Skip to content

feat(appo): restore Go2 Motrix owner on tensor runtime - #2077

Merged
TATP-233 merged 1 commit into
mainfrom
task/update-unilab-rl-1.4.8
Oct 8, 2026
Merged

TATP-233 merged 1 commit into
mainfrom
task/update-unilab-rl-1.4.8

Conversation

@TATP-233

@TATP-233 TATP-233 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • restore appo/go2_joystick_flat/motrix as a configured tensor-runtime owner
  • register Go2JoystickFlat for Motrix
  • upgrade the optional/dev unilab-rl dependency to 1.4.8
  • follow the upstream SAC package rename (fast_sac → sac, FastSACLearner → SACLearner)
  • route APPO through the same tensor EnvFactory as off-policy training, with the tensor NaN-guard factory
  • regenerate bilingual support-matrix evidence and update Motrix scope docs
  • add a bounded Motrix APPO slow smoke that asserts model_1.pt

Dependency changes

unilab-rl 1.4.8 includes:

  • tensor-native APPO collector, rollout IPC, and staging (upstream chore: set up github collaboration workflow #84)
  • Torch action/reset/state env contract
  • Gaussian std metrics derived directly from distribution parameters
  • CPU-authoritative env + CUDA collector bootstrap correction fix
  • SAC package/learner rename with no compatibility layer

This PR removes the need for a consumer-side APPO NumPy adapter or custom MLPModel subclass.

Validation

Local gates:

make check
# passed

make test
# 1893 passed, 28 skipped, 552 deselected

make test-all
# passed, including benchmark smoke 35/35 modules and 36/36 entrypoints

Focused runtime tests:

uv run pytest -q \
  tests/scripts/test_train_script_configs.py \
  tests/algos/test_appo_runner.py \
  tests/algos/test_offpolicy_double_buffer_runner.py \
  tests/algos/test_offpolicy_dp_sync.py \
  -m 'slow or not slow'
# 70 passed, 2 skipped

Bounded real training through the public CLI route:

MuJoCo:
  status=completed
  completed_iterations=1
  total_env_steps=40
  checkpoint=model_1.pt

Motrix:
  status=completed
  completed_iterations=1
  total_env_steps=16
  checkpoint=model_1.pt

Both runs use the published unilab-rl==1.4.8 package.

@TATP-233
TATP-233 requested a review from caozx1110 as a code owner October 8, 2026 11:30
Restore the APPO Go2 flat Motrix owner using the canonical tensor reset
events and register the task backend. Upgrade unilab-rl to 1.4.8 for the
tensor-native APPO collector and parameter-backed Gaussian metrics, and
follow the upstream SAC package rename.

Commands routed through appo/go2_joystick_flat/motrix now train without a
consumer-side NumPy adapter or model workaround. Bounded MuJoCo and Motrix
smokes produce model_1.pt and update the generated support evidence.
@TATP-233
TATP-233 force-pushed the task/update-unilab-rl-1.4.8 branch from c467e54 to acba2f3 Compare October 8, 2026 11:35
@TATP-233
TATP-233 merged commit 215341e into main Oct 8, 2026
8 checks passed
@TATP-233
TATP-233 deleted the task/update-unilab-rl-1.4.8 branch October 8, 2026 11:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant