Skip to content

fix(training): keep MuJoCo collector tensors on CPU by default - #2101

Merged
TATP-233 merged 1 commit into
mainfrom
fix/issue-2100-mujoco-collector-tensor-device
Oct 9, 2026
Merged

TATP-233 merged 1 commit into
mainfrom
fix/issue-2100-mujoco-collector-tensor-device

Conversation

@TATP-233

@TATP-233 TATP-233 commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Added training.collector_tensor_device: cpu | cuda to the SAC, FlashSAC, PPO, and APPO training owners, with cpu as the default.
  • Routed that explicit choice into the existing Manager owner field (manager_torch_device). An unindexed cuda request is resolved by the environment to the current rank-local Torch CUDA ordinal.
  • Prevents automatic learner CUDA selection from implicitly moving a MuJoCo HOST_BRIDGE collector/Manager carrier to CUDA. Non-MuJoCo CUDA requests fail closed; DEVICE_RESIDENT backends continue to own their GPU placement.
  • Updated ADR-0012 and the maintained English/Chinese installation and support-matrix docs.

Linked Work

Validation

  • make test-all passed on the final local head before this PR was created
  • Additional task-specific validation listed below

Commands actually run:

uv run pytest tests/base/backend/test_process_device.py
# 32 passed

uv run pytest tests/envs/test_manager_based_rl_env.py -k 'host_bridge or manager_tensor'
# 6 passed, 80 deselected

uv run pytest tests/scripts/test_train_scripts.py -k 'collector_tensor_device or host_bridge'
# Exit code 5: 137 collected, 137 deselected. This file is entirely marked slow by existing config.

uv run pytest --override-ini "addopts=--tb=short" tests/scripts/test_train_scripts.py -k 'collector_tensor_device or host_bridge'
# 7 passed, 130 deselected

make check
# Passed: ruff format/check, mypy, pyright, and focused test lint

make test
# 1857 passed, 43 skipped, 495 deselected

make test-all
# Passed, including 1857 passed / 43 skipped and benchmark import smoke 34/34 module-mode and 35/35 script-mode

Complete default SAC validation:

uv run train --algo sac --task g1_walk_flat --sim mujoco \
  algo.max_iterations=160 training.no_play=true training.export_onnx=false \
  training.nan_guard.enabled=false training.logger=tensorboard

Result:

  • Completed 160/160 iterations.
  • run_summary.json tail FPS: 93,873.5; TensorBoard last-80 median FPS: 97,666.7.
  • TensorBoard last-80 median apply_action: 0.279 ms.
  • Manifest: collector_tensor_native=false, env_public_device=cpu, CPU inference transport, and learner_device=cuda:0.
  • Local artifact: logs/sac/G1WalkFlat/2026-10-09_13-49-45_mujoco/run_summary.json.

Explicit CUDA smoke:

uv run train --algo sac --task g1_walk_flat --sim mujoco \
  training.collector_tensor_device=cuda algo.max_iterations=10 \
  training.no_play=true training.export_onnx=false \
  training.nan_guard.enabled=false training.logger=tensorboard

Result: completed 10/10 iterations with collector_tensor_native=true, env_public_device=cuda:0, and CUDA inference transport.

Remote CI route:

Impact

  • Backend impact: mujoco
  • Platform impact: Linux CUDA training behavior; docs also cover ROCm's CUDA Torch namespace
  • Training effect expected: yes — MuJoCo defaults to CPU Manager/collector carriers while retaining an explicit CUDA opt-in

Artifacts

  • W&B: none
  • benchmark result: local run_summary.json files cited above
  • video / screenshot: not applicable
  • ONNX / checkpoint: local smoke checkpoints were disabled or incidental and are not attached

Checklist

  • Added or updated tests where needed
  • Updated docs if behavior or workflow changed
  • Linked the driving issue
  • Noted any follow-up work explicitly

Follow-up: none required by this scoped change. Motrix/Drake/Superdex remain intentionally outside the #2100 backend boundary.

@TATP-233
TATP-233 requested a review from caozx1110 as a code owner October 9, 2026 06:26
@TATP-233
TATP-233 merged commit cbafb50 into main Oct 9, 2026
8 checks passed
@TATP-233
TATP-233 deleted the fix/issue-2100-mujoco-collector-tensor-device branch October 9, 2026 06:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant