Skip to content

fix(runtime): honor host-bridge manager Torch devices - #2084

Merged
TATP-233 merged 1 commit into
mainfrom
fix/host-bridge-manager-tensors
Oct 8, 2026
Merged

TATP-233 merged 1 commit into
mainfrom
fix/host-bridge-manager-tensors

Conversation

@TATP-233

@TATP-233 TATP-233 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • Derive ManagerBasedRlEnv/TorchEnv placement from the backend's declared HOST_BRIDGE Torch-device capability. A capable host bridge now uses the current GPU (including ROCm PyTorch's cuda namespace); CPU-only capabilities stay on CPU even when a GPU is visible.
  • Keep locomotion gait terms, selected packed resets, policy I/O, and related tests on the negotiated Manager device. Packed tensor-reset event commits now honor the host bridge's declared device contract instead of forcing selected rows through CPU first.
  • Update bilingual installation/support documentation and ADR-0012/ADR-0006. The support-matrix generator and its slow test now express that CPU-authoritative host bridges support ROCm Torch buffers without making a CUDA-physics claim.
  • Correct the optional Drake startup-boundary assertion to accept the current DrakeEnvPool batch extension wording emitted by this UniSim profile.

CPU-authoritative physics is unchanged. This affects MuJoCo/Motrix/Drake/SuperDex host-bridge execution on Linux CUDA/ROCm and CPU; device-resident backends retain their existing CUDA-only lifecycle.

Linked Work

Validation

  • make test-all passed on the final local head before this PR was created or updated
  • Additional task-specific validation listed below

Commands actually run:

UV_NO_SYNC=1 uv sync --all-extras --dev
UV_NO_SYNC=1 make check
UV_NO_SYNC=1 uv run pytest -q \
  tests/envs/test_manager_based_rl_env.py \
  tests/envs/locomotion/test_manager_gait_terms.py \
  tests/envs/locomotion/go2/test_manager_based_cfg.py \
  tests/base/test_entity_scene_consumer.py \
  tests/managers/test_isaac_lab_migration_fixture.py \
  tests/managers/test_mjlab_migration_fixture.py \
  tests/managers/test_stage_curriculum_demo.py
UV_NO_SYNC=1 uv run pytest -q \
  tests/envs/test_go2_superdex.py \
  tests/envs/locomotion/go2/test_manager_based_cfg.py::test_go2_flat_drake_factory_is_registered_before_optional_runtime_load \
  tests/scripts/test_check_docs.py::test_documentation_files_match_current_repo_contracts
UV_NO_SYNC=1 uv run scripts/generate_support_matrix.py --write
UV_NO_SYNC=1 uv run pytest -q -m slow tests/scripts/test_support_matrix.py
UV_NO_SYNC=1 make test-all

Final local result:

  • make check: passed (ruff format, ruff check, mypy, pyright, focused test lint).
  • make test-all: passed (1854 passed, 17 skipped, 489 deselected, benchmark smoke imports 34/34 module-mode and 35/35 script-mode).
  • Focused GPU coverage exercised CUDA host-bridge placement and SuperDex CUDA packed resets on the local GPU. CPU-only capability fallback is also covered.

Remote CI route:

  • Base main: current-head remote CI requested below; status will be recorded after it reports.

Impact

  • Backend impact: mujoco / motrix / drake / superdex (HOST_BRIDGE only; device-resident paths unchanged)
  • Platform impact: Linux (CPU, CUDA, ROCm) / macOS CPU-authoritative host bridge; no new CUDA-only physics claim
  • Training effect expected: yes — GPU host-bridge Manager/policy I/O stays on the negotiated GPU instead of being forced to CPU; dimensions and policy contracts are unchanged

Artifacts

  • W&B: N/A
  • benchmark result: N/A (benchmark smoke included in make test-all)
  • video / screenshot: N/A
  • ONNX / checkpoint: N/A

Checklist

  • Added or updated tests where needed
  • Updated docs if behavior or workflow changed
  • Linked the driving issue
  • Noted any follow-up work explicitly

No unresolved follow-up is known. The remote CUDA/ROCm-specific hardware coverage remains dependent on CI runners.

@TATP-233

TATP-233 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

CI link for the current head: https://github.com/Motphys/UniLab/actions/runs/37776382218

@TATP-233
TATP-233 force-pushed the fix/host-bridge-manager-tensors branch from 9551cb2 to fa89b76 Compare October 8, 2026 12:24
@TATP-233

TATP-233 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Final-head CI for fa89b76e: https://github.com/Motphys/UniLab/actions/runs/37779801809

@TATP-233

TATP-233 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Correction: final-head CI run is https://github.com/Motphys/UniLab/actions/runs/37776693104 for fa89b76e.

@TATP-233
TATP-233 force-pushed the fix/host-bridge-manager-tensors branch from fa89b76 to c3e0d59 Compare October 8, 2026 12:35
@TATP-233

TATP-233 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Current-head CI for c3e0d593: https://github.com/Motphys/UniLab/actions/runs/37778027767 (previous CPU-only-Torch CI failure was fixed).

@TATP-233
TATP-233 force-pushed the fix/host-bridge-manager-tensors branch from c3e0d59 to ae7396b Compare October 8, 2026 12:47
@TATP-233

TATP-233 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Current-head CI for ae7396b3: https://github.com/Motphys/UniLab/actions/runs/37779358883. CPU-only Torch builds now skip the fake-GPU negative branch.

@TATP-233

TATP-233 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

@TATP-233
TATP-233 merged commit 5e7114f into main Oct 8, 2026
8 checks passed
@TATP-233
TATP-233 deleted the fix/host-bridge-manager-tensors branch October 8, 2026 12:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant