Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -61,14 +61,16 @@ owns its explicit packed conversion boundary. MuJoCo remains a production backen
canonical in-process host-bridge implementation; it is not a test-only fallback.

`env.tensor_runtime` and `env.tensor_runtime_device` are removed. Tensor execution is an invariant
of the Manager runtime. Placement is derived from declared backend data plane and rank topology.
No compatibility alias, dual-mode field, or hidden fallback is retained.

For a `HOST_BRIDGE` backend, CPU-authoritative physics does not imply CPU Manager tensors.
When PyTorch exposes an available GPU and the backend accepts its current `cuda` ordinal,
the Manager uses that GPU, including ROCm PyTorch's `cuda` namespace. CPU-only capabilities
or hosts without an available GPU retain CPU placement. `DEVICE_RESIDENT` adapters keep
their own CUDA-only platform requirements; accepting HIP Torch buffers on a host bridge
of the Manager runtime. Placement is derived from the declared backend data plane and, for a
`HOST_BRIDGE`, the process's explicit carrier request. No NumPy/dual-runtime compatibility alias
or hidden fallback is retained.

For a `HOST_BRIDGE` backend, CPU-authoritative physics keeps the default Manager/TorchEnv
carriers on CPU. A training process may explicitly request `cuda`/`cuda:<ordinal>` through
`manager_torch_device` (routed from its resolved learner device), including ROCm PyTorch's
`cuda` namespace; UniLab validates the request against the backend's declared Torch devices.
GPU visibility alone is not a placement request. `DEVICE_RESIDENT` adapters keep their own
CUDA-only placement and platform requirements; accepting HIP Torch buffers on a host bridge
does not make those physics engines ROCm-compatible.

### Manager-owned RNG
Expand Down
11 changes: 7 additions & 4 deletions docs/sphinx/source/en/1-getting_started/2-installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -243,10 +243,13 @@ ROCm notes:
profile, use `uv run --no-sync` (or `UV_NO_SYNC=1 make ...`) for validation;
automatic synchronization would reinstall that profile's CUDA wheels.
- With ROCm PyTorch and an available GPU, a `HOST_BRIDGE` backend accepting
`cuda` buffers runs the Manager/TorchEnv tensors on the current GPU. MuJoCo
physics remains on CPU; packed transfers connect it to GPU observations,
actions, rewards, and resets. CUDA-only physics backends remain unsupported
on ROCm. See {doc}`/adr/ADR-0012-sole-tensor-manager-and-scoped-backends`.
`cuda` buffers can use the current GPU for Manager/TorchEnv tensors when the
training process explicitly requests that learner device (which routes
`manager_torch_device`). The direct environment default remains CPU. MuJoCo
physics stays on CPU; packed transfers connect it to explicitly requested GPU
observations, actions, rewards, and resets. CUDA-only physics backends remain
unsupported on ROCm. See
{doc}`/adr/ADR-0012-sole-tensor-manager-and-scoped-backends`.
- When installing from PyPI instead of a source checkout, `make sync-rocm` does
not apply. Install the torch build validated by the repository from the
PyTorch ROCm index first, then `unilab`. The published dependency range is
Expand Down
17 changes: 9 additions & 8 deletions docs/sphinx/source/en/5-reference/5-support_matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,12 +74,12 @@ This table is derived from UniSim's SDK-free public static inventory. It describ

| Backend | Execution / process / data plane | Torch devices | CUDA runtime | Linux+CUDA | macOS | ROCm | Worker | Reset randomization | Fixed variants | Host callbacks | Packed bridge |
|---|---|---|---|---|---|---|---|---|---|---|---|
| `mujoco` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`) | In-process; no external Python worker | unknown | unknown | Unsupported | Exact |
| `motrix` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`) | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Exact |
| `drake` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`) | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Exact |
| `mujoco` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`) | In-process; no external Python worker | unknown | unknown | Unsupported | Exact |
| `motrix` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`) | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Exact |
| `drake` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`) | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Exact |
| `mjwarp` | Device-resident / in-process / direct | CUDA | Required for the entire tensor lifecycle | Supported: Linux CUDA only | Unsupported; no CPU, MPS, or ROCm fallback | Unsupported; no CPU, MPS, or ROCm fallback | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Unsupported |
| `newton` | Device-resident / in-process / direct | CUDA | Required for the entire tensor lifecycle | Supported: Linux CUDA only | Unsupported; no CPU, MPS, or ROCm fallback | Unsupported; no CPU, MPS, or ROCm fallback | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Unsupported |
| `superdex` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`) | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Exact |
| `superdex` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`) | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Exact |
| `genesis` | Device-resident / in-process / direct | CUDA | Required for the entire tensor lifecycle | Supported: Linux CUDA only | Unsupported; no CPU, MPS, or ROCm fallback | Unsupported; no CPU, MPS, or ROCm fallback | In-process; no external Python worker | Unsupported | Unsupported | Unsupported | Unsupported |

### Entrypoint x Task Owner
Expand Down Expand Up @@ -125,7 +125,8 @@ configuration. If `CUDA_VISIBLE_DEVICES` is set, backend ordinals address that
remapped namespace, not host-global physical indices.

On macOS and ROCm, use a CPU-authoritative host-bridge backend. On ROCm,
Manager/TorchEnv uses the current GPU when ROCm PyTorch is available and the
backend accepts that `cuda` device; CPU physics does not require CPU Manager
tensors. GPU Torch buffers on a host bridge do not imply GPU physics or a
device-resident backend lifecycle.
Manager/TorchEnv uses the current GPU only when the training process explicitly
requests that learner device and the backend accepts the routed `cuda`
`manager_torch_device`; direct environment construction remains CPU. GPU Torch
buffers on a host bridge do not imply GPU physics or a device-resident backend
lifecycle.
10 changes: 6 additions & 4 deletions docs/sphinx/source/zh_CN/1-getting_started/2-installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -218,10 +218,12 @@ ROCm 说明:
- 如果手工安装了 ROCm wheel,但项目仍使用默认依赖配置档,验证时使用
`uv run --no-sync`(或 `UV_NO_SYNC=1 make ...`);自动同步会重新安装默认配置档的
CUDA wheel。
- 安装 ROCm PyTorch 且 GPU 可用时,接受 `cuda` buffer 的 `HOST_BRIDGE` 后端会让
Manager/TorchEnv tensor 使用当前 GPU。MuJoCo 物理仿真仍在 CPU 上,通过 packed
传输连接 GPU 上的 observation、action、reward 与 reset。CUDA-only 物理后端仍不
支持 ROCm。见 {doc}`/adr/ADR-0012-sole-tensor-manager-and-scoped-backends`。
- 安装 ROCm PyTorch 且 GPU 可用时,接受 `cuda` buffer 的 `HOST_BRIDGE` 后端只有在
训练进程显式请求该 learner 设备(并路由 `manager_torch_device`)时才会让
Manager/TorchEnv tensor 使用当前 GPU;直接构造环境的默认仍是 CPU。MuJoCo 物理
仿真仍在 CPU 上,通过 packed 传输连接显式请求的 GPU observation、action、reward
与 reset。CUDA-only 物理后端仍不支持 ROCm。见
{doc}`/adr/ADR-0012-sole-tensor-manager-and-scoped-backends`。
- 从 PyPI 安装(不克隆仓库)时,`make sync-rocm` 不适用;先从 PyTorch ROCm 索引
安装仓库验证过的 torch build,再安装 `unilab`。发布的依赖范围是
`torch>=2.9,<2.15`,pip 会保留已安装的 ROCm build,不会替换为 CUDA wheel:
Expand Down
17 changes: 9 additions & 8 deletions docs/sphinx/source/zh_CN/5-reference/5-support_matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,12 +65,12 @@ uv run scripts/generate_support_matrix.py --write

| Backend | Execution / process / data plane | Torch devices | CUDA runtime | Linux+CUDA | macOS | ROCm | Worker | Reset randomization | Fixed variants | Host callbacks | Packed bridge |
|---|---|---|---|---|---|---|---|---|---|---|---|
| `mujoco` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`) | In-process; no external Python worker | unknown | unknown | 不支持 | 支持 |
| `motrix` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`) | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 支持 |
| `drake` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`) | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 支持 |
| `mujoco` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`) | In-process; no external Python worker | unknown | unknown | 不支持 | 支持 |
| `motrix` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`) | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 支持 |
| `drake` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`) | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 支持 |
| `mjwarp` | Device-resident / in-process / direct | CUDA | Required for the entire tensor lifecycle | Supported: Linux CUDA only | Unsupported; no CPU, MPS, or ROCm fallback | Unsupported; no CPU, MPS, or ROCm fallback | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 不支持 |
| `newton` | Device-resident / in-process / direct | CUDA | Required for the entire tensor lifecycle | Supported: Linux CUDA only | Unsupported; no CPU, MPS, or ROCm fallback | Unsupported; no CPU, MPS, or ROCm fallback | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 不支持 |
| `superdex` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`) | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 支持 |
| `superdex` | Host bridge / in-process / host bridge | CPU / CUDA | Required only when the learner requests CUDA state/control buffers | Supported: CPU-authoritative physics with optional CUDA Torch buffers | CPU-authoritative host bridge only; no CUDA physics claim | CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`) | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 支持 |
| `genesis` | Device-resident / in-process / direct | CUDA | Required for the entire tensor lifecycle | Supported: Linux CUDA only | Unsupported; no CPU, MPS, or ROCm fallback | Unsupported; no CPU, MPS, or ROCm fallback | In-process; no external Python worker | 不支持 | 不支持 | 不支持 | 不支持 |

### Entrypoint x Task Owner
Expand Down Expand Up @@ -114,7 +114,8 @@ NVIDIA 驱动与容器运行时不匹配,再调整任务配置。设置了
`CUDA_VISIBLE_DEVICES` 时,后端序号指向重映射后的命名空间,而不是宿主机
全局物理索引。

在 macOS 与 ROCm 上请使用 CPU-authoritative host-bridge 后端。在 ROCm 上,
ROCm PyTorch 的 GPU 可用且后端接受当前 `cuda` 设备时,Manager/TorchEnv 使用
当前 GPU;CPU 物理仿真不要求 Manager tensor 使用 CPU。host bridge 使用 GPU
Torch buffer 不代表 GPU physics,也不代表 device-resident 后端生命周期。
在 macOS 与 ROCm 上请使用 CPU-authoritative host-bridge 后端。在 ROCm 上,只有
训练进程显式请求该 learner 设备且后端接受被路由的 `cuda`
`manager_torch_device` 时,Manager/TorchEnv 才使用当前 GPU;直接构造环境仍为
CPU。host bridge 使用 GPU Torch buffer 不代表 GPU physics,也不代表
device-resident 后端生命周期。
7 changes: 4 additions & 3 deletions scripts/generate_support_matrix.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,16 +25,17 @@

_SUPPORT_LEVEL_ENUM = re.compile(r"\bSupportLevel\.([A-Z][A-Z_]+)\b")

_ROCM_HOST_BRIDGE_LABEL = "Supported: CPU-authoritative physics with ROCm Torch buffers (`cuda`)"
_ROCM_HOST_BRIDGE_LABEL = "CPU-authoritative host bridge with explicit ROCm Torch buffers (`cuda`)"


def _with_support_level_values(content: str) -> str:
"""Render hardened UniSim lifecycle enum values as their public strings."""
rendered = _SUPPORT_LEVEL_ENUM.sub(lambda match: match.group(1).lower(), content)
# UniSim's static inventory is deliberately conservative: it only states
# that no ROCm-native CUDA-only fallback exists. Host bridges accepting
# `cuda` Torch buffers also accept ROCm PyTorch's CUDA-namespace API while
# physics remains CPU-authoritative; make that repository guarantee visible.
# `cuda` Torch buffers also accept ROCm PyTorch's CUDA-namespace API,
# but only when the process explicitly requests those Manager carriers;
# make that opt-in repository guarantee visible.
return rendered.replace(
"CPU-authoritative host bridge only; no ROCm CUDA-only fallback",
_ROCM_HOST_BRIDGE_LABEL,
Expand Down
40 changes: 38 additions & 2 deletions src/unilab/base/process_device.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,8 +16,9 @@
from typing import Any, cast

# These backends consume an explicit integer device id while materializing
# their simulator. MuJoCo/Motrix/Drake either run on the host or own their
# device selection internally and must not receive a synthetic override.
# their simulator. MuJoCo/Motrix/Drake run CPU-authoritative physics; their
# optional accelerator Manager buffers are selected separately through
# ``manager_torch_device``.
BACKEND_ENV_DEVICE_FIELDS: dict[str, str] = {
"isaacgym": "isaacgym_device_id",
"isaacsim": "isaacsim_device_id",
Expand All @@ -41,6 +42,11 @@
# integer payload sent to that worker before construction.
_EXTERNAL_CUDA_IPC_BACKENDS = {"isaacgym", "isaacsim"}

# HOST_BRIDGE backends accept optional accelerator Torch carriers across their packed
# boundary. A process may explicitly request those carriers on its learner Torch
# device; unset non-accelerator requests retain the owner/default CPU placement.
_HOST_BRIDGE_TORCH_BACKENDS = {"mujoco", "motrix", "drake", "superdex"}


# Set once ``bind_genesis_process_device`` has pinned CUDA_VISIBLE_DEVICES for
# this process. Genesis/Quadrants binds its CUDA runtime to the first visible
Expand Down Expand Up @@ -170,6 +176,35 @@ def apply_backend_env_device_override(
return result


def apply_manager_torch_device_override(
env_cfg_override: Mapping[str, Any] | None,
backend_type: str,
*,
learner_device: str | None = None,
) -> dict[str, Any]:
"""Return an env override selecting optional HOST_BRIDGE Torch carriers.

The input mapping is never mutated. Explicit CPU requests force CPU carriers.
Unset and non-CUDA learner requests retain owner/default placement. A CUDA request
is validated against backend capabilities when the environment binds.
DEVICE_RESIDENT backends own placement and never receive this synthetic field.
"""

result = dict(env_cfg_override) if env_cfg_override is not None else {}
if _normalize_backend(backend_type) not in _HOST_BRIDGE_TORCH_BACKENDS:
return result
if learner_device is None:
return result
device = str(learner_device).strip()
if not device or device.lower() == "cpu":
result["manager_torch_device"] = "cpu"
return result
if device.split(":", 1)[0].lower() != "cuda":
return result
result["manager_torch_device"] = device
return result


def resolve_backend_process_device(backend_type: str, learner_device: str | None) -> str | None:
backend = _normalize_backend(backend_type)
if (
Expand Down Expand Up @@ -399,6 +434,7 @@ def _reset_genesis_device_pin_for_tests() -> None:
__all__ = [
"BACKEND_ENV_DEVICE_FIELDS",
"apply_backend_env_device_override",
"apply_manager_torch_device_override",
"bind_backend_process_device",
"bind_backend_process_device_for_backend",
"bind_genesis_process_device",
Expand Down
Loading
Loading