Repository navigation
NpEnv removal P1: tensor-native Manager-Based TorchEnv runtime #1703
Description
Activity
P1 first slice is merged in PR #1708 (merge commit 6aaf0c4): durable
TorchEnv/TorchEnvState, tensor step/autoreset/final-observation/training-state contracts, capability/device/fail-closed validation, shared timing schema owner, and focused lifecycle tests. #1703 remains open becauseManagerBasedRlEnvstill derives fromNpEnv; the next slice must switch its lifecycle toTorchEnvthrough the explicit NumPy Manager host boundary and migrate its tests.Core P1 lifecycle migration is merged in PR #1709 (merge commit 1bfaacb), with UniSim public-width dependency unisim#322 pinned at c63f5562.
ManagerBasedRlEnvnow inheritsTorchEnv; public actions/reset rows/state/backend control are Torch, and its temporary Manager term execution uses named NumPy host boundaries. Direct RSL-RL is tensor-first and the G1 direct runtime no longer retains_cpu_env. #1703 remains open for device-resident Manager term compute, RNG provenance, and completing the all-tensor API boundary before broad task parity.First Manager term-family carrier migration is merged in PR #1710 (merge commit c7bf511): MetricsManager accumulators/counters/substep buffers are Torch tensors on the environment device, NumPy term results cross one explicit host boundary, Torch term results remain on-device, and reset logging is the bounded host publication. ADR-0011 also updates the Manager package import boundary to permit its target Torch carrier while continuing to reject mjlab at runtime. #1703 remains open for observation/reward/termination/command/event carriers and RNG provenance.
Reward and termination Manager carriers are now Torch tensors after PR #1711 (merge commit 49bd2b3). Reward aggregation/per-term rates and termination failure/timeout/per-term buffers run on env.device; already-tensor metric/reward/termination terms require exact dtype/device, legacy NumPy terms cross the named temporary boundary, and reset/logging remain bounded host publication points. #1703 remains open for action, observation, command, event, Entity state/reset APIs, and RNG provenance.
ActionManager top-level history is now Torch after PR #1712 (merge commit e5ed471).
action,prev_action, andprev_prev_actionlive on env.device; process input is strict contiguous float32 exact-device Torch; selected-row reset/history semantics are preserved. Existing ActionTerms temporarily receive one full-batch NumPy copy through the named term boundary. Action-rate/acceleration rewards consume Torch history, whilelast_actionexplicitly publishes NumPy to the still-host ObservationManager. #1703 remains open because BaseAction/Entity control writes, observation/command/event carriers, Entity tensor reads, reset tensor commits, and RNG provenance are pending.ObservationManager public output carrier is now Torch after PR #1713 (merge commit ae2db05).
compute()/compute_group()publish contiguous float32 tensors on env.device, the cache and Managerobs_bufstay Torch, andTorchEnvState.obsconsumes them without a per-step Torch↔NumPy round trip. Terms, clip/scale, NaN handling, noise, delay, and history still run their explicit NumPy pipeline; temporal buffer migration remains open.Merged #1715 (
099a1663): BaseAction raw/processed buffers, cold affine/clip tensors, action capability routing, and rough/motion action math now execute on-device. The remaining Entity write boundary is intentionally explicit: apply still publishes one host copy atapply_actions(), and encoder/default/current-joint-pos math plus Motion.targetremain migration-only NumPy boundaries.Merged #1716 (
9749919a): AllegroIncrementalPositionAction now declares the tensor action path, keeps raw/clipped/integrated targets and bounds as float32 tensors on env.device, and executes clipping/integration/limiting/selected-row reset on-device. Entity writes and the still-NumPy Allegro observation remain explicit migration-only host publications. Next blocker for Stewart/state-feedback actions is a scene-owned device-aware state/sensor read facade with phase-scoped consistency.Merged #1717 (
7af18e44): added the publicEntity.joint_tensor_view(device)facade. It resolves entity joints into backend qpos/qvel columns, validates capabilities/device/layout/dtype/finiteness, and returns entity-ordered Torch state.RelativeJointPositionActionnow consumes it for on-device state-feedback math before the explicit Entity host write. Next: extend the facade to named sensors/body state and migrate Stewart, then observation non-temporal processing.Merged #1719 (
cf857971): ObservationManager now accepts exact-device float32 Torch terms and executes non-temporal clip/scale/NaN handling/row selection/concatenation with Torch on env.device. Scale tensors are cold-path persistent. NumPy terms retain the existing host boundary; noise/delay/history remain explicitly NumPy migration paths. Remaining priorities: named sensor/body tensor facade, Stewart action, command/event/reset tensor transactions, then legacy NpEnv deletion.Summary
Completes the next #1703 P1 slice on top of #1720:
- Adds a scene-owned packed tensor read plan phase boundary and wires state-feedback actions to sensor/body tensor reads.
- Refreshes the packed scene packet across Manager phases and after bounded in-phase mutations.
- Pairs selected-row reset with one public
HostBridgeTransferPlan.apply_reset()commit and one selected post-reset packed H2D read where native/public qpos and qvel widths agree. MotrixSim's current native/public reset-width mismatch explicitly falls back to the existing publicset_state()boundary rather than guessing a private projection. - Makes
ManagerBasedRlEnv._controlthe authoritative float32 Torch tensor. Entity control writes can now publish directly into that tensor; mapped-entity reset control restoration uses the publicctrlstate view. The Stewart action no longer detaches to NumPy at its Entity write boundary.
Validation
UNILAB_LOCAL_UNISIM=/home/user/ws/unilabsim/unisim uv run pytest -q \ tests/base tests/envs/mdp tests/envs/test_manager_based_rl_env.py \ tests/envs/test_stewart.py tests/envs/test_multi_entity_consumer.py \ tests/tasks/test_tensor_runtime_components.py 834 passed, 19 skipped, 219 deselected make check passed (mypy clean; pyright 0 errors, existing optional drake import warning) make test-all 1831 passed, 26 skipped, 617 deselected benchmark module-mode 35/35; script-mode 36/36Boundary notes
- Mapped entity patches, reset randomization, mocap writes, and CUDA-device packed reset commits still fail closed and remain explicit migration boundaries.
- MotrixSim packed reset requires a UniSim follow-up to unify native/public reset widths. This PR intentionally does not implement that projection in UniLab.
Closes nothing; #1703 remains open for observation term consumption, temporal observation paths, command/event transactions, and RNG provenance.
PR #1721 当前 head
8e2330b4的 CI 已完成:Focused CPU tensor profile: SUCCESSGated single-GPU tensor profile: SKIPPED(按预期,无 gated GPU runner)- PR 状态:OPEN / MERGEABLE,等待 review;遵守 review-first,不 self-merge。
PR #1721 已 review 合入
develop/tensor-runtime(merge boundary:1ca0ab5b)。已在 #1721 的进度评论基础上完成该 merge 边界的追踪。继续推进下一个 P1 slice:temporal observation buffers tensorization,PR #1722:
CircularBuffer现在以 Torch tensor 持有 history 和 per-batch counters,设备绑定到首次 append 的 observation 设备;DelayBuffer.append/backfill/compute/peek返回/消费 Torch tensor,lag RNG 仍保持 env-owned NumPy RNG,避免本 slice 改变 RNG stream;ObservationManager删除 delay/history 前的obs.detach().cpu().numpy()host detour;- 新增 no-host-detour 测试覆盖 delay + history 的 device/dtype/语义。
本地验证:
focused: 202 passed, 1 deselected make check: passed make test-all: 1832 passed, 26 skipped; benchmark 35/35 + 36/36剩余 observation boundary:
NoiseCfg/NoiseModel仍为 NumPy owner;这是下一个需要和 RNG provenance 一起处理的边界。Command manager boundary slice is ready in PR #1723 (
feat/issue-1703-command-tensors, headb6ef45e7):CommandTerm.time_left/command_counterare runtime-device Torch tensors; scalar, NumPy row-wise, and Torch row-wisedtare supported, while resampling still consumes env-owned NumPy RNG so this slice does not change the RNG stream.CommandManager.get_command()now returns a validated device tensor. Existing NumPy command terms cross one explicit conversion boundary; locomotion command reward consumers convert at their own named NumPy boundary.MotionCommandmetrics and its Numba kernel remain explicitly NumPy-owned and fail closed if that carrier changes prematurely.- Local gates:
make checkpass; focused suite 221 passed / 37 deselected;make test-all1833 passed / 26 skipped and benchmark imports 35/35 + 36/36.
Remaining #1703 priorities are unchanged: observation noise/RNG provenance, command term internals, event tensor transactions, selected-row reset semantics, and MotrixSim native/public reset-width follow-up in UniSim.
EventManager scheduling-carrier slice is ready as stacked PR #1724 (
feat/issue-1703-event-scheduling-tensors, head83538abe, base #1722):- interval
time_left, reset-throttling step IDs, and trigger-once flags are now runtime-device Torch tensors (float64,int64,bool); - all interval/reset sampling still uses env-owned NumPy RNG, so the RNG stream is unchanged in this slice;
- event terms and randomization transactions remain at their existing explicit NumPy term boundary.
Local gates: focused 124 passed;
make checkpass;make test-all1832 passed / 26 skipped + benchmark imports 35/35 and 36/36.- interval
PR #1724 CPU validation is confirmed: manually dispatched
Tensor Runtime CIon head83538abecompleted successfully (Focused CPU tensor profile, run 36556322932). The gated single-GPU job was not requested. PR remains OPEN/MERGEABLE awaiting review.Observation noise carrier slice is ready as stacked PR #1740 (
feat/issue-1703-observation-noise-tensors, headb92714d8, base #1722): Constant/Uniform/Gaussian noise configs and additive-bias noise models accept NumPy or Torch observations and preserve the input carrier/device; ObservationManager removes the pre-noise.detach().cpu().numpy()full-observation detour. Env-owned NumPy RNG remains authoritative and consumes the same stream; tensor random noise performs one sampled unit-buffer H2D copy, while observation arithmetic/output stays on-device. Additive bias state materializes on the active observation device. Local gates: focused noise/partial-reset 30 passed; focused observation buffers 33 passed; related manager/env/base suites 189 passed;make checkpass;make test1841 passed / 26 skipped; slow train selection has only the unchanged unisim#291 failure. Tensor Runtime CI run 36571563620 is green.PR #1740 closed automatically when its base #1722 merged. It has been recreated as direct PR #1741 (
feat/issue-1703-observation-noise-tensors, rebased head13a9cc3e, basedevelop/tensor-runtime) with the same changes and validation summary. Retargeted-head Tensor Runtime CI run 36573120024 is dispatched.PR #1741 retargeted-head validation is green: Focused CPU profile passed; gated single-GPU profile skipped as expected (run 36573120024). PR is OPEN/MERGEABLE awaiting review.
Merge-boundary update: #1722/#1723/#1724 have been review-merged into
develop/tensor-runtime. This closes the temporal observation carrier, command carrier, and event scheduling-carrier slices. The next #1703 P1 slice is direct PR #1741 (feat/issue-1703-observation-noise-tensors, head13a9cc3e), which removes the remaining observation-noise full-observation host detour while keeping the env-owned NumPy RNG authoritative. Remaining #1703 priorities: event term/randomization transactions, command term internals, RNG provenance formalization, selected-row reset semantics, and representative CUDA/CPU backend gates in P3.P1 merge update: observation temporal buffers (#1722), command carriers (#1723), event scheduling carriers (#1724), and observation noise carriers (#1741) are merged. Remaining #1703 implementation work is concentrated in event term/randomization tensor transactions, command term internals, explicit RNG provenance, and selected-row reset semantics; representative backend validation then belongs to #1705.
Parent: #1701
Problem
ManagerBasedRlEnvcurrently derives fromNpEnv, while the reviewed tensor lifecycle is task-owned and scoped. RemovingNpEnvrequires the Manager-Based runtime itself to become tensor-native.Outcome
Introduce the durable
TorchEnvbase and migrate the Manager-Based runtime lifecycle to Torch without relying on a new compatibility adapter.Work
TorchEnvandTorchEnvState.SimBackendtensor APIs and fail closed on unsupported capabilities.Acceptance
NpEnvbase for normal execution.