Skip to content

NpEnv removal P1: tensor-native Manager-Based TorchEnv runtime #1703

Description

@TATP-233

Parent: #1701

Problem

ManagerBasedRlEnv currently derives from NpEnv, while the reviewed tensor lifecycle is task-owned and scoped. Removing NpEnv requires the Manager-Based runtime itself to become tensor-native.

Outcome

Introduce the durable TorchEnv base and migrate the Manager-Based runtime lifecycle to Torch without relying on a new compatibility adapter.

Work

  • Add the owner module for TorchEnv and TorchEnvState.
  • Move step/reset lifecycle, state buffers, autoreset, termination/truncation, final observations, episode bookkeeping, and finite checks to Torch.
  • Keep observation groups as dictionaries and preserve policy/critic dimensions.
  • Move Manager-Based observation, reward, termination, command, event, curriculum, and metric compute to device-resident Torch operations.
  • Make RNG seed provenance and selected-row reset semantics explicit.
  • Preserve explicit host-bridge H2D/D2H boundaries for CPU-authoritative physics.
  • Consume only public SimBackend tensor APIs and fail closed on unsupported capabilities.
  • Add focused lifecycle and partial-reset tests before broad task migration.

Acceptance

  • The core Manager-Based lifecycle no longer calls the NpEnv base for normal execution.
  • CPU Torch buffers and CUDA Torch buffers follow the declared device contract without hidden NumPy conversion.
  • Existing legacy code may remain temporarily for not-yet-migrated callers, but the new path is not an alias or fallback for the final contract.

Activity

  1. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    P1 first slice is merged in PR #1708 (merge commit 6aaf0c4): durable TorchEnv/TorchEnvState, tensor step/autoreset/final-observation/training-state contracts, capability/device/fail-closed validation, shared timing schema owner, and focused lifecycle tests. #1703 remains open because ManagerBasedRlEnv still derives from NpEnv; the next slice must switch its lifecycle to TorchEnv through the explicit NumPy Manager host boundary and migrate its tests.

  2. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Core P1 lifecycle migration is merged in PR #1709 (merge commit 1bfaacb), with UniSim public-width dependency unisim#322 pinned at c63f5562. ManagerBasedRlEnv now inherits TorchEnv; public actions/reset rows/state/backend control are Torch, and its temporary Manager term execution uses named NumPy host boundaries. Direct RSL-RL is tensor-first and the G1 direct runtime no longer retains _cpu_env. #1703 remains open for device-resident Manager term compute, RNG provenance, and completing the all-tensor API boundary before broad task parity.

  3. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    First Manager term-family carrier migration is merged in PR #1710 (merge commit c7bf511): MetricsManager accumulators/counters/substep buffers are Torch tensors on the environment device, NumPy term results cross one explicit host boundary, Torch term results remain on-device, and reset logging is the bounded host publication. ADR-0011 also updates the Manager package import boundary to permit its target Torch carrier while continuing to reject mjlab at runtime. #1703 remains open for observation/reward/termination/command/event carriers and RNG provenance.

  4. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Reward and termination Manager carriers are now Torch tensors after PR #1711 (merge commit 49bd2b3). Reward aggregation/per-term rates and termination failure/timeout/per-term buffers run on env.device; already-tensor metric/reward/termination terms require exact dtype/device, legacy NumPy terms cross the named temporary boundary, and reset/logging remain bounded host publication points. #1703 remains open for action, observation, command, event, Entity state/reset APIs, and RNG provenance.

  5. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    ActionManager top-level history is now Torch after PR #1712 (merge commit e5ed471). action, prev_action, and prev_prev_action live on env.device; process input is strict contiguous float32 exact-device Torch; selected-row reset/history semantics are preserved. Existing ActionTerms temporarily receive one full-batch NumPy copy through the named term boundary. Action-rate/acceleration rewards consume Torch history, while last_action explicitly publishes NumPy to the still-host ObservationManager. #1703 remains open because BaseAction/Entity control writes, observation/command/event carriers, Entity tensor reads, reset tensor commits, and RNG provenance are pending.

  6. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    ObservationManager public output carrier is now Torch after PR #1713 (merge commit ae2db05). compute() / compute_group() publish contiguous float32 tensors on env.device, the cache and Manager obs_buf stay Torch, and TorchEnvState.obs consumes them without a per-step Torch↔NumPy round trip. Terms, clip/scale, NaN handling, noise, delay, and history still run their explicit NumPy pipeline; temporal buffer migration remains open.

  7. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Merged #1715 (099a1663): BaseAction raw/processed buffers, cold affine/clip tensors, action capability routing, and rough/motion action math now execute on-device. The remaining Entity write boundary is intentionally explicit: apply still publishes one host copy at apply_actions(), and encoder/default/current-joint-pos math plus Motion .target remain migration-only NumPy boundaries.

  8. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Merged #1716 (9749919a): AllegroIncrementalPositionAction now declares the tensor action path, keeps raw/clipped/integrated targets and bounds as float32 tensors on env.device, and executes clipping/integration/limiting/selected-row reset on-device. Entity writes and the still-NumPy Allegro observation remain explicit migration-only host publications. Next blocker for Stewart/state-feedback actions is a scene-owned device-aware state/sensor read facade with phase-scoped consistency.

  9. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Merged #1717 (7af18e44): added the public Entity.joint_tensor_view(device) facade. It resolves entity joints into backend qpos/qvel columns, validates capabilities/device/layout/dtype/finiteness, and returns entity-ordered Torch state. RelativeJointPositionAction now consumes it for on-device state-feedback math before the explicit Entity host write. Next: extend the facade to named sensors/body state and migrate Stewart, then observation non-temporal processing.

  10. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Merged #1719 (cf857971): ObservationManager now accepts exact-device float32 Torch terms and executes non-temporal clip/scale/NaN handling/row selection/concatenation with Torch on env.device. Scale tensors are cold-path persistent. NumPy terms retain the existing host boundary; noise/delay/history remain explicitly NumPy migration paths. Remaining priorities: named sensor/body tensor facade, Stewart action, command/event/reset tensor transactions, then legacy NpEnv deletion.

  11. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Summary

    Completes the next #1703 P1 slice on top of #1720:

    • Adds a scene-owned packed tensor read plan phase boundary and wires state-feedback actions to sensor/body tensor reads.
    • Refreshes the packed scene packet across Manager phases and after bounded in-phase mutations.
    • Pairs selected-row reset with one public HostBridgeTransferPlan.apply_reset() commit and one selected post-reset packed H2D read where native/public qpos and qvel widths agree. MotrixSim's current native/public reset-width mismatch explicitly falls back to the existing public set_state() boundary rather than guessing a private projection.
    • Makes ManagerBasedRlEnv._control the authoritative float32 Torch tensor. Entity control writes can now publish directly into that tensor; mapped-entity reset control restoration uses the public ctrl state view. The Stewart action no longer detaches to NumPy at its Entity write boundary.

    Validation

    UNILAB_LOCAL_UNISIM=/home/user/ws/unilabsim/unisim uv run pytest -q \
      tests/base tests/envs/mdp tests/envs/test_manager_based_rl_env.py \
      tests/envs/test_stewart.py tests/envs/test_multi_entity_consumer.py \
      tests/tasks/test_tensor_runtime_components.py
    834 passed, 19 skipped, 219 deselected
    
    make check
    passed (mypy clean; pyright 0 errors, existing optional drake import warning)
    
    make test-all
    1831 passed, 26 skipped, 617 deselected
    benchmark module-mode 35/35; script-mode 36/36
    

    Boundary notes

    • Mapped entity patches, reset randomization, mocap writes, and CUDA-device packed reset commits still fail closed and remain explicit migration boundaries.
    • MotrixSim packed reset requires a UniSim follow-up to unify native/public reset widths. This PR intentionally does not implement that projection in UniLab.

    Closes nothing; #1703 remains open for observation term consumption, temporal observation paths, command/event transactions, and RNG provenance.

  12. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    PR #1721 当前 head 8e2330b4 的 CI 已完成:

    • Focused CPU tensor profile: SUCCESS
    • Gated single-GPU tensor profile: SKIPPED(按预期,无 gated GPU runner)
    • PR 状态:OPEN / MERGEABLE,等待 review;遵守 review-first,不 self-merge。
  13. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    PR #1721 已 review 合入 develop/tensor-runtime(merge boundary:1ca0ab5b)。已在 #1721 的进度评论基础上完成该 merge 边界的追踪。

    继续推进下一个 P1 slice:temporal observation buffers tensorization,PR #1722:

    • CircularBuffer 现在以 Torch tensor 持有 history 和 per-batch counters,设备绑定到首次 append 的 observation 设备;
    • DelayBuffer.append/backfill/compute/peek 返回/消费 Torch tensor,lag RNG 仍保持 env-owned NumPy RNG,避免本 slice 改变 RNG stream;
    • ObservationManager 删除 delay/history 前的 obs.detach().cpu().numpy() host detour;
    • 新增 no-host-detour 测试覆盖 delay + history 的 device/dtype/语义。

    本地验证:

    focused: 202 passed, 1 deselected
    make check: passed
    make test-all: 1832 passed, 26 skipped; benchmark 35/35 + 36/36
    

    剩余 observation boundary:NoiseCfg/NoiseModel 仍为 NumPy owner;这是下一个需要和 RNG provenance 一起处理的边界。

  14. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Command manager boundary slice is ready in PR #1723 (feat/issue-1703-command-tensors, head b6ef45e7):

    • CommandTerm.time_left/command_counter are runtime-device Torch tensors; scalar, NumPy row-wise, and Torch row-wise dt are supported, while resampling still consumes env-owned NumPy RNG so this slice does not change the RNG stream.
    • CommandManager.get_command() now returns a validated device tensor. Existing NumPy command terms cross one explicit conversion boundary; locomotion command reward consumers convert at their own named NumPy boundary.
    • MotionCommand metrics and its Numba kernel remain explicitly NumPy-owned and fail closed if that carrier changes prematurely.
    • Local gates: make check pass; focused suite 221 passed / 37 deselected; make test-all 1833 passed / 26 skipped and benchmark imports 35/35 + 36/36.

    Remaining #1703 priorities are unchanged: observation noise/RNG provenance, command term internals, event tensor transactions, selected-row reset semantics, and MotrixSim native/public reset-width follow-up in UniSim.

  15. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    EventManager scheduling-carrier slice is ready as stacked PR #1724 (feat/issue-1703-event-scheduling-tensors, head 83538abe, base #1722):

    • interval time_left, reset-throttling step IDs, and trigger-once flags are now runtime-device Torch tensors (float64, int64, bool);
    • all interval/reset sampling still uses env-owned NumPy RNG, so the RNG stream is unchanged in this slice;
    • event terms and randomization transactions remain at their existing explicit NumPy term boundary.

    Local gates: focused 124 passed; make check pass; make test-all 1832 passed / 26 skipped + benchmark imports 35/35 and 36/36.

  16. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    PR #1724 CPU validation is confirmed: manually dispatched Tensor Runtime CI on head 83538abe completed successfully (Focused CPU tensor profile, run 36556322932). The gated single-GPU job was not requested. PR remains OPEN/MERGEABLE awaiting review.

  17. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Observation noise carrier slice is ready as stacked PR #1740 (feat/issue-1703-observation-noise-tensors, head b92714d8, base #1722): Constant/Uniform/Gaussian noise configs and additive-bias noise models accept NumPy or Torch observations and preserve the input carrier/device; ObservationManager removes the pre-noise .detach().cpu().numpy() full-observation detour. Env-owned NumPy RNG remains authoritative and consumes the same stream; tensor random noise performs one sampled unit-buffer H2D copy, while observation arithmetic/output stays on-device. Additive bias state materializes on the active observation device. Local gates: focused noise/partial-reset 30 passed; focused observation buffers 33 passed; related manager/env/base suites 189 passed; make check pass; make test 1841 passed / 26 skipped; slow train selection has only the unchanged unisim#291 failure. Tensor Runtime CI run 36571563620 is green.

  18. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    PR #1740 closed automatically when its base #1722 merged. It has been recreated as direct PR #1741 (feat/issue-1703-observation-noise-tensors, rebased head 13a9cc3e, base develop/tensor-runtime) with the same changes and validation summary. Retargeted-head Tensor Runtime CI run 36573120024 is dispatched.

  19. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    PR #1741 retargeted-head validation is green: Focused CPU profile passed; gated single-GPU profile skipped as expected (run 36573120024). PR is OPEN/MERGEABLE awaiting review.

  20. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    Merge-boundary update: #1722/#1723/#1724 have been review-merged into develop/tensor-runtime. This closes the temporal observation carrier, command carrier, and event scheduling-carrier slices. The next #1703 P1 slice is direct PR #1741 (feat/issue-1703-observation-noise-tensors, head 13a9cc3e), which removes the remaining observation-noise full-observation host detour while keeping the env-owned NumPy RNG authoritative. Remaining #1703 priorities: event term/randomization transactions, command term internals, RNG provenance formalization, selected-row reset semantics, and representative CUDA/CPU backend gates in P3.

  21. TATP-233 commented on Sep 29, 2026

    @TATP-233
    CollaboratorAuthor

    P1 merge update: observation temporal buffers (#1722), command carriers (#1723), event scheduling carriers (#1724), and observation noise carriers (#1741) are merged. Remaining #1703 implementation work is concentrated in event term/randomization tensor transactions, command term internals, explicit RNG provenance, and selected-row reset semantics; representative backend validation then belongs to #1705.

  22. TATP-233 commented on Oct 8, 2026

    @TATP-233
    CollaboratorAuthor

    Closing as superseded rather than completed-as-written: #1811 completed and explicitly superseded the remaining P1 scope, with the tensor-only Manager architecture integrated through #2023/#2044. The remaining lifecycle-throughput mechanism is tracked by #2043.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions