Skip to content

feat(training): enable single-host multi-GPU MPS - #2069

Merged
TATP-233 merged 2 commits into
develop/tensor-runtimefrom
feature/dp-cuda-mps-multi-gpu
Oct 8, 2026
Merged

TATP-233 merged 2 commits into
develop/tensor-runtimefrom
feature/dp-cuda-mps-multi-gpu

Conversation

@TATP-233

@TATP-233 TATP-233 commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • Completes the off-policy single-host multi-GPU parent launch wiring: a parent mask with multiple opaque CUDA_VISIBLE_DEVICES entries infers one rank per entry, rank 0 is remapped to its one-entry local namespace before CUDA initialization, and DpRankSupervisor still sees the full parent mask while spawning sibling ranks.
  • Exports the inferred UNILAB_DP_WORLD_SIZE before runner assembly so rank 0 constructs the same NCCL DpParameterSync as spawned ranks.
  • Extends CUDA MPS validation from strict single-rank to valid per-rank topology on one host. Each rank still independently validates its rank-local learner/collector GPU UUID and daemon access.
  • Adds rank and world_size to the MPS manifest evidence.

This is an implementation/gate change, not a performance claim. See the benchmark result below: DP + MPS currently regresses and must not yet be promoted as recommended.

Linked Work

Validation

  • Focused tests passed on the implementation branch
  • make test-all was not run for this experimental follow-up branch

Commands actually run:

uv run --no-sync pyright \
  src/unilab/scripts/train_offpolicy.py \
  src/unilab/training/cuda_process_sharing.py \
  tests/training/test_cuda_process_sharing.py

uv run --no-sync pytest \
  tests/training/test_cuda_process_sharing.py \
  tests/algos/test_offpolicy_double_buffer_runner.py -q
# 53 passed

uv run --no-sync pytest tests/scripts/test_train_scripts.py \
  -q -m slow -k 'offpolicy_parent or offpolicy_supervisor'
# 3 passed

Two-GPU hardware validation on server 592 with exact pinned siblings plus unilab_rl PR #79:

SAC / G1 Walk Flat / MJWarp
RC=0; status=completed; iterations=2/2; world_size=2; NCCL DP;
normal_completion; cleanup errors=[]; replay published==released;
final occupancy=0; dropped batches=0.

FlashSAC / G1 Motion Tracking / MJWarp
RC=0; status=completed; iterations=2/2; world_size=2; NCCL DP;
normal_completion; cleanup errors=[]; replay published==released;
final occupancy=0; dropped batches=0.

Two-GPU MPS smoke:

Both tasks: RC=0, status=completed, effective=mps, validated=true,
rank=0, world_size=2, learner/collector UUID match, server PID recorded,
NCCL DP active, normal shutdown, no replay drops.

Benchmark evidence

One 300-iteration arm per task, final-100 mean, RTX 5090 ×2:

Task DP base steps/s DP MPS steps/s Delta
FlashSAC / G1 Motion Tracking 66,809.27 57,840.75 -13.42%
SAC / G1 Walk Flat 31,523.51 30,132.91 -4.41%

All four 300-iteration runs completed normally with world_size=2, no replay drops, final occupancy zero, and clean shutdown. This explicitly fails the no-regression direction discussed in #2063; the work is not ready for production support promotion.

Likely mechanisms match #2063:

  • unilab_rl PR Task:Go2 handstand #79 restores DP by leaving the whole-cycle graph and using eager fallback;
  • NCCL all-reduce synchronizes ranks, amplifying per-rank jitter;
  • learner-side kernels under MPS compete more aggressively.

Impact

  • Backend impact: MJWarp CUDA path
  • Platform impact: Linux/NVIDIA CUDA
  • Training behavior: single-rank default remains unchanged. Multi-entry parent visibility now enters DP as already documented. DP + MPS is functional but not recommended based on current evidence.

Artifacts

  • W&B: none
  • benchmark result: server-local ~/unilabsim/UniLab/benchmarks/dp-mps-20261005/ on 592; not committed
  • video/screenshot: none
  • ONNX/checkpoint: none

Checklist

  • Added or updated tests where needed
  • Updated docs/ADR for supported DP+MPS status
  • Linked prerequisite PRs/discussions
  • Follow-up recorded

Follow-up required before claiming DP+MPS support:

  1. implement per-rank evidence aggregation plus one host-level server record;
  2. investigate whole-cycle NCCL-compatible capture rather than eager DP fallback;
  3. meet or set an evidence-based CUDA MPS under multi-GPU DP: evidence gaps and design constraints #2063 no-regression gate;
  4. add cross-rank config agreement and atomic failure;
  5. document single-host daemon ownership and concurrent-run policy.

@TATP-233
TATP-233 requested a review from caozx1110 as a code owner October 7, 2026 16:09
@TATP-233
TATP-233 merged commit fc8626c into develop/tensor-runtime Oct 8, 2026
2 checks passed
@TATP-233
TATP-233 deleted the feature/dp-cuda-mps-multi-gpu branch October 8, 2026 05:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant