You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(training): let trainer use the sole live uni-cumps daemon without shell eval #2093
uni-cumps now owns the explicit daemon lifecycle, but the trainer still requires the user to export the daemon environment into the shell:
uv run uni-cumps start
uv run --extra mjwarp train --algo flashsac \
--task g1_motion_tracking \
--sim mjwarp \
training.cuda_process_sharing=mps
With only start completed, the trainer probes /tmp/nvidia-mps/control and fails:
ValueError: training.cuda_process_sharing='mps' could not reach the control daemon
through /tmp/nvidia-mps/control.
The working command currently requires this extra shell integration:
eval"$(uv run uni-cumps env)"
Issue #2079 explicitly deferred this launcher integration: Phase 1 only improved diagnostics and taught users the explicit daemon command. That phase has now landed, so the follow-up can make the normal workflow start-and-train.
Deliverable
For the supported single-host, single-rank MJWarp off-policy path, let the trainer resolve a UniLab-recorded daemon without mutating the parent shell.
When training.cuda_process_sharing=mps and CUDA_MPS_PIPE_DIRECTORY is absent:
select the sole live daemon whose canonical GPU UUID matches the rank-local learner/collector GPU;
set CUDA_MPS_PIPE_DIRECTORY and CUDA_MPS_LOG_DIRECTORY only in the trainer process environment before the fail-closed probe and before environment/learner/collector construction;
keep using an explicit caller-provided CUDA_MPS_PIPE_DIRECTORY when present;
fail closed when there is no live matching daemon, multiple matching live daemons, or host/UID mismatch;
include host-level daemon identity additively in the existing producer diagnostics without making it a stable public scalar contract.
The trainer must not start, stop, repair, or lease-refcount a daemon.
Scope and delivery boundaries
In this work item:
trainer-side selection of one already-running, already-recorded daemon;
single-GPU MJWarp SAC/FlashSAC path;
tests and production-guide updates;
improved errors that distinguish “no live daemon” from “daemon exists but failed validation”.
Separate outcomes:
starting/stopping or leasing a daemon from the trainer;
multi-GPU DP daemon selection and cross-rank aggregation;
task-per-GPU scheduling/refcounts;
multi-node daemon discovery;
training.cuda_process_sharing=auto.
Affected owner layers and contracts
Training assembly: src/unilab/scripts/train_offpolicy.py
Daemon record owner: src/unilab/training/cuda_mps_cli.py
A design choice is needed for selecting among multiple matching daemons if a stable training.cuda_mps_daemon owner key is preferred over requiring explicit environment in that case.
Proposed owner
UniLab training runtime maintainers
Validation plan
Unit tests with fake daemon records and fake process identity.
Builder test proving selected daemon environment reaches the existing fail-closed probe.
Explicit-environment precedence test.
No-daemon and ambiguous-daemon fail-closed tests before env construction.
Linux/NVIDIA integration: uni-cumps start, then the documented trainer command without shell eval.
Problem
uni-cumpsnow owns the explicit daemon lifecycle, but the trainer still requires the user to export the daemon environment into the shell:With only
startcompleted, the trainer probes/tmp/nvidia-mps/controland fails:The working command currently requires this extra shell integration:
Issue #2079 explicitly deferred this launcher integration: Phase 1 only improved diagnostics and taught users the explicit daemon command. That phase has now landed, so the follow-up can make the normal workflow start-and-train.
Deliverable
For the supported single-host, single-rank MJWarp off-policy path, let the trainer resolve a UniLab-recorded daemon without mutating the parent shell.
When
training.cuda_process_sharing=mpsandCUDA_MPS_PIPE_DIRECTORYis absent:CUDA_MPS_PIPE_DIRECTORYandCUDA_MPS_LOG_DIRECTORYonly in the trainer process environment before the fail-closed probe and before environment/learner/collector construction;CUDA_MPS_PIPE_DIRECTORYwhen present;The trainer must not start, stop, repair, or lease-refcount a daemon.
Scope and delivery boundaries
In this work item:
Separate outcomes:
training.cuda_process_sharing=auto.Affected owner layers and contracts
src/unilab/scripts/train_offpolicy.pysrc/unilab/training/cuda_mps_cli.pysrc/unilab/training/cuda_process_sharing.pyunilab cuda mpsmanagement CLI (initial single-GPU implementation) #2079 follow-up boundaryDefinition of done
uv run uni-cumps start, the documented training command succeeds withouteval "$(uv run uni-cumps env)".CUDA_MPS_PIPE_DIRECTORYbehavior remains unchanged.uni-cumps startand the required GPU.make check,make test, andmake test-allpass on the final head.Dependencies and blockers
training.cuda_mps_daemonowner key is preferred over requiring explicit environment in that case.Proposed owner
UniLab training runtime maintainers
Validation plan
uni-cumps start, then the documented trainer command without shelleval.