Repository navigation
feat(training): enable single-host multi-GPU MPS - #2069
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
CUDA_VISIBLE_DEVICESentries infers one rank per entry, rank 0 is remapped to its one-entry local namespace before CUDA initialization, andDpRankSupervisorstill sees the full parent mask while spawning sibling ranks.UNILAB_DP_WORLD_SIZEbefore runner assembly so rank 0 constructs the same NCCLDpParameterSyncas spawned ranks.rankandworld_sizeto the MPS manifest evidence.This is an implementation/gate change, not a performance claim. See the benchmark result below: DP + MPS currently regresses and must not yet be promoted as recommended.
Linked Work
develop/tensor-runtimeValidation
make test-allwas not run for this experimental follow-up branchCommands actually run:
Two-GPU hardware validation on server 592 with exact pinned siblings plus unilab_rl PR #79:
Two-GPU MPS smoke:
Benchmark evidence
One 300-iteration arm per task, final-100 mean, RTX 5090 ×2:
All four 300-iteration runs completed normally with world_size=2, no replay drops, final occupancy zero, and clean shutdown. This explicitly fails the no-regression direction discussed in #2063; the work is not ready for production support promotion.
Likely mechanisms match #2063:
Impact
Artifacts
~/unilabsim/UniLab/benchmarks/dp-mps-20261005/on 592; not committedChecklist
Follow-up required before claiming DP+MPS support: