Repository navigation
fix(cli): keep default MPS runtime within socket limits - #2092
Merged
Merged
Conversation
TATP-233
force-pushed
the
fix/issue-2091-cuda-mps-default-path
branch
from
October 8, 2026 14:20
7748cb9 to
a1084ce
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Driving issue
Closes #2091.
Problem
The zero-argument
uni-cumpsworkflow could strand itself after a failed or exited NVIDIA control daemon:control_lockand apipe/logFIFO;controlandcontrol_privilegedsockets;startthen reported “Refusing to attach to an existing control path,” whileenv/stophad no daemon record to use.Measured on CUDA MPS 13.0 / driver 580:
Fix
Default runtime paths are now short and private:
All default runtime directories are created with mode
0700and owned by the invoking user.The default daemon name now uses the complete canonical GPU UUID:
rather than a shortened prefix, avoiding daemon-name collisions.
Explicit
--pipe-dirvalues are checked before launch. A resulting control path over 95 bytes fails with an actionable message recommending a short path such as/tmp/<short-name>/pipe.Failed-launch recovery is narrow:
control_lockand thepipe/logFIFO) are removed after a failed control launch;control/control_privilegedsockets are removed when no PID file exists;A nonzero control exit now appends the control-log tail, making silent NVIDIA failures diagnosable.
Final workflow
Verified on the available Linux/NVIDIA host:
All four completed successfully;
doctorexited 0. A secondstartwhile live correctly reported the existing live daemon. Restarting after a dead control process also succeeded after safe orphan cleanup.The immediate
start→doctorpair was additionally repeated 20 times in one shell; all 20 doctors returnedvalid=Trueand exit 0. The orphan check now validates PID liveness and the/proc/<pid>/statprocess name rather than treating the mere presence of a stale PID file as live ownership.Validation
Commands run on final local head
a1084ce8:Results:
1845 passed, 43 skipped, 488 deselected;34/34, script-mode35/35.Focused checks:
Result:
57 passed.New coverage includes:
Scope boundaries