Repository navigation
test: Deflake L0_socket reuse-port request distribution check - #9001
Draft
devin-ai-integration[bot] wants to merge 1 commit into
Draft
devin-ai-integration[bot] wants to merge 1 commit into
devin-ai-integration[bot] wants to merge 1 commit into
Conversation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Contributor
Author
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does the PR do?
L0_socketstarts 3 servers sharing the HTTP/gRPC port (--reuse-*-port=true), sends a single batch of 11 client processes, and fails if any server'snv_inference_request_successis 0. The kernel uses a hash to pick whichSO_REUSEPORTlistener gets each new connection, so with 11 connections there's roughly a 3·(2/3)^11 ≈ 5% chance per protocol that one server gets nothing. That matches the intermittent pseudo-nightly failures: gRPC counts were 0/16/6 on 10-06 and 8/14/0 on 10-02, and the runs in between passed.Change: send batches of 11 clients until all three servers have served at least one request, up to
MAX_CLIENT_ROUNDS=10. If a server still has 0 after 10 rounds, the test fails as before. A listener that really never receives connections is still caught; only the chance-based miss goes away. The test also prints the per-round counts and waits on each client PID individually. Before,wait $pidsonly returned the exit status of the last client, so failures in other clients went unnoticed.Checklist
<commit_type>: <Title>Commit Type:
Related PRs:
None
Where should the reviewer start?
qa/L0_socket/test.sh, the reuse-port distribution loop near the end of the file.Test plan:
bash -nand pre-commit pass. I tested the loop logic locally with a fake client that assigns requests at random. I haven't run it against a real server. It needs anL0_socket--baseCI run.Caveats:
In the worst case the test takes up to 10x as many client runs. In practice it almost always stops after round 1.
Background
The pseudo-nightly CI health report flagged
L0_socket--baseas newly broken (dl/dgx/tritonserver job 473179531).Related Issues:
L0_socket--baseflake (no ticket yet)Link to Devin session: https://nvidia-cloud.devinenterprise.com/sessions/42b5577b240344db884012e785abfcc6
Open in Devin Desktop: https://nvidia-cloud.devinenterprise.com/desktop/session/42b5577b240344db884012e785abfcc6?variant=devin