Repository navigation
Conversation
Add --grpc-stream-accept-prefetch (default 16) to keep several stream accept requests outstanding.
|
|
Thought we decided not to expose |
| import test_util as tu | ||
| import tritonclient.grpc as grpcclient | ||
| from tritonclient.utils import InferenceServerException | ||
| import os # noqa: E402 |
There was a problem hiding this comment.
why al of the # noqa labeling?
| grpc_options_.push_back( | ||
| {OPTION_GRPC_STREAM_ACCEPT_PREFETCH, "grpc-stream-accept-prefetch", | ||
| Option::ArgInt, | ||
| "The number of accept requests each gRPC streaming inference handler " |
There was a problem hiding this comment.
Let's clean this up a bit. Perhaps:
"The maximum number of accept requests each gRPC streaming inference handler keeps outstanding. "
"Up to this many new streams can be matched immediately. "
"Must be in the range 1 to 128, inclusive. Default is 16."
is "maximum" the correct term to use above?
also, what is an "accept request"?
| * `--grpc-stream-accept-prefetch`: 16 by default. | ||
| The number of accept requests each streaming inference handler keeps outstanding, so up to this many new streams can be matched immediately. |
There was a problem hiding this comment.
maybe reword as
The number of accept requests each stream inference handler keeps outstanding.
Up to this many new streams can be matched immediately.
Valid range is `1` to `128`, inclusive.
Default value is `16`.
Legacy versions of Triton default value was `1`.
fix: Prevent silent gRPC stream cancellations under backpressure
What does the PR do?
Under ensemble
max_inflight_requestsbackpressure, new gRPC streams can be cancelled by gRPC before Triton ever sees them. The client gets a bareCANCELLEDwith no responses, and nothing appears in Triton's logs or metrics.The streaming handler keeps only one
RequestModelStreamInfer(accept request) outstanding and re-posts it after handling each new stream. Under backpressure,TRITONSERVER_ServerInferAsyncblocks that same thread for seconds per request. New streams in a burst are therefore matched about one per pipeline pass. gRPC cancels any call left unmatched longer thanGRPC_ARG_SERVER_MAX_UNREQUESTED_TIME_IN_SERVER_SECONDS(30 s by default).This PR keeps K accept requests outstanding instead of one, so a burst is matched right away. Matched streams are not subject to that deadline.
ModelStreamInferHandler::StartNewRequest()posts K accept requests on its first call. Each accepted stream then posts one replacement, so K stay outstanding.--grpc-stream-accept-prefetch(range 1–128, default 16). A value of1restores the previous behavior.stream_accept_prefetchoption in the in-process Python frontend (KServeGrpc.Options).inference_protocols.md.Checklist
<commit_type>: <Title>Commit Type:
Check the conventional commit type
box here and add the label to the github PR.
Related PRs:
Where should the reviewer start?
src/grpc/stream_infer_handler.cc:StartNewRequest()/PostAcceptRequest()src/grpc/grpc_server.cc: handler construction and thestream_accept_prefetchkey inGetOptionssrc/command_line_parser.cc: new option and validationqa/L0_simple_ensemble/test.sh: new--grpc-stream-accept-prefetchblockTest plan:
L0_simple_ensemble:EnsembleBackpressureTest(16 concurrent streams × 8 responses) runs at the default with no flag. Without this change it cancels 5 of 16 streams atmax_inflight_requests: 1.L0_simple_ensemble: new block that rejects0,129andabc. For the default,1and128, it checks that exactly that many accept requests are posted at startup.L0_simple_ensemble: the backpressure test now asserts errors before response counts, so a cancelled stream reportsCANCELLEDinstead of "expected 8, got 0".L0_python_api:stream_accept_prefetchrange checks.Failures were only gRPC
CANCELLEDat ~31 s on streams that never reached Triton. Startup memory (37–39 MB), startup time (~3 s) and shutdown time (~4–5 s) did not change with K from 1 to 128.All pre-commit hooks pass.
CI Pipeline ID: [72460154]
Caveats:
InferAsyncstill blocks the streaming handler thread undermax_inflight_requestsbackpressure. Streams are now admitted, but they are still served at the pipeline's pace. That needs a separate change.KServeGrpctests inL0_python_apiarexfail(run=False), so the Python option is covered by option-level tests only.Background
This became visible when Triton's pinned gRPC moved from 1.54.3 to 1.81.1; the same streams survived on 1.54.3. A prototype with 16 accept requests on the default single completion queue passed the backpressure test in CI (pipeline 67201458).
Raising
grpc.server_max_unrequested_time_in_serverinstead gave 5, 1 and 0 cancellations at 30, 45 and 60 s. The safe value depends on the workload, so it is only an emergency mitigation.Related Issues: (use one of the action keywords Closes / Fixes / Resolves / Relates to)