Skip to content

fix: Prevent silent gRPC stream cancellations under backpressure - #9004

Open
akhilraj9 wants to merge 1 commit into
mainfrom
asaraswathi/tri-1711-stream-accept-prefetch
Open

akhilraj9 wants to merge 1 commit into
mainfrom
asaraswathi/tri-1711-stream-accept-prefetch

Conversation

@akhilraj9

@akhilraj9 akhilraj9 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

fix: Prevent silent gRPC stream cancellations under backpressure

What does the PR do?

Under ensemble max_inflight_requests backpressure, new gRPC streams can be cancelled by gRPC before Triton ever sees them. The client gets a bare CANCELLED with no responses, and nothing appears in Triton's logs or metrics.

The streaming handler keeps only one RequestModelStreamInfer (accept request) outstanding and re-posts it after handling each new stream. Under backpressure, TRITONSERVER_ServerInferAsync blocks that same thread for seconds per request. New streams in a burst are therefore matched about one per pipeline pass. gRPC cancels any call left unmatched longer than GRPC_ARG_SERVER_MAX_UNREQUESTED_TIME_IN_SERVER_SECONDS (30 s by default).

This PR keeps K accept requests outstanding instead of one, so a burst is matched right away. Matched streams are not subject to that deadline.

  • ModelStreamInferHandler::StartNewRequest() posts K accept requests on its first call. Each accepted stream then posts one replacement, so K stay outstanding.
  • New option --grpc-stream-accept-prefetch (range 1–128, default 16). A value of 1 restores the previous behavior.
  • Matching stream_accept_prefetch option in the in-process Python frontend (KServeGrpc.Options).
  • Docs: new "GRPC Streaming Accept Prefetch" subsection in inference_protocols.md.
  • No change to completion-queue layout, threading, the scheduler or backends.

Checklist

  • PR title reflects the change and is of format <commit_type>: <Title>
  • Changes are described in the pull request.
  • Related issues are referenced.
  • Populated github labels field
  • Added test plan and verified test passes.
  • Verified that the PR passes existing CI.
  • Verified copyright is correct on all changed files.
  • Added succinct git squash message before merging ref.
  • All template sections are filled out.
  • Optional: Additional screenshots for behavior/output changes with before/after.

Commit Type:

Check the conventional commit type
box here and add the label to the github PR.

  • build
  • ci
  • docs
  • feat
  • fix
  • perf
  • refactor
  • revert
  • style
  • test

Related PRs:

Where should the reviewer start?

  • src/grpc/stream_infer_handler.cc: StartNewRequest() / PostAcceptRequest()
  • src/grpc/grpc_server.cc: handler construction and the stream_accept_prefetch key in GetOptions
  • src/command_line_parser.cc: new option and validation
  • qa/L0_simple_ensemble/test.sh: new --grpc-stream-accept-prefetch block

Test plan:

  • L0_simple_ensemble: EnsembleBackpressureTest (16 concurrent streams × 8 responses) runs at the default with no flag. Without this change it cancels 5 of 16 streams at max_inflight_requests: 1.

  • L0_simple_ensemble: new block that rejects 0, 129 and abc. For the default, 1 and 128, it checks that exactly that many accept requests are posted at startup.

  • L0_simple_ensemble: the backpressure test now asserts errors before response counts, so a cancelled stream reports CANCELLED instead of "expected 8, got 0".

  • L0_python_api: stream_accept_prefetch range checks.

    Failures were only gRPC CANCELLED at ~31 s on streams that never reached Triton. Startup memory (37–39 MB), startup time (~3 s) and shutdown time (~4–5 s) did not change with K from 1 to 128.

  • All pre-commit hooks pass.

  • CI Pipeline ID: [72460154]

Caveats:

  • The root cause is not fixed here: InferAsync still blocks the streaming handler thread under max_inflight_requests backpressure. Streams are now admitted, but they are still served at the pipeline's pace. That needs a separate change.
  • How large a burst the default absorbs depends on pipeline speed. In the L0 test (~4 s per request) the default absorbs about 40 new streams at once. Larger bursts need a larger value.
  • Each outstanding accept request holds one pre-allocated handler state. These count in the gRPC graceful-shutdown connection count until they are drained.
  • The live KServeGrpc tests in L0_python_api are xfail(run=False), so the Python option is covered by option-level tests only.

Background

This became visible when Triton's pinned gRPC moved from 1.54.3 to 1.81.1; the same streams survived on 1.54.3. A prototype with 16 accept requests on the default single completion queue passed the backpressure test in CI (pipeline 67201458).

Raising grpc.server_max_unrequested_time_in_server instead gave 5, 1 and 0 cancellations at 30, 45 and 60 s. The safe value depends on the workload, so it is only an emergency mitigation.

Related Issues: (use one of the action keywords Closes / Fixes / Resolves / Relates to)

  • Relates to TRI-1711

Add --grpc-stream-accept-prefetch (default 16) to keep several stream
accept requests outstanding.
@akhilraj9
akhilraj9 requested review from Vinya567, pskiran1, whoisj and yinggeh and removed request for Vinya567, pskiran1 and whoisj October 8, 2026 15:37
@greptile-apps

greptile-apps Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Adds tunable parameter for gRPC stream accept request prefetching.

The PR appears safe to merge; no actionable defects were found.

What we checked:

  • Startup checks wait for accepts: InferHandler::Start() waits until StartNewRequest() finishes. gRPC starts before HTTP, whose readiness the test waits for.

Summary

Adds configurable gRPC stream accept prefetch so bursts can match waiting accepts while the handler is busy.

  • Defaults to 16, with a range of 1–128, through both the command line and KServeGrpc.Options.
  • Adds startup-count and range checks, documentation, and clearer backpressure test errors.
  • No actionable issues found. Tests were not run during this review.

Acknowledged limitations supplied by akhilraj9: InferAsync still blocks under backpressure; bursts beyond the available accepts can still be cancelled; pending accepts count toward graceful shutdown; live Python frontend tests remain disabled. These were explicitly described as known or deferred.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Command-line or Python option] --> B[Prefetch count: default 16]
  B --> C[Handler posts K accepts at startup]
  C --> D[gRPC matches a new stream]
  D --> E[Handler posts one replacement]
  E --> F[Handler reads requests]
  F --> G[Triton runs inference]
Loading

Reviews (1) · Last reviewed commit: "fix: Prevent silent gRPC stream cancella..." · Reviewed by Greptile

@yinggeh

yinggeh commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Thought we decided not to expose --grpc-stream-accept-prefetch?

import test_util as tu
import tritonclient.grpc as grpcclient
from tritonclient.utils import InferenceServerException
import os # noqa: E402

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why al of the # noqa labeling?

grpc_options_.push_back(
{OPTION_GRPC_STREAM_ACCEPT_PREFETCH, "grpc-stream-accept-prefetch",
Option::ArgInt,
"The number of accept requests each gRPC streaming inference handler "

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's clean this up a bit. Perhaps:

       "The maximum number of accept requests each gRPC streaming inference handler keeps outstanding. "
       "Up to this many new streams can be matched immediately. "
       "Must be in the range 1 to 128, inclusive. Default is 16."

is "maximum" the correct term to use above?

also, what is an "accept request"?

Comment on lines +159 to +160
* `--grpc-stream-accept-prefetch`: 16 by default.
The number of accept requests each streaming inference handler keeps outstanding, so up to this many new streams can be matched immediately.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe reword as

  The number of accept requests each stream inference handler keeps outstanding.
  Up to this many new streams can be matched immediately.
  Valid range is `1` to `128`, inclusive.
  Default value is `16`.
  Legacy versions of Triton default value was `1`.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants