Repository navigation
Conversation
The prometheus-adapter rule averaged per-model queue-time ratios, so idle models diluted the pod's reported queue time, and the 1+delta denominator divided by n+1 requests instead of n. Sum both per-model counters over the pod before dividing, guard the zero-request case with clamp_min, and use increase instead of delta since both series are counters that reset on server restart. Verified with promtool: one pod, two models, traffic to one (100 req/30s, 1000us queued each) — the old rule reports 495.05, the new rule 1000. Fixes triton-inference-server#8979 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
100-JM
added a commit
to 100-JM/100-JM
that referenced
this pull request
Oct 10, 2026
- kubeflow/trainer#4159 moved to Merged (7 merged total) - triton-inference-server/server#8995 added to In review (10 total) - tech badges and a one-line contribution summary Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #8979.
What
The prometheus-adapter rule behind the
avg_time_queue_usHPA metric indeploy/k8s-onprem/values.yamlunder-reports the mean queue time, delaying scale-up:avg(...) by (pod)averages per-model ratios, so every idle model contributes a 0 and dilutes the result — a pod serving 1 of N loaded models reports ~1/N of the true queue wait. Triton loads every model in the repository, and the chart's own quickstart repository holds two models, so a stock install halves the metric.1 + delta(...)divides the accumulated queue time by n+1 requests instead of n.Change
Aggregate both per-model counters over the pod before dividing, guard the zero-request case with
clamp_min(which, unlike1 + x, does not shift the result), and useincreaseinstead ofdeltasince both series are counters that reset on server restart:This follows the issue's "mean across the pod" form, which matches the README's description of scaling "based on the average queue time". The recording rules in
triton-inference-server/tutorials(TensorRT-LLM autoscaling guide) already useclamp_minthe same way.Verification
Evaluated both expanded queries with
promtool test rules(promtool 3.5.0) on the scenario from #8979 — one pod, two models, traffic to one: 100 requests per 30s window, each queued 1000us:An added idle third model leaves the fixed rule at 1000 while dragging the old rule down further (330.03). The zero-traffic case returns 0 rather than NaN thanks to
clamp_min.Not included:
seriesQuerystill hardcodesnamespace="default"(breaks discovery in other namespaces, likely the cause of #6247) — happy to address that here too if preferred, but it changes discovery behavior so I kept this PR to the metric computation.🤖 Generated with Claude Code