Skip to content

fix: avg_time_queue_us under-reports queue time in k8s-onprem autoscaling rule - #8995

Open
100-JM wants to merge 1 commit into
triton-inference-server:mainfrom
100-JM:fix-k8s-onprem-queue-metric
Open

100-JM wants to merge 1 commit into
triton-inference-server:mainfrom
100-JM:fix-k8s-onprem-queue-metric

Conversation

@100-JM

@100-JM 100-JM commented Oct 3, 2026

Copy link
Copy Markdown

Fixes #8979.

What

The prometheus-adapter rule behind the avg_time_queue_us HPA metric in deploy/k8s-onprem/values.yaml under-reports the mean queue time, delaying scale-up:

  1. avg(...) by (pod) averages per-model ratios, so every idle model contributes a 0 and dilutes the result — a pod serving 1 of N loaded models reports ~1/N of the true queue wait. Triton loads every model in the repository, and the chart's own quickstart repository holds two models, so a stock install halves the metric.
  2. 1 + delta(...) divides the accumulated queue time by n+1 requests instead of n.

Change

Aggregate both per-model counters over the pod before dividing, guard the zero-request case with clamp_min (which, unlike 1 + x, does not shift the result), and use increase instead of delta since both series are counters that reset on server restart:

metricsQuery: 'sum(increase(nv_inference_queue_duration_us{<<.LabelMatchers>>}[30s])) by (<<.GroupBy>>) / clamp_min(sum(increase(nv_inference_request_success{<<.LabelMatchers>>}[30s])) by (<<.GroupBy>>), 1)'

This follows the issue's "mean across the pod" form, which matches the README's description of scaling "based on the average queue time". The recording rules in triton-inference-server/tutorials (TensorRT-LLM autoscaling guide) already use clamp_min the same way.

Verification

Evaluated both expanded queries with promtool test rules (promtool 3.5.0) on the scenario from #8979 — one pod, two models, traffic to one: 100 requests per 30s window, each queued 1000us:

query reported mean queue time
current rule 495.05 us (fails the 1000us assertion)
this PR 1000 us (exact)

An added idle third model leaves the fixed rule at 1000 while dragging the old rule down further (330.03). The zero-traffic case returns 0 rather than NaN thanks to clamp_min.

Not included: seriesQuery still hardcodes namespace="default" (breaks discovery in other namespaces, likely the cause of #6247) — happy to address that here too if preferred, but it changes discovery behavior so I kept this PR to the metric computation.

🤖 Generated with Claude Code

The prometheus-adapter rule averaged per-model queue-time ratios, so
idle models diluted the pod's reported queue time, and the 1+delta
denominator divided by n+1 requests instead of n. Sum both per-model
counters over the pod before dividing, guard the zero-request case with
clamp_min, and use increase instead of delta since both series are
counters that reset on server restart.

Verified with promtool: one pod, two models, traffic to one (100 req/30s,
1000us queued each) — the old rule reports 495.05, the new rule 1000.

Fixes triton-inference-server#8979

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@greptile-apps

greptile-apps Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Changes Kubernetes autoscaling metric calculation.

The PR appears safe to merge.

Summary

The PR changes the autoscaling metric from an average of per-model queue-time ratios to a per-pod ratio of counter increases, avoiding dilution by idle models and accounting for counter resets.

Reviews (1) · Last reviewed commit: "Fix avg_time_queue_us under-reporting in..."

100-JM added a commit to 100-JM/100-JM that referenced this pull request Oct 10, 2026
- kubeflow/trainer#4159 moved to Merged (7 merged total)
- triton-inference-server/server#8995 added to In review (10 total)
- tech badges and a one-line contribution summary

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

k8s-onprem autoscaling metric avg_time_queue_us under-reports queue time by roughly the number of loaded models

2 participants