Skip to content

fix: preserve model selection in GenAI-Perf sweeps - #471

Open
1fanwang wants to merge 2 commits into
triton-inference-server:mainfrom
1fanwang:1fannnw/fix-multi-model-sweep
Open

1fanwang wants to merge 2 commits into
triton-inference-server:mainfrom
1fanwang:1fannnw/fix-multi-model-sweep

Conversation

@1fanwang

Copy link
Copy Markdown

Summary

A GenAI-Perf sweep with multiple models crashes before profiling with KeyError: 'llama2:7b'. The generator expects a search space for every model, but analyze supplies one shared search space for the mixed workload.

Generate sweep points from the supplied search spaces, including their debug descriptions. Requests keep both models and the existing selection strategy. This does not add independent per-model sweeps.

Fixes the crash reported in #428.

Testing Done

The regression in this PR runs through configuration parsing, analyze initialization, sweep generation, and the real request-file writer. It reads JSONL prompts and serves an empty metrics endpoint through a local HTTP server. The normal CI suite runs it without mocking those paths.

I ran it on macOS arm64 with Python 3.12.9. It does not execute native Perf Analyzer, model inference, or a GPU benchmark.

Source Before After
Current main The second model raises KeyError in normal and debug modes Both concurrency values retain both request models
r26.08 The same failure The identical runtime patch produces the same result

Both unmodified sources report:

E   KeyError: 'llama2:7b'

With the patch, both sources emit these request-file observations:

models=['mistral:7b', 'llama2:7b'] log_level=INFO
concurrency=1 request_models=['mistral:7b', 'llama2:7b']
concurrency=2 request_models=['mistral:7b', 'llama2:7b']
Commands and remaining raw output

From the checkout root, named worktree in this run, these commands run the committed CI test against both baselines and both patched sources. The Python environment has the package dependencies. Importlib mode keeps the test location from replacing the runtime selected by PYTHONPATH.

set -euo pipefail
git fetch https://github.com/triton-inference-server/perf_analyzer.git r26.08
mkdir -p ../baseline-source ../release-baseline-source ../release-source
git archive ee519c63382237fadd8b92216d667b65fa503d35 genai-perf | tar -x -C ../baseline-source
git archive a218ec10e31310ecd8a7d08b2d1df75bca7c973e genai-perf | tar -x -C ../release-baseline-source
git archive a218ec10e31310ecd8a7d08b2d1df75bca7c973e genai-perf | tar -x -C ../release-source
git diff ee519c63382237fadd8b92216d667b65fa503d35 HEAD -- genai-perf/genai_perf/config/generate/sweep_objective_generator.py | (cd ../release-source && git apply)
for tree in baseline-source release-baseline-source worktree release-source; do
  case "$tree" in
    baseline-source|release-baseline-source) expected=1 ;;
    *) expected=0 ;;
  esac
  result=0
  PYTHONPATH="../$tree/genai-perf" PYTHONDONTWRITEBYTECODE=1 \
    python -m pytest --import-mode=importlib -q -s --tb=short \
    --basetemp="../evidence/public-$tree" \
    genai-perf/tests/test_sweep_objective_generator.py::test_analyze_sweeps_model_selection \
    > "../public-$tree.log" 2>&1 || result=$?
  printf 'source=%s exit=%s\n' "$tree" "$result"
  cat "../public-$tree.log"
  test "$result" -eq "$expected"
done

The test also records the debug-mode result:

models=['mistral:7b', 'llama2:7b'] log_level=DEBUG
concurrency=1 request_models=['mistral:7b', 'llama2:7b']
concurrency=2 request_models=['mistral:7b', 'llama2:7b']

The single-model control emits the same values before and after:

models=['mistral:7b'] log_level=INFO
concurrency=1 request_models=['mistral:7b', 'mistral:7b']
concurrency=2 request_models=['mistral:7b', 'mistral:7b']
models=['mistral:7b'] log_level=DEBUG
concurrency=1 request_models=['mistral:7b', 'mistral:7b']
concurrency=2 request_models=['mistral:7b', 'mistral:7b']

The baseline invocations exited with status 1; the patched invocations exited with status 0.

The existing W503 warning and type-check diagnostics are unchanged from the base.

Signed-off-by: 1fanwang <1fannnw@gmail.com>
Signed-off-by: 1fanwang <1fannnw@gmail.com>
@greptile-apps

greptile-apps Bot commented Sep 20, 2026

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

The PR appears safe to merge with no actionable correctness, security, or repository-rule issues identified.

Summary

The PR fixes mixed-model GenAI-Perf sweeps by generating objectives from the supplied shared search spaces rather than incorrectly expecting one search space per configured request model.

  • Preserves all configured request models and the existing model-selection strategy at every sweep point.
  • Aligns sweep generation, configuration counting, and debug descriptions around the same search-space mapping.
  • Adds an integration-style regression test for single- and mixed-model sweeps in normal and debug modes.
  • Documents that multi-model analysis sweeps a shared mixed-model workload rather than independent per-model workloads.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
    A[Configured model list] --> B[Model-selection strategy]
    A --> C[Shared sweep search space]
    C --> D[Sweep objective generator]
    D --> E[Stimulus sweep points]
    E --> F[Request-file generation]
    B --> F
    A --> F
    F --> G[Mixed-model workload per sweep point]
Loading

Reviews (1) · Last reviewed commit: "test: exercise mixed-model sweeps with a..."

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant