Skip to content

Add an optional benchmarks page for E2E eval results - #43

Merged
joeldickson merged 7 commits into
mainfrom
feature/benchmarks-page
Oct 1, 2026
Merged

joeldickson merged 7 commits into
mainfrom
feature/benchmarks-page

Conversation

@PunchBui

@PunchBui PunchBui commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds an optional Benchmarks page: a models × skills matrix of E2E eval results on main, with a per-test breakdown for each skill. Off by default; enable with DROPMCP_BENCHMARKS=true.

Page

  • An Overall row gives each model's average score, pass count and test coverage; an All models column gives each skill's average.
  • Latest shows the newest result per test and model. 60-day average shows the mean of every run in the window and colours a cell by whether that average meets the threshold. ▲/▼ compare the newest result with the window average, and expanded tests show a trend line of recent scores.
  • Columns are ranked by the selected average. Compact and a model picker (with All / Top 5 presets) keep wide comparisons readable. The view is kept in the URL (metric, sort, compact, models) so it can be shared.
  • The header link is shown only to identified users.

API and data

  • The page reports eval_results_project only. /api/benchmarks takes no project parameter, so callers cannot select another project's data.
  • The data source is supplied by the deployment as benchmark_results_store. Enabling the page without a project or a valid store fails at startup. No query, host or credential defaults ship in the library, in line with Stop shipping deployment-specific StarRocks queries #41.
  • MySQLBenchmarkResultsStore (in eval_results_mysql.py, mysql extra) runs caller-supplied SQL bound as (project, datadate). Credentials and the datadate format are chosen by the caller, so partition keys that are not ISO dates prune correctly. Query failures propagate so the page reports an outage rather than an empty result.
  • /api/benchmarks answers 401 without the identity header and never returns reasoning or error text. Responses are no-store, nosniff and no-referrer.
  • Every run in the lookback window is kept, so history needs no extra query.

Things to know

  • The identity header is trusted as sent, exactly as /api/me and the subscription routes already do. The 401 and the hidden link are a usability gate, not access control; put an authenticating proxy in front of the server and deny /benchmarks and /api/benchmarks to everyone else. The README says so.
  • Payload size grows with tests × models (about 390 KB uncompressed for 46 tests × 20 models, about 70 KB gzipped). Loading a skill's per-test detail on expand would be the fix if that becomes a problem; not done here.
  • History mixes commits. A real improvement to a skill shifts the average, so read it with the ▲/▼ and the run count.
  • New dependency: @tanstack/react-query (MIT, about 12 kB gzipped). The page loads its data with useQuery and a single QueryClientProvider at the root, instead of useEffect plus hand-rolled loading and error state. It is the only consumer; the other pages still fetch in useEffect. It is its own commit (efa338c), so it is easy to drop or to migrate the rest over if you would rather not add the dependency in this PR.
  • The existing _datadate_cutoff in eval_results_mysql.py is untouched. It still binds an ISO date; changing it would change which rows /api/telemetry returns.

Test plan

  • ruff check src
  • pytest: 249 passed (203 before this change)
  • npm run lint and npm run build
  • npx playwright test tests/benchmarks.spec.ts: 27 passed. The existing screenshot tests only have Linux baselines, so they were not run locally; this change adds no screenshot snapshots.
  • Tests were checked against deliberate breakages (auth check, cache header, ordering, colouring, partition format, connection cleanup) and fail as expected.
  • Ran the server locally against a fake store with 20 models and looked at light, dark, compact, history and mobile layouts.
  • Not run against a real database.

🤖 Generated with Claude Code

A models x skills matrix of E2E results on main, with a per-test breakdown
for each skill. Off by default; enable with DROPMCP_BENCHMARKS=true.

Page
- An Overall row gives each model's average score, pass count and test
  coverage; an All models column gives each skill's average.
- Latest shows the newest result per test and model. 60-day average shows the
  mean of every run in the window, coloured by whether it meets the threshold.
  Arrows compare the newest result with the window average, and expanded tests
  show a trend line of recent scores.
- Columns are ranked by the selected average. Compact drops the detail lines
  and a model picker (with All and Top 5 presets) hides columns, for wide
  comparisons. The view is kept in the URL so it can be shared.
- The header link is shown only to identified users.

API and data
- The page reports eval_results_project only. /api/benchmarks takes no
  project parameter, so callers cannot select another project's data.
- The data source is supplied by the deployment as benchmark_results_store.
  Enabling the page without a project or a valid store fails at startup. No
  query, host or credential defaults ship in the library.
- MySQLBenchmarkResultsStore runs caller-supplied SQL bound as
  (project, datadate), with credentials and the datadate format chosen by
  the caller. Query failures propagate so the page can report an outage
  instead of an empty result.
- /api/benchmarks answers 401 without the identity header and never returns
  reasoning or error text. Responses are no-store, nosniff and no-referrer.
  The identity header is trusted as sent, so this is a usability gate;
  access control belongs in front of the server.
- Every run in the lookback window is kept, so history needs no extra query.
  Payload grows with tests x models.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@PunchBui PunchBui self-assigned this Sep 30, 2026
@PunchBui
PunchBui requested review from dicko2 and joeldickson and removed request for joeldickson September 30, 2026 11:28
Dachawatsathon and others added 5 commits October 1, 2026 09:58
- Read the view straight from the URL during render instead of memoising a
  cheap parse; there is no React Compiler here, so the memo bought nothing.
- Update the URL with a functional setSearchParams so the handler never
  closes over a stale copy of the params, and drop the useCallback that
  existed only to carry that closure into a component that is not memoised.
- Render the status banner with an explicit ternary rather than && on a
  string.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Fetch the page's data with useQuery instead of hand-rolled useState and
useEffect for data, loading and error. This drops the effect and its three
pieces of state, and gives the page request deduplication and a cache for
free, so returning to the page within a minute shows the last result
immediately.

- Adds @tanstack/react-query (MIT) and a single QueryClientProvider at the
  root. The benchmarks page is its only consumer; the other pages still fetch
  in useEffect.
- Retries are off so a 401 reports "Sign in" straight away instead of after
  several backoffs.
- Results are fresh for a minute, in line with the server-side cache.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Move data loading and URL view state into hooks, give the page, results,
controls, model picker, matrix, skill rows and cells one visible region
each, and pull Sparkline and Pill into their own files. Rendered markup
and styles are unchanged; tests/benchmarks.spec.ts passes untouched.
Split BenchmarkCells into one file per cell (Empty, Score, Overall,
Skill, Test) and move the header row, overall row, test row and legend
out of their parents. No markup or style changes.
react-component-layout: replace the 60-80 and 15 line limits with a
structural rule. A file names one thing on screen and exports one
component; parts stay with their region, regions (rows, panels,
sections, pickers, anything another file needs) get their own file,
and files grouped by kind (Cells, Rows, helpers) are buckets.

Apply it across the client:
- RepoFeedbackPage: toolbar, list, card, card header, details panel and
  triage row move to components/.
- Shared feedback filters (search field, option filter, value filter),
  FeedbackField and FeedbackResolutionLink replace the copies in
  FeedbackToolbar and RepoFeedbackPage.
- CatalogGrid's extra exports become CatalogSkeletonGrid,
  CatalogNoMatches and ErrorState; CatalogPage's content states become
  CatalogResults.
- SearchToolbar's group-subscription row becomes
  GroupSubscriptionFilter; InstallPanel splits into tab list, tab
  panels and CopyableSnippet; DetailPage's hero, header and sections
  get their own files on a shared DetailSection.
- FeedbackCard header and FeedbackDetailsPanel's artifact list get
  their own files.

No markup or style changes; CSS rules move with their components.
react-component-layout gains two sections: folders follow the screen
(one folder per area, shared components at the top, a sub-folder only
for a region with its own family, two levels at most) and keep the
component tree shallow (compose at the parent, no pass-through
wrappers).

Apply it to the client: components/ now holds catalog/ (install/),
detail/ (resources/, telemetry/), feedback/ (skill/, repo/) and
benchmarks/ (controls/, matrix/), with Header, Footer, FeedbackHeader
and ErrorState shared at the top. ScreenshotsSection only wrapped
DetailSection around ScreenshotGallery, so DetailPage composes them
directly. No markup or style changes.
@joeldickson
joeldickson merged commit 31d8cfe into main Oct 1, 2026
4 checks passed
joeldickson added a commit to agoda-com/Local-Dev-Telemetry-Manager that referenced this pull request Oct 1, 2026
…e limits (#26)

* react-component-layout: replace line limits with one file, one region.

Drop the 60-80 and 15 line limits. A file names one thing on screen and
exports one component; parts stay with their region, regions (rows,
panels, sections, pickers, anything another file needs) get their own
file, and files grouped by kind (Cells, Rows, helpers) are buckets.
Also adds the view-model section from the canonical copy.
Same text as agoda-com/dropmcp#43.

* react-component-layout: folders follow the screen, keep the tree shallow.

Add two sections: one folder per area under components/, shared
components at the top (no shared/ or common/), a sub-folder only for a
region with its own family, two levels at most, no one-file folders,
folders named after the screen; and keep the component tree shallow by
composing at the parent instead of chaining pass-through wrappers.

---------

Co-authored-by: Joel Dickson <joel.dickson@agoda.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants