Add an optional benchmarks page for E2E eval results - #43
Merged
Merged
Conversation
A models x skills matrix of E2E results on main, with a per-test breakdown for each skill. Off by default; enable with DROPMCP_BENCHMARKS=true. Page - An Overall row gives each model's average score, pass count and test coverage; an All models column gives each skill's average. - Latest shows the newest result per test and model. 60-day average shows the mean of every run in the window, coloured by whether it meets the threshold. Arrows compare the newest result with the window average, and expanded tests show a trend line of recent scores. - Columns are ranked by the selected average. Compact drops the detail lines and a model picker (with All and Top 5 presets) hides columns, for wide comparisons. The view is kept in the URL so it can be shared. - The header link is shown only to identified users. API and data - The page reports eval_results_project only. /api/benchmarks takes no project parameter, so callers cannot select another project's data. - The data source is supplied by the deployment as benchmark_results_store. Enabling the page without a project or a valid store fails at startup. No query, host or credential defaults ship in the library. - MySQLBenchmarkResultsStore runs caller-supplied SQL bound as (project, datadate), with credentials and the datadate format chosen by the caller. Query failures propagate so the page can report an outage instead of an empty result. - /api/benchmarks answers 401 without the identity header and never returns reasoning or error text. Responses are no-store, nosniff and no-referrer. The identity header is trusted as sent, so this is a usability gate; access control belongs in front of the server. - Every run in the lookback window is kept, so history needs no extra query. Payload grows with tests x models. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
PunchBui
requested review from
dicko2 and
joeldickson
and removed request for
joeldickson
September 30, 2026 11:28
- Read the view straight from the URL during render instead of memoising a cheap parse; there is no React Compiler here, so the memo bought nothing. - Update the URL with a functional setSearchParams so the handler never closes over a stale copy of the params, and drop the useCallback that existed only to carry that closure into a component that is not memoised. - Render the status banner with an explicit ternary rather than && on a string. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Fetch the page's data with useQuery instead of hand-rolled useState and useEffect for data, loading and error. This drops the effect and its three pieces of state, and gives the page request deduplication and a cache for free, so returning to the page within a minute shows the last result immediately. - Adds @tanstack/react-query (MIT) and a single QueryClientProvider at the root. The benchmarks page is its only consumer; the other pages still fetch in useEffect. - Retries are off so a 401 reports "Sign in" straight away instead of after several backoffs. - Results are fresh for a minute, in line with the server-side cache. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Move data loading and URL view state into hooks, give the page, results, controls, model picker, matrix, skill rows and cells one visible region each, and pull Sparkline and Pill into their own files. Rendered markup and styles are unchanged; tests/benchmarks.spec.ts passes untouched.
Split BenchmarkCells into one file per cell (Empty, Score, Overall, Skill, Test) and move the header row, overall row, test row and legend out of their parents. No markup or style changes.
react-component-layout: replace the 60-80 and 15 line limits with a structural rule. A file names one thing on screen and exports one component; parts stay with their region, regions (rows, panels, sections, pickers, anything another file needs) get their own file, and files grouped by kind (Cells, Rows, helpers) are buckets. Apply it across the client: - RepoFeedbackPage: toolbar, list, card, card header, details panel and triage row move to components/. - Shared feedback filters (search field, option filter, value filter), FeedbackField and FeedbackResolutionLink replace the copies in FeedbackToolbar and RepoFeedbackPage. - CatalogGrid's extra exports become CatalogSkeletonGrid, CatalogNoMatches and ErrorState; CatalogPage's content states become CatalogResults. - SearchToolbar's group-subscription row becomes GroupSubscriptionFilter; InstallPanel splits into tab list, tab panels and CopyableSnippet; DetailPage's hero, header and sections get their own files on a shared DetailSection. - FeedbackCard header and FeedbackDetailsPanel's artifact list get their own files. No markup or style changes; CSS rules move with their components.
react-component-layout gains two sections: folders follow the screen (one folder per area, shared components at the top, a sub-folder only for a region with its own family, two levels at most) and keep the component tree shallow (compose at the parent, no pass-through wrappers). Apply it to the client: components/ now holds catalog/ (install/), detail/ (resources/, telemetry/), feedback/ (skill/, repo/) and benchmarks/ (controls/, matrix/), with Header, Footer, FeedbackHeader and ErrorState shared at the top. ScreenshotsSection only wrapped DetailSection around ScreenshotGallery, so DetailPage composes them directly. No markup or style changes.
joeldickson
added a commit
to agoda-com/Local-Dev-Telemetry-Manager
that referenced
this pull request
Oct 1, 2026
…e limits (#26) * react-component-layout: replace line limits with one file, one region. Drop the 60-80 and 15 line limits. A file names one thing on screen and exports one component; parts stay with their region, regions (rows, panels, sections, pickers, anything another file needs) get their own file, and files grouped by kind (Cells, Rows, helpers) are buckets. Also adds the view-model section from the canonical copy. Same text as agoda-com/dropmcp#43. * react-component-layout: folders follow the screen, keep the tree shallow. Add two sections: one folder per area under components/, shared components at the top (no shared/ or common/), a sub-folder only for a region with its own family, two levels at most, no one-file folders, folders named after the screen; and keep the component tree shallow by composing at the parent instead of chaining pass-through wrappers. --------- Co-authored-by: Joel Dickson <joel.dickson@agoda.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds an optional Benchmarks page: a models × skills matrix of E2E eval results on
main, with a per-test breakdown for each skill. Off by default; enable withDROPMCP_BENCHMARKS=true.Page
metric,sort,compact,models) so it can be shared.API and data
eval_results_projectonly./api/benchmarkstakes no project parameter, so callers cannot select another project's data.benchmark_results_store. Enabling the page without a project or a valid store fails at startup. No query, host or credential defaults ship in the library, in line with Stop shipping deployment-specific StarRocks queries #41.MySQLBenchmarkResultsStore(ineval_results_mysql.py,mysqlextra) runs caller-supplied SQL bound as(project, datadate). Credentials and thedatadateformat are chosen by the caller, so partition keys that are not ISO dates prune correctly. Query failures propagate so the page reports an outage rather than an empty result./api/benchmarksanswers401without the identity header and never returns reasoning or error text. Responses areno-store,nosniffandno-referrer.Things to know
/api/meand the subscription routes already do. The401and the hidden link are a usability gate, not access control; put an authenticating proxy in front of the server and deny/benchmarksand/api/benchmarksto everyone else. The README says so.@tanstack/react-query(MIT, about 12 kB gzipped). The page loads its data withuseQueryand a singleQueryClientProviderat the root, instead ofuseEffectplus hand-rolled loading and error state. It is the only consumer; the other pages still fetch inuseEffect. It is its own commit (efa338c), so it is easy to drop or to migrate the rest over if you would rather not add the dependency in this PR._datadate_cutoffineval_results_mysql.pyis untouched. It still binds an ISO date; changing it would change which rows/api/telemetryreturns.Test plan
ruff check srcpytest: 249 passed (203 before this change)npm run lintandnpm run buildnpx playwright test tests/benchmarks.spec.ts: 27 passed. The existing screenshot tests only have Linux baselines, so they were not run locally; this change adds no screenshot snapshots.🤖 Generated with Claude Code