Symgliph's golden-set format separates retrieval, evidence, answer constraints, and provenance into linked tables. It is intentionally compatible with the Hugging Face dataset viewer and MTEB-style retrieval evaluation without reducing the benchmark to query/response pairs.
Every record uses symgliph.semantic-routing-gold/v1. The package contains six
JSONL configurations:
| Configuration | Key | Purpose |
|---|---|---|
corpus |
artifact_id |
Canonical pattern, tool, document, or symbol text |
queries |
query_id |
Blind user goals, split, language, and annotation tier |
qrels |
query + artifact | Graded relevance labels |
evidence |
query + artifact | Exact byte ranges, source digests, constraints, and references |
negatives |
query + artifact | Incorrect but plausible mined candidates |
provenance |
source_id |
Immutable revisions, licenses, and transformations |
manifest.json commits to the canonical rows with BLAKE3 and records a human
collection tier. VALIDATION.json keeps collection value, structural validity,
publication readiness, and expert-adjudication readiness as separate claims.
The Rust validator rejects:
- duplicate artifact, query, qrel, or negative identifiers;
- dangling references between tables;
- qrels without digest-bound evidence;
- stale source digests and invalid byte ranges;
- blind goals that contain normalized target names;
- a target simultaneously labeled as a hard negative;
- duplicate goals or source families crossing split boundaries; and
- mismatched schemas or split labels.
An unresolved source license does not corrupt local evaluation, so it is a
warning that forces publish_ready: false. Missing expert review independently
forces gold_ready: false.
datasets/semantic-routing-gold is generated from Fabric commit
befdfeefb2402db706f4b54165b8a47ef2cedbad and the passing blind-discovery
report.
| Table | Rows |
|---|---|
| Corpus | 24 |
| Blind queries | 6 |
| Qrels | 6 |
| Exact evidence | 6 |
| Hard negatives | 12 |
| Provenance sources | 1 |
Its dataset root is
ecb156cfec7ce8c60eb1fda1819b7f9480be392ca58f1160d800a6493d5afe3f.
It is designated collection_tier: gold, structurally valid, and publishable
under the recorded MIT provenance. gold_ready: false separately records that
the six mappings and constraint lists still need independent double review and
adjudication before making an expert-gold claim.
Rebuild it with:
cargo run --features golden --bin symgliph-golden -- build-fabric \
--subset .symglyph/fabric-subsets/befdfeefb2402db706f4b54165b8a47ef2cedbad-24 \
--spec benchmarks/fabric-discovery.json \
--proof .symglyph/fabric-discovery-openrouter-proof-befdfeefb240.json \
--output datasets/semantic-routing-goldscripts/toolret-golden.sh downloads immutable revisions of
hugFu/ToolRet-Queries and mangopy/ToolRet-Tools, then normalizes their
Parquet tables in Rust. The importer handles the upstream use of bare NaN
inside JSON-encoded label documents by converting only non-string non-finite
tokens to JSON null; target IDs and relevance grades are unchanged.
The resulting ignored local package contains 44,453 tools, 7,961 queries, and
14,106 qrel/evidence pairs. It validates, but both upstream Hub repositories
omit an aggregate license. Its records therefore retain NOASSERTION and the
package cannot be published by this workflow.
cargo run --features golden --bin symgliph-golden -- merge \
datasets/semantic-routing-gold \
.symglyph/golden/toolret-candidate \
--output .symglyph/golden/combined-candidateThe current combined root is
e1f39920a4c513a368f13aeb62280c9e4c3c87e019383d4d91f3e2c28e2210db,
covering 44,477 artifacts, 7,967 queries, 14,112 qrels, 14,112 exact evidence
records, and 12 hard negatives. It remains local because ToolRet's licensing
is unresolved.
- Resolve and record every upstream component license before redistribution.
- Mine hard negatives with a pinned retriever, then have reviewers reject false negatives.
- Have two independent reviewers label target relevance, evidence ranges, ambiguity, acceptable alternatives, and answer constraints.
- Send disagreements to a third adjudicator and mark only accepted records as
expert_adjudicated. - Split by entire source family, never random rows, to prevent near-duplicate tools from crossing train and test.
- Keep final labels in a private or gated evaluation repository; publish only evaluation inputs and the scoring harness to reduce benchmark contamination.
- Report retrieval, exact selection, abstention, evidence recall, constraint preservation, digest integrity, token/cost reduction, and latency separately.
The generated Hugging Face card already defines every table as a configuration.
Publishing is deliberately not automated: inspect VALIDATION.json, create the
target dataset repository, and upload only a package with
publish_ready: true after an explicit release decision.
After authenticating hf with a write-scoped token, the fail-closed publisher
revalidates the schema, immutable root, gold designation, and publication flag
before creating or updating the requested dataset:
scripts/publish-gold-hf.sh codetestcode/semantic-routing-gold