Skip to content

Latest commit

 

History

History
132 lines (105 loc) · 5.39 KB

File metadata and controls

132 lines (105 loc) · 5.39 KB

Semantic Routing Gold

Symgliph's golden-set format separates retrieval, evidence, answer constraints, and provenance into linked tables. It is intentionally compatible with the Hugging Face dataset viewer and MTEB-style retrieval evaluation without reducing the benchmark to query/response pairs.

Contract

Every record uses symgliph.semantic-routing-gold/v1. The package contains six JSONL configurations:

Configuration Key Purpose
corpus artifact_id Canonical pattern, tool, document, or symbol text
queries query_id Blind user goals, split, language, and annotation tier
qrels query + artifact Graded relevance labels
evidence query + artifact Exact byte ranges, source digests, constraints, and references
negatives query + artifact Incorrect but plausible mined candidates
provenance source_id Immutable revisions, licenses, and transformations

manifest.json commits to the canonical rows with BLAKE3 and records a human collection tier. VALIDATION.json keeps collection value, structural validity, publication readiness, and expert-adjudication readiness as separate claims.

Fail-closed validation

The Rust validator rejects:

  • duplicate artifact, query, qrel, or negative identifiers;
  • dangling references between tables;
  • qrels without digest-bound evidence;
  • stale source digests and invalid byte ranges;
  • blind goals that contain normalized target names;
  • a target simultaneously labeled as a hard negative;
  • duplicate goals or source families crossing split boundaries; and
  • mismatched schemas or split labels.

An unresolved source license does not corrupt local evaluation, so it is a warning that forces publish_ready: false. Missing expert review independently forces gold_ready: false.

Current generated packages

Publishable Fabric seed

datasets/semantic-routing-gold is generated from Fabric commit befdfeefb2402db706f4b54165b8a47ef2cedbad and the passing blind-discovery report.

Table Rows
Corpus 24
Blind queries 6
Qrels 6
Exact evidence 6
Hard negatives 12
Provenance sources 1

Its dataset root is ecb156cfec7ce8c60eb1fda1819b7f9480be392ca58f1160d800a6493d5afe3f. It is designated collection_tier: gold, structurally valid, and publishable under the recorded MIT provenance. gold_ready: false separately records that the six mappings and constraint lists still need independent double review and adjudication before making an expert-gold claim.

Rebuild it with:

cargo run --features golden --bin symgliph-golden -- build-fabric \
  --subset .symglyph/fabric-subsets/befdfeefb2402db706f4b54165b8a47ef2cedbad-24 \
  --spec benchmarks/fabric-discovery.json \
  --proof .symglyph/fabric-discovery-openrouter-proof-befdfeefb240.json \
  --output datasets/semantic-routing-gold

Local ToolRet candidate

scripts/toolret-golden.sh downloads immutable revisions of hugFu/ToolRet-Queries and mangopy/ToolRet-Tools, then normalizes their Parquet tables in Rust. The importer handles the upstream use of bare NaN inside JSON-encoded label documents by converting only non-string non-finite tokens to JSON null; target IDs and relevance grades are unchanged.

The resulting ignored local package contains 44,453 tools, 7,961 queries, and 14,106 qrel/evidence pairs. It validates, but both upstream Hub repositories omit an aggregate license. Its records therefore retain NOASSERTION and the package cannot be published by this workflow.

Combined candidate

cargo run --features golden --bin symgliph-golden -- merge \
  datasets/semantic-routing-gold \
  .symglyph/golden/toolret-candidate \
  --output .symglyph/golden/combined-candidate

The current combined root is e1f39920a4c513a368f13aeb62280c9e4c3c87e019383d4d91f3e2c28e2210db, covering 44,477 artifacts, 7,967 queries, 14,112 qrels, 14,112 exact evidence records, and 12 hard negatives. It remains local because ToolRet's licensing is unresolved.

Path to an expert golden release

  1. Resolve and record every upstream component license before redistribution.
  2. Mine hard negatives with a pinned retriever, then have reviewers reject false negatives.
  3. Have two independent reviewers label target relevance, evidence ranges, ambiguity, acceptable alternatives, and answer constraints.
  4. Send disagreements to a third adjudicator and mark only accepted records as expert_adjudicated.
  5. Split by entire source family, never random rows, to prevent near-duplicate tools from crossing train and test.
  6. Keep final labels in a private or gated evaluation repository; publish only evaluation inputs and the scoring harness to reduce benchmark contamination.
  7. Report retrieval, exact selection, abstention, evidence recall, constraint preservation, digest integrity, token/cost reduction, and latency separately.

The generated Hugging Face card already defines every table as a configuration. Publishing is deliberately not automated: inspect VALIDATION.json, create the target dataset repository, and upload only a package with publish_ready: true after an explicit release decision.

After authenticating hf with a write-scoped token, the fail-closed publisher revalidates the schema, immutable root, gold designation, and publication flag before creating or updating the requested dataset:

scripts/publish-gold-hf.sh codetestcode/semantic-routing-gold