PR #29 gets its semantic data from a build script that writes one 329 KB JSON blob into ui/public/ — variable library concepts, ontology terms, curated edges and participant counts in a single object, with the UI reducing over it client-side. Fine for a demo, generalizes to nothing.
We need to define the packet the UI actually receives. Scope is deliberately narrow: terms and variables only — what exists, what it means, what it relates to, who collects it, what shape its values take. Participant counts and row-level results are a query response and a separate contract.
It should be general enough that a filter panel, a visualization builder and consumers we haven't thought of read the same thing. The UI asks about a concept, decides what it needs, asks again — so the shape has to support successive queries rather than one round trip. Sessions don't outlive a graph release; a reload rebuilds.
What to settle:
- Graph-shaped neighbourhood. Nodes and edges anchored at seed terms, bounded by hop depth and by requested node and edge categories — genes matter to some consumer eventually and to a filter panel never. Never the whole graph.
- A declared boundary. Each packet says where it was truncated. Without that the UI can't tell "no such edges exist" from "you didn't ask for them", and successive queries become guesswork.
- Self-contained terms. Every CURIE the packet mentions resolves inside it — no vocabulary knowledge, no second call to render a label.
- Provenance on every relation. Predicate, readable reason, and tier: ontology hierarchy, harvested KG edge, curated and unreviewed. A curated assertion must never look like a published one.
- Variable typing. Value type, units, permissible values, BDCHM binding — so a UI can build a filter for a variable it has never seen instead of waiting for data to infer its shape.
- One envelope at every scale. A search hit, a concept detail and a suggested relation are the same shape with different fill, not three schemas.
The schema should be LinkML with the TypeScript types generated from it — the source-of-truth principle applied to our own interface. Traversal helpers generated alongside are nearly free and save every consumer rewriting the descendant walk.
Defining the contract, not building it. How we reach VarLib and Monarch, and where curated edges live, is separate work.
PR #29 gets its semantic data from a build script that writes one 329 KB JSON blob into
ui/public/— variable library concepts, ontology terms, curated edges and participant counts in a single object, with the UI reducing over it client-side. Fine for a demo, generalizes to nothing.We need to define the packet the UI actually receives. Scope is deliberately narrow: terms and variables only — what exists, what it means, what it relates to, who collects it, what shape its values take. Participant counts and row-level results are a query response and a separate contract.
It should be general enough that a filter panel, a visualization builder and consumers we haven't thought of read the same thing. The UI asks about a concept, decides what it needs, asks again — so the shape has to support successive queries rather than one round trip. Sessions don't outlive a graph release; a reload rebuilds.
What to settle:
The schema should be LinkML with the TypeScript types generated from it — the source-of-truth principle applied to our own interface. Traversal helpers generated alongside are nearly free and save every consumer rewriting the descendant walk.
Defining the contract, not building it. How we reach VarLib and Monarch, and where curated edges live, is separate work.