ui/src/demoData.ts is 160 lines of invented labels — "Coronary artery disease", "Echocardiography" — plain strings with no vocabulary behind them. The synthetic corpus (#25) produces the same shapes with real MONDO, HP, and OMOP concepts and real BDCHM structure, so the demo could be exercising the vocabulary a portal will actually meet rather than a parallel invention.
That matters for the same reason the corpus goes through dm-bip instead of emitting BDCHM directly: work built against invented labels needs rework when it meets real data.
Size is not a constraint. Flattening one cohort's participants to JSON comes out at 56KB raw and 1KB gzipped — CURIEs repeat heavily and compress away. Both cohorts land around 110KB raw.
This is one change covering two contexts: docs/index.html links to study-palette.netlify.app, so the Pages "live demo" and the PR previews are the same app in different deploy contexts.
Three things need deciding.
Labels are the actual work. Shipping CURIEs alone gives you a donut segment labelled MONDO:0005068. Sex, race, and ethnicity are BDCHM enum members carrying descriptions we can read locally, so those are free. Condition concepts resolve through reachable_from MONDO and HP, so there are no local labels — those need an ontology lookup at build time, cached into the fixture rather than fetched in the browser.
The corpus has no procedures. The second donut is Procedures, and the corpus covers conditions, measurements, drug exposures, visits, and demography. Either add Procedure records upstream — probably worth doing regardless, since a portal will need them and it helps coverage — or re-point that chart at drug exposures, which we do have.
Where the derived fixture lives. Netlify builds from git and will not run dm-bip, so the flattened JSON is either committed or fetched from a release during the build. Committing ~110KB is a fixture rather than a corpus and is much simpler, but it does soften the rule set in #25 that generated data stays out of git — worth restating the line explicitly as "corpus not committed, small derived fixture is" rather than letting it blur. Fetching keeps the rule clean at the cost of making PR previews depend on release plumbing that does not exist yet.
One property in favour: the flattening step is the index-assembly job in miniature, so the demo path would exercise the same shape as the real pipeline instead of being a separate invention.
Depends on #25.
ui/src/demoData.tsis 160 lines of invented labels —"Coronary artery disease","Echocardiography"— plain strings with no vocabulary behind them. The synthetic corpus (#25) produces the same shapes with real MONDO, HP, and OMOP concepts and real BDCHM structure, so the demo could be exercising the vocabulary a portal will actually meet rather than a parallel invention.That matters for the same reason the corpus goes through dm-bip instead of emitting BDCHM directly: work built against invented labels needs rework when it meets real data.
Size is not a constraint. Flattening one cohort's participants to JSON comes out at 56KB raw and 1KB gzipped — CURIEs repeat heavily and compress away. Both cohorts land around 110KB raw.
This is one change covering two contexts:
docs/index.htmllinks tostudy-palette.netlify.app, so the Pages "live demo" and the PR previews are the same app in different deploy contexts.Three things need deciding.
Labels are the actual work. Shipping CURIEs alone gives you a donut segment labelled
MONDO:0005068. Sex, race, and ethnicity are BDCHM enum members carrying descriptions we can read locally, so those are free. Condition concepts resolve throughreachable_fromMONDO and HP, so there are no local labels — those need an ontology lookup at build time, cached into the fixture rather than fetched in the browser.The corpus has no procedures. The second donut is Procedures, and the corpus covers conditions, measurements, drug exposures, visits, and demography. Either add
Procedurerecords upstream — probably worth doing regardless, since a portal will need them and it helps coverage — or re-point that chart at drug exposures, which we do have.Where the derived fixture lives. Netlify builds from git and will not run dm-bip, so the flattened JSON is either committed or fetched from a release during the build. Committing ~110KB is a fixture rather than a corpus and is much simpler, but it does soften the rule set in #25 that generated data stays out of git — worth restating the line explicitly as "corpus not committed, small derived fixture is" rather than letting it blur. Fetching keeps the rule clean at the cost of making PR previews depend on release plumbing that does not exist yet.
One property in favour: the flattening step is the index-assembly job in miniature, so the demo path would exercise the same shape as the real pipeline instead of being a separate invention.
Depends on #25.