A schema that has never been populated is a guess. Once the BDCHM extension from the sibling issue exists, the fastest way to find out whether it holds is to build an instance of it from the synthetic corpus.
The corpus is the right test subject: it is BDCHM-conformant, transformed through dm-bip against a pinned BDCHM, already typed to Parquet by synthetic/schema.py, and it is the only end-to-end data we control. It also spans the shapes that tend to break assumptions — nested MeasurementObservationSet, temporal structure, and concepts drawn from more than one source vocabulary.
What this should produce:
- Summary statistics generated from the corpus, conforming to the draft schema
- A record of where the schema did not fit — slots that had nowhere to put a value, or values with no slot
The second item is the point. This exists to feed corrections back into the draft before anything is built on it, not to produce a publishable artifact.
Depends on the schema draft. Scoped deliberately small; if the sprint tightens, this is the one to drop, since the draft is what tis-lab/BDC-Portal#62 asks for.
A schema that has never been populated is a guess. Once the BDCHM extension from the sibling issue exists, the fastest way to find out whether it holds is to build an instance of it from the synthetic corpus.
The corpus is the right test subject: it is BDCHM-conformant, transformed through dm-bip against a pinned BDCHM, already typed to Parquet by
synthetic/schema.py, and it is the only end-to-end data we control. It also spans the shapes that tend to break assumptions — nestedMeasurementObservationSet, temporal structure, and concepts drawn from more than one source vocabulary.What this should produce:
The second item is the point. This exists to feed corrections back into the draft before anything is built on it, not to produce a publishable artifact.
Depends on the schema draft. Scoped deliberately small; if the sprint tightens, this is the one to drop, since the draft is what tis-lab/BDC-Portal#62 asks for.