Skip to content

Generate summary statistics from the synthetic corpus against the draft schema #48

Description

@amc-corey-cox

A schema that has never been populated is a guess. Once the BDCHM extension from the sibling issue exists, the fastest way to find out whether it holds is to build an instance of it from the synthetic corpus.

The corpus is the right test subject: it is BDCHM-conformant, transformed through dm-bip against a pinned BDCHM, already typed to Parquet by synthetic/schema.py, and it is the only end-to-end data we control. It also spans the shapes that tend to break assumptions — nested MeasurementObservationSet, temporal structure, and concepts drawn from more than one source vocabulary.

What this should produce:

  • Summary statistics generated from the corpus, conforming to the draft schema
  • A record of where the schema did not fit — slots that had nowhere to put a value, or values with no slot

The second item is the point. This exists to feed corrections back into the draft before anything is built on it, not to produce a publishable artifact.

Depends on the schema draft. Scoped deliberately small; if the sprint tightens, this is the one to drop, since the draft is what tis-lab/BDC-Portal#62 asks for.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Metadata Source of TruthLinkML metadata, semantic bindings, query engine

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions