Skip to content

Make the synthetic corpus reflect real study heterogeneity #27

Description

@amc-corey-cox

The corpus in #25 produces two structurally identical cohorts. Real BDC studies differ from each other in specific, checkable ways, and the portal's job is largely to reconcile those differences — so a corpus without them hides the problems it exists to exercise.

The principle

Harmonization aligns on the concept CURIE and preserves value-level granularity. The corpus-wide vocabulary is the union across studies; each study contributes a subset of it, never a reduction of what it actually collected. Race in the real specs shows this exactly:

JHS      OMOP:8516                                                       1 category
CARDIA   OMOP:8516  OMOP:8527                                            2
MESA     OMOP:8515  OMOP:8516  OMOP:8527  UNKNOWN                        4
WHI      OMOP:8515  OMOP:8516  OMOP:8527  OMOP:8552  OMOP:8657           5
CHS      + OMOP:45880900 (more than one race)                            6

Every study's "White" is OMOP:8527. But CHS keeps more than one race even though most studies have no such category, and JHS is single-race because that is the cohort. Nobody is pushed to a common denominator and nobody's extra granularity is folded into "Other".

The alignment axis is the CURIE-valued slot that says what a record is about: observation_type on MeasurementObservation, condition_concept on Condition.

What to add

1. method_type, throughout. We emit none. It appears in 469 of 837 real specs, with a large vocabulary — blood assay (723), anthropometry (345), Electrocardiogram (235), calculated (154), spirometry (132), auscultatory method (104), Seated random-zero average (38). This is what "align without over-condensing" looks like in practice: BMI is one concept in one unit everywhere, but how it was obtained survives. A portal filtering "BMI by anthropometry" has nothing to work with today.

2. Coverage as subsets, not substitutions. A study should be missing concepts entirely — no HDL, no WBC — rather than having a different version of them. Derive the pattern from what VarLib already records about which studies contribute which concepts.

3. Granularity depth. One study records heart failure with a single MONDO term; another distinguishes two because it asked a finer question. Same vocabulary, different depth into it — the CHS more than one race pattern.

4. Chunky sparsity. We sprinkle 0.5% of cells independently. Real gaps are structural: a whole visit with no labs drawn, a measurement absent for early exams.

What NOT to do

Recorded because both were proposed during design and both are wrong:

  • Do not vary units between studies. Harmonization normalizes them — BMI is kg/m2 in all 11 real studies. Varying units would show the portal a problem that harmonization exists to remove.
  • Do not make concept choice a study-level style. ARIC uses both heart-failure MONDO terms within a single spec file, mapped from different source variables, because they capture different things. Concept follows source semantics, not study preference.

Deliberately out of scope

The raw tables are effectively pre-harmonized — they carry target vocabulary directly, so no spec uses value_mappings (against 257 real spec files that do). That weakens the corpus as a transform-layer test case, but it is invisible to the portal, which consumes harmonized output either way. Worth doing eventually for pipeline-test value; not what this issue is about.

Depends on #25.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions