The corpus in #25 produces two structurally identical cohorts. Real BDC studies differ from each other in specific, checkable ways, and the portal's job is largely to reconcile those differences — so a corpus without them hides the problems it exists to exercise.
The principle
Harmonization aligns on the concept CURIE and preserves value-level granularity. The corpus-wide vocabulary is the union across studies; each study contributes a subset of it, never a reduction of what it actually collected. Race in the real specs shows this exactly:
JHS OMOP:8516 1 category
CARDIA OMOP:8516 OMOP:8527 2
MESA OMOP:8515 OMOP:8516 OMOP:8527 UNKNOWN 4
WHI OMOP:8515 OMOP:8516 OMOP:8527 OMOP:8552 OMOP:8657 5
CHS + OMOP:45880900 (more than one race) 6
Every study's "White" is OMOP:8527. But CHS keeps more than one race even though most studies have no such category, and JHS is single-race because that is the cohort. Nobody is pushed to a common denominator and nobody's extra granularity is folded into "Other".
The alignment axis is the CURIE-valued slot that says what a record is about: observation_type on MeasurementObservation, condition_concept on Condition.
What to add
1. method_type, throughout. We emit none. It appears in 469 of 837 real specs, with a large vocabulary — blood assay (723), anthropometry (345), Electrocardiogram (235), calculated (154), spirometry (132), auscultatory method (104), Seated random-zero average (38). This is what "align without over-condensing" looks like in practice: BMI is one concept in one unit everywhere, but how it was obtained survives. A portal filtering "BMI by anthropometry" has nothing to work with today.
2. Coverage as subsets, not substitutions. A study should be missing concepts entirely — no HDL, no WBC — rather than having a different version of them. Derive the pattern from what VarLib already records about which studies contribute which concepts.
3. Granularity depth. One study records heart failure with a single MONDO term; another distinguishes two because it asked a finer question. Same vocabulary, different depth into it — the CHS more than one race pattern.
4. Chunky sparsity. We sprinkle 0.5% of cells independently. Real gaps are structural: a whole visit with no labs drawn, a measurement absent for early exams.
What NOT to do
Recorded because both were proposed during design and both are wrong:
- Do not vary units between studies. Harmonization normalizes them — BMI is
kg/m2 in all 11 real studies. Varying units would show the portal a problem that harmonization exists to remove.
- Do not make concept choice a study-level style. ARIC uses both heart-failure MONDO terms within a single spec file, mapped from different source variables, because they capture different things. Concept follows source semantics, not study preference.
Deliberately out of scope
The raw tables are effectively pre-harmonized — they carry target vocabulary directly, so no spec uses value_mappings (against 257 real spec files that do). That weakens the corpus as a transform-layer test case, but it is invisible to the portal, which consumes harmonized output either way. Worth doing eventually for pipeline-test value; not what this issue is about.
Depends on #25.
The corpus in #25 produces two structurally identical cohorts. Real BDC studies differ from each other in specific, checkable ways, and the portal's job is largely to reconcile those differences — so a corpus without them hides the problems it exists to exercise.
The principle
Harmonization aligns on the concept CURIE and preserves value-level granularity. The corpus-wide vocabulary is the union across studies; each study contributes a subset of it, never a reduction of what it actually collected. Race in the real specs shows this exactly:
Every study's "White" is
OMOP:8527. But CHS keeps more than one race even though most studies have no such category, and JHS is single-race because that is the cohort. Nobody is pushed to a common denominator and nobody's extra granularity is folded into "Other".The alignment axis is the CURIE-valued slot that says what a record is about:
observation_typeonMeasurementObservation,condition_conceptonCondition.What to add
1.
method_type, throughout. We emit none. It appears in 469 of 837 real specs, with a large vocabulary —blood assay(723),anthropometry(345),Electrocardiogram(235),calculated(154),spirometry(132),auscultatory method(104),Seated random-zero average(38). This is what "align without over-condensing" looks like in practice: BMI is one concept in one unit everywhere, but how it was obtained survives. A portal filtering "BMI by anthropometry" has nothing to work with today.2. Coverage as subsets, not substitutions. A study should be missing concepts entirely — no HDL, no WBC — rather than having a different version of them. Derive the pattern from what VarLib already records about which studies contribute which concepts.
3. Granularity depth. One study records heart failure with a single MONDO term; another distinguishes two because it asked a finer question. Same vocabulary, different depth into it — the CHS more than one race pattern.
4. Chunky sparsity. We sprinkle 0.5% of cells independently. Real gaps are structural: a whole visit with no labs drawn, a measurement absent for early exams.
What NOT to do
Recorded because both were proposed during design and both are wrong:
kg/m2in all 11 real studies. Varying units would show the portal a problem that harmonization exists to remove.Deliberately out of scope
The raw tables are effectively pre-harmonized — they carry target vocabulary directly, so no spec uses
value_mappings(against 257 real spec files that do). That weakens the corpus as a transform-layer test case, but it is invisible to the portal, which consumes harmonized output either way. Worth doing eventually for pipeline-test value; not what this issue is about.Depends on #25.