The hypertension and Type 2 diabetes code sets are compiled and in use — 13 HTN subtypes and 11 for T2D encoded in synthetic/vocab.py, each row carrying MONDO, HPO, SNOMED CT and ICD-10-CM together, drawn on by the corpus (#33, #34). What does not exist is a document that defines the cohorts: inclusion and exclusion rules, and a code breakdown table someone can read without opening the generator.
The current documentation is prose in synthetic/README.md aimed at explaining the corpus, not at defining a cohort. Three things learned while encoding the code sets belong in that document:
- BDCHM has no slot for source terminology.
Condition.condition_concept is single-valued over ConditionConceptEnum, which is MONDO ∪ HPO, so SNOMED and ICD-10-CM cannot be carried in harmonized output at all. In this corpus they stay in the raw dbGaP-style tables. Any cohort definition keyed on ICD has to say which layer it applies to.
- T2D granularity lives in ICD-10-CM, not MONDO. The code-set reference gives
MONDO:0005148 for nearly every T2D row and puts the complication detail in the ICD column. A definition at complication level cannot be written in MONDO alone.
- Some reference cells name a hierarchy rather than a term (
"44054006 hierarchy + renal complication descendants", "E11.3*"). Those were treated as absent rather than guessed at. Two rows are unused: pregnancy-related hypertension, since both cohorts enrol at 45-78, and T2D with ophthalmic complications, where a wildcard ICD plus the shared root MONDO leaves it indistinguishable from any other T2D record.
Worth deciding as part of this: whether the document defines cohorts against the raw layer, the harmonized layer, or both, since the available codes differ between them.
The hypertension and Type 2 diabetes code sets are compiled and in use — 13 HTN subtypes and 11 for T2D encoded in
synthetic/vocab.py, each row carrying MONDO, HPO, SNOMED CT and ICD-10-CM together, drawn on by the corpus (#33, #34). What does not exist is a document that defines the cohorts: inclusion and exclusion rules, and a code breakdown table someone can read without opening the generator.The current documentation is prose in
synthetic/README.mdaimed at explaining the corpus, not at defining a cohort. Three things learned while encoding the code sets belong in that document:Condition.condition_conceptis single-valued overConditionConceptEnum, which is MONDO ∪ HPO, so SNOMED and ICD-10-CM cannot be carried in harmonized output at all. In this corpus they stay in the raw dbGaP-style tables. Any cohort definition keyed on ICD has to say which layer it applies to.MONDO:0005148for nearly every T2D row and puts the complication detail in the ICD column. A definition at complication level cannot be written in MONDO alone."44054006 hierarchy + renal complication descendants","E11.3*"). Those were treated as absent rather than guessed at. Two rows are unused: pregnancy-related hypertension, since both cohorts enrol at 45-78, and T2D with ophthalmic complications, where a wildcard ICD plus the shared root MONDO leaves it indistinguishable from any other T2D record.Worth deciding as part of this: whether the document defines cohorts against the raw layer, the harmonized layer, or both, since the available codes differ between them.