Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
82 changes: 79 additions & 3 deletions synthetic/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,8 +42,16 @@ which slots nest, because the model alone does not decide that: BDCHM gives
as a uuid5 string while `value_quantity` is nested inline.

It also reports where the model and the data disagree on cardinality. That is
not hypothetical — BDCHM declares `identity` multivalued and every
transformation spec, RTI's included, emits a scalar.
not hypothetical — BDCHM declares `identity` and `Condition.associated_evidence`
multivalued, and every transformation spec, RTI's included, emits a scalar for
both.

The `associated_evidence` case has a sharp edge worth knowing about: in
linkml-map a `value:` derivation emits a scalar, while `populated_from` with
`value_mappings` honours the multivalued declaration and emits a list. Using
both for one slot leaves it carrying two shapes, and `schema.py` types a column
from the first record it sees — so it would get the rest wrong. The specs here
branch with `case()` instead, which stays scalar and matches what RTI emits.

Then, from a dm-bip checkout:

Expand Down Expand Up @@ -83,13 +91,79 @@ without running the pipeline. Regenerate it with `python sample.py` after a run.
| Race | 50/30/10/10 white, black, Asian, American Indian | 60/20/10/10 white, black, Middle Eastern, Native Hawaiian |
| BMI | continuous | categorical |

Both carry the same five conditions — heart failure, family history of stroke,
hypertension, Type 2 diabetes, and myocardial infarction — and the same
measurements, including fasting glucose and HbA1c.

5% of Study Two's participants are the same individuals as in Study One. They
share a `dbGaP_Subject_ID`, so harmonization resolves them to one `Person` with
two `Participant` records — the same way real cross-study participation appears.

One visit per participant is `TELEHEALTH`; the rest are `STUDY_SITE_VISIT`.
Deceased participants stop attending, so nobody is measured after they die.

## Time

Conditions carry `age_at_condition_start`, `age_at_condition_end` where they
resolved, and the `associated_visit` that recorded them. Measurements carry
`age_at_observation` and their own `associated_visit` — including the systolic
and diastolic observations nested inside a blood pressure set, which are
queryable on their own and should not need a join back through the set to be
placed in time.

The reason to populate those slots is that a quarter of diagnoses are **incident**:
made between two visits rather than before enrolment. A condition only makes
the measurements it explains abnormal from its diagnosis onwards, and stops
once it resolves, so the same participant's results fall into two distinct
distributions:

```
glucose at or above 126 mg/dL, participants diagnosed with T2D mid-study
before the diagnosis 10.3% (the corpus-wide rate)
after the diagnosis 52.2% (the diabetic rate)
```

A corpus where every diagnosis predates every visit answers "measurements after
diagnosis" with the whole cohort, and the temporal slots are decoration.
`validate.py` checks the separation rather than merely checking the slots are
non-null.

Family history is the exception: `age_at_condition_start` is defined as the
*participant's* age, which means nothing for a condition a parent had, so
family-history records carry the recording visit and no ages.

## Concepts and source terminology

Hypertension and Type 2 diabetes are coded from the BDC cohort-readiness
code-set reference — 13 hypertension subtypes and 11 for T2D, against the
single `HP:0000822` and no diabetes at all that the corpus started with. A
participant screened and found negative is coded to the hierarchy root, since a
questionnaire asks about hypertension in general; a subtype is recorded once
there is a diagnosis.

SNOMED CT and ICD-10-CM appear **only in the raw dbGaP-style tables**. BDCHM has
no slot for source terminology anywhere — `condition_concept` is single-valued
over `ConditionConceptEnum`, which is MONDO ∪ HPO — so mapping source codes into
harmonized output would mean inventing a place to put them. Leaving them in the
raw layer is also what real data looks like, and it means the transformation
exercises the code-mapping path rather than assuming it away.

That matters most for T2D, where the reference gives `MONDO:0005148` for nearly
every row and puts the complication-level detail in ICD-10-CM. The granularity
this corpus offers for diabetes complications lives in the source codes.

Two reference rows are deliberately not used, and one class of cell is not
guessed at:

- **Pregnancy-related hypertension** — both cohorts enrol at 45-78.
- **T2D with ophthalmic complications** — its ICD-10-CM cell is the bare
wildcard `E11.3*` and its MONDO is the shared root, so it would be
indistinguishable from every other T2D record.
- Cells naming a hierarchy rather than a term (`"44054006 hierarchy + renal
complication descendants"`) or a wildcard rather than a code become `None`
rather than a guess. A subtype with no concrete source code emits none, which
is a shape real data has too.

## Known shapes that surprise people

**`cause_of_death` is present for living participants**, multivalued, with a null
Expand Down Expand Up @@ -124,7 +198,9 @@ has to render:
- `associated_assay` referencing a CBC instance (WBC)
- `qualifier` marking a value as an average (BUN)
- `Condition.relationship_to_participant` for family history
- `associated_evidence` on study-record-sourced conditions
- `associated_evidence` distinguishing an ECG-backed infarction from self-report
- `age_at_condition_start` / `_end` and `associated_visit` on conditions
- `age_at_observation` on every measurement, nested ones included
- `exposure_status` distinguishing absent from present drug exposures
- Continuous and categorical presentations of the same concept across cohorts

Expand Down
63 changes: 50 additions & 13 deletions synthetic/generate.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,29 +26,58 @@
7,
),
"visit": (["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "VISIT_TYPE", "AGE_DAYS"], 4),
# AGE_DAYS repeats the visit's age on every measurement table. It is
# derivable by joining to the visit table, but real dbGaP measurement tables
# carry the age alongside the result, and it is what age_at_observation is
# populated from without requiring the transformation to join.
"clinical": (
["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "HEIGHT_CM", "WEIGHT_KG", "BMI", "SBP", "DBP"],
7,
["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "AGE_DAYS",
"HEIGHT_CM", "WEIGHT_KG", "BMI", "SBP", "DBP"],
8,
),
"labs": (
["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "HDL", "HDL_OPERATOR", "BUN", "WBC"],
6,
["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "AGE_DAYS",
"HDL", "HDL_OPERATOR", "BUN", "WBC", "GLUCOSE", "HBA1C"],
9,
),
# One row per participant, so the visit that recorded each condition is a
# per-condition column rather than the table-level VISIT_NUM the measurement
# tables carry. SNOMED and ICD-10-CM appear only here: BDCHM has no slot for
# source terminology, so they stay in the raw layer, which is where they sit
# in real dbGaP tables too.
"conditions": (
[
"dbGaP_Subject_ID",
"SUBJECT_ID",
"HEART_FAILURE",
"HF_CONCEPT",
"HF_AGE_START",
"HF_VISIT",
"FAM_STROKE",
"FS_CONCEPT",
"FS_RELATIVE",
"FS_VISIT",
"HYPERTENSION",
"HTN_CONCEPT",
"HTN_SNOMED",
"HTN_ICD10",
"HTN_AGE_START",
"HTN_AGE_END",
"HTN_VISIT",
"DIABETES",
"DM_CONCEPT",
"DM_SNOMED",
"DM_ICD10",
"DM_AGE_START",
"DM_AGE_END",
"DM_VISIT",
"HEART_ATTACK",
"HA_CONCEPT",
"HA_SOURCE",
"HA_AGE_START",
"HA_VISIT",
],
10,
28,
),
"meds": (["dbGaP_Subject_ID", "SUBJECT_ID", "CCB_STATUS", "CCB_CONCEPT"], 3),
}
Expand Down Expand Up @@ -98,25 +127,33 @@ def build_rows(participants, study):
for visit in p.visits:
rows["visit"].append([gid, sid, visit.number, visit.category, visit.age_days])
rows["clinical"].append([
gid, sid, visit.number,
gid, sid, visit.number, visit.age_days,
visit.height_cm, visit.weight_kg,
visit.bmi_category if study.bmi_categorical else visit.bmi,
visit.systolic, visit.diastolic,
])
rows["labs"].append([
gid, sid, visit.number,
gid, sid, visit.number, visit.age_days,
visit.hdl, visit.hdl_operator, visit.bun, visit.wbc,
visit.glucose, visit.hba1c,
])

c = p.conditions
hf, fs, htn, dm, ha = (
c["heart_failure"], c["family_stroke"], c["hypertension"],
c["diabetes"], c["heart_attack"],
)
rows["conditions"].append([
gid, sid,
c["heart_failure"]["status"], c["heart_failure"]["concept"],
c["family_stroke"]["status"], c["family_stroke"]["concept"],
c["family_stroke"]["relationship"],
c["hypertension"]["status"],
c["heart_attack"]["status"], c["heart_attack"]["concept"],
"STUDY_RECORD" if c["heart_attack"]["from_study_record"] else "SELF_REPORT",
hf["status"], hf["concept"], hf["age_start"], hf["visit"],
fs["status"], fs["concept"], fs["relationship"], fs["visit"],
htn["status"], htn["concept"], htn["snomed"], htn["icd10"],
htn["age_start"], htn["age_end"], htn["visit"],
dm["status"], dm["concept"], dm["snomed"], dm["icd10"],
dm["age_start"], dm["age_end"], dm["visit"],
ha["status"], ha["concept"],
"STUDY_RECORD" if ha["from_study_record"] else "SELF_REPORT",
ha["age_start"], ha["visit"],
])

rows["meds"].append([
Expand Down
Loading
Loading