Repository navigation
Add synthetic corpus generator for two example cohorts #25
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from 4 commits
Commits
Show all changes
10 commits
Select commit
Hold shift + click to select a range
5bc5825
Add synthetic corpus generator for two example cohorts
amc-corey-cox 3a91718
Fix nested identities, assay slots, and visit provenance in specs
amc-corey-cox 42e9830
Route categorical BMI to value_concept instead of value_decimal
amc-corey-cox 49cbbcf
Pin BDCHM fetch, wire synthetic checks into CI, satisfy lint
amc-corey-cox ae611b8
Address PR review feedback
amc-corey-cox 9795eb0
Commit a feature-demonstrating sample of harmonized output
amc-corey-cox 61f1718
Document the cause_of_death shape as faithful to the real transformation
amc-corey-cox c3354db
End the blood pressure template on a newline rather than an escaped q…
amc-corey-cox 40acdfb
Emit JSONL alongside YAML and build Parquet with BDCHM-derived types
amc-corey-cox 824729e
Draw measurements conditional on diagnosis so cohorts separate
amc-corey-cox File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,61 @@ | ||
| name: Test Synthetic Corpus | ||
|
|
||
| on: | ||
| pull_request: | ||
| paths: | ||
| - "synthetic/**" | ||
| - ".github/workflows/test-synthetic.yaml" | ||
| push: | ||
| branches: [main] | ||
| paths: | ||
| - "synthetic/**" | ||
| - ".github/workflows/test-synthetic.yaml" | ||
|
|
||
| jobs: | ||
| test-synthetic: | ||
| runs-on: ubuntu-latest | ||
| steps: | ||
| - uses: actions/checkout@v4 | ||
|
|
||
| - name: Set up Python | ||
| uses: actions/setup-python@v5 | ||
| with: | ||
| python-version: "3.12" | ||
|
|
||
| # The corpus is a build product, not a committed artifact: it is | ||
| # regenerated here and checked, rather than diffed against something | ||
| # stored in the repo. | ||
| - name: Generate raw tables | ||
| working-directory: synthetic | ||
| run: python generate.py | ||
|
|
||
| - name: Generate transformation specs | ||
| working-directory: synthetic | ||
| run: python specs.py | ||
|
|
||
| - name: Validate distributions and invariants | ||
| working-directory: synthetic | ||
| run: python validate.py | ||
|
|
||
| - name: Check specs are well-formed YAML | ||
| working-directory: synthetic | ||
| run: | | ||
| pip install --quiet pyyaml | ||
| python - <<'PY' | ||
| import glob | ||
| import sys | ||
|
|
||
| import yaml | ||
|
|
||
| bad = [] | ||
| paths = sorted(glob.glob("specs/*/*.yaml")) | ||
| for path in paths: | ||
| try: | ||
| yaml.safe_load(open(path)) | ||
| except yaml.YAMLError as exc: | ||
| bad.append(f"{path}: {exc}") | ||
| print(f"{len(paths) - len(bad)}/{len(paths)} specs parse") | ||
| if bad: | ||
| print("\n".join(bad)) | ||
| sys.exit(1) | ||
| PY |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,6 @@ | ||
| # Generated corpus — reproducible from SEED, distributed as release assets | ||
| data/ | ||
| output/ | ||
| # Vendored upstream schema, fetched not committed | ||
| bdchm.yaml | ||
| __pycache__/ |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,89 @@ | ||
| # Synthetic corpus | ||
|
|
||
| Two fictional cohorts, structurally faithful to harmonized BDC data, for teams | ||
| building against the portal without touching participant data. | ||
|
|
||
| Nothing here derives from real participants. Every coded value — CURIEs, enum | ||
| members, units — is taken from RTI's `priority_variables_transform` specs or | ||
| from BDCHM itself, so the corpus resolves against the same vocabulary as real | ||
| harmonized data. | ||
|
|
||
| ## Why it goes through dm-bip | ||
|
|
||
| The generator emits **dbGaP-style raw tables**, not harmonized output, and those | ||
| are transformed by dm-bip against BDCHM. Emitting BDCHM directly would be a | ||
| second implementation of the transformation, free to drift from the real one in | ||
| ways nobody would notice until a portal built on it met real data. | ||
|
|
||
| ``` | ||
| generate.py -> data/raw/*.txt.gz -> dm-bip map-data -> harmonized BDCHM | ||
| specs.py -> specs/*/*.yaml -> ^ | ||
| ``` | ||
|
|
||
| ## Running it | ||
|
|
||
| ```bash | ||
| python generate.py # raw tables, per study | ||
| python specs.py # BDCHM-targeted transformation specs | ||
| python validate.py # distributions and invariants | ||
| ``` | ||
|
|
||
| Then, from a dm-bip checkout: | ||
|
|
||
| ```bash | ||
| SYNTH=/path/to/synthetic | ||
| make pipeline CONFIG=$SYNTH/pipeline/example_study_one.mk \ | ||
| SYNTH_DIR=$SYNTH SYNTH_OUTPUT_DIR=$SYNTH/output/study_one | ||
| ``` | ||
|
|
||
| BDCHM is fetched rather than vendored, pinned to a release so an upstream change | ||
| is a deliberate bump here rather than a silent change in what the pipeline | ||
| produces: | ||
|
|
||
| ```bash | ||
| ./fetch-bdchm.sh # currently v1.3.0 | ||
| ``` | ||
|
|
||
| ## What is and isn't committed | ||
|
|
||
| The generator is committed; the corpus is not. Generation is deterministic given | ||
| `SEED`, so the same seed reproduces the same corpus byte for byte — which makes | ||
| committing several megabytes of generated `.gz` pointless, and keeps diffs | ||
| reviewable. Built corpora are distributed as release assets. | ||
|
|
||
| ## The cohorts | ||
|
|
||
| | | Example Study One | Example Study Two | | ||
| |---|---|---| | ||
| | Participants | 500 | 500 | | ||
| | Visits | 3 | 5 | | ||
| | Sex skew | slightly male | slightly female | | ||
| | Race | 50/30/10/10 white, black, Asian, American Indian | 60/20/10/10 white, black, Middle Eastern, Native Hawaiian | | ||
| | BMI | continuous | categorical | | ||
|
|
||
| 5% of Study Two's participants are the same individuals as in Study One. They | ||
| share a `dbGaP_Subject_ID`, so harmonization resolves them to one `Person` with | ||
| two `Participant` records — the same way real cross-study participation appears. | ||
|
|
||
| One visit per participant is `TELEHEALTH`; the rest are `STUDY_SITE_VISIT`. | ||
| Deceased participants stop attending, so nobody is measured after they die. | ||
|
|
||
| ## Structural features exercised | ||
|
|
||
| Beyond the clinical values, the corpus is meant to exercise the shapes a portal | ||
| has to render: | ||
|
|
||
| - `MeasurementObservationSet` with nested systolic and diastolic observations | ||
| - `Quantity.operator` for results censored below an assay's detection limit | ||
| - `Assay` with `lower_limit_of_detection` / `upper_limit_of_detection` (HDL) | ||
| - `associated_assay` referencing a CBC instance (WBC) | ||
| - `qualifier` marking a value as an average (BUN) | ||
| - `Condition.relationship_to_participant` for family history | ||
| - `associated_evidence` on study-record-sourced conditions | ||
| - `exposure_status` distinguishing absent from present drug exposures | ||
| - Continuous and categorical presentations of the same concept across cohorts | ||
|
|
||
| Coverage against BDCHM's full slot inventory — the percentage of slots appearing | ||
| at least once — is intended as the acceptance number for broadening the corpus | ||
| beyond this first pass. **That measurement is not built yet.** This pass covers | ||
| the brief only. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,16 @@ | ||
| #!/usr/bin/env bash | ||
| # Fetch the pinned BDCHM schema used as the transformation target. | ||
| # | ||
| # Pinned rather than tracking main so the corpus is reproducible: a schema | ||
| # change upstream should be a deliberate bump here, visible in review, not | ||
| # something that silently alters what the pipeline produces. | ||
| set -euo pipefail | ||
|
|
||
| BDCHM_VERSION="v1.3.0" | ||
| REPO="RTIInternational/NHLBI-BDC-DMC-HM" | ||
| URL="https://raw.githubusercontent.com/${REPO}/${BDCHM_VERSION}/src/bdchm/schema/bdchm.yaml" | ||
|
|
||
| DEST="$(cd "$(dirname "$0")" && pwd)/bdchm.yaml" | ||
|
|
||
| curl -fsSL "$URL" -o "$DEST" | ||
| echo "Fetched BDCHM ${BDCHM_VERSION} -> ${DEST}" | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,157 @@ | ||
| # ruff: noqa: S311 | ||
| """ | ||
| Emit the synthetic corpus as dbGaP-style raw tables. | ||
|
|
||
| These are inputs to dm-bip, not harmonized output. Running them through the | ||
| pipeline with BDCHM as the target schema is what makes the result structurally | ||
| identical to real harmonized data, rather than a second implementation of the | ||
| transformation that can quietly drift. | ||
|
|
||
| python synthetic/generate.py [--out DIR] | ||
| """ | ||
|
|
||
| import argparse | ||
| import gzip | ||
| from pathlib import Path | ||
|
|
||
| import population as pop | ||
|
|
||
| CITATION = "Synthetic corpus for portal development. Not derived from participant data." | ||
|
|
||
| # Column layouts. The phv accessions are fictional but well-formed, and are | ||
| # what the transformation specs reference. | ||
| LAYOUTS = { | ||
| "subject": ( | ||
| ["dbGaP_Subject_ID", "SUBJECT_ID", "SEX", "RACE", "ETHNICITY", "AGE", "VITAL_STATUS", "DEATH_CAUSE"], | ||
| 7, | ||
| ), | ||
| "visit": (["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "VISIT_TYPE", "AGE_DAYS"], 4), | ||
| "clinical": ( | ||
| ["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "HEIGHT_CM", "WEIGHT_KG", "BMI", "SBP", "DBP"], | ||
| 7, | ||
| ), | ||
| "labs": ( | ||
| ["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "HDL", "HDL_OPERATOR", "BUN", "WBC"], | ||
| 6, | ||
| ), | ||
| "conditions": ( | ||
| [ | ||
| "dbGaP_Subject_ID", | ||
| "SUBJECT_ID", | ||
| "HEART_FAILURE", | ||
| "HF_CONCEPT", | ||
| "FAM_STROKE", | ||
| "FS_CONCEPT", | ||
| "FS_RELATIVE", | ||
| "HYPERTENSION", | ||
| "HEART_ATTACK", | ||
| "HA_CONCEPT", | ||
| "HA_SOURCE", | ||
| ], | ||
| 10, | ||
| ), | ||
| "meds": (["dbGaP_Subject_ID", "SUBJECT_ID", "CCB_STATUS", "CCB_CONCEPT"], 3), | ||
| } | ||
|
|
||
|
|
||
| def phv_ids(study, table, count): | ||
| """Fictional but well-formed phv accessions, unique per study and table.""" | ||
| base = int(study.phs[-3:]) * 100000 + int(study.tables[table][-3:]) * 100 | ||
| return [f"phv{base + i:08d}.v1" for i in range(count)] | ||
|
|
||
|
|
||
| def write_table(out_dir, study, table, rows): | ||
| """Write one table in dbGaP raw format, returning its path.""" | ||
| columns, phv_count = LAYOUTS[table] | ||
| pht = study.tables[table] | ||
| path = out_dir / f"{study.phs}.v1.{pht}.v1.p1.c1.ex0_1s.HMB.txt.gz" | ||
|
|
||
| with gzip.open(path, "wt", newline="") as fh: | ||
| fh.write(f"# Study accession: {study.phs}.v1.p1\n") | ||
| fh.write(f"# Table accession: {pht}\n") | ||
| fh.write("# Consent group: All subjects\n") | ||
| fh.write(f"# Citation: {CITATION}\n") | ||
| fh.write("#\n") | ||
| fh.write("##\t" + "\t".join(phv_ids(study, table, phv_count)) + "\n") | ||
| fh.write("\t".join(columns) + "\n") | ||
| fh.write("\n") | ||
| for row in rows: | ||
| fh.write("\t".join("" if c is None else str(c) for c in row) + "\n") | ||
| return path | ||
|
|
||
|
|
||
| def build_rows(participants, study): | ||
| """Shape the population into the six per-study tables.""" | ||
| rows = {k: [] for k in LAYOUTS} | ||
|
|
||
| for p in (x for x in participants if x.study is study): | ||
| gid, sid = p.person.dbgap_id, p.subject_id | ||
| person = p.person | ||
|
|
||
| rows["subject"].append([ | ||
| gid, sid, person.sex, person.race, person.ethnicity, | ||
| person.baseline_age, | ||
| 1 if person.deceased else 0, | ||
| person.death_cause, | ||
| ]) | ||
|
|
||
| for visit in p.visits: | ||
| rows["visit"].append([gid, sid, visit.number, visit.category, visit.age_days]) | ||
| rows["clinical"].append([ | ||
| gid, sid, visit.number, | ||
| visit.height_cm, visit.weight_kg, | ||
| visit.bmi_category if study.bmi_categorical else visit.bmi, | ||
| visit.systolic, visit.diastolic, | ||
| ]) | ||
| rows["labs"].append([ | ||
| gid, sid, visit.number, | ||
| visit.hdl, visit.hdl_operator, visit.bun, visit.wbc, | ||
| ]) | ||
|
|
||
| c = p.conditions | ||
| rows["conditions"].append([ | ||
| gid, sid, | ||
| c["heart_failure"]["status"], c["heart_failure"]["concept"], | ||
| c["family_stroke"]["status"], c["family_stroke"]["concept"], | ||
| c["family_stroke"]["relationship"], | ||
| c["hypertension"]["status"], | ||
| c["heart_attack"]["status"], c["heart_attack"]["concept"], | ||
| "STUDY_RECORD" if c["heart_attack"]["from_study_record"] else "SELF_REPORT", | ||
| ]) | ||
|
|
||
| rows["meds"].append([ | ||
| gid, sid, | ||
| "PRESENT" if p.ccb else "ABSENT", | ||
| p.ccb, | ||
| ]) | ||
|
|
||
| return rows | ||
|
|
||
|
|
||
| def main(): | ||
| """Generate the raw tables for both studies.""" | ||
| parser = argparse.ArgumentParser(description=__doc__) | ||
| parser.add_argument( | ||
| "--out", | ||
| type=Path, | ||
| default=Path(__file__).parent / "data" / "raw", | ||
| help="directory to write the dbGaP-style tables into", | ||
| ) | ||
| args = parser.parse_args() | ||
|
|
||
| _, participants = pop.build() | ||
|
|
||
| written = [] | ||
| for study, slug in ((pop.STUDY_ONE, "study_one"), (pop.STUDY_TWO, "study_two")): | ||
| out_dir = args.out / slug | ||
| out_dir.mkdir(parents=True, exist_ok=True) | ||
| rows = build_rows(participants, study) | ||
| for table in LAYOUTS: | ||
| written.append((study.name, table, write_table(out_dir, study, table, rows[table]), len(rows[table]))) | ||
|
|
||
| for name, table, path, count in written: | ||
| print(f"{name:20} {table:12} {count:6} rows {path.name}") | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| main() |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,14 @@ | ||
| # Pipeline config for the synthetic Example Study One cohort. | ||
| # Target schema is BDCHM itself, so the harmonized output is structurally | ||
| # identical to real harmonized data rather than a simplified stand-in. | ||
|
|
||
| DM_RAW_SOURCE := $(SYNTH_DIR)/data/raw/study_one | ||
| DM_SCHEMA_NAME := ExampleStudyOne | ||
| DM_OUTPUT_DIR := $(or $(SYNTH_OUTPUT_DIR),output/ExampleStudyOne) | ||
| DM_INPUT_DIR := $(DM_OUTPUT_DIR)/prepared | ||
| DM_TRANS_SPEC_DIR := $(SYNTH_DIR)/specs/example_study_one | ||
| DM_MAPPING_SPEC := $(DM_TRANS_SPEC_DIR) | ||
| DM_MAP_TARGET_SCHEMA := $(SYNTH_DIR)/bdchm.yaml | ||
| DM_MAPPING_PREFIX := SYNTH1 | ||
| DM_MAPPING_POSTFIX := -data | ||
| DM_MAP_STRICT := false |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,14 @@ | ||
| # Pipeline config for the synthetic Example Study Two cohort. | ||
| # Target schema is BDCHM itself, so the harmonized output is structurally | ||
| # identical to real harmonized data rather than a simplified stand-in. | ||
|
|
||
| DM_RAW_SOURCE := $(SYNTH_DIR)/data/raw/study_two | ||
| DM_SCHEMA_NAME := ExampleStudyTwo | ||
| DM_OUTPUT_DIR := $(or $(SYNTH_OUTPUT_DIR),output/ExampleStudyTwo) | ||
| DM_INPUT_DIR := $(DM_OUTPUT_DIR)/prepared | ||
| DM_TRANS_SPEC_DIR := $(SYNTH_DIR)/specs/example_study_two | ||
| DM_MAPPING_SPEC := $(DM_TRANS_SPEC_DIR) | ||
| DM_MAP_TARGET_SCHEMA := $(SYNTH_DIR)/bdchm.yaml | ||
| DM_MAPPING_PREFIX := SYNTH2 | ||
| DM_MAPPING_POSTFIX := -data | ||
| DM_MAP_STRICT := false |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.