Skip to content
Merged
Show file tree
Hide file tree
Changes from 4 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 61 additions & 0 deletions .github/workflows/test-synthetic.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
name: Test Synthetic Corpus

on:
pull_request:
paths:
- "synthetic/**"
- ".github/workflows/test-synthetic.yaml"
push:
branches: [main]
paths:
- "synthetic/**"
- ".github/workflows/test-synthetic.yaml"

jobs:
test-synthetic:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"

# The corpus is a build product, not a committed artifact: it is
# regenerated here and checked, rather than diffed against something
# stored in the repo.
- name: Generate raw tables
working-directory: synthetic
run: python generate.py

- name: Generate transformation specs
working-directory: synthetic
run: python specs.py

- name: Validate distributions and invariants
working-directory: synthetic
run: python validate.py

- name: Check specs are well-formed YAML
working-directory: synthetic
run: |
pip install --quiet pyyaml
python - <<'PY'
import glob
import sys

import yaml

bad = []
paths = sorted(glob.glob("specs/*/*.yaml"))
for path in paths:
try:
yaml.safe_load(open(path))
except yaml.YAMLError as exc:
bad.append(f"{path}: {exc}")
print(f"{len(paths) - len(bad)}/{len(paths)} specs parse")
if bad:
print("\n".join(bad))
sys.exit(1)
PY
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ target-version = "py312"
line-length = 120
include = [
"api/**/*.py",
"synthetic/**/*.py",
]

[tool.ruff.lint]
Expand Down
6 changes: 6 additions & 0 deletions synthetic/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
# Generated corpus — reproducible from SEED, distributed as release assets
data/
output/
# Vendored upstream schema, fetched not committed
bdchm.yaml
__pycache__/
89 changes: 89 additions & 0 deletions synthetic/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,89 @@
# Synthetic corpus

Two fictional cohorts, structurally faithful to harmonized BDC data, for teams
building against the portal without touching participant data.

Nothing here derives from real participants. Every coded value — CURIEs, enum
members, units — is taken from RTI's `priority_variables_transform` specs or
from BDCHM itself, so the corpus resolves against the same vocabulary as real
harmonized data.

## Why it goes through dm-bip

The generator emits **dbGaP-style raw tables**, not harmonized output, and those
are transformed by dm-bip against BDCHM. Emitting BDCHM directly would be a
second implementation of the transformation, free to drift from the real one in
ways nobody would notice until a portal built on it met real data.

```
generate.py -> data/raw/*.txt.gz -> dm-bip map-data -> harmonized BDCHM
specs.py -> specs/*/*.yaml -> ^
```

## Running it

```bash
python generate.py # raw tables, per study
python specs.py # BDCHM-targeted transformation specs
python validate.py # distributions and invariants
```

Then, from a dm-bip checkout:

```bash
SYNTH=/path/to/synthetic
make pipeline CONFIG=$SYNTH/pipeline/example_study_one.mk \
SYNTH_DIR=$SYNTH SYNTH_OUTPUT_DIR=$SYNTH/output/study_one
```

BDCHM is fetched rather than vendored, pinned to a release so an upstream change
is a deliberate bump here rather than a silent change in what the pipeline
produces:

```bash
./fetch-bdchm.sh # currently v1.3.0
```

## What is and isn't committed

The generator is committed; the corpus is not. Generation is deterministic given
`SEED`, so the same seed reproduces the same corpus byte for byte — which makes
committing several megabytes of generated `.gz` pointless, and keeps diffs
reviewable. Built corpora are distributed as release assets.

## The cohorts

| | Example Study One | Example Study Two |
|---|---|---|
| Participants | 500 | 500 |
| Visits | 3 | 5 |
| Sex skew | slightly male | slightly female |
| Race | 50/30/10/10 white, black, Asian, American Indian | 60/20/10/10 white, black, Middle Eastern, Native Hawaiian |
| BMI | continuous | categorical |

5% of Study Two's participants are the same individuals as in Study One. They
share a `dbGaP_Subject_ID`, so harmonization resolves them to one `Person` with
two `Participant` records — the same way real cross-study participation appears.

One visit per participant is `TELEHEALTH`; the rest are `STUDY_SITE_VISIT`.
Deceased participants stop attending, so nobody is measured after they die.

## Structural features exercised

Beyond the clinical values, the corpus is meant to exercise the shapes a portal
has to render:

- `MeasurementObservationSet` with nested systolic and diastolic observations
- `Quantity.operator` for results censored below an assay's detection limit
- `Assay` with `lower_limit_of_detection` / `upper_limit_of_detection` (HDL)
- `associated_assay` referencing a CBC instance (WBC)
- `qualifier` marking a value as an average (BUN)
- `Condition.relationship_to_participant` for family history
- `associated_evidence` on study-record-sourced conditions
- `exposure_status` distinguishing absent from present drug exposures
- Continuous and categorical presentations of the same concept across cohorts

Coverage against BDCHM's full slot inventory — the percentage of slots appearing
at least once — is intended as the acceptance number for broadening the corpus
beyond this first pass. **That measurement is not built yet.** This pass covers
the brief only.
16 changes: 16 additions & 0 deletions synthetic/fetch-bdchm.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
#!/usr/bin/env bash
# Fetch the pinned BDCHM schema used as the transformation target.
#
# Pinned rather than tracking main so the corpus is reproducible: a schema
# change upstream should be a deliberate bump here, visible in review, not
# something that silently alters what the pipeline produces.
set -euo pipefail

BDCHM_VERSION="v1.3.0"
REPO="RTIInternational/NHLBI-BDC-DMC-HM"
URL="https://raw.githubusercontent.com/${REPO}/${BDCHM_VERSION}/src/bdchm/schema/bdchm.yaml"

DEST="$(cd "$(dirname "$0")" && pwd)/bdchm.yaml"

curl -fsSL "$URL" -o "$DEST"
Comment thread
amc-corey-cox marked this conversation as resolved.
Outdated
echo "Fetched BDCHM ${BDCHM_VERSION} -> ${DEST}"
157 changes: 157 additions & 0 deletions synthetic/generate.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,157 @@
# ruff: noqa: S311
"""
Emit the synthetic corpus as dbGaP-style raw tables.

These are inputs to dm-bip, not harmonized output. Running them through the
pipeline with BDCHM as the target schema is what makes the result structurally
identical to real harmonized data, rather than a second implementation of the
transformation that can quietly drift.

python synthetic/generate.py [--out DIR]
"""

import argparse
import gzip
from pathlib import Path

import population as pop

CITATION = "Synthetic corpus for portal development. Not derived from participant data."

# Column layouts. The phv accessions are fictional but well-formed, and are
# what the transformation specs reference.
LAYOUTS = {
"subject": (
["dbGaP_Subject_ID", "SUBJECT_ID", "SEX", "RACE", "ETHNICITY", "AGE", "VITAL_STATUS", "DEATH_CAUSE"],
7,
),
"visit": (["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "VISIT_TYPE", "AGE_DAYS"], 4),
"clinical": (
["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "HEIGHT_CM", "WEIGHT_KG", "BMI", "SBP", "DBP"],
7,
),
"labs": (
["dbGaP_Subject_ID", "SUBJECT_ID", "VISIT_NUM", "HDL", "HDL_OPERATOR", "BUN", "WBC"],
6,
),
"conditions": (
[
"dbGaP_Subject_ID",
"SUBJECT_ID",
"HEART_FAILURE",
"HF_CONCEPT",
"FAM_STROKE",
"FS_CONCEPT",
"FS_RELATIVE",
"HYPERTENSION",
"HEART_ATTACK",
"HA_CONCEPT",
"HA_SOURCE",
],
10,
),
"meds": (["dbGaP_Subject_ID", "SUBJECT_ID", "CCB_STATUS", "CCB_CONCEPT"], 3),
}


def phv_ids(study, table, count):
"""Fictional but well-formed phv accessions, unique per study and table."""
base = int(study.phs[-3:]) * 100000 + int(study.tables[table][-3:]) * 100
return [f"phv{base + i:08d}.v1" for i in range(count)]


def write_table(out_dir, study, table, rows):
"""Write one table in dbGaP raw format, returning its path."""
columns, phv_count = LAYOUTS[table]
pht = study.tables[table]
path = out_dir / f"{study.phs}.v1.{pht}.v1.p1.c1.ex0_1s.HMB.txt.gz"

with gzip.open(path, "wt", newline="") as fh:
fh.write(f"# Study accession: {study.phs}.v1.p1\n")
fh.write(f"# Table accession: {pht}\n")
fh.write("# Consent group: All subjects\n")
fh.write(f"# Citation: {CITATION}\n")
fh.write("#\n")
fh.write("##\t" + "\t".join(phv_ids(study, table, phv_count)) + "\n")
fh.write("\t".join(columns) + "\n")
fh.write("\n")
for row in rows:
fh.write("\t".join("" if c is None else str(c) for c in row) + "\n")
return path


def build_rows(participants, study):
"""Shape the population into the six per-study tables."""
rows = {k: [] for k in LAYOUTS}

for p in (x for x in participants if x.study is study):
gid, sid = p.person.dbgap_id, p.subject_id
person = p.person

rows["subject"].append([
gid, sid, person.sex, person.race, person.ethnicity,
person.baseline_age,
1 if person.deceased else 0,
person.death_cause,
])

for visit in p.visits:
rows["visit"].append([gid, sid, visit.number, visit.category, visit.age_days])
rows["clinical"].append([
gid, sid, visit.number,
visit.height_cm, visit.weight_kg,
visit.bmi_category if study.bmi_categorical else visit.bmi,
visit.systolic, visit.diastolic,
])
rows["labs"].append([
gid, sid, visit.number,
visit.hdl, visit.hdl_operator, visit.bun, visit.wbc,
])

c = p.conditions
rows["conditions"].append([
gid, sid,
c["heart_failure"]["status"], c["heart_failure"]["concept"],
c["family_stroke"]["status"], c["family_stroke"]["concept"],
c["family_stroke"]["relationship"],
c["hypertension"]["status"],
c["heart_attack"]["status"], c["heart_attack"]["concept"],
"STUDY_RECORD" if c["heart_attack"]["from_study_record"] else "SELF_REPORT",
])

rows["meds"].append([
gid, sid,
"PRESENT" if p.ccb else "ABSENT",
p.ccb,
])

return rows


def main():
"""Generate the raw tables for both studies."""
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--out",
type=Path,
default=Path(__file__).parent / "data" / "raw",
help="directory to write the dbGaP-style tables into",
)
args = parser.parse_args()

_, participants = pop.build()

written = []
for study, slug in ((pop.STUDY_ONE, "study_one"), (pop.STUDY_TWO, "study_two")):
out_dir = args.out / slug
out_dir.mkdir(parents=True, exist_ok=True)
rows = build_rows(participants, study)
for table in LAYOUTS:
written.append((study.name, table, write_table(out_dir, study, table, rows[table]), len(rows[table])))

for name, table, path, count in written:
print(f"{name:20} {table:12} {count:6} rows {path.name}")


if __name__ == "__main__":
main()
14 changes: 14 additions & 0 deletions synthetic/pipeline/example_study_one.mk
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# Pipeline config for the synthetic Example Study One cohort.
# Target schema is BDCHM itself, so the harmonized output is structurally
# identical to real harmonized data rather than a simplified stand-in.

DM_RAW_SOURCE := $(SYNTH_DIR)/data/raw/study_one
DM_SCHEMA_NAME := ExampleStudyOne
DM_OUTPUT_DIR := $(or $(SYNTH_OUTPUT_DIR),output/ExampleStudyOne)
DM_INPUT_DIR := $(DM_OUTPUT_DIR)/prepared
DM_TRANS_SPEC_DIR := $(SYNTH_DIR)/specs/example_study_one
DM_MAPPING_SPEC := $(DM_TRANS_SPEC_DIR)
DM_MAP_TARGET_SCHEMA := $(SYNTH_DIR)/bdchm.yaml
DM_MAPPING_PREFIX := SYNTH1
DM_MAPPING_POSTFIX := -data
DM_MAP_STRICT := false
14 changes: 14 additions & 0 deletions synthetic/pipeline/example_study_two.mk
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# Pipeline config for the synthetic Example Study Two cohort.
# Target schema is BDCHM itself, so the harmonized output is structurally
# identical to real harmonized data rather than a simplified stand-in.

DM_RAW_SOURCE := $(SYNTH_DIR)/data/raw/study_two
DM_SCHEMA_NAME := ExampleStudyTwo
DM_OUTPUT_DIR := $(or $(SYNTH_OUTPUT_DIR),output/ExampleStudyTwo)
DM_INPUT_DIR := $(DM_OUTPUT_DIR)/prepared
DM_TRANS_SPEC_DIR := $(SYNTH_DIR)/specs/example_study_two
DM_MAPPING_SPEC := $(DM_TRANS_SPEC_DIR)
DM_MAP_TARGET_SCHEMA := $(SYNTH_DIR)/bdchm.yaml
DM_MAPPING_PREFIX := SYNTH2
DM_MAPPING_POSTFIX := -data
DM_MAP_STRICT := false
Loading
Loading