One cohort is spelled five different ways across synthetic/:
| Thing |
Spelling |
| pipeline config |
pipeline/example_study_one.mk |
| raw data |
data/raw/study_one |
| transformation specs |
specs/example_study_one |
DM_SCHEMA_NAME |
ExampleStudyOne |
| output directory |
output/study_one |
Three renderings of one identity (example_study_one, study_one, ExampleStudyOne),
distributed so that no two adjacent directories agree.
Why it matters
DM_OUTPUT_DIR := $(or $(SYNTH_OUTPUT_DIR),output/ExampleStudyOne) makes
SYNTH_OUTPUT_DIR a free parameter — no cohort name reaches it from the config. The
caller has to pair it with CONFIG by hand:
make pipeline CONFIG=$SYNTH/pipeline/example_study_one.mk \
SYNTH_DIR=$SYNTH SYNTH_OUTPUT_DIR=$SYNTH/output/study_one
Nothing connects the two halves of that command, and the naming spread gives the eye
nothing to check against. A mismatched pair writes one cohort's products into the other's
directory without erroring; see
linkml/dm-bip#357, which proposes catching
that on the dm-bip side.
The default is also wrong in a quiet way: omit SYNTH_OUTPUT_DIR and output lands in
output/ExampleStudyOne relative to the dm-bip checkout the command is run from,
not in synthetic/.
Proposed
Let each config own its output directory, so SYNTH_OUTPUT_DIR stops being something the
caller must supply correctly:
DM_OUTPUT_DIR := $(or $(SYNTH_OUTPUT_DIR),$(SYNTH_DIR)/output/study_one)
The default then lands in the right place, the override survives for deliberate use, and
the documented invocation collapses to:
make pipeline CONFIG=$SYNTH/pipeline/example_study_one.mk SYNTH_DIR=$SYNTH
Settling on one spelling per cohort is the larger half of the cleanup and touches
directory names, so it is worth deciding whether it is in scope here or a follow-up.
dm-bip#357 is defence in depth on the tool side; this is the fix on ours.
Follow-on
synthetic/README.md currently spends several lines documenting the pairing rule and the
surprising default. Both paragraphs go away if this lands.
One cohort is spelled five different ways across
synthetic/:pipeline/example_study_one.mkdata/raw/study_onespecs/example_study_oneDM_SCHEMA_NAMEExampleStudyOneoutput/study_oneThree renderings of one identity (
example_study_one,study_one,ExampleStudyOne),distributed so that no two adjacent directories agree.
Why it matters
DM_OUTPUT_DIR := $(or $(SYNTH_OUTPUT_DIR),output/ExampleStudyOne)makesSYNTH_OUTPUT_DIRa free parameter — no cohort name reaches it from the config. Thecaller has to pair it with
CONFIGby hand:Nothing connects the two halves of that command, and the naming spread gives the eye
nothing to check against. A mismatched pair writes one cohort's products into the other's
directory without erroring; see
linkml/dm-bip#357, which proposes catching
that on the dm-bip side.
The default is also wrong in a quiet way: omit
SYNTH_OUTPUT_DIRand output lands inoutput/ExampleStudyOnerelative to the dm-bip checkout the command is run from,not in
synthetic/.Proposed
Let each config own its output directory, so
SYNTH_OUTPUT_DIRstops being something thecaller must supply correctly:
The default then lands in the right place, the override survives for deliberate use, and
the documented invocation collapses to:
Settling on one spelling per cohort is the larger half of the cleanup and touches
directory names, so it is worth deciding whether it is in scope here or a follow-up.
dm-bip#357is defence in depth on the tool side; this is the fix on ours.Follow-on
synthetic/README.mdcurrently spends several lines documenting the pairing rule and thesurprising default. Both paragraphs go away if this lands.