Repository navigation
Port subworkflow prepare_genome to nf-core structure - #290
Conversation
Add meta.yml and nf-tests for the local PREPARE_GENOME subworkflow. There is no prepare_genome subworkflow in nf-core/modules to install in its place. The tests cover building every index from a FASTA and a GTF, compressed inputs with a GFF3 annotation and supplied indices, a non fastq source that needs no index, and a stub run. Also: - Take source, aligner, pseudo_aligner and skip_alignment as inputs instead of reading params inside the subworkflow. - Emit salmon_index and suppa_tpm as bare paths whatever the input. An uncompressed --salmon_index or --suppa_tpm was emitted as [ [:], path ], which SALMON_QUANT and SUPPA cannot use, while the .tar.gz and .gz inputs came out as a path. - Give GFFREAD a meta id taken from the GFF file name, so the converted annotation is no longer called null.gtf. - Pass every path to the modules through file(), add the MODULE section headers and format with nextflow lint -format. Generated by Claude Opus 5.5 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
❌ nf-test failed with latest Nextflow versionNote Tests with Nextflow's latest version failed but it will not cause a CI workflow failure.
See the full run for details. |
erikrikarddaniel
left a comment
There was a problem hiding this comment.
Reviewed by Claude Code.
Traced both fixes to their consumers rather than taking them on trust. The suppa_tpm one would have thrown rather than merely emitted an odd shape: workflows/rnasplice.nf does ch_suppa_tpm.map { tpm_psi -> [[ id: tpm_psi.baseName ], tpm_psi] } (lines 313 and 386), and baseName does not exist on [[:], path]. Same story for the Salmon index reaching SALMON_QUANT's plain path index. The emitted shapes are uniform now, the take: list matches the call in main.nf, and the meta.yml enums agree with nextflow_schema.json.
Two non-blocking points.
Four of the six uncompress branches are never exercised. The tests cover a gzipped genome FASTA and a gzipped DEXSeq GFF, so GUNZIP_GTF, GUNZIP_GFF, GUNZIP_TRANSCRIPT_FASTA and GUNZIP_SUPPA_TPM never run. All six branches have the same shape and were rewritten in the same pass, so one of them holding a wrong variable would go green today. The gzipped GFF3 path is the one I would most want covered, since it is where GUNZIP_GFF feeds the new [[ id: gff_file.baseName ], gff_file] map — the null.gtf fix is only tested on an uncompressed GFF3. Switching the second test's GFF3 and transcript FASTA inputs to their .gz variants would cover three of the four for the price of a snapshot update.
The second test writes its stand-in inputs into ${outputDir}. Suggested inline: ${workDir} keeps the output directory holding only outputs.
Nothing else came up. nf-test is green on my side of the diff too, and the CHANGELOG entries are in the right sections and in the order this repo uses.
|
Thanks @erikrikarddaniel ! I'll fix those before merging |
Review feedback on nf-core#290. The tests only ran GUNZIP_FASTA and GUNZIP_GFF_DEXSEQ, so a wrong variable in the GTF, GFF3, transcript FASTA or SUPPA TPM branch would have gone unnoticed. The test datasets carry no gzipped copy of those files, so the tests gzip them on the fly: the second test now takes a gzipped GFF3, transcript FASTA and TPM table, and the genome_bam test a gzipped GTF and an uncompressed TPM table, which keeps the uncompressed --suppa_tpm fix covered. The gzipped GFF3 also covers the GFFREAD naming fix on the path that goes through GUNZIP_GFF. The stand-in inputs are written to workDir instead of outputDir, so the output directory holds only outputs. Generated by Claude Opus 5.5 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013jvD13HCjdAWqMnRvEzVH6
Description
Moves the local
PREPARE_GENOMEsubworkflow to the nf-core subworkflow template, the same wayALIGN_STAR,DEXSEQ_DEU,EDGER_DEUandLEAFCUTTERwere ported. nf-core/modules has noprepare_genomesubworkflow to install instead, so it stays local.Changes
meta.yml. Authors are taken from the git history of the subworkflow (@asmaali98, @bensouthgate, @jma1991, @valentinoruggieri), plus @piplus2 as author and maintainer.params.--source,--aligner,--pseudo_alignerand--skip_alignmentare explicit inputs, passed in frommain.nf.file(..., checkIfExists: true)before reaching a module, the MODULE section headers are added and the file is formatted withnextflow lint -format.Fixes
--salmon_indexor--suppa_tpmwas emitted as[ [:], path ], while the.tar.gzand.gzinputs came out as a bare path.SALMON_QUANTandSUPPAexpect the path, so only the compressed form worked. Both are now emitted as paths whatever the input. This came in with 495ff9c.--gffwas namednull.gtf, sinceGFFREADnames its output aftermeta.idand the meta map was empty. It is now named after the GFF file.Testing
.tar.gz) and Salmon (directory) indices plus a DEXSeq GFF and a SUPPA TPM table, agenome_bamsource that needs no index, and a stub.genome.gff3test file makes gffread fail on a malformed strand column, so the GFF3 test usesreference/genes_chrX.gff3from the rnasplice test-datasets.nf-core pipelines lint(tools 4.1.0) 0 failures,prekclean.nf-core subworkflows lintwarns thatpreprocess/transcripts/fasta/gencodeandstar/genomeparams/upgradeare missing frommeta.yml. They are listed under their real names; the linter splits local module names on underscores, the same false positive as forstrand_junctionsinLEAFCUTTER.Generated by Claude Opus 5.5
PR checklist
nf-core pipelines lint).nextflow run . -profile test,docker --outdir <OUTDIR>).nextflow run . -profile debug,test,docker --outdir <OUTDIR>).docs/usage.mdis updated.docs/output.mdis updated.CHANGELOG.mdis updated.README.mdis updated (including new tool citations and authors/contributors).🤖 Generated with Claude Code