Skip to content

Port subworkflow prepare_genome to nf-core structure - #290

Merged
piplus2 merged 2 commits into
nf-core:devfrom
piplus2:prepare-genome-nfcore
Sep 23, 2026
Merged

piplus2 merged 2 commits into
nf-core:devfrom
piplus2:prepare-genome-nfcore

Conversation

@piplus2

@piplus2 piplus2 commented Sep 23, 2026

Copy link
Copy Markdown

Description

Moves the local PREPARE_GENOME subworkflow to the nf-core subworkflow template, the same way ALIGN_STAR, DEXSEQ_DEU, EDGER_DEU and LEAFCUTTER were ported. nf-core/modules has no prepare_genome subworkflow to install instead, so it stays local.

Changes

  • New meta.yml. Authors are taken from the git history of the subworkflow (@asmaali98, @bensouthgate, @jma1991, @valentinoruggieri), plus @piplus2 as author and maintainer.
  • The subworkflow no longer reads params. --source, --aligner, --pseudo_aligner and --skip_alignment are explicit inputs, passed in from main.nf.
  • Every path goes through file(..., checkIfExists: true) before reaching a module, the MODULE section headers are added and the file is formatted with nextflow lint -format.

Fixes

  • An uncompressed --salmon_index or --suppa_tpm was emitted as [ [:], path ], while the .tar.gz and .gz inputs came out as a bare path. SALMON_QUANT and SUPPA expect the path, so only the compressed form worked. Both are now emitted as paths whatever the input. This came in with 495ff9c.
  • The GTF converted from --gff was named null.gtf, since GFFREAD names its output after meta.id and the meta map was empty. It is now named after the GFF file.

Testing

  • 4 new nf-tests for the subworkflow: building every index from a FASTA and a GTF, gzipped inputs with a GFF3 annotation and supplied STAR (.tar.gz) and Salmon (directory) indices plus a DEXSeq GFF and a SUPPA TPM table, a genome_bam source that needs no index, and a stub.
  • The nf-core genome.gff3 test file makes gffread fail on a malformed strand column, so the GFF3 test uses reference/genes_chrX.gff3 from the rnasplice test-datasets.
  • Full suite green: 73/73 nf-tests, nf-core pipelines lint (tools 4.1.0) 0 failures, prek clean.
  • nf-core subworkflows lint warns that preprocess/transcripts/fasta/gencode and star/genomeparams/upgrade are missing from meta.yml. They are listed under their real names; the linter splits local module names on underscores, the same false positive as for strand_junctions in LEAFCUTTER.

Generated by Claude Opus 5.5

PR checklist

  • This comment contains a description of changes (with reason).
  • If you've fixed a bug or added code that should be tested, add tests!
  • If you've added a new tool - have you followed the pipeline conventions in the contribution docs
  • If necessary, also make a PR on the nf-core/rnasplice branch on the nf-core/test-datasets repository.
  • Make sure your code lints (nf-core pipelines lint).
  • Ensure the test suite passes (nextflow run . -profile test,docker --outdir <OUTDIR>).
  • Check for unexpected warnings in debug mode (nextflow run . -profile debug,test,docker --outdir <OUTDIR>).
  • Usage Documentation in docs/usage.md is updated.
  • Output Documentation in docs/output.md is updated.
  • CHANGELOG.md is updated.
  • README.md is updated (including new tool citations and authors/contributors).

🤖 Generated with Claude Code

Add meta.yml and nf-tests for the local PREPARE_GENOME subworkflow. There is
no prepare_genome subworkflow in nf-core/modules to install in its place.
The tests cover building every index from a FASTA and a GTF, compressed
inputs with a GFF3 annotation and supplied indices, a non fastq source that
needs no index, and a stub run.

Also:

- Take source, aligner, pseudo_aligner and skip_alignment as inputs instead
  of reading params inside the subworkflow.

- Emit salmon_index and suppa_tpm as bare paths whatever the input. An
  uncompressed --salmon_index or --suppa_tpm was emitted as [ [:], path ],
  which SALMON_QUANT and SUPPA cannot use, while the .tar.gz and .gz inputs
  came out as a path.

- Give GFFREAD a meta id taken from the GFF file name, so the converted
  annotation is no longer called null.gtf.

- Pass every path to the modules through file(), add the MODULE section
  headers and format with nextflow lint -format.

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

nf-core pipelines lint overall result: Passed ✅ ⚠️

Posted for pipeline commit d4dd27d

+| ✅ 311 tests passed       |+
#| ❔   5 tests were ignored |#
#| ❔   1 tests had warnings |#
!| ❗  12 tests had warnings |!
Details

❗ Test warnings:

  • readme - README contains the placeholder zenodo.XXXXXXX. This should be replaced with the zenodo doi (after the first release).
  • pipeline_todos - TODO string in CHANGELOG.md: ## v1.1.0dev - [unreleased replace with date on release ]
  • pipeline_todos - TODO string in main.nf.test: define inputs of the process here. Example:
  • pipeline_todos - TODO string in awsfulltest.yml: You can customise AWS full pipeline tests as required
  • pipeline_todos - TODO string in methods_description_template.yml: #Update the HTML below to your preferred methods description, e.g. add publication citation for this pipeline
  • pipeline_todos - TODO string in nextflow.config: Specify any additional parameters here
  • pipeline_todos - TODO string in CONTRIBUTING.md: Add any pipeline specific contribution guidelines here, such as coding styles, procedures, checklists etc.
  • schema_params - Schema param fasta not found from nextflow config
  • schema_params - Schema param gtf not found from nextflow config
  • schema_params - Schema param gff not found from nextflow config
  • schema_params - Schema param star_index not found from nextflow config
  • schema_params - Schema param salmon_index not found from nextflow config

❔ Tests ignored:

  • files_unchanged - File ignored due to lint config: .github/PULL_REQUEST_TEMPLATE.md
  • files_unchanged - File ignored due to lint config: .github/workflows/branch.yml
  • files_unchanged - File ignored due to lint config: .github/workflows/linting.yml
  • files_unchanged - File ignored due to lint config: assets/nf-core-rnasplice_logo_light.png
  • files_unchanged - File ignored due to lint config: docs/images/nf-core-rnasplice_logo_dark.png

❔ Tests fixed:

✅ Tests passed:

Run details

  • nf-core/tools version 4.1.0
  • Run at 2026-09-23 15:43:42

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

❌ nf-test failed with latest Nextflow version

Note

Tests with Nextflow's latest version failed but it will not cause a CI workflow failure.
Please check if the failure is expected with newer (edge-)releases of Nextflow or if it needs fixing.

  • ❌ docker | latest-everything | Shard 7/8

See the full run for details.

@erikrikarddaniel erikrikarddaniel left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed by Claude Code.

Traced both fixes to their consumers rather than taking them on trust. The suppa_tpm one would have thrown rather than merely emitted an odd shape: workflows/rnasplice.nf does ch_suppa_tpm.map { tpm_psi -> [[ id: tpm_psi.baseName ], tpm_psi] } (lines 313 and 386), and baseName does not exist on [[:], path]. Same story for the Salmon index reaching SALMON_QUANT's plain path index. The emitted shapes are uniform now, the take: list matches the call in main.nf, and the meta.yml enums agree with nextflow_schema.json.

Two non-blocking points.

Four of the six uncompress branches are never exercised. The tests cover a gzipped genome FASTA and a gzipped DEXSeq GFF, so GUNZIP_GTF, GUNZIP_GFF, GUNZIP_TRANSCRIPT_FASTA and GUNZIP_SUPPA_TPM never run. All six branches have the same shape and were rewritten in the same pass, so one of them holding a wrong variable would go green today. The gzipped GFF3 path is the one I would most want covered, since it is where GUNZIP_GFF feeds the new [[ id: gff_file.baseName ], gff_file] map — the null.gtf fix is only tested on an uncompressed GFF3. Switching the second test's GFF3 and transcript FASTA inputs to their .gz variants would cover three of the four for the price of a snapshot update.

The second test writes its stand-in inputs into ${outputDir}. Suggested inline: ${workDir} keeps the output directory holding only outputs.

Nothing else came up. nf-test is green on my side of the diff too, and the CHANGELOG entries are in the right sections and in the order this repo uses.

Comment thread subworkflows/local/prepare_genome/tests/main.nf.test Outdated
Comment thread subworkflows/local/prepare_genome/tests/main.nf.test Outdated
@piplus2

piplus2 commented Sep 23, 2026

Copy link
Copy Markdown
Author

Thanks @erikrikarddaniel ! I'll fix those before merging

Review feedback on nf-core#290. The tests only ran GUNZIP_FASTA and
GUNZIP_GFF_DEXSEQ, so a wrong variable in the GTF, GFF3, transcript FASTA
or SUPPA TPM branch would have gone unnoticed. The test datasets carry no
gzipped copy of those files, so the tests gzip them on the fly: the second
test now takes a gzipped GFF3, transcript FASTA and TPM table, and the
genome_bam test a gzipped GTF and an uncompressed TPM table, which keeps
the uncompressed --suppa_tpm fix covered. The gzipped GFF3 also covers the
GFFREAD naming fix on the path that goes through GUNZIP_GFF.

The stand-in inputs are written to workDir instead of outputDir, so the
output directory holds only outputs.

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013jvD13HCjdAWqMnRvEzVH6
@piplus2
piplus2 merged commit b789036 into nf-core:dev Sep 23, 2026
22 checks passed
@piplus2 piplus2 mentioned this pull request Sep 24, 2026
5 of 11 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants