Skip to content

Port rMATS to the nf-core rmats/prep module and refactor RMATS_POST to the nf-core module structure - #285

Merged
piplus2 merged 3 commits into
nf-core:devfrom
piplus2:rmats-nfcore
Sep 22, 2026
Merged

piplus2 merged 3 commits into
nf-core:devfrom
piplus2:rmats-nfcore

Conversation

@piplus2

@piplus2 piplus2 commented Sep 22, 2026

Copy link
Copy Markdown

Summary

Port the rMATS modules to the nf-core structure. There is an nf-core rmats/prep module but no rmats/post, so:

  • RMATS_PREP is replaced by the nf-core rmats/prep module, installed with nf-core modules install and left untouched. It preps one BAM file at a time, so each BAM is read once whatever the number of contrasts it takes part in, instead of every BAM of a contrast in one go. The RMATS subworkflow gathers the .rmats files of the samples of each contrast by sample id and hands them to the post step with the bam lists.
  • RMATS_POST is refactored to the nf-core module template: [ meta, rmats, bam_list1, bam_list2 ], [ meta2, gtf ] and read_length inputs, per file type emits (mats, from_gtf, raw_input, summary, log), environment.yml, meta.yml, a stub and nf-tests. It also reports the PAIRADISE version.
  • --cstat, --paired-stats and --novelSS with --mil/--mel move to ext.args in modules.config, together with the library options of the prep step (-t, --libType, --variable-read-length, --allow-clipping). --statoff stays in the module, since it follows from the absence of a second bam list.
  • Authors from the git history of both modules: @asmaali98, @bensouthgate, @jma1991 and @piplus2, maintainer @piplus2.

Why the split is safe

rmats.py --task post never reads the BAM files. It walks --tmp for .rmats files and matches each one to the --b1/--b2 lists by the BAM file name recorded on its first line (split_sg_files_by_bam in rmatspipeline.pyx), ignoring .rmats files of BAMs that are not in the lists. CREATE_BAMLIST and the nf-core prep module both write the staged file name, so they agree. The hashes of the rMATS result tables in the pipeline snapshots are unchanged by the port.

Output layout

  • rmats/prep/{sample}.rmats and rmats/prep/{sample}_read_outcomes_by_bam.txt (was rmats/{contrast}/rmats_temp/* with timestamped names, plus a rmats_prep.log)
  • rmats/{contrast}{_paired}/* and rmats/{contrast}{_paired}.log (was rmats/{contrast}/rmats_post{_paired}/* and .log)
  • the intermediate tmp/ folder of the post step is no longer published

docs/output.md, tests/.nftignore and the stable_path ignore lists of the pipeline tests are updated accordingly.

One thing to keep an eye on

The nf-core rmats/prep module is process_single (1 CPU, 6 GB), where the old all-in-one RMATS_PREP was process_high. One BAM per task needs far less than all of them at once, but a large BAM may still want more memory; that is a withName: RMATS_PREP resource override in modules.config if it turns out to be needed.

Testing

  • nf-test test modules/local/rmats_post --profile +docker: 4/4 pass (unpaired, paired, single condition, stub), snapshot stable on rerun
  • pipeline nf-tests default, genome_bam and multiple_runs regenerated; the rMATS result table hashes are the same as before
  • full nf-test test tests/ suite: 7/7 pass
  • nf-core pipelines lint: 0 failures; nf-core modules lint --local: only the generic warnings every local module here gets; prek clean

Generated by Claude Opus 5

PR checklist

  • This comment contains a description of changes (with reason).
  • If you've fixed a bug or added code that should be tested, add tests!
  • If you've added a new tool - have you followed the pipeline conventions in the contribution docs
  • Make sure your code lints (nf-core pipelines lint).
  • Ensure the test suite passes (nf-test test tests/ --profile +docker).
  • Output Documentation in docs/output.md is updated.
  • CHANGELOG.md is updated.
  • README.md is updated (including new tool citations and authors/contributors).

🤖 Generated with Claude Code

https://claude.ai/code/session_019SiaVrAVHBtxZcGgoiwQun

piplus2 and others added 2 commits September 22, 2026 10:03
Replace the local RMATS_PREP module with the nf-core rmats/prep module. It
preps one BAM file at a time, so each BAM is read once whatever the number
of contrasts it takes part in, instead of every BAM of a contrast in one
go. rMATS matches the .rmats files to the --b1/--b2 lists by BAM file name
and never reads the BAMs again in the post step, which is what makes the
split possible: the RMATS subworkflow gathers the .rmats files of the
samples of each contrast by sample id and hands them to RMATS_POST with
the bam lists.

Refactor RMATS_POST to the nf-core module template: [ meta, rmats,
bam_list1, bam_list2 ], [ meta2, gtf ] and read_length inputs, per file
type emits, environment.yml, meta.yml, a stub and nf-tests (unpaired,
paired, single condition, stub). --cstat, --paired-stats and --novelSS
with --mil/--mel move to ext.args in modules.config, together with the
library options of the prep step, and the paired model is marked by the
prefix instead of a nested folder. --statoff stays in the module, since
it follows from the absence of a second bam list. The module also reports
the PAIRADISE version.

The .rmats files are published to rmats/prep/{sample}.rmats and the
results of a contrast to rmats/{contrast}{_paired}/, with the log next to
the folder. The intermediate tmp/ folder of the post step is no longer
published. The pipeline snapshots change accordingly; the hashes of the
rMATS result tables are unchanged.

Generated by Claude Opus 5

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SiaVrAVHBtxZcGgoiwQun
@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

nf-core pipelines lint overall result: Passed ✅ ⚠️

Posted for pipeline commit e91fe6f

+| ✅ 312 tests passed       |+
#| ❔   5 tests were ignored |#
#| ❔   1 tests had warnings |#
!| ❗  12 tests had warnings |!
Details

❗ Test warnings:

  • readme - README contains the placeholder zenodo.XXXXXXX. This should be replaced with the zenodo doi (after the first release).
  • pipeline_todos - TODO string in CHANGELOG.md: ## v1.1.0dev - [unreleased replace with date on release ]
  • pipeline_todos - TODO string in nextflow.config: Specify any additional parameters here
  • pipeline_todos - TODO string in CONTRIBUTING.md: Add any pipeline specific contribution guidelines here, such as coding styles, procedures, checklists etc.
  • pipeline_todos - TODO string in main.nf.test: define inputs of the process here. Example:
  • pipeline_todos - TODO string in methods_description_template.yml: #Update the HTML below to your preferred methods description, e.g. add publication citation for this pipeline
  • pipeline_todos - TODO string in awsfulltest.yml: You can customise AWS full pipeline tests as required
  • schema_params - Schema param fasta not found from nextflow config
  • schema_params - Schema param gtf not found from nextflow config
  • schema_params - Schema param gff not found from nextflow config
  • schema_params - Schema param star_index not found from nextflow config
  • schema_params - Schema param salmon_index not found from nextflow config

❔ Tests ignored:

  • files_unchanged - File ignored due to lint config: .github/PULL_REQUEST_TEMPLATE.md
  • files_unchanged - File ignored due to lint config: .github/workflows/branch.yml
  • files_unchanged - File ignored due to lint config: .github/workflows/linting.yml
  • files_unchanged - File ignored due to lint config: assets/nf-core-rnasplice_logo_light.png
  • files_unchanged - File ignored due to lint config: docs/images/nf-core-rnasplice_logo_dark.png

❔ Tests fixed:

✅ Tests passed:

Run details

  • nf-core/tools version 4.1.0
  • Run at 2026-09-22 09:55:49

@erikrikarddaniel erikrikarddaniel left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude's comments:

I checked the part the pipeline snapshots can't cover. The test data is paired-end and unstranded, which are rMATS's defaults for -t and --libType. So unchanged hashes alone can't show that post is safe without -t, --libType, --variable-read-length and --allow-clipping. I ran rmats.py 4.3.0 from the module's container on two of the chrX test BAMs, with prep set to stranded paired-end and, separately, to single-end. In both cases post's output was byte-identical with and without those four options (about 1500 events). The prep settings do change the result (1499 vs 1520 events), so the options sit on the right step.

One question inline, about duplicate sample ids with --source genome_bam.

A small naming thought, not blocking: modules/local/rmats/post would sit next to modules/nf-core/rmats/prep and match the nf-core name if RMATS_POST ever goes upstream.

// The post step of a contrast takes the `.rmats` files of every sample in its bam
// lists. The bam lists carry the BAM file names, which is what rMATS matches the
// `.rmats` files by, so the samples of a contrast are looked up by sample id here
ch_rmats_by_sample = RMATS_PREP.out.rmats

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can a --source genome_bam samplesheet have two rows with the same sample? The comment at line 33 mentions technical replicates sharing a sample id, and assets/schema_input_genome_bam.json has no uniqueEntries. Unlike the fastq source, the genome_bam branch of PIPELINE_INITIALISATION doesn't groupTuple either.

If it can, each BAM gets its own RMATS_PREP task with the same prefix, so both write ${meta.id}.rmats. Here both then match the same id, and RMATS_POST stages them into rmats_tmp/*, which fails. I checked that part with a minimal process:

Process `P` input file name collision -- There are multiple input files for each of the following file names: rmats_tmp/S1.rmats

The two published rmats/prep/{sample}.rmats would also overwrite each other. The old all-in-one prep took both BAMs in one list, so it didn't hit this. I didn't run the pipeline with such a samplesheet.

If duplicates aren't meant to be supported, "uniqueEntries": ["sample"] in the schema would say so up front. If they are, naming the prep output after the BAM instead of the sample would keep them apart, e.g. ext.prefix = { genome_bam.baseName } on RMATS_PREP. The bam lists already carry the BAM names that rMATS matches on.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, and the answer is that duplicates were never viable for this source, so I went with the first option. With --source genome_bam every BAM goes through BAM_SORT_STATS_SAMTOOLS first, whose prefix is ${meta.id}_sorted (conf/modules.config), so two rows of the same sample already produced two S1_sorted.bam: the old all-in-one prep staged both into one task (the same name collision you saw) and CREATE_BAMLIST would have listed the name twice, which rMATS rejects as duplicate input bam files. That also rules out naming the prep output after the BAM, since by then the BAM is already named after the sample.

e91fe6f adds "uniqueEntries": ["sample"] to the genome_bam, transcriptome_bam and salmon_results schemas, the three sources without a merge step, and says so in docs/usage.md. Checked with a fifth row repeating ERR188383 on the test_genome_bam samplesheet:

The following errors have been detected in dup.csv:
-> Entry 5: Detected duplicate entries: [sample:ERR188383]

The three source tests still pass with unique names.

Only the fastq source merges the rows of a sample. With --source
genome_bam two rows of the same sample already collided in
BAM_SORT_STATS_SAMTOOLS, which names the sorted BAM after the sample id,
and from there in CREATE_BAMLIST and rMATS's own duplicate BAM check. The
per sample RMATS_PREP now adds a `${meta.id}.rmats` collision to that
list, so the schemas of the genome_bam, transcriptome_bam and
salmon_results sources declare `uniqueEntries` on `sample` and the
samplesheet fails validation up front instead. Documented in usage.md.

Raised in review of nf-core#285.

Generated by Claude Opus 5

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SiaVrAVHBtxZcGgoiwQun
@piplus2

piplus2 commented Sep 22, 2026

Copy link
Copy Markdown
Author

Thanks for checking the prep/post option split against stranded and single-end data, that was the part the unstranded paired test set could not tell.

On the name: agreed that modules/local/rmats/post reads better next to modules/nf-core/rmats/prep. I have left it at modules/local/rmats_post in this PR to keep the diff to the port itself, and would do the move together with the other local modules that still use the flat name (create_bamlist, strand_junctions, misopysettings) so the directory follows one rule.

Generated by Claude Opus 5

@piplus2
piplus2 merged commit 4a8604e into nf-core:dev Sep 22, 2026
23 checks passed
@piplus2

piplus2 commented Sep 22, 2026

Copy link
Copy Markdown
Author

thanks @erikrikarddaniel !

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants