Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
148 changes: 148 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,148 @@
# nf-core: agents

This is the main AI context file for nf-core pipelines. All AI agents and coding assistants **MUST** follow the rules contained in this document.

## Natural language

All comments and documentation **MUST** be written in English with British spelling. Documentation files **SHOULD** additionally follow the style guide at https://nf-co.re/docs/developing/documentation/style-guide.
Never use emdashes in prose text, be succinct and to the point. Avoid telltale LLM phrasing such as "Not X, but Y", and excessive use of bold formatting.

## Key nf-core terms

- Module: a single process that achieves a single, well defined task (e.g. aligning reads to a genome)
- Subworkflow: a sequence of chained modules that achieve a specific objective (e.g. FASTQ cleanup and quality check)
- Workflow: a complete sequence of modules and subworkflows that performs a specific analysis (e.g. bulk RNA-seq analysis)
- Pipeline: a complete, executable Nextflow project that defines workflow logic, input handling, and output publishing

## Nextflow pitfalls

- Nextflow supports 2 ways to publish files to the output directory: workflow outputs (modern) and `publishDir` configuration directives in modules.config (legacy). You **SHOULD** publish output consistently with the existing code.
- nf-core tools commands may fail. If that happens, ask the user for help. You **MUST NOT** generate any file that is supposed to be generated by nf-core tools.
- Certain very old pipelines might be using Nextflow DSL1 syntax (with the entire workflow in a single file and channel from/to keywords). This syntax is now deprecated. You **MUST NOT** attempt to work on those pipelines.

## nf-core template structure

The directory you are working on was created with the nf-core pipeline template. Key features of the template are demonstrated below:

```
.

├── conf // directory containing Nextflow configurations for the pipeline (see "Configuration files" below)
├── main.nf // core Nextflow script, may need editing if input structure changes
├── modules // Nextflow DSL2 modules
│ ├── local // local modules (see "Modules" below)
| | └── mymodule // each module must be in a separate directory
│ └── nf-core // nf-core modules (see "Modules" below)
├── nextflow_schema.json // JSON schema describing pipeline parameters
├── subworkflows // Nextflow subworkflows (see "Subworkflows" below)
│ ├── local // local subworkflows
| | └── myswf // each subworkflow must be in a separate directory
│ └── nf-core // nf-core subworkflows
├── tests // nf-test end-to-end tests for the pipeline
│ └── default.nf.test // main test script, must exist
└── workflows // do not add files
└── {pipeline-name}.nf // Nextflow file containing main pipeline logic
```

The pipeline also contains other files and directories. If a file does not follow the treemap above, you **MUST** verify with the user before editing it.

## Modules

- You **SHOULD** use existing nf-core modules for the tools you need, where available.
- You can find available modules and install modules with nf-core tools (see "nf-core tools" section below).
- You **SHOULD NOT** edit nf-core modules in the pipeline modules directory. If unavoidable, you **MAY** edit their `main.nf` if necessary, and if target pipeline logic cannot be achieved with the existing module code. If you do it, you **MUST** run `nf-core modules patch {name}` afterwards, and flag that a PR will be needed to upstream the change.
- The pipeline has a local modules directory. If a script is only useful within this pipeline, you **MAY** create a local module for it.
- Use nf-core tools (see below) to create local module boilerplate and then edit the files.
- Use `ext.args` to pass any command-line arguments (except input files) to the underlying tool.
- Use `ext.prefix` to customize the name of the output files. To include runtime variables in those arguments, use Groovy-style closures, for example: `ext.prefix = { "${meta.id}_filtered" }`, usually through the modules.config file.

## Subworkflows

- You **SHOULD** use nf-core subworkflows that are relevant to the pipeline tasks.
- If none is applicable, create a local subworkflow when it thematically makes sense.

## Pipeline structure

An nf-core pipeline contains 3 main parts called by the root `workflow` block in `main.nf`:

- initialisation workflow (defined in `subworkflows/local/utils_nfcore_{name}_pipeline/main.nf`): handles input processing and validation
- main workflow (defined in `workflows/{name}.nf`): contains the main analysis, including generation of all output files (some nf-core pipelines contain more than one workflow)
- completion workflow (defined in `subworkflows/local/utils_nfcore_{name}_pipeline/main.nf`): handles sending completion notifications

## Configuration files

- You **MUST NOT** edit `base.config`, `igenomes.config`, and `igenomes_ignored.config`
- Set `ext.args` and `ext.prefix` for modules in `modules.config`, using `withName` blocks
- The test in `test.config` **SHOULD** take a few minutes and only test the basic functionality with minimal input
- The test in `test_full.config` **SHOULD** use input and parameters that trigger all pipeline functionality

## Meta map

The meta map is a Nextflow map passed along with each file that contains sample-specific information. The map is created during input processing and passed through modules.

- The meta map **MUST** contain an `id` field with a unique identifier.
- nf-core modules and subworkflows may **only** access `id` and `single_end` fields.
- Local modules, subworkflows, and workflows may create and access any meta fields that are useful for the pipeline.
- Channel operations **MUST** preserve the meta map when present; they **MAY** add, remove, or modify specific keys as required.

## nf-core tools

nf-core provides a CLI toolkit for working with the nf-core template. The core command is `nf-core`. You **SHOULD** always use the tools instead of creating files manually.

Use `nf-core --help` to obtain information about nf-core commands. You can also use `--help` for subcommands.

Write the names of subtool modules in commands with a slash, like `samtools/sort`.

## nf-test and testing

- Each pipeline **MUST** have at least 1 test case.
- Tests have a standardized syntax, with setup (optional), input ("when"), and assertion ("then") sections.
- Tests at a path can be executed with `nf-test test {path}`.
- Most tests create at least 1 snapshot file. You **MUST NOT** edit snapshots manually.
- If you expect the output to change (e.g. after a tool update), update the snapshot with `nf-test test --profile +{docker|singularity|conda} --update-snapshot`. Only regenerate snapshots on the same CPU architecture as CI.
- If a new output file has unstable content, add it to `.nftignore`.

## git and branch policy

This repository has at least 3 git branches: `main` (or `master`), `dev`, and `TEMPLATE`.

- You **MUST NOT** switch or write to the TEMPLATE branch.
- You **MUST NOT** write any code to `main`; use a pull request instead.
- Always create a new branch with a meaningful name for each feature, then open a pull request to `dev`.
- You **MUST NOT** commit feature work directly to `dev`, even on a fork.
- If you work on multiple features in parallel, you **SHOULD** use a separate worktree for each task to prevent clobber.

## Commit rules and routine

- Each commit **SHOULD** contain one logical change.
- Commit title **SHOULD** be concise and written in imperative mood.
- If the commit consists only of installing or updating an nf-core module or subworkflow, limit the commit title to `Install/update nf-core module/subworkflow {name}`.
- Before each commit, you **MUST** stage changes and then run `prek`. Resolve all errors and all possible warnings. Repeat until there are no solvable outstanding issues.

## Push routine

- You should only push to GitHub after implementing some meaningful changes and if the code is working.
- Before pushing, you **MUST** run `nf-core pipelines lint`, resolve all errors and all possible warnings. Repeat until there are no solvable outstanding issues.
- If you are preparing a release (PR to main), use `nf-core pipelines lint --release` instead.
- You **MUST** also run `nf-test test tests/`. If the pipeline fails, resolve the underlying issues. If the test fails due to mismatching snapshots, update them if permitted (see "nf-test and testing" above). Otherwise, fix the issue that caused the unexpected change.
- If you know the code will cause issues or you intend to push more changes, you **SHOULD** add `[skip ci]` at the end of the commit title. You **SHOULD** omit this tag for final review-ready commits.

## PR procedure

- A PR **SHOULD** contain a single feature.
- You **SHOULD** add a line in the relevant section in CHANGELOG.md, listing contributors and the expected PR number.
- The PR **MUST** use and follow the nf-core PR template, including the checklist.
- The PR message **SHOULD** start with a brief explanation of the changes made and the motivation.
- Each PR requires reviews (1 for dev, 2 for main) and passing CI before merging.
- A human can request PR reviews on Slack.

## Agent self-disclosure

- If you generated a majority of the code in a commit, you **MUST** add "Generated by {your name}" at the end of the commit message body.
- If you open a PR autonomously, you **MUST** add "Generated by {your name}" at the end of the PR message (above the checklist).

## References

- Nextflow documentation: https://docs.seqera.io/nextflow
- nf-core tools documentation: https://nf-co.re/docs/nf-core-tools/
- nf-test documentation: https://www.nf-test.com/docs/getting-started/
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- #270 - Move the `LEAFCUTTER` subworkflow to the nf-core subworkflow template. Its `juncs` output is now the per sample `[ meta, junc ]` tuples, instead of a single list flattening the meta maps in with the paths (by @piplus2)
- #272 - `--leafcutter` now works with `--source genome_bam`, whatever the strandedness, since the junction strand no longer has to come from the alignment. For unstranded libraries the clusters differ slightly from LeafCutter's documented STAR route, see the LeafCutter section of `docs/usage.md` (by @piplus2)
- #281 - Refactor `ISOFORMSWITCHANALYZER` to the nf-core module template and update `IsoformSwitchAnalyzeR` 2.2.0 -> 2.12.0 (R 4.3 -> 4.5). `bin/run_isoformswitchanalyzer.R` is now a module template (by @piplus2)
- #284 - Refactor `PREPROCESS_TRANSCRIPTS_FASTA_GENCODE` to the nf-core module template, with `environment.yml`, `meta.yml`, a stub and nf-tests. It reports the `coreutils` version of `cut`, which does the work, instead of `sed` (by @piplus2)

### Fixed

Expand Down Expand Up @@ -87,6 +88,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- #258 - Fix `ISOFORMSWITCHANALYZER` receiving the `.tar.gz` archives instead of the extracted Salmon directories with `--source salmon_results`. `TX2GENE_TXIMPORT` extracts them and now emits them (by @piplus2)
- #261 - Fix the paired rMATS sample order when a sample spans several samplesheet rows (by @piplus2)
- #261 - Fix single condition rMATS runs, which aborted before producing any output (by @piplus2)
- #284 - Fix `--gencode` with `--transcript_fasta` aborting in `PREPARE_GENOME`, as `PREPROCESS_TRANSCRIPTS_FASTA_GENCODE` dropped the meta map the downstream channel expects (by @piplus2)
- #261 - Fix samples split over several samplesheet rows, whose fastq files were passed on without being concatenated (by @piplus2)
- #263 - **Breaking change**: `DEXSEQ_COUNT` no longer counts BAM input as forward stranded, it follows the samplesheet `strandedness`, so DEXSeq results change for BAM samplesheets that do not set the column (reported by @albamasmalavila, fix by @piplus2)
- #264 - Fix the `DEXSEQ_DTU` stub, which wrote file names the process outputs did not match, so a stub run of the DTU path failed (by @piplus2)
Expand Down
12 changes: 12 additions & 0 deletions modules/local/preprocess_transcripts_fasta_gencode/environment.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
---
# yaml-language-server: $schema=https://raw.githubusercontent.com/nf-core/modules/master/modules/environment-schema.json
channels:
- conda-forge
- bioconda
dependencies:
- conda-forge::coreutils=9.5
- conda-forge::grep=3.11
- conda-forge::gzip=1.13
- conda-forge::lbzip2=2.5
- conda-forge::sed=4.8
- conda-forge::tar=1.34
28 changes: 18 additions & 10 deletions modules/local/preprocess_transcripts_fasta_gencode/main.nf
Original file line number Diff line number Diff line change
@@ -1,26 +1,34 @@
process PREPROCESS_TRANSCRIPTS_FASTA_GENCODE {
tag "$fasta"
tag "${meta.id ?: fasta.baseName}"
label 'process_single'

conda "conda-forge::sed=4.7"
container "${ workflow.containerEngine == 'singularity' && !task.ext.singularity_pull_docker_container ?
'https://depot.galaxyproject.org/singularity/ubuntu:20.04' :
'nf-core/ubuntu:20.04' }"
conda "${moduleDir}/environment.yml"
container "${workflow.containerEngine in ['singularity', 'apptainer'] && !task.ext.singularity_pull_docker_container
? 'https://community-cr-prod.seqera.io/docker/registry/v2/blobs/sha256/52/52ccce28d2ab928ab862e25aae26314d69c8e38bd41ca9431c67ef05221348aa/data'
: 'community.wave.seqera.io/library/coreutils_grep_gzip_lbzip2_pruned:838ba80435a629f8'}"

input:
tuple val(meta), path(fasta)

output:
path "*.fa" , emit: fasta
tuple val("${task.process}"), val('sed'), eval('sed --version 2>&1 | head -1 | sed "s/sed (GNU sed) //g"'), topic: versions, emit: versions_sed
tuple val(meta), path("*.fixed.fa"), emit: fasta
tuple val("${task.process}"), val('cut'), eval('cut --version 2>&1 | head -1 | sed "s/^.*coreutils) //"'), topic: versions, emit: versions_cut

when:
task.ext.when == null || task.ext.when

script:
def gzipped = fasta.toString().endsWith('.gz')
def outfile = gzipped ? file(fasta.baseName).baseName : fasta.baseName
def gzipped = fasta.extension == 'gz'
def prefix = task.ext.prefix ?: (gzipped ? file(fasta.baseName).baseName : fasta.baseName)
def command = gzipped ? 'zcat' : 'cat'
"""
${command} ${fasta} | cut -d "|" -f1 > ${outfile}.fixed.fa
${command} ${fasta} | cut -d "|" -f1 > ${prefix}.fixed.fa
"""

stub:
def gzipped = fasta.extension == 'gz'
def prefix = task.ext.prefix ?: (gzipped ? file(fasta.baseName).baseName : fasta.baseName)
"""
touch ${prefix}.fixed.fa
"""
}
72 changes: 72 additions & 0 deletions modules/local/preprocess_transcripts_fasta_gencode/meta.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
# yaml-language-server: $schema=https://raw.githubusercontent.com/nf-core/modules/master/modules/meta-schema.json
name: "preprocess_transcripts_fasta_gencode"
description: Strip the GENCODE transcript FASTA headers down to the transcript id,
so that they match the transcript ids of the GTF
keywords:
- fasta
- gencode
- transcriptome
- header
tools:
- "cut":
description: "GNU coreutils cut, which removes sections from each line of a file"
homepage: "https://www.gnu.org/software/coreutils/"
documentation: "https://www.gnu.org/software/coreutils/manual/html_node/cut-invocation.html"
licence: ["GPL-3.0-or-later"]
identifier: ""

input:
- - meta:
type: map
description: |
Groovy Map containing sample information
e.g. [ id:'test' ]
- fasta:
type: file
description: GENCODE transcript FASTA, whose headers hold the transcript id,
gene id, names, length and biotype separated by `|`, optionally gzipped
pattern: "*.{fa,fasta,fa.gz,fasta.gz}"
ontologies:
- edam: "http://edamontology.org/format_1929" # FASTA

output:
fasta:
- - meta:
type: map
description: |
Groovy Map containing sample information
e.g. [ id:'test' ]
- "*.fixed.fa":
type: file
description: Uncompressed transcript FASTA with the headers cut down to
the transcript id, the first `|` separated field
pattern: "*.fixed.fa"
ontologies:
- edam: "http://edamontology.org/format_1929" # FASTA
versions_cut:
- - ${task.process}:
type: string
description: The process the versions were collected from
- cut:
type: string
description: The tool name
- cut --version 2>&1 | head -1 | sed "s/^.*coreutils) //":
type: eval
description: The expression to obtain the version of the tool

topics:
versions:
- - ${task.process}:
type: string
description: The process the versions were collected from
- cut:
type: string
description: The tool name
- cut --version 2>&1 | head -1 | sed "s/^.*coreutils) //":
type: eval
description: The expression to obtain the version of the tool
authors:
- "@drpatelh"
- "@piplus2"
maintainers:
- "@piplus2"
Loading
Loading