nanometanf is an nf-core compliant Nextflow pipeline for Oxford Nanopore long-read sequencing data analysis. It serves as the computational backend for Nanometa Live and covers quality control (Chopper, FASTP, NanoPlot), taxonomic classification (Kraken2), and validation (BLAST, minimap2). Real-time and batch execution modes are supported.
Capabilities:
- Real-time FASTQ monitoring during a sequencing run, using Nextflow
watchPath - Streaming Kraken2 architecture (v1.5+) with per-sample parallelism, append-only batch storage, and incremental taxid counting
- Pre-demultiplexed barcode directory support (flat or per-barcode layouts)
- Adaptive batching and configurable concurrency for high-throughput runs
- Tool-agnostic canonical output layer (
outdir/canonical/) consumed by Nanometa Live and other frontends; schema indocs/development/canonical_output_specification.md - Optional runtime-metrics instrumentation (
--runtime_metrics_interval_seconds) emitting periodic[runtime-metrics]log lines plus the standard Nextflowexecution_trace_<ts>.txtwith queue-waiting fields for backpressure analysis - nf-core compliance with an extensive nf-test suite (59 pipeline-level + 23
module-internal tests; CI matrix runs Nextflow 26.04.0 and
latest-everythingunder both Docker and the platform-profile stub-sanity matrix)
The recommended environment is the nf-core conda environment with Nextflow
26.04.0 or later.
QC only:
nextflow run foi-bioinformatics/nanometanf \
--input samplesheet.csv \
--outdir results \
-profile condaFull analysis with classification:
nextflow run foi-bioinformatics/nanometanf \
--input samplesheet.csv \
--kraken2_db /path/to/kraken2_db \
--outdir results \
-profile condaCI smoke test (Docker is used in CI; locally use conda):
nextflow run foi-bioinformatics/nanometanf -profile test,docker --outdir test_resultsReal-time monitoring examples
# Watch a directory during sequencing
nextflow run foi-bioinformatics/nanometanf \
--realtime_mode \
--nanopore_output_dir /path/to/fastq \
--kraken2_db /path/to/db \
--outdir results \
-profile conda
# High-throughput real-time with the streaming classifier (v1.5+)
nextflow run foi-bioinformatics/nanometanf \
--realtime_mode \
--kraken2_enable_incremental true \
--max_classification_forks 8 \
--nanopore_output_dir /path/to/fastq \
--kraken2_db /path/to/db \
--outdir results \
-profile conda
# Same, with periodic runtime-metrics snapshots written to .nextflow.log
nextflow run foi-bioinformatics/nanometanf \
--realtime_mode \
--runtime_metrics_interval_seconds 60 \
--nanopore_output_dir /path/to/fastq \
--kraken2_db /path/to/db \
--outdir results \
-profile condaFor runs with many barcodes (more than ten), the pipeline uses a streaming Kraken2 architecture with the following properties:
- Per-sample parallelism. Merger and report modules no longer use
maxForks 1globally; samples are processed independently. - Append-only batch storage. Each batch is written to its own file with an atomic JSON index. Cumulative output is rebuilt only at end of session, giving O(1) per batch instead of O(n) cumulative rewrites.
- Incremental taxid counting. Per-sample state files accumulate counts without re-reading prior outputs. Per-sample state is released as soon as the final batch of a sample is processed, so the head-process heap does not grow with run length.
- Concurrency cap.
--max_classification_forksis a global cap on in-flight classification tasks (default 8).--max_concurrent_batchesis reserved for a future per-barcode throttle and is currently advisory only -- see the audit P2.9 note inCLAUDE.mdfor the workaround.
Internal benchmarks measure roughly four to five times higher throughput on runs with twelve or more barcodes compared with the previous architecture.
nextflow run foi-bioinformatics/nanometanf \
--realtime_mode \
--kraken2_enable_incremental true \
--max_classification_forks 8 \
...Hardware-specific profiles set resource defaults for different sequencer
platforms and deployment scenarios. Combine with an execution-engine profile
(conda, docker, or singularity).
| Profile | Use case | Description |
|---|---|---|
test |
CI and validation | Minimal dataset with reduced resources for fast tests |
minion |
MinION / Mk1C | Conservative memory and CPU allocation |
promethion |
PromethION (standard) | Higher throughput defaults |
promethion_8 |
PromethION (8-barcode) | Tuned for multiplexed PromethION runs |
field |
Field deployments | Reduced resource footprint for laptops |
Examples:
# PromethION run on a workstation with conda
nextflow run foi-bioinformatics/nanometanf -profile promethion,conda --input samplesheet.csv
# MinION field run with conda and real-time monitoring
nextflow run foi-bioinformatics/nanometanf -profile minion,conda \
--input samplesheet.csv --realtime_mode
# CI smoke test
nextflow run foi-bioinformatics/nanometanf -profile test,dockerFor additional resource tuning, use one of the platform profiles above, or
pass your own file with -c my.config.
Note:
conf/production.config,conf/cloud.configandconf/cluster.configare unregistered and currently do not parse (Nextflow strict syntax). They are kept for reference only -- do not pass them with-cuntil they are fixed.
User documentation:
- Usage guide -- parameters and execution modes
- Output reference -- directory layout and file formats
- Quick start -- scenario-based walkthrough
- Real-time processing
- Performance tuning
- Troubleshooting
Development documentation:
- Development guide
- Canonical output specification -- source of truth for
outdir/canonical/ - Testing guide -- nf-test conventions
- CLAUDE.md -- developer notes for AI-assisted work
Release information:
Input -> QC -> Classification -> Validation -> Reports
| | | | |
FASTQ Chopper Kraken2 BLAST MultiQC
FASTP (streaming) minimap2 Canonical JSON
NanoPlot Taxpasta HTML
Supported inputs (mutually exclusive; enforced at schema level via a oneOf
constraint):
--inputFASTQ samplesheet (standard batch analysis)--input_dirdirectory scan (auto-detects barcode subdirectories vs. flat layout; replaces the deprecated--barcode_input_dir)--realtime_mode+--nanopore_output_dirfor live FASTQ monitoring
See the usage guide for parameter details.
nanometanf was originally written by Andreas Sjodin (@andreassjodin).
The pipeline is built on the nf-core framework. We thank the nf-core community and the developers of the integrated tools.
If you use nanometanf in your work, please cite:
nanometanf: an Oxford Nanopore sequencing analysis pipeline.
Andreas Sjodin. GitHub 2025. DOI to be assigned upon Zenodo deposit.
A full list of tool references is available in CITATIONS.md.
The nf-core framework reference:
The nf-core framework for community-curated bioinformatics pipelines.
Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso, Sven Nahnsen.
Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.
Released under the MIT License. See LICENSE.