Skip to content

About

Backend of Nanometa-Live

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

517 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

foi-bioinformatics/nanometanf

GitHub Actions CI Status GitHub Actions Linting Status nf-test

Nextflow nf-core template version run with conda run with docker run with singularity Launch on Seqera Platform

Introduction

nanometanf is an nf-core compliant Nextflow pipeline for Oxford Nanopore long-read sequencing data analysis. It serves as the computational backend for Nanometa Live and covers quality control (Chopper, FASTP, NanoPlot), taxonomic classification (Kraken2), and validation (BLAST, minimap2). Real-time and batch execution modes are supported.

Capabilities:

  • Real-time FASTQ monitoring during a sequencing run, using Nextflow watchPath
  • Streaming Kraken2 architecture (v1.5+) with per-sample parallelism, append-only batch storage, and incremental taxid counting
  • Pre-demultiplexed barcode directory support (flat or per-barcode layouts)
  • Adaptive batching and configurable concurrency for high-throughput runs
  • Tool-agnostic canonical output layer (outdir/canonical/) consumed by Nanometa Live and other frontends; schema in docs/development/canonical_output_specification.md
  • Optional runtime-metrics instrumentation (--runtime_metrics_interval_seconds) emitting periodic [runtime-metrics] log lines plus the standard Nextflow execution_trace_<ts>.txt with queue-waiting fields for backpressure analysis
  • nf-core compliance with an extensive nf-test suite (59 pipeline-level + 23 module-internal tests; CI matrix runs Nextflow 26.04.0 and latest-everything under both Docker and the platform-profile stub-sanity matrix)

Quick start

The recommended environment is the nf-core conda environment with Nextflow 26.04.0 or later.

QC only:

nextflow run foi-bioinformatics/nanometanf \
    --input samplesheet.csv \
    --outdir results \
    -profile conda

Full analysis with classification:

nextflow run foi-bioinformatics/nanometanf \
    --input samplesheet.csv \
    --kraken2_db /path/to/kraken2_db \
    --outdir results \
    -profile conda

CI smoke test (Docker is used in CI; locally use conda):

nextflow run foi-bioinformatics/nanometanf -profile test,docker --outdir test_results
Real-time monitoring examples
# Watch a directory during sequencing
nextflow run foi-bioinformatics/nanometanf \
    --realtime_mode \
    --nanopore_output_dir /path/to/fastq \
    --kraken2_db /path/to/db \
    --outdir results \
    -profile conda

# High-throughput real-time with the streaming classifier (v1.5+)
nextflow run foi-bioinformatics/nanometanf \
    --realtime_mode \
    --kraken2_enable_incremental true \
    --max_classification_forks 8 \
    --nanopore_output_dir /path/to/fastq \
    --kraken2_db /path/to/db \
    --outdir results \
    -profile conda

# Same, with periodic runtime-metrics snapshots written to .nextflow.log
nextflow run foi-bioinformatics/nanometanf \
    --realtime_mode \
    --runtime_metrics_interval_seconds 60 \
    --nanopore_output_dir /path/to/fastq \
    --kraken2_db /path/to/db \
    --outdir results \
    -profile conda

Streaming classification architecture (v1.5+)

For runs with many barcodes (more than ten), the pipeline uses a streaming Kraken2 architecture with the following properties:

  • Per-sample parallelism. Merger and report modules no longer use maxForks 1 globally; samples are processed independently.
  • Append-only batch storage. Each batch is written to its own file with an atomic JSON index. Cumulative output is rebuilt only at end of session, giving O(1) per batch instead of O(n) cumulative rewrites.
  • Incremental taxid counting. Per-sample state files accumulate counts without re-reading prior outputs. Per-sample state is released as soon as the final batch of a sample is processed, so the head-process heap does not grow with run length.
  • Concurrency cap. --max_classification_forks is a global cap on in-flight classification tasks (default 8). --max_concurrent_batches is reserved for a future per-barcode throttle and is currently advisory only -- see the audit P2.9 note in CLAUDE.md for the workaround.

Internal benchmarks measure roughly four to five times higher throughput on runs with twelve or more barcodes compared with the previous architecture.

nextflow run foi-bioinformatics/nanometanf \
    --realtime_mode \
    --kraken2_enable_incremental true \
    --max_classification_forks 8 \
    ...

Deployment profiles

Hardware-specific profiles set resource defaults for different sequencer platforms and deployment scenarios. Combine with an execution-engine profile (conda, docker, or singularity).

Profile Use case Description
test CI and validation Minimal dataset with reduced resources for fast tests
minion MinION / Mk1C Conservative memory and CPU allocation
promethion PromethION (standard) Higher throughput defaults
promethion_8 PromethION (8-barcode) Tuned for multiplexed PromethION runs
field Field deployments Reduced resource footprint for laptops

Examples:

# PromethION run on a workstation with conda
nextflow run foi-bioinformatics/nanometanf -profile promethion,conda --input samplesheet.csv

# MinION field run with conda and real-time monitoring
nextflow run foi-bioinformatics/nanometanf -profile minion,conda \
    --input samplesheet.csv --realtime_mode

# CI smoke test
nextflow run foi-bioinformatics/nanometanf -profile test,docker

For additional resource tuning, use one of the platform profiles above, or pass your own file with -c my.config.

Note: conf/production.config, conf/cloud.config and conf/cluster.config are unregistered and currently do not parse (Nextflow strict syntax). They are kept for reference only -- do not pass them with -c until they are fixed.

Documentation

Documentation index.

User documentation:

Development documentation:

Release information:

Pipeline summary

Input -> QC -> Classification -> Validation -> Reports
  |       |          |                |            |
FASTQ  Chopper    Kraken2          BLAST       MultiQC
       FASTP    (streaming)       minimap2     Canonical JSON
       NanoPlot Taxpasta                       HTML

Supported inputs (mutually exclusive; enforced at schema level via a oneOf constraint):

  1. --input FASTQ samplesheet (standard batch analysis)
  2. --input_dir directory scan (auto-detects barcode subdirectories vs. flat layout; replaces the deprecated --barcode_input_dir)
  3. --realtime_mode + --nanopore_output_dir for live FASTQ monitoring

See the usage guide for parameter details.

Credits

nanometanf was originally written by Andreas Sjodin (@andreassjodin).

The pipeline is built on the nf-core framework. We thank the nf-core community and the developers of the integrated tools.

Citations

If you use nanometanf in your work, please cite:

nanometanf: an Oxford Nanopore sequencing analysis pipeline.

Andreas Sjodin. GitHub 2025. DOI to be assigned upon Zenodo deposit.

A full list of tool references is available in CITATIONS.md.

The nf-core framework reference:

The nf-core framework for community-curated bioinformatics pipelines.

Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso, Sven Nahnsen.

Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.

License

Released under the MIT License. See LICENSE.

About

Backend of Nanometa-Live

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages