Skip to content

2. Modules

jaclew edited this page Jan 7, 2025 · 19 revisions

repgenr

This is the entry point to the software, where subcommands for modules are specified. To list available modules, run repgenr without arguments or use repgenr --help. Every module also accepts --help to display additional information.

metadata

The metadata module retrieves the GTDB metadata table based on specified input criteria. It outputs the NCBI accession numbers for all samples at the requested taxonomic level. A random outgroup sample annotated as a "representative genome" in GTDB is selected at one taxonomic level higher than the specified level. This outgroup sample is later used in a module that computes phylogeny. The outgroup can optionally be specified by the user with an NCBI assembly accession number.

Metadata information is stored in the specified work directory as follows:

  • metadata_level.txt: Stores the specified output taxonomic level.
  • metadata_selected.tsv: Contains GTDB metadata for the output samples.
  • metadata_summary.tsv: Summarizes the number of samples per taxonomic level.
  • metadata_summary_number_in_level.png: Bar plot showing sample abundance at each taxonomic level.
  • metadata_summary_number_per_level.png: Bar plot of abundance across taxonomic levels.
  • outgroup_accession.txt: Stores the NCBI accession number of the outgroup.

genome

The genome module uses the accession numbers obtained from the metadata command to download genomes from NCBI. The genomes are saved in a folder named genomes, with files named <family>_<genus>_<species>_<NCBI-accession-number>.fasta. A list of downloaded accessions is stored in the file ncbi_acc_download_list.txt. The outgroup sample is stored in a separate folder named outgroup.

glance

The glance module quickly computes an average nucleotide identity (ANI) dendrogram for the downloaded genomes. It generates a PDF file named glance_clustering_dendrogram.pdf in the specified working directory. This module helps overview genome similarity and inform dereplication settings, such as ANI values.

derep

The derep module clusters genomes based on average nucleotide identity (ANI) using the dRep software. Users can adjust the ANI threshold for final clustering (--secondary_ani, default=0.99) and for initial clustering (--primary_ani, default=0.9). The chosen algorithm for ANI comparison can be set with --S_algorithm (default: fastANI). This module requires significant computational resources and time, especially for large datasets. To improve performance, large datasets can be divided into chunks (--process_size) and processed in parallel (--num_processes), reducing the number of all-vs-all comparisons. After initial chunk processing, dRep is executed on their outputs. Optionally, the process can be paused for review after the first step with --prompt. If chunking is used, a second step is performed on the chunk outputs using --secondary_only.

The results are stored as follows:

  • genomes_derep_representants: Folder containing dereplicated genomes.
  • derep_genomes_status.tsv: the columns show (1) file-name, (2) representative genome (true or false), (3) QC-information.
  • derep_genomeInformation.tsv: the columns show genome (1) file-name, (2) completeness, (3) contamination, (4) strain heterogeneity, (5) length, (6) N50, and (7) centrality. For information about these metrics, see dRep documentation.
  • derep_potentially_discarded_genomes.tsv: Lists genomes that fall below default thresholds for completeness, contamination, and length.
  • derep_genomes_summary.tsv: the columns show (1) representative-genome, (2) number of dereplicated-and-removed genomes, and (3) the names of dereplicated-and-removed genomes.
  • derep_genomes_summary2.tsv: presents the previous file in a different way, showing columns as (1) representative-genome, (2) contained genome, and (3) boolean, if the contained genome is the same as the representative-genome.
  • derep_clustered_genomes.tsv: the columns show (1) dereplication representative-genome and (2) dereplicated-and-removed genome. If chunking is used, derep_chunks_clustered_genomes.tsv provides the same information with an additional column indicating the chunk.

Intermediary files of genome-quality reported by dRep is saved in the folder dereplication_intermediary_files and is used if the derep-module is executed multiple times, greatly reducing computation time. Intermediary files can be ignored by --skip_intermediary. The parameters used during dereplication is saved in the text-file named derep_parameters.txt and includes ANI-thresholds and the chunk size used. The folder dereplication_workdir is a working-directory for the derep module and unless --keep_files is applied it is removed after execution.

derep_stocker

The derep_stocker module facilitates the management of multiple derep runs, useful when testing different settings. Completed runs are named (--name) and archived (--pack), while previous runs are loaded with --unpack. Runs can be listed (--list) and names deleted (--delete). If error messages about missing files appear, ensure derep_genomes_summary.tsv and derep_genomes_summary2.tsv are in your work directory; if absent, they can be created with derep_summarize.py --workdir <path_to_workdir>.

phylo

The phylo-module computes a phylogenetic tree based on the representative genome sequences that result from the derep module. Alternatively, phylogeny is computed for all genome sequences by invoking --all_genomes. The module can be run either in a fast and rough mode or a slow but accurate mode, specified by parameter --mode. Optionally, the outgroup is ignored when computing the tree by invoking --no_outgroup. The module outputs the tree in newick format to the file genomes_derep_representants.dnd. The multiple sequence alignment-file used to generate the tree is saved by applying --keep_msa (only applicable to --mode accurate) and phylogeny calculation can be ignored if applying --halt_after_msa in case a custom approach is preferred. To keep temporary files when using --mode accurate then apply --keep_files.

tree2tax

The tree2tax module generates files compatible with downstream applications such as FlexTaxD, enabling the initiation or modification of a taxonomic database. Users have the option to assign custom basenames to nodes, which are automatically enumerated (e.g., <basename>_1, <basename>_2, …, <basename>_N). In the absence of user-defined basenames, each node is assigned a unique MD5 hash identifier derived from its descending leaves, ensuring uniqueness within any non-redundant database.

When integrating RepGenR output with an existing FlexTaxD database to replace a taxonomic branch, the --root_name parameter must be used to designate the appropriate parent node as the root. Optionally, the --remove_outgroup parameter can be invoked to exclude the outgroup from the final output. For analyses including all genomes instead of just dereplicated representatives, the --all_genomes flag should be specified. To include dereplicated (redundant) genomes alongside the representant genome, use flag ´--include_dereplicated´.

The module outputs two key files:

  • derep_genomes_tree2tax.tsv: This file presents a parent-child relationship mapping of the phylogeny, with two columns indicating the names of the child (1) and parent (2) nodes.
  • derep_genomes_map.tsv: This file provides the path to genome files, showing the accession number (1) and the corresponding genome name (2).

These outputs are instrumental for users aiming to enrich or update their taxonomic databases with precise phylogenetic information.

[NEW] Virus workflow

RepGenR now supports the download and organization of virus genomes using BV-BRC.org. While still under development, the documentation for the viral workflow is limited. For a demonstration, please see the virus-example: https://github.com/FOI-Bioinformatics/repgenr/wiki/3c.-RepGenR-virus-workflow

A short note on segmented viral sequences: The automated estimate of genome lengths can not handle multiple segments and should be manually set in vgenome using --length_range. Similarly, the automated determination of an outgroup may be flawed and --no_outgroup is recommended in the vgenome step.

vmetadata

The virus-metadata module retrieves the fasta for the input virus group, e.g. a virus family. Available viruses from BV-BRC.org are listed by --list. Using data in the downloaded sequence headers, additional taxonomic information is retrieved from NCBI using Entrez. Metadata files are produced in the specified work directory as follows:

  • virus_metadata_base.tsv: contains information parsed from the BV-BRC.org fasta file. Entries are grouped by NCBI taxonomy ID.
  • virus_metadata_ncbi.tsv: contains additional metadata from NCBI such as taxonomy by name and ID.
  • virus_metadata_ncbi_summary.tsv: summarizes the number of datasets at taxonomic levels.

vgenome

The virus-genome module lets the user select taxa to unpack from the downloaded fasta file. The metadata file metadata_ncbi_summary.tsv is useful to overview the downloaded dataset and select taxonomy.

derep

After executing the vmetadata and vgenome modules then the derep module should be run using --virus flag to apply settings described here: https://drep.readthedocs.io/en/latest/choosing_parameters.html#comparing-and-dereplicating-non-bacterial-genomes

Clone this wiki locally