Detailed workflow and tool citations

This page follows a nucleotide assembly through MetaPathways, from reference preparation to connected annotation, abundance and pathway tables. It identifies the software responsible for each calculation and explains how Nextflow schedules it. For a shorter introduction, see the conceptual overview, local/HPC architecture and data-flow overview.

If you want pathway/genome databases (PGDBs), complete the Pathway Tools installation guide first. Annotation, read abundance and reporting can run with --skip_ptools.

How to cite your analysis

Cite MetaPathways, Nextflow, the biological tools used by your selected stages, and the reference databases you searched. Add Pathway Tools, MetaCyc, MAGSplitter and Camelot when those components are used. Cite the container runtime and Slurm when applicable. The tables below connect each component to its publication or official project; the full references include DOI links and websites.

Download the software and database bibliography. Papers describe methods, not the precise software or database versions in your run: also record the MP version, environment, reference releases, image identity, settings and execution logs. See reproducibility. For the bundled test data, use the separate CAMI and CAMI II citations.

1. Prepare references and the optional Pathway Tools image

These are setup commands, run before sample analysis. An existing compatible MPDB and SIF can be reused across samples and computers; they are not rebuilt for each sample.

Detailed workflow 1

Zoom diagram · Mermaid source

build_db prepares selected public references and the supporting MPDB structure. FAST uses fastdb; BLAST+ uses makeblastdb. rRNA searches require nucleotide BLAST indexes. MP’s own preparation scripts build the lookup tables consumed by annotation and reporting. Database options and custom references are described in database construction.

build_pt builds and validates a private Apptainer image from the licensed installer. With its optional MPDB destination, MP extracts the matching MetaCyc protein FASTA and companion flat files, builds the selected search index, and derives the EC/reaction/pathway mappings. This licensed preparation is separate from the public-reference build. See MetaCyc preparation.

The BLAST+ installation inside the Pathway Tools image supports Pathway Tools’ own sequence databases. It is distinct from the FAST/BLAST choice for MP’s functional annotation searches. Docker is an alternative way to run the MP application image; it is not required to build the Pathway Tools SIF.

2. Plan and schedule work

analysis_wf validates the manifest or automatically matched input layout, resolves sample IDs, and builds a task dependency graph. The Python controller writes modular Nextflow definitions plus task specifications. Nextflow runs eligible tasks locally or submits them to Slurm; it does not submit every downstream task before its inputs are ready. See input organization and execution.

Each worker checks its task receipt and tracked inputs before either reusing valid outputs or invoking the recorded command. Tasks write biological outputs into the sample’s output directory and retain command output, status and resource diagnostics. Nextflow’s trace, report and timeline describe the invocation; MP’s receipts also distinguish reused work and optional entity outcomes.

Scheduling level

What can run together

Across samples

Ready stages from different samples can overlap. Samples are not processed one at a time.

Within an annotation stage

Independent searches or parses for different reference databases can overlap. The following stage waits for that group.

After Pathway Tools input preparation

Read abundance, the community PGDB, and MAG splitting can proceed independently. Each MAG PGDB waits for splitting.

Within a task

Capable tools receive the requested thread budget. pProdigal and ptRNAscan distribute work across their underlying predictors. Serial stages and each Pathway Tools task reserve one CPU.

Across executors

Local execution fits aggregate CPU/memory limits. Slurm uses per-job requests plus submission/queue limits; dependencies remain managed by Nextflow.

--threads controls capable tools, not the total workflow concurrency. --max_cpus, --max_memory, --max_tasks and Slurm submission settings are explained in resources. Multiple processes shown inside one task, such as CoverM, SAMtools and featureCounts, run as subcommands of that task; they are not separate Slurm jobs.

3. Predict and annotate features

This diagram follows the stage order for nucleotide FASTA input. Boxes grouping multiple MP transformations are abbreviated for readability. The reference searches within a group can run concurrently; this does not imply that all RNA prediction and protein search stages run independently within a sample. Protein-only inputs follow the compatible reduced path described in the stage reference.

Detailed workflow 2

Zoom diagram · Mermaid source

pProdigal parallelizes Prodigal gene prediction. MP derives and filters protein sequences before searching the selected protein references with FAST or BLAST+. MP computes sequence-based reference scores and then applies its configured search thresholds and score-ratio rules. FAST is a threaded implementation derived from LAST; citing FAST does not mean the separate LAST executable was run.

Barrnap predicts rRNA features; BLAST+ searches their sequences against the chosen rRNA references. ptRNAscan partitions the input and starts single-threaded tRNAscan-SE workers within the assigned CPU budget. Barrnap and tRNAscan-SE use their own underlying RNA search software/models, including nhmmer/HMMER and Infernal, respectively. These are not additional MP-level workflow stages.

MP combines protein and RNA evidence, performs feature-overlap processing through pybedtools/BEDTools, and writes annotation tables and annotated sequence products. An annotation row’s taxonomy comes from that row’s reference database. Hit taxonomy and the within-database lowest common ancestor (LCA) remain distinct; MP does not borrow a taxon assignment from another database to fill the row. References without supported taxonomy produce the documented unavailable status. See annotation interpretation.

4. Measure abundance, build PGDBs and assemble reports

Solid arrows show the downstream processing paths; dotted arrows connect existing products to reporting. Additional inputs are named inside the relevant boxes. The abundance branch requires reads. The PGDB branch requires Pathway Tools; genome-specific PGDBs additionally require a contig-to-genome map.

Detailed workflow 3

Zoom diagram · Mermaid source

CoverM generates contig abundance and the BAM used for feature counting. MP invokes coverm contig without overriding its mapper, so the mapper follows the installed CoverM version’s default. Check that version and its logged command/backend when recording methods; a directory called bwa/ is not evidence that BWA performed the mapping. Backend citations are listed below for use when applicable.

SAMtools name-sorts the expected CoverM BAM. featureCounts, distributed with Subread, uses the annotated feature coordinates; MP’s abund_calc.py derives the feature abundance table. Contig-level coverage and feature-level counts are separate measurements. Reads are not required to predict features or infer pathways.

MAGSplitter reuses the community’s existing annotations and the supplied membership map; it does not bin contigs, rerun protein searches or infer genome quality. MP stages sequence-backed inputs for both community and genome entities, preserving feature identities and recording input normalization. This provides Pathway Tools with sequence and coordinate information as well as functional annotations. See PGDB input preparation.

Pathway Tools/PathoLogic builds each PGDB using the MetaCyc reference in the selected installation. Taxonomic scope/pruning and transport inference are analysis settings; they do not change MP’s upstream taxonomic annotation tables. After construction, Pathway Tools exports the database and Camelot/MP extracts pathway-to-gene associations. Independent containerized entities can run concurrently, with one CPU per Pathway Tools task.

A required community PGDB failure stops dependent work. Optional MAG PGDB failures are recorded and do not become successful empty pathway predictions. With --skip_ptools, the PGDB branch is omitted. With compact results enabled, per-sample cleanup waits for that sample’s terminal tasks and retains reporting inputs and diagnostics. See execution and cleanup.

After successful workflow completion, the MP controller builds the combined report and SQLite index. This is controller-side postprocessing, not a separate Slurm biological job. The explorer follows sample-scoped contig, feature and entity relationships and exports selected tables; it does not rerun annotation or pathway inference. See the report tutorial and results schema.

Tool and wrapper index

“MP” in the diagrams means code shipped with MetaPathways. These transformations are covered by the MetaPathways citation. A repository link is supplied where a separate paper has not been established; it is not a claim that a paper does not exist.

Component

Role in this workflow

Citation and official project

MetaPathways

Validation, planning, QC, parsing, LCA, annotation integration, PGDB staging, abundance calculations and reporting

Publication; GitHub

Nextflow

Task dependencies, execution, trace and scheduling

Publication; website

pProdigal / Prodigal

Parallel wrapper / protein-coding gene prediction

pProdigal; Prodigal paper; Prodigal

FAST

Protein search and reference indexing (fastal, fastdb)

GitHub; LAST method ancestry

NCBI BLAST+

Alternative protein searches, rRNA searches and database indexing

Publication; website

barrnap / nhmmer

rRNA prediction / underlying profile search

barrnap; nhmmer paper; HMMER

ptRNAscan / tRNAscan-SE / Infernal

Parallel wrapper / tRNA prediction / underlying RNA search

MP wrapper source; tRNAscan-SE paper, website; Infernal paper, website

pybedtools / BEDTools

Feature interval comparison and overlap processing

pybedtools paper, GitHub; BEDTools paper, GitHub

CoverM

Read mapping orchestration and contig coverage measurements

Publication; GitHub

SAMtools

BAM processing for read abundance

Publication; website

featureCounts / Subread

Assign alignments to annotated features

Publication; website

MAGSplitter

Split annotated features using contig-to-genome membership

GitHub

Pathway Tools / PathoLogic

Build, save and export pathway/genome databases

Publication; website

Camelot (camelot-frs)

Load PGDB flat-file frames and query pathway/gene relationships

Bitbucket

Minimap2, BWA, Strobealign

Read-mapping backends; cite the backend actually used by CoverM

Minimap2, GitHub; BWA, GitHub; Strobealign, version-specific citation guidance

The BWA reference below describes the original BWA method; for a run using BWA-MEM, also cite Li (2013), Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM, arXiv:1303.3997. Follow Strobealign’s version-specific citation guidance rather than assuming one article covers every backend release. Installed backend packages do not imply that all of them ran.

Reference databases

Choose citations for the references actually used, and report their release dates or versions. Searching a database is not the same as running its authors’ annotation software: for example, MP’s eggNOG reference search does not invoke eggNOG-mapper.

Reference

Use

Citation and official website

UniProtKB/Swiss-Prot

Protein function and supported taxon identifiers

UniProt; uniprot.org

UniRef

Clustered protein reference sequences and supported taxonomy

UniRef; UniRef

CAZy

Carbohydrate-active enzyme reference annotations

CAZy; cazy.org; dbCAN download service used by the MP reference builder

eggNOG

Orthology-based functional reference, when supplied/selected

eggNOG; eggnog.embl.de

SILVA

Ribosomal RNA reference searches

SILVA; arb-silva.de

NCBI Taxonomy

Taxon names and lineage tree for supported reference hits/LCA

NCBI Taxonomy; NCBI

ENZYME / ExPASy

EC nomenclature and supporting enzyme tables

ENZYME; enzyme.expasy.org

MetaCyc

Licensed optional protein reference and Pathway Tools’ pathway knowledge base

MetaCyc; metacyc.org

MetaCyc has two roles: selecting it as an MP protein annotation database is optional, while Pathway Tools uses its own installed MetaCyc knowledge base for inference. Omitting the protein search does not remove that inference reference. Custom references require their own provenance and citations.

Supporting software and distribution

These components support execution, parsing, packaging or presentation; they are not all standalone biological tasks. Exact transitive dependencies vary with the resolved environment. The environment specification and exported package list are the version inventory for a particular run.

Component

Role and reference

Apptainer

Private SIF build/execution. Project, GitHub, recommended foundational Singularity paper. The project recommends that paper under its former name; record the actual Apptainer version separately.

Slurm

Optional cluster scheduler. Publication, official documentation.

Mamba, Conda, Bioconda

Environment/package installation. Mamba, Conda, Bioconda paper, Bioconda.

Docker and Quay

Alternative application container runtime and image distribution. Docker, Quay.

Python, Java/OpenJDK, Groovy

MP and Nextflow implementation/runtime stack. Python, OpenJDK, Groovy.

pandas, NumPy, SciPy

Tabular/numerical processing and supporting dependencies. pandas paper, project; NumPy paper, project; SciPy paper, project.

pyfastx, pysam

Sequence access and alignment-format support in the software environment. pyfastx paper, GitHub; pysam.

sexpdata, html2text

Parse Lisp-style data and convert HTML text during PGDB extraction. sexpdata, html2text.

tqdm, peppy

Supporting progress/configuration dependencies. tqdm, peppy.

SQLite

Disk-backed relational report index through Python’s SQLite interface. sqlite.org.

Cython, setuptools, pip

Build/install support. Cython paper, project; setuptools, pip.

curl, GNU Wget, urllib3

Download/HTTP support. curl, Wget, urllib3.

Xvfb/X.Org, xauth, ncurses, OpenSSL, libxml2

Headless Pathway Tools display and container/runtime support. X.Org, ncurses, OpenSSL, libxml2.

Sphinx, MyST, Read the Docs theme, Mermaid

Documentation build and these diagrams. Sphinx, MyST, theme, sphinxcontrib-mermaid, Mermaid, Read the Docs.

Historical Snakefiles, older mapping helpers and bundled legacy executables are not evidence that those routes ran. The supported controller uses Nextflow, and the abundance route above uses CoverM followed by featureCounts. Cite software from the actual commands and environment used for your analysis.

Full references

The references below identify methods and resources; publication versions are not pins for the installed software or MPDB release. Wrappers without a verified separate publication are linked in the tool index above.

MetaPathways

McLaughlin RJ, Liu TX, Altman T, et al. (2024). MetaPathways v3.5: Modularity and Scalability Improvements for Pathway Inference from Environmental Genomes. bioRxiv. doi:10.1101/2024.06.04.597460. Official project.

Nextflow

Di Tommaso P, Chatzou M, Floden EW, et al. (2017). Nextflow enables reproducible computational workflows. Nature Biotechnology 35(4): 316–319. doi:10.1038/nbt.3820. Official project.

Prodigal

Hyatt D, Chen GL, LoCascio PF, et al. (2010). Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 11(1): 119. doi:10.1186/1471-2105-11-119. Official project.

LAST

Kiełbasa SM, Wan R, Sato K, et al. (2011). Adaptive seeds tame genomic sequence comparison. Genome Research 21(3): 487–493. doi:10.1101/gr.113985.110. Official project.

BLAST+

Camacho C, Coulouris G, Avagyan V, et al. (2009). BLAST+: architecture and applications. BMC Bioinformatics 10(1): 421. doi:10.1186/1471-2105-10-421. Official project.

nhmmer / HMMER

Wheeler TJ, Eddy SR. (2013). nhmmer: DNA homology search with profile HMMs. Bioinformatics 29(19): 2487–2489. doi:10.1093/bioinformatics/btt403. Official project.

Infernal

Nawrocki EP, Eddy SR. (2013). Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29(22): 2933–2935. doi:10.1093/bioinformatics/btt509. Official project.

tRNAscan-SE

Chan PP, Lin BY, Mak AJ, et al. (2021). tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Research 49(16): 9077–9096. doi:10.1093/nar/gkab688. Official project.

pybedtools

Dale RK, Pedersen BS, Quinlan AR. (2011). Pybedtools: a flexible Python library for manipulating genomic datasets and annotations. Bioinformatics 27(24): 3423–3424. doi:10.1093/bioinformatics/btr539. Official project.

BEDTools

Quinlan AR, Hall IM. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26(6): 841–842. doi:10.1093/bioinformatics/btq033. Official project.

CoverM

Aroney STN, Newell RJP, Nissen JN, et al. (2025). CoverM: read alignment statistics for metagenomics. Bioinformatics 41(4): btaf147. doi:10.1093/bioinformatics/btaf147. Official project.

SAMtools

Danecek P, Bonfield JK, Liddle J, et al. (2021). Twelve years of SAMtools and BCFtools. GigaScience 10(2): giab008. doi:10.1093/gigascience/giab008. Official project.

featureCounts

Liao Y, Smyth GK, Shi W. (2014). featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30(7): 923–930. doi:10.1093/bioinformatics/btt656. Official project.

Pathway Tools

Karp PD, Latendresse M, Paley SM, et al. (2016). Pathway Tools version 19.0 update: software for pathway/genome informatics and systems biology. Briefings in Bioinformatics 17(5): 877–890. doi:10.1093/bib/bbv079. Official project.

Minimap2

Li H. (2018). Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34(18): 3094–3100. doi:10.1093/bioinformatics/bty191. Official project.

BWA

Li H, Durbin R. (2009). Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics 25(14): 1754–1760. doi:10.1093/bioinformatics/btp324. Official project.

Strobealign

Sahlin K. (2022). Strobealign: flexible seed size enables ultra-fast and accurate read alignment. Genome Biology 23(1): 260. doi:10.1186/s13059-022-02831-7. Official project.

Apptainer / Singularity

Kurtzer GM, Sochat V, Bauer MW. (2017). Singularity: Scientific containers for mobility of compute. PLOS ONE 12(5): e0177459. doi:10.1371/journal.pone.0177459. Official project.

Slurm

Yoo AB, Jette MA, Grondona M. (2003). SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Processing. Lecture Notes in Computer Science 2862: 44–60. doi:10.1007/10968987_3. Official project.

Bioconda

The Bioconda Team, Grüning B, Dale R, et al. (2018). Bioconda: sustainable and comprehensive software distribution for the life sciences. Nature Methods 15(7): 475–476. doi:10.1038/s41592-018-0046-7. Official project.

pandas

McKinney W. (2010). Data Structures for Statistical Computing in Python. Proceedings of the Python in Science Conference: 56–61. doi:10.25080/majora-92bf1922-00a. Official project.

NumPy

Harris CR, Millman KJ, van der Walt SJ, et al. (2020). Array programming with NumPy. Nature 585(7825): 357–362. doi:10.1038/s41586-020-2649-2. Official project.

SciPy

Virtanen P, Gommers R, Oliphant TE, et al. (2020). SciPy 1.0: fundamental algorithms for scientific computing in Python. Nature Methods 17(3): 261–272. doi:10.1038/s41592-019-0686-2. Official project.

pyfastx

Du L, Liu Q, Fan Z, et al. (2021). Pyfastx: a robust Python package for fast random access to sequences from plain and gzipped FASTA/Q files. Briefings in Bioinformatics 22(4): bbaa368. doi:10.1093/bib/bbaa368. Official project.

Cython

Behnel S, Bradshaw R, Citro C, et al. (2011). Cython: The Best of Both Worlds. Computing in Science & Engineering 13(2): 31–39. doi:10.1109/mcse.2010.118. Official project.

UniProt

The UniProt Consortium, Bateman A, Martin MJ, et al. (2023). UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Research 51(D1): D523–D531. doi:10.1093/nar/gkac1052. Official project.

UniRef

Suzek BE, Huang H, McGarvey P, et al. (2007). UniRef: comprehensive and non-redundant UniProt reference clusters. Bioinformatics 23(10): 1282–1288. doi:10.1093/bioinformatics/btm098. Official project.

CAZy

Drula E, Garron ML, Dogan S, et al. (2022). The carbohydrate-active enzyme database: functions and literature. Nucleic Acids Research 50(D1): D571–D577. doi:10.1093/nar/gkab1045. Official project.

eggNOG

Huerta-Cepas J, Szklarczyk D, Heller D, et al. (2019). eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Research 47(D1): D309–D314. doi:10.1093/nar/gky1085. Official project.

SILVA

Quast C, Pruesse E, Yilmaz P, et al. (2013). The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Research 41(D1): D590–D596. doi:10.1093/nar/gks1219. Official project.

NCBI Taxonomy

Schoch CL, Ciufo S, Domrachev M, et al. (2020). NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database 2020: baaa062. doi:10.1093/database/baaa062. Official project.

ENZYME

Bairoch A. (2000). The ENZYME database in 2000. Nucleic Acids Research 28(1): 304–305. doi:10.1093/nar/28.1.304. Official project.

MetaCyc

Caspi R, Billington R, Keseler IM, et al. (2020). The MetaCyc database of metabolic pathways and enzymes - a 2019 update. Nucleic Acids Research 48(D1): D445–D453. doi:10.1093/nar/gkz862. Official project.