Detailed workflow and tool citations
This page follows a nucleotide assembly through MetaPathways, from reference preparation to connected annotation, abundance and pathway tables. It identifies the software responsible for each calculation and explains how Nextflow schedules it. For a shorter introduction, see the conceptual overview, local/HPC architecture and data-flow overview.
If you want pathway/genome databases (PGDBs), complete the Pathway Tools installation guide first. Annotation, read abundance and reporting can run with --skip_ptools.
How to cite your analysis
Cite MetaPathways, Nextflow, the biological tools used by your selected stages, and the reference databases you searched. Add Pathway Tools, MetaCyc, MAGSplitter and Camelot when those components are used. Cite the container runtime and Slurm when applicable. The tables below connect each component to its publication or official project; the full references include DOI links and websites.
Download the software and database bibliography. Papers describe methods, not the precise software or database versions in your run: also record the MP version, environment, reference releases, image identity, settings and execution logs. See reproducibility. For the bundled test data, use the separate CAMI and CAMI II citations.
1. Prepare references and the optional Pathway Tools image
These are setup commands, run before sample analysis. An existing compatible MPDB and SIF can be reused across samples and computers; they are not rebuilt for each sample.
build_db prepares selected public references and the supporting MPDB structure. FAST uses fastdb; BLAST+ uses makeblastdb. rRNA searches require nucleotide BLAST indexes. MP’s own preparation scripts build the lookup tables consumed by annotation and reporting. Database options and custom references are described in database construction.
build_pt builds and validates a private Apptainer image from the licensed installer. With its optional MPDB destination, MP extracts the matching MetaCyc protein FASTA and companion flat files, builds the selected search index, and derives the EC/reaction/pathway mappings. This licensed preparation is separate from the public-reference build. See MetaCyc preparation.
The BLAST+ installation inside the Pathway Tools image supports Pathway Tools’ own sequence databases. It is distinct from the FAST/BLAST choice for MP’s functional annotation searches. Docker is an alternative way to run the MP application image; it is not required to build the Pathway Tools SIF.
2. Plan and schedule work
analysis_wf validates the manifest or automatically matched input layout, resolves sample IDs, and builds a task dependency graph. The Python controller writes modular Nextflow definitions plus task specifications. Nextflow runs eligible tasks locally or submits them to Slurm; it does not submit every downstream task before its inputs are ready. See input organization and execution.
Each worker checks its task receipt and tracked inputs before either reusing valid outputs or invoking the recorded command. Tasks write biological outputs into the sample’s output directory and retain command output, status and resource diagnostics. Nextflow’s trace, report and timeline describe the invocation; MP’s receipts also distinguish reused work and optional entity outcomes.
Scheduling level |
What can run together |
|---|---|
Across samples |
Ready stages from different samples can overlap. Samples are not processed one at a time. |
Within an annotation stage |
Independent searches or parses for different reference databases can overlap. The following stage waits for that group. |
After Pathway Tools input preparation |
Read abundance, the community PGDB, and MAG splitting can proceed independently. Each MAG PGDB waits for splitting. |
Within a task |
Capable tools receive the requested thread budget. pProdigal and ptRNAscan distribute work across their underlying predictors. Serial stages and each Pathway Tools task reserve one CPU. |
Across executors |
Local execution fits aggregate CPU/memory limits. Slurm uses per-job requests plus submission/queue limits; dependencies remain managed by Nextflow. |
--threads controls capable tools, not the total workflow concurrency. --max_cpus, --max_memory, --max_tasks and Slurm submission settings are explained in resources. Multiple processes shown inside one task, such as CoverM, SAMtools and featureCounts, run as subcommands of that task; they are not separate Slurm jobs.
3. Predict and annotate features
This diagram follows the stage order for nucleotide FASTA input. Boxes grouping multiple MP transformations are abbreviated for readability. The reference searches within a group can run concurrently; this does not imply that all RNA prediction and protein search stages run independently within a sample. Protein-only inputs follow the compatible reduced path described in the stage reference.
pProdigal parallelizes Prodigal gene prediction. MP derives and filters protein sequences before searching the selected protein references with FAST or BLAST+. MP computes sequence-based reference scores and then applies its configured search thresholds and score-ratio rules. FAST is a threaded implementation derived from LAST; citing FAST does not mean the separate LAST executable was run.
Barrnap predicts rRNA features; BLAST+ searches their sequences against the chosen rRNA references. ptRNAscan partitions the input and starts single-threaded tRNAscan-SE workers within the assigned CPU budget. Barrnap and tRNAscan-SE use their own underlying RNA search software/models, including nhmmer/HMMER and Infernal, respectively. These are not additional MP-level workflow stages.
MP combines protein and RNA evidence, performs feature-overlap processing through pybedtools/BEDTools, and writes annotation tables and annotated sequence products. An annotation row’s taxonomy comes from that row’s reference database. Hit taxonomy and the within-database lowest common ancestor (LCA) remain distinct; MP does not borrow a taxon assignment from another database to fill the row. References without supported taxonomy produce the documented unavailable status. See annotation interpretation.
4. Measure abundance, build PGDBs and assemble reports
Solid arrows show the downstream processing paths; dotted arrows connect existing products to reporting. Additional inputs are named inside the relevant boxes. The abundance branch requires reads. The PGDB branch requires Pathway Tools; genome-specific PGDBs additionally require a contig-to-genome map.
CoverM generates contig abundance and the BAM used for feature counting. MP invokes coverm contig without overriding its mapper, so the mapper follows the installed CoverM version’s default. Check that version and its logged command/backend when recording methods; a directory called bwa/ is not evidence that BWA performed the mapping. Backend citations are listed below for use when applicable.
SAMtools name-sorts the expected CoverM BAM. featureCounts, distributed with Subread, uses the annotated feature coordinates; MP’s abund_calc.py derives the feature abundance table. Contig-level coverage and feature-level counts are separate measurements. Reads are not required to predict features or infer pathways.
MAGSplitter reuses the community’s existing annotations and the supplied membership map; it does not bin contigs, rerun protein searches or infer genome quality. MP stages sequence-backed inputs for both community and genome entities, preserving feature identities and recording input normalization. This provides Pathway Tools with sequence and coordinate information as well as functional annotations. See PGDB input preparation.
Pathway Tools/PathoLogic builds each PGDB using the MetaCyc reference in the selected installation. Taxonomic scope/pruning and transport inference are analysis settings; they do not change MP’s upstream taxonomic annotation tables. After construction, Pathway Tools exports the database and Camelot/MP extracts pathway-to-gene associations. Independent containerized entities can run concurrently, with one CPU per Pathway Tools task.
A required community PGDB failure stops dependent work. Optional MAG PGDB failures are recorded and do not become successful empty pathway predictions. With --skip_ptools, the PGDB branch is omitted. With compact results enabled, per-sample cleanup waits for that sample’s terminal tasks and retains reporting inputs and diagnostics. See execution and cleanup.
After successful workflow completion, the MP controller builds the combined report and SQLite index. This is controller-side postprocessing, not a separate Slurm biological job. The explorer follows sample-scoped contig, feature and entity relationships and exports selected tables; it does not rerun annotation or pathway inference. See the report tutorial and results schema.
Tool and wrapper index
“MP” in the diagrams means code shipped with MetaPathways. These transformations are covered by the MetaPathways citation. A repository link is supplied where a separate paper has not been established; it is not a claim that a paper does not exist.
Component |
Role in this workflow |
Citation and official project |
|---|---|---|
MetaPathways |
Validation, planning, QC, parsing, LCA, annotation integration, PGDB staging, abundance calculations and reporting |
|
Nextflow |
Task dependencies, execution, trace and scheduling |
|
pProdigal / Prodigal |
Parallel wrapper / protein-coding gene prediction |
|
FAST |
Protein search and reference indexing ( |
|
NCBI BLAST+ |
Alternative protein searches, rRNA searches and database indexing |
|
barrnap / nhmmer |
rRNA prediction / underlying profile search |
|
ptRNAscan / tRNAscan-SE / Infernal |
Parallel wrapper / tRNA prediction / underlying RNA search |
|
pybedtools / BEDTools |
Feature interval comparison and overlap processing |
|
CoverM |
Read mapping orchestration and contig coverage measurements |
|
SAMtools |
BAM processing for read abundance |
|
featureCounts / Subread |
Assign alignments to annotated features |
|
MAGSplitter |
Split annotated features using contig-to-genome membership |
|
Pathway Tools / PathoLogic |
Build, save and export pathway/genome databases |
|
Camelot ( |
Load PGDB flat-file frames and query pathway/gene relationships |
|
Minimap2, BWA, Strobealign |
Read-mapping backends; cite the backend actually used by CoverM |
Minimap2, GitHub; BWA, GitHub; Strobealign, version-specific citation guidance |
The BWA reference below describes the original BWA method; for a run using BWA-MEM, also cite Li (2013), Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM, arXiv:1303.3997. Follow Strobealign’s version-specific citation guidance rather than assuming one article covers every backend release. Installed backend packages do not imply that all of them ran.
Reference databases
Choose citations for the references actually used, and report their release dates or versions. Searching a database is not the same as running its authors’ annotation software: for example, MP’s eggNOG reference search does not invoke eggNOG-mapper.
Reference |
Use |
Citation and official website |
|---|---|---|
UniProtKB/Swiss-Prot |
Protein function and supported taxon identifiers |
|
UniRef |
Clustered protein reference sequences and supported taxonomy |
|
CAZy |
Carbohydrate-active enzyme reference annotations |
CAZy; cazy.org; dbCAN download service used by the MP reference builder |
eggNOG |
Orthology-based functional reference, when supplied/selected |
|
SILVA |
Ribosomal RNA reference searches |
|
NCBI Taxonomy |
Taxon names and lineage tree for supported reference hits/LCA |
|
ENZYME / ExPASy |
EC nomenclature and supporting enzyme tables |
|
MetaCyc |
Licensed optional protein reference and Pathway Tools’ pathway knowledge base |
MetaCyc has two roles: selecting it as an MP protein annotation database is optional, while Pathway Tools uses its own installed MetaCyc knowledge base for inference. Omitting the protein search does not remove that inference reference. Custom references require their own provenance and citations.
Supporting software and distribution
These components support execution, parsing, packaging or presentation; they are not all standalone biological tasks. Exact transitive dependencies vary with the resolved environment. The environment specification and exported package list are the version inventory for a particular run.
Component |
Role and reference |
|---|---|
Apptainer |
Private SIF build/execution. Project, GitHub, recommended foundational Singularity paper. The project recommends that paper under its former name; record the actual Apptainer version separately. |
Slurm |
Optional cluster scheduler. Publication, official documentation. |
Mamba, Conda, Bioconda |
Environment/package installation. Mamba, Conda, Bioconda paper, Bioconda. |
Docker and Quay |
Alternative application container runtime and image distribution. Docker, Quay. |
Python, Java/OpenJDK, Groovy |
MP and Nextflow implementation/runtime stack. Python, OpenJDK, Groovy. |
pandas, NumPy, SciPy |
Tabular/numerical processing and supporting dependencies. pandas paper, project; NumPy paper, project; SciPy paper, project. |
pyfastx, pysam |
Sequence access and alignment-format support in the software environment. pyfastx paper, GitHub; pysam. |
sexpdata, html2text |
Parse Lisp-style data and convert HTML text during PGDB extraction. sexpdata, html2text. |
tqdm, peppy |
Supporting progress/configuration dependencies. tqdm, peppy. |
SQLite |
Disk-backed relational report index through Python’s SQLite interface. sqlite.org. |
Cython, setuptools, pip |
Build/install support. Cython paper, project; setuptools, pip. |
curl, GNU Wget, urllib3 |
|
Xvfb/X.Org, xauth, ncurses, OpenSSL, libxml2 |
Headless Pathway Tools display and container/runtime support. X.Org, ncurses, OpenSSL, libxml2. |
Sphinx, MyST, Read the Docs theme, Mermaid |
Documentation build and these diagrams. Sphinx, MyST, theme, sphinxcontrib-mermaid, Mermaid, Read the Docs. |
Historical Snakefiles, older mapping helpers and bundled legacy executables are not evidence that those routes ran. The supported controller uses Nextflow, and the abundance route above uses CoverM followed by featureCounts. Cite software from the actual commands and environment used for your analysis.
Full references
The references below identify methods and resources; publication versions are not pins for the installed software or MPDB release. Wrappers without a verified separate publication are linked in the tool index above.
MetaPathways
McLaughlin RJ, Liu TX, Altman T, et al. (2024). MetaPathways v3.5: Modularity and Scalability Improvements for Pathway Inference from Environmental Genomes. bioRxiv. doi:10.1101/2024.06.04.597460. Official project.
Nextflow
Di Tommaso P, Chatzou M, Floden EW, et al. (2017). Nextflow enables reproducible computational workflows. Nature Biotechnology 35(4): 316–319. doi:10.1038/nbt.3820. Official project.
Prodigal
Hyatt D, Chen GL, LoCascio PF, et al. (2010). Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 11(1): 119. doi:10.1186/1471-2105-11-119. Official project.
LAST
Kiełbasa SM, Wan R, Sato K, et al. (2011). Adaptive seeds tame genomic sequence comparison. Genome Research 21(3): 487–493. doi:10.1101/gr.113985.110. Official project.
BLAST+
Camacho C, Coulouris G, Avagyan V, et al. (2009). BLAST+: architecture and applications. BMC Bioinformatics 10(1): 421. doi:10.1186/1471-2105-10-421. Official project.
nhmmer / HMMER
Wheeler TJ, Eddy SR. (2013). nhmmer: DNA homology search with profile HMMs. Bioinformatics 29(19): 2487–2489. doi:10.1093/bioinformatics/btt403. Official project.
Infernal
Nawrocki EP, Eddy SR. (2013). Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29(22): 2933–2935. doi:10.1093/bioinformatics/btt509. Official project.
tRNAscan-SE
Chan PP, Lin BY, Mak AJ, et al. (2021). tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Research 49(16): 9077–9096. doi:10.1093/nar/gkab688. Official project.
pybedtools
Dale RK, Pedersen BS, Quinlan AR. (2011). Pybedtools: a flexible Python library for manipulating genomic datasets and annotations. Bioinformatics 27(24): 3423–3424. doi:10.1093/bioinformatics/btr539. Official project.
BEDTools
Quinlan AR, Hall IM. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26(6): 841–842. doi:10.1093/bioinformatics/btq033. Official project.
CoverM
Aroney STN, Newell RJP, Nissen JN, et al. (2025). CoverM: read alignment statistics for metagenomics. Bioinformatics 41(4): btaf147. doi:10.1093/bioinformatics/btaf147. Official project.
SAMtools
Danecek P, Bonfield JK, Liddle J, et al. (2021). Twelve years of SAMtools and BCFtools. GigaScience 10(2): giab008. doi:10.1093/gigascience/giab008. Official project.
featureCounts
Liao Y, Smyth GK, Shi W. (2014). featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30(7): 923–930. doi:10.1093/bioinformatics/btt656. Official project.
Pathway Tools
Karp PD, Latendresse M, Paley SM, et al. (2016). Pathway Tools version 19.0 update: software for pathway/genome informatics and systems biology. Briefings in Bioinformatics 17(5): 877–890. doi:10.1093/bib/bbv079. Official project.
Minimap2
Li H. (2018). Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34(18): 3094–3100. doi:10.1093/bioinformatics/bty191. Official project.
BWA
Li H, Durbin R. (2009). Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics 25(14): 1754–1760. doi:10.1093/bioinformatics/btp324. Official project.
Strobealign
Sahlin K. (2022). Strobealign: flexible seed size enables ultra-fast and accurate read alignment. Genome Biology 23(1): 260. doi:10.1186/s13059-022-02831-7. Official project.
Apptainer / Singularity
Kurtzer GM, Sochat V, Bauer MW. (2017). Singularity: Scientific containers for mobility of compute. PLOS ONE 12(5): e0177459. doi:10.1371/journal.pone.0177459. Official project.
Slurm
Yoo AB, Jette MA, Grondona M. (2003). SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Processing. Lecture Notes in Computer Science 2862: 44–60. doi:10.1007/10968987_3. Official project.
Bioconda
The Bioconda Team, Grüning B, Dale R, et al. (2018). Bioconda: sustainable and comprehensive software distribution for the life sciences. Nature Methods 15(7): 475–476. doi:10.1038/s41592-018-0046-7. Official project.
pandas
McKinney W. (2010). Data Structures for Statistical Computing in Python. Proceedings of the Python in Science Conference: 56–61. doi:10.25080/majora-92bf1922-00a. Official project.
NumPy
Harris CR, Millman KJ, van der Walt SJ, et al. (2020). Array programming with NumPy. Nature 585(7825): 357–362. doi:10.1038/s41586-020-2649-2. Official project.
SciPy
Virtanen P, Gommers R, Oliphant TE, et al. (2020). SciPy 1.0: fundamental algorithms for scientific computing in Python. Nature Methods 17(3): 261–272. doi:10.1038/s41592-019-0686-2. Official project.
pyfastx
Du L, Liu Q, Fan Z, et al. (2021). Pyfastx: a robust Python package for fast random access to sequences from plain and gzipped FASTA/Q files. Briefings in Bioinformatics 22(4): bbaa368. doi:10.1093/bib/bbaa368. Official project.
Cython
Behnel S, Bradshaw R, Citro C, et al. (2011). Cython: The Best of Both Worlds. Computing in Science & Engineering 13(2): 31–39. doi:10.1109/mcse.2010.118. Official project.
UniProt
The UniProt Consortium, Bateman A, Martin MJ, et al. (2023). UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Research 51(D1): D523–D531. doi:10.1093/nar/gkac1052. Official project.
UniRef
Suzek BE, Huang H, McGarvey P, et al. (2007). UniRef: comprehensive and non-redundant UniProt reference clusters. Bioinformatics 23(10): 1282–1288. doi:10.1093/bioinformatics/btm098. Official project.
CAZy
Drula E, Garron ML, Dogan S, et al. (2022). The carbohydrate-active enzyme database: functions and literature. Nucleic Acids Research 50(D1): D571–D577. doi:10.1093/nar/gkab1045. Official project.
eggNOG
Huerta-Cepas J, Szklarczyk D, Heller D, et al. (2019). eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Research 47(D1): D309–D314. doi:10.1093/nar/gky1085. Official project.
SILVA
Quast C, Pruesse E, Yilmaz P, et al. (2013). The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Research 41(D1): D590–D596. doi:10.1093/nar/gks1219. Official project.
NCBI Taxonomy
Schoch CL, Ciufo S, Domrachev M, et al. (2020). NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database 2020: baaa062. doi:10.1093/database/baaa062. Official project.
ENZYME
Bairoch A. (2000). The ENZYME database in 2000. Nucleic Acids Research 28(1): 304–305. doi:10.1093/nar/28.1.304. Official project.
MetaCyc
Caspi R, Billington R, Keseler IM, et al. (2020). The MetaCyc database of metabolic pathways and enzymes - a 2019 update. Nucleic Acids Research 48(D1): D445–D453. doi:10.1093/nar/gkz862. Official project.