PhyloProcessR is a modular R toolkit for converting raw targeted sequence-capture reads into curated, analysis-ready phylogenomic datasets. Its functions can be called independently, assembled into project-specific pipelines, or used through the supplied complete workflows. It combines domain-specific R functions with established bioinformatics programs for stages that require external executables.
The package and supplied workflows support:
- Organizing raw read data.
- Removing adapters, quality filtering, normalizing, and merging paired-end reads.
- Removing contaminant reads.
- Assembling cleaned reads and recovering targeted loci.
- Mapping reads, calling variants, and generating IUPAC or haplotype consensuses.
- Aligning, filtering, and trimming recovered loci.
- Assessing capture, missing data, depth, paralogy, and alignment quality.
- Separating supported two-copy marker groups for review and analysis.
- Integrating legacy sequences and discovering shared novel loci.
- Concatenating loci by target or gene for downstream analyses.
The R package can be installed and tested without every external program. Individual workflow stages require only the command-line tools they invoke; the provided container and Conda environment install the complete environment.
PhyloProcessR requires R 4.0 or later. Its R package dependencies are declared
in DESCRIPTION and installed by standard R package installers. Depending on
the workflow, external tools can include fastp, BWA, HISAT2, SPAdes, MEGAHIT,
CAP3, BLAST, LAST, MAFFT, trimAl, IQ-TREE, GATK, and Samtools. See the workflow
configuration files for the programs used by each stage.
Clone the repository to obtain the container recipes, Conda environment, and workflow configuration files:
git clone https://github.com/PhyloForge/PhyloProcessR.gitThen change to the setup directory:
cd PhyloProcessR/setup-files/The recommended way to run the full pipeline is with Docker for local Linux or macOS use, or Apptainer/Singularity on an HPC system. This keeps the R and bioinformatics dependencies in a reproducible environment.
Pull the published image from Docker Hub.
For Docker:
docker pull chutter/phyloprocessr:latestMount the current working directory and run a workflow script:
docker run -v $(pwd):/app -w /app chutter/phyloprocessr:latest Rscript workflow-5_trimming.RFor Apptainer or Singularity, convert the same image to a local .sif file:
apptainer pull phyloprocessr.sif docker://chutter/phyloprocessr:latestThen run a workflow:
apptainer exec phyloprocessr.sif Rscript workflow-5_trimming.RInside either container, tools are on PATH; workflow configurations can use
conda.env = "" or /opt/conda/bin/ where a tool-directory argument is
required.
Use this option when you need to modify the supplied environment.
Build Docker from the setup-files directory. On Apple Silicon, the
linux/amd64 platform maintains compatibility with tools distributed only for
x86_64 Linux.
docker build --platform linux/amd64 -t chutter/phyloprocessr:1.0.0 -t chutter/phyloprocessr:latest .Build an Apptainer image from the supplied definition:
apptainer build phyloprocessr.sif Apptainer.defAfter installing a Conda-compatible package manager, create the environment
from setup-files/environment.yml:
conda env create -f environment.yml -n PhyloProcessRActivate it before running a workflow:
conda activate PhyloProcessRInstall the stable version 1.0.0 from GitHub:
install.packages("remotes")
remotes::install_github("PhyloForge/PhyloProcessR@v1.0.0")Install the current development version when you need changes made after the release:
remotes::install_github("PhyloForge/PhyloProcessR")When working inside the supplied Conda environment or container, its pinned R dependencies are already present. Install a local checkout without upgrading them:
remotes::install_local(".", upgrade = "never", dependencies = FALSE)Load the package with:
library(PhyloProcessR)For reproducible analyses, record the installed package version or Git commit
rather than reinstalling the moving development branch in every script.
Supplied workflow configuration files therefore set
install.latest.github = FALSE by default. Developers who intentionally want
the newest beta code for debugging can set it to TRUE; the workflow will then
install the current GitHub version with remotes before loading PhyloProcessR.
To check external programs in the active Conda environment:
setupCheck(anaconda.environment = Sys.getenv("CONDA_PREFIX"))The seven-sample laptop example provides a compact, redistribution-ready dataset for testing PhyloProcessR without a cluster. It retains both single- and multilane libraries and uses a reduced 40-locus target panel to demonstrate how individual package functions can be composed into a project-specific workflow for preprocessing, assembly, target recovery, annotation, alignment, trimming, and quality control.
The example runs in three documented stages and produces a ranked set of 20 alignments together with expected summaries and final alignments for comparison. It defaults to two CPU threads and a 4 GB memory limit. See the complete instructions and expected results and the example-data license.
PhyloProcessR supplies reference workflows for distinct analysis stages.
Configuration files and R scripts are in workflows/. You can run a complete
workflow or compose the underlying functions in a custom R script.
| Workflow | Script | Description |
|---|---|---|
| Workflow X0 | workflow-X0_read-screening.R |
Quickly screen raw reads, cleaning loss, and capture success one sample at a time |
| Workflow 1 | workflow-1_preprocess.R |
Organize, clean, assess, decontaminate, and merge raw reads |
| Workflow 2 | workflow-2_assembly.R |
De novo assembly and target matching; optional additive per-locus recovery and contig curation |
| Workflow 3 | workflow-3_variant-calling.R |
Call variants and make IUPAC or haplotype consensus contigs |
| Workflow 4 | workflow-4_alignment.R |
Annotate contigs and align target markers |
| Workflow 5 | workflow-5_trimming.R |
Trim alignments, concatenate genes, build unlinked dataset; optionally include novel markers from Workflow X4 |
| Workflow X1 | workflow-X1_joint-genotype_VCF.R |
Map samples to a common reference and make a joint VCF file |
| Workflow X2 | workflow-X2_capture-assessment.R |
Assess capture efficiency for cleaned reads from workflow 1 |
| Workflow X3 | workflow-X3_legacy-integration.R |
Integrate Sanger/GenBank legacy alignments into the capture dataset; supports NEXUS conversion and mitochondrial loci |
| Workflow X4 | workflow-X4_novel-loci.R |
Recover, assemble, and align novel shared genomic regions |
| Workflow X5 | workflow-X5_paralog-analysis.R |
Assess saved copies, separate supported two-copy groups, and export filtered markers |
Each workflow has a matching configuration file. Set all project parameters in that file before you run the script.
The tutorials are versioned in the main repository:
- Install PhyloProcessR
- Configure a project
- Run the workflows
- Assess the results
- Integrate legacy data
The tutorial index explains the documentation structure and simplified technical English conventions.
The archived version 1.0.0 release is available at
doi:10.5281/zenodo.22134845. Use the
concept DOI to refer to the software
across releases. Machine-readable citation metadata are provided in
CITATION.cff.
Contributions are welcome; see CONTRIBUTING.md and
CODE_OF_CONDUCT.md. PhyloProcessR is distributed under
the GNU General Public License, version 3 or later.