Skip to content

Repository files navigation

PhyloProcessR

R CMD check License: GPL v3+ DOI

PhyloProcessR is a modular R toolkit for converting raw targeted sequence-capture reads into curated, analysis-ready phylogenomic datasets. Its functions can be called independently, assembled into project-specific pipelines, or used through the supplied complete workflows. It combines domain-specific R functions with established bioinformatics programs for stages that require external executables.

The package and supplied workflows support:

  1. Organizing raw read data.
  2. Removing adapters, quality filtering, normalizing, and merging paired-end reads.
  3. Removing contaminant reads.
  4. Assembling cleaned reads and recovering targeted loci.
  5. Mapping reads, calling variants, and generating IUPAC or haplotype consensuses.
  6. Aligning, filtering, and trimming recovered loci.
  7. Assessing capture, missing data, depth, paralogy, and alignment quality.
  8. Separating supported two-copy marker groups for review and analysis.
  9. Integrating legacy sequences and discovering shared novel loci.
  10. Concatenating loci by target or gene for downstream analyses.

The R package can be installed and tested without every external program. Individual workflow stages require only the command-line tools they invoke; the provided container and Conda environment install the complete environment.

Prerequisites

PhyloProcessR requires R 4.0 or later. Its R package dependencies are declared in DESCRIPTION and installed by standard R package installers. Depending on the workflow, external tools can include fastp, BWA, HISAT2, SPAdes, MEGAHIT, CAP3, BLAST, LAST, MAFFT, trimAl, IQ-TREE, GATK, and Samtools. See the workflow configuration files for the programs used by each stage.

Complete workflow environment

Clone the repository to obtain the container recipes, Conda environment, and workflow configuration files:

git clone https://github.com/PhyloForge/PhyloProcessR.git

Then change to the setup directory:

cd PhyloProcessR/setup-files/

The recommended way to run the full pipeline is with Docker for local Linux or macOS use, or Apptainer/Singularity on an HPC system. This keeps the R and bioinformatics dependencies in a reproducible environment.

Option 1: Run a pre-built container (recommended)

Pull the published image from Docker Hub.

For Docker:

docker pull chutter/phyloprocessr:latest

Mount the current working directory and run a workflow script:

docker run -v $(pwd):/app -w /app chutter/phyloprocessr:latest Rscript workflow-5_trimming.R

For Apptainer or Singularity, convert the same image to a local .sif file:

apptainer pull phyloprocessr.sif docker://chutter/phyloprocessr:latest

Then run a workflow:

apptainer exec phyloprocessr.sif Rscript workflow-5_trimming.R

Inside either container, tools are on PATH; workflow configurations can use conda.env = "" or /opt/conda/bin/ where a tool-directory argument is required.

Option 2: Build a container from the recipes

Use this option when you need to modify the supplied environment.

Build Docker from the setup-files directory. On Apple Silicon, the linux/amd64 platform maintains compatibility with tools distributed only for x86_64 Linux.

docker build --platform linux/amd64 -t chutter/phyloprocessr:1.0.0 -t chutter/phyloprocessr:latest .

Build an Apptainer image from the supplied definition:

apptainer build phyloprocessr.sif Apptainer.def

Option 3: Create the Conda environment

After installing a Conda-compatible package manager, create the environment from setup-files/environment.yml:

conda env create -f environment.yml -n PhyloProcessR

Activate it before running a workflow:

conda activate PhyloProcessR

Install the R package

Install the stable version 1.0.0 from GitHub:

install.packages("remotes")
remotes::install_github("PhyloForge/PhyloProcessR@v1.0.0")

Install the current development version when you need changes made after the release:

remotes::install_github("PhyloForge/PhyloProcessR")

When working inside the supplied Conda environment or container, its pinned R dependencies are already present. Install a local checkout without upgrading them:

remotes::install_local(".", upgrade = "never", dependencies = FALSE)

Load the package with:

library(PhyloProcessR)

For reproducible analyses, record the installed package version or Git commit rather than reinstalling the moving development branch in every script. Supplied workflow configuration files therefore set install.latest.github = FALSE by default. Developers who intentionally want the newest beta code for debugging can set it to TRUE; the workflow will then install the current GitHub version with remotes before loading PhyloProcessR.

To check external programs in the active Conda environment:

setupCheck(anaconda.environment = Sys.getenv("CONDA_PREFIX"))

Laptop-scale reproducible example

The seven-sample laptop example provides a compact, redistribution-ready dataset for testing PhyloProcessR without a cluster. It retains both single- and multilane libraries and uses a reduced 40-locus target panel to demonstrate how individual package functions can be composed into a project-specific workflow for preprocessing, assembly, target recovery, annotation, alignment, trimming, and quality control.

The example runs in three documented stages and produces a ranked set of 20 alignments together with expected summaries and final alignments for comparison. It defaults to two CPU threads and a 4 GB memory limit. See the complete instructions and expected results and the example-data license.

Workflows

PhyloProcessR supplies reference workflows for distinct analysis stages. Configuration files and R scripts are in workflows/. You can run a complete workflow or compose the underlying functions in a custom R script.

Workflow Script Description
Workflow X0 workflow-X0_read-screening.R Quickly screen raw reads, cleaning loss, and capture success one sample at a time
Workflow 1 workflow-1_preprocess.R Organize, clean, assess, decontaminate, and merge raw reads
Workflow 2 workflow-2_assembly.R De novo assembly and target matching; optional additive per-locus recovery and contig curation
Workflow 3 workflow-3_variant-calling.R Call variants and make IUPAC or haplotype consensus contigs
Workflow 4 workflow-4_alignment.R Annotate contigs and align target markers
Workflow 5 workflow-5_trimming.R Trim alignments, concatenate genes, build unlinked dataset; optionally include novel markers from Workflow X4
Workflow X1 workflow-X1_joint-genotype_VCF.R Map samples to a common reference and make a joint VCF file
Workflow X2 workflow-X2_capture-assessment.R Assess capture efficiency for cleaned reads from workflow 1
Workflow X3 workflow-X3_legacy-integration.R Integrate Sanger/GenBank legacy alignments into the capture dataset; supports NEXUS conversion and mitochondrial loci
Workflow X4 workflow-X4_novel-loci.R Recover, assemble, and align novel shared genomic regions
Workflow X5 workflow-X5_paralog-analysis.R Assess saved copies, separate supported two-copy groups, and export filtered markers

Each workflow has a matching configuration file. Set all project parameters in that file before you run the script.

Tutorials

The tutorials are versioned in the main repository:

  1. Install PhyloProcessR
  2. Configure a project
  3. Run the workflows
  4. Assess the results
  5. Integrate legacy data

The tutorial index explains the documentation structure and simplified technical English conventions.

Citation, contributing, and license

The archived version 1.0.0 release is available at doi:10.5281/zenodo.22134845. Use the concept DOI to refer to the software across releases. Machine-readable citation metadata are provided in CITATION.cff.

Contributions are welcome; see CONTRIBUTING.md and CODE_OF_CONDUCT.md. PhyloProcessR is distributed under the GNU General Public License, version 3 or later.

About

R package for processing high-throughput sequencing data from targeted sequence capture for phylogenomic analyses

Resources

Code of conduct

Contributing

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages