Skip to content

Repository files navigation

HillYeah: Coverage-Standardized Biodiversity Comparison

Important

Preliminary data analysis tool, subject to ongoing development. Results should be interpreted with appropriate caution and validated with statistical rigor.

This pipeline uses iNEXT (Hsieh, Ma & Chao) to compare biodiversity across sequencing methods while controlling for differences in sequencing depth and sampling completeness. Rather than relying on arbitrary rarefaction depths, all diversity estimates are standardized to equal sample coverage, ensuring that observed differences reflect community structure rather than sequencing effort.

Live Application

Click the image below to launch the interactive iNEXT tool:

HillYeah App

R Script Workflow Overview

As an alternative to the interactive application, the R script provides R users with a reproducible pipeline for calculating size- and coverage-based rarefaction/extrapolation analyses on their own data.

1. Data Preparation

Imports OTU/ESV tables from multiple sequencing or clustering methods. Cleans abundance matrices (removes zero-sum taxa, handles missing values). Assigns unique OTU identifiers across methods. Splits data into method × sample assemblages for downstream analysis.

2. Coverage-Standardized Diversity Estimation

Computes sample coverage for all assemblages. Identifies the minimum shared coverage across samples. Estimates Hill diversity (q = 0, 1, 2) at this common coverage using estimateD. Results are saved for reproducibility.

3. Visualization of Hill Numbers

Generates violin + boxplots of coverage-standardized diversity. Diversity is compared across methods for each Hill order. Log-scaled y-axis highlights differences across orders of magnitude.

4. Statistical Testing

Friedman tests assess overall method effects while accounting for paired samples. Paired Wilcoxon tests with FDR correction identify pairwise differences. Effect sizes are calculated for all pairwise comparisons.

5. Rarefaction and Extrapolation Analyses

A. Sample-Based Rarefaction

Converts data to presence/absence matrices. Uses incidence-based iNEXT to generate sample accumulation curves. Estimates asymptotic species richness across methods.

B. Read-Depth Rarefaction

Pools samples within each method. Performs abundance-based rarefaction/extrapolation across sequencing depth. Visualizes richness accumulation as a function of reads.

6. Output Figures

Hill diversity comparison plots Sample accumulation curves Read-depth rarefaction/extrapolation curves Combined multi-panel figures exported as publication-ready PDFs

Sample and sequence depth rarefaction

About

calculation of hill diversity numbers from metabarcoding OTU tables

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages