A professional, end-to-end bioinformatics pipeline for bulk RNA-seq data analysis. This project covers the entire workflow from raw FASTQ files to advanced systems biology analysis (WGCNA).
This repository provides a standardized framework for:
- Raw data quality control and trimming.
- Efficient read alignment and post-processing.
- Gene-level quantification.
- Differential expression analysis (DEA) with high-quality visualizations.
- Functional interpretation via GO, KEGG, and GSEA.
- Gene co-expression network construction (WGCNA).
scripts/: Contains all modular Bash and R scripts.README.md: Documentation and usage guide.LICENSE: MIT License.
All upstream steps are automated in shell scripts located in /scripts.
- Environment Setup: Configuration using Conda and SRA-Toolkit.
- Preprocessing: Quality control with FastQC and adapter trimming with Cutadapt.
- Script:
scripts/01_preprocessing.sh
- Script:
- Alignment: Supports Hisat2 and STAR for mapping reads to a reference genome.
- Script:
scripts/02_alignment.sh
- Script:
- Quantification: Counting reads per gene using featureCounts.
- Script:
scripts/03_quantification.sh
- Script:
Downstream statistical analysis is performed using R (v4.0+).
- Differential Expression: Utilizing
DESeq2for normalization and identifying DEGs.- Script:
scripts/04_deseq2_analysis.R
- Script:
- Functional Enrichment: GO/KEGG Over-representation Analysis (ORA) and Gene Set Enrichment Analysis (GSEA).
- Script:
scripts/05_enrichment.R
- Script:
- Network Analysis: WGCNA for identifying highly correlated gene modules.
- Script:
scripts/06_wgcna.R
- Script:
Ensure you have Conda installed. Create the environment:
conda create -n rnaseq_env bioconda::sra-tools bioconda::fastqc bioconda::cutadapt bioconda::hisat2 bioconda::subread
conda activate rnaseq_env