Skip to content

About

A comprehensive and standard pipeline for RNA-seq data analysis, from raw reads preprocessing to downstream WGCNA.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

RNA-seq Analysis Pipeline

A professional, end-to-end bioinformatics pipeline for bulk RNA-seq data analysis. This project covers the entire workflow from raw FASTQ files to advanced systems biology analysis (WGCNA).

Description

This repository provides a standardized framework for:

  • Raw data quality control and trimming.
  • Efficient read alignment and post-processing.
  • Gene-level quantification.
  • Differential expression analysis (DEA) with high-quality visualizations.
  • Functional interpretation via GO, KEGG, and GSEA.
  • Gene co-expression network construction (WGCNA).

Repository Structure

  • scripts/: Contains all modular Bash and R scripts.
  • README.md: Documentation and usage guide.
  • LICENSE: MIT License.

Workflow Overview

1. Upstream Processing (Shell Scripts)

All upstream steps are automated in shell scripts located in /scripts.

  • Environment Setup: Configuration using Conda and SRA-Toolkit.
  • Preprocessing: Quality control with FastQC and adapter trimming with Cutadapt.
    • Script: scripts/01_preprocessing.sh
  • Alignment: Supports Hisat2 and STAR for mapping reads to a reference genome.
    • Script: scripts/02_alignment.sh
  • Quantification: Counting reads per gene using featureCounts.
    • Script: scripts/03_quantification.sh

2. Downstream Analysis (R Scripts)

Downstream statistical analysis is performed using R (v4.0+).

  • Differential Expression: Utilizing DESeq2 for normalization and identifying DEGs.
    • Script: scripts/04_deseq2_analysis.R
  • Functional Enrichment: GO/KEGG Over-representation Analysis (ORA) and Gene Set Enrichment Analysis (GSEA).
    • Script: scripts/05_enrichment.R
  • Network Analysis: WGCNA for identifying highly correlated gene modules.
    • Script: scripts/06_wgcna.R

Quick Start

Prerequisites

Ensure you have Conda installed. Create the environment:

conda create -n rnaseq_env bioconda::sra-tools bioconda::fastqc bioconda::cutadapt bioconda::hisat2 bioconda::subread
conda activate rnaseq_env

About

A comprehensive and standard pipeline for RNA-seq data analysis, from raw reads preprocessing to downstream WGCNA.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages