Skip to content

Repository files navigation

Variant Prioritization: Multi-Evidence Target Discovery for IBD

A high-throughput target prioritization pipeline for Inflammatory Bowel Disease (IBD), developed as a one-day project for the YC AI X Bio Hackathon (March 8th, 2026).

This toolkit integrates genetic associations, functional genomics, and pharmacological metadata to rank potential therapeutic targets for Crohn's Disease and Ulcerative Colitis.


🚀 Overview

Target discovery in complex diseases like IBD requires reconciling disparate evidence layers. This project automates the synthesis of:

  • Genetics: GWAS signals and local fine-mapping (FINEMAP/Wakefield ABF).
  • Functional Genomics: eQTL colocalization (COLOC) across relevant tissues (Colon, Whole Blood).
  • Biological Context: Cell-specificity and pathway relevance.
  • Tractability: Druggability scores derived from Open Targets.
  • Population Robustness: Stability of LD patterns across major populations (AFR, AMR, EAS, EUR, SAS).

🏆 Top Candidates

The pipeline identified several high-confidence targets, benchmarking successfully against known positive controls:

  1. IL23R (rs11209026) - Score: 0.794
    • Potent genetics support; established druggable cytokine pathway.
  2. CARD9 (rs10781499) - Score: 0.652
    • Immune signaling adaptor with strong rare-variant and eQTL support.
  3. NOD2 (rs2066844) - Score: 0.649
    • High disease relevance in Paneth cell biology; historical Crohn's benchmark.
  4. TYK2 (rs34536443) - Score: 0.643
    • JAK-family kinase; high therapeutic tractability.

🛠 Project Structure

.
├── run_job.py              # Main execution engine for scoring and ranking
├── deliverable/            # Analysis outputs and visualization artifacts
│   ├── ranked_targets.csv  # Final prioritized gene list
│   ├── report.md           # Detailed run summary
│   └── streamlit_app/      # Interactive dashboard for evidence exploration
├── scripts/                # Utility scripts for data fetching and env setup
├── variants/               # Input seed variants (e.g., ibd_top20.csv)
├── analysis_env/           # (Local) Scaffold for LD references and results
└── environment.ld-finemap.yml # Conda environment definition

⚙️ Installation & Usage

1. Environment Setup

The pipeline requires a specific environment for local fine-mapping and colocalization.

conda env create -f environment.ld-finemap.yml
conda activate ld-finemap

2. Running a Prioritization Job

Execute the main pipeline using run_job.py. It coordinates live API enrichment (Open Targets, GWAS Catalog) and local data processing.

python run_job.py --job codex_ibd_job.json

3. Interactive Visualization

Explore the results using the built-in Streamlit app:

cd deliverable/streamlit_app
streamlit run app.py

🧬 Methodology

The prioritization engine uses a composite scoring model with the following default weights:

  • Genetics (30%): Fine-mapped PIPs and GWAS p-values.
  • eQTL (25%): Colocalization posterior probabilities (PP4).
  • Cell Specificity (20%): Expression enrichment in relevant cells.
  • Druggability (15%): Open Targets tractability modalities.
  • Population Robustness (10%): Cross-ancestry LD stability.

Scores are reconciled via a provenance-aware logic that prioritizes exact local support over proxy annotations.


⚖️ License & Acknowledgements

Developed during the YC AI X Bio Hackathon (2026). Intended for research use.

  • Data Sources: GWAS Catalog, Open Targets, GTEx, Ensembl VEP.

About

Repo for YC BioXAI Hackathon Project (03082026)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages