A high-throughput target prioritization pipeline for Inflammatory Bowel Disease (IBD), developed as a one-day project for the YC AI X Bio Hackathon (March 8th, 2026).
This toolkit integrates genetic associations, functional genomics, and pharmacological metadata to rank potential therapeutic targets for Crohn's Disease and Ulcerative Colitis.
Target discovery in complex diseases like IBD requires reconciling disparate evidence layers. This project automates the synthesis of:
- Genetics: GWAS signals and local fine-mapping (FINEMAP/Wakefield ABF).
- Functional Genomics: eQTL colocalization (COLOC) across relevant tissues (Colon, Whole Blood).
- Biological Context: Cell-specificity and pathway relevance.
- Tractability: Druggability scores derived from Open Targets.
- Population Robustness: Stability of LD patterns across major populations (AFR, AMR, EAS, EUR, SAS).
The pipeline identified several high-confidence targets, benchmarking successfully against known positive controls:
- IL23R (rs11209026) - Score: 0.794
- Potent genetics support; established druggable cytokine pathway.
- CARD9 (rs10781499) - Score: 0.652
- Immune signaling adaptor with strong rare-variant and eQTL support.
- NOD2 (rs2066844) - Score: 0.649
- High disease relevance in Paneth cell biology; historical Crohn's benchmark.
- TYK2 (rs34536443) - Score: 0.643
- JAK-family kinase; high therapeutic tractability.
.
├── run_job.py # Main execution engine for scoring and ranking
├── deliverable/ # Analysis outputs and visualization artifacts
│ ├── ranked_targets.csv # Final prioritized gene list
│ ├── report.md # Detailed run summary
│ └── streamlit_app/ # Interactive dashboard for evidence exploration
├── scripts/ # Utility scripts for data fetching and env setup
├── variants/ # Input seed variants (e.g., ibd_top20.csv)
├── analysis_env/ # (Local) Scaffold for LD references and results
└── environment.ld-finemap.yml # Conda environment definition
The pipeline requires a specific environment for local fine-mapping and colocalization.
conda env create -f environment.ld-finemap.yml
conda activate ld-finemapExecute the main pipeline using run_job.py. It coordinates live API enrichment (Open Targets, GWAS Catalog) and local data processing.
python run_job.py --job codex_ibd_job.jsonExplore the results using the built-in Streamlit app:
cd deliverable/streamlit_app
streamlit run app.pyThe prioritization engine uses a composite scoring model with the following default weights:
- Genetics (30%): Fine-mapped PIPs and GWAS p-values.
- eQTL (25%): Colocalization posterior probabilities (PP4).
- Cell Specificity (20%): Expression enrichment in relevant cells.
- Druggability (15%): Open Targets tractability modalities.
- Population Robustness (10%): Cross-ancestry LD stability.
Scores are reconciled via a provenance-aware logic that prioritizes exact local support over proxy annotations.
Developed during the YC AI X Bio Hackathon (2026). Intended for research use.
- Data Sources: GWAS Catalog, Open Targets, GTEx, Ensembl VEP.