Skip to content

Repository files navigation

ConvergentDecidua

A reproducible comparative atlas for the evolution of decidualization.

Background

Decidualization may have originated as a maternal wound-healing and immune-regulatory response to trophoblast invasion, and embryos subsequently evolved to exploit and shape that response. The modern decidua is therefore a derived organ built from ancient tissue-repair circuitry, analogous in some respects to how parasites induce host structures such as insect galls.

The repeated emergence of menstruation-associated spontaneous decidualization in distantly related mammals suggests a striking case of convergent evolution in reproductive timing. In humans and other menstruating species, endometrial stromal cells acquire a decidual phenotype as part of the reproductive cycle itself, before embryo implantation is confirmed. This differs from the condition in most mammals, where decidualization is induced locally by implantation or embryo-derived cues.

This shift altered the temporal relationship between maternal tissues and embryonic invasion. Rather than responding only after implantation begins, the uterus in spontaneously decidualizing species enters a hormonally regulated anticipatory stromal state. In many menstruating mammals, that cyclic preconditioning is associated with progesterone-dependent stromal differentiation, endometrial breakdown and repair, immune remodeling, and, in several lineages, invasive implantation. These traits are related but not interchangeable, and their evolutionary relationships remain part of the question.

One influential model proposes that spontaneous decidualization evolved as a maternal adaptation to increasingly invasive trophoblast behavior. Under this view, maternal tissues gained the ability to prepare for implantation in advance, regulate trophoblast invasion, and evaluate embryo quality before extensive placental integration. Such a transition would likely require changes not only in decidual marker genes, but in the regulatory architecture controlling progesterone responsiveness, stromal differentiation, inflammatory signaling, immune interaction, and reproductive-cycle timing.

The key comparative test is not simply human versus mouse, but trait-positive lineages versus closely related trait-negative controls. If similar timing phenotypes evolved more than once, then those lineages may share cell-state and regulatory features absent from related species that retain implantation-induced decidualization.

This project therefore asks whether lineage-specific regulatory sequence changes plausibly shifted conserved stromal programs from embryo-induced activation toward more autonomous cyclic timing. Single-cell transcriptomics identifies the stromal cell states and gene programs whose deployment differs across species, while comparative genomics examines whether those expression differences track with candidate promoter, enhancer, or other noncoding changes predicted to alter transcription factor binding or chromatin activity.

Rather than focusing solely on genes associated with decidual identity, the emphasis here is on the evolution of timing itself: the transition from embryo-dependent activation to autonomous cyclical initiation. The near-term goal is to build a reproducible evidence chain from cell-state-specific expression differences to candidate regulatory mutations that can be prioritized for downstream validation.

Ultimately, understanding how spontaneous decidualization evolved may illuminate broader principles governing evolutionary changes in developmental timing, maternal-fetal conflict, reproductive immunology, and the evolution of complex endocrine-regulated cellular states.

CLI: wombat · Visualization: DecidualAtlas (Streamlit) · Current milestone: MVR 0.1

What this project does

ConvergentDecidua builds a cross-species atlas and comparative genomics workflow to investigate how spontaneous decidualization evolved in menstruating mammals. The pipeline ingests public scRNA-seq, scATAC-seq, and bulk RNA-seq datasets from human and mouse, maps orthologs, integrates stromal cell populations, and scores 8 decidualization-related gene modules. Its near-term goal is to connect cell-state-specific expression differences to candidate lineage-specific regulatory sequence changes associated with shifts in decidual timing.

For the full scientific background — hypothesis, species rationale, dataset targets, AI model strategy (DeciduaAI), and long-term roadmap — see BACKGROUND.md.

Current status (MVR 0.1)

The pipeline has been executed end-to-end on real data:

Step Result
Ortholog backbone 25,439 human↔mouse pairs (16,168 Tier 1 + 9,271 Tier 2) via Ensembl Compara
Datasets fetched GSE127918 (human scRNA), GSE111976 (human scRNA), GSE226429 (mouse bulk)
QC GSE127918 → 9,292 cells; GSE111976 → 1,578 cells; GSE226429 → 6 samples
Integration 9,065 human stromal cells via Harmony with UMAP embedding
Cell types 4 subtypes: 8,466 fibroblast, 394 decidual, 129 pre-decidual, 76 senescent
Scoring 8 modules scored; decidual_score highest in decidual_stromal (0.74)
Reports Methods, coverage, QC, ortholog, scoring (heatmap + violin plots), manifest

See PLAN.md for detailed status, known gaps, and the implementation plan.

Known gaps

  • No mouse scRNA integrated yet — GSE226417 requires R/Seurat (RData format), E-MTAB-11491 has 644 individual files
  • No scATAC data — GSE183771 is a 7GB+ tar archive, deferred
  • Integration is human-only — blocked by the mouse data gap above

Repository structure

wombat/              CLI and orchestration (Click commands, config loader)
src/                 Analysis modules
  ingest/              GEO/ArrayExpress download → AnnData
  metadata/            .obs harmonization
  qc/                  scRNA, scATAC, bulk QC pipelines
  orthologs/           Ensembl BioMart / Compara ortholog mapping
  cell_states/         Marker-based annotation + Harmony/scVI integration
  scoring/             8 decidualization modules via scanpy.tl.score_genes
  reports/             Methods, coverage, QC, manifest generation
configs/             YAML configs (datasets, species, markers)
decidual_atlas/      Streamlit visualization app
workflows/           Snakemake rules
tests/               pytest suite
results/             Pipeline outputs (.gitignored)
  orthologs/           backbone.parquet, orthogroups
  raw/                 Downloaded source files
  processed/           Standardized h5ad per dataset
  qc/                  QC-filtered h5ad
  integrated/          Harmony/scVI joint embeddings
  scored/              Decidualization-scored h5ad
  reports/             Generated reports and figures

Installation

# Clone
git clone https://github.com/BioNanomics/ConvergentDecidua.git
cd ConvergentDecidua

# Create virtual environment
python -m venv .venv
source .venv/bin/activate

# Install with all optional dependencies
pip install -e ".[all]"

# Or install specific groups
pip install -e ".[dev]"       # pytest, ruff, pre-commit
pip install -e ".[atlas]"     # streamlit, plotly
pip install -e ".[ingest]"    # GEOparse
pip install -e ".[workflow]"  # snakemake

Requires Python ≥3.9, <3.13. Additional runtime dependencies: tabulate, harmonypy.

Wombat CLI

wombat is the command-line interface that drives the entire pipeline. Each step can be run independently or chained via Snakemake.

wombat [OPTIONS] COMMAND [ARGS]

Options:
  -v, --verbose    Increase verbosity (-v, -vv)

Commands:
  init              Validate that all required configs exist
  validate-config   Validate all YAML configuration files
  build-registry    Export dataset registry to Parquet and CSV
  fetch             Download datasets and convert to standardized AnnData
  qc                Run quality control on processed datasets
  orthologs build   Build ortholog backbone and orthogroup tables
  integrate         Integrate datasets across species
  score-decidua     Compute decidualization scores across all 8 modules
  generate-reports  Generate all pipeline reports and release manifest
  serve-atlas       Launch the DecidualAtlas Streamlit app

Typical workflow

# 1. Validate configuration
wombat validate-config

# 2. Build ortholog backbone (human↔mouse)
wombat orthologs build

# 3. Fetch and QC datasets
wombat fetch --dataset GSE127918
wombat fetch --all-datasets
wombat qc --species human
wombat qc --species mouse

# 4. Integrate stromal cells
wombat integrate --mode stromal --method harmony

# 5. Score decidualization modules
wombat score-decidua

# 6. Generate reports
wombat generate-reports

# 7. Launch interactive viewer
wombat serve-atlas --port 8501

Decidualization scoring modules

The pipeline scores each cell on 8 gene-set modules defined in configs/markers.yaml:

Module What it captures
decidual_score Core decidualization signature (PRL, IGFBP1, FOXO1, etc.)
progesterone_response_score Progesterone receptor pathway activity
estrogen_response_score Estrogen receptor pathway activity
stress_response_score Oxidative/cellular stress response
senescence_score Cellular senescence markers
immune_interface_score Stromal–immune crosstalk genes
ECM_remodeling_score Extracellular matrix remodeling
angiogenesis_score Angiogenesis and vascular remodeling

Configuration

All pipeline parameters are defined in YAML files under configs/:

  • datasets.yaml — Dataset registry (accession, species, assay, priority)
  • species.yaml — Species metadata (genome build, Ensembl release, decidualization mode)
  • markers.yaml — Cell-type markers and decidualization gene sets

Config loading uses wombat.config.load_config(name) — never hardcode paths.

How we got here

The project was built in two phases:

Phase 1: Code scaffolding (Phases 1–10)

Starting from an empty repo, 10 implementation phases built the full pipeline code:

  1. Project skeleton — pyproject.toml, Click CLI, CI, pre-commit
  2. Configuration — YAML schemas for datasets, species, markers
  3. Data ingestion — GEO + ArrayExpress downloaders → standardized AnnData
  4. Metadata harmonization — Consistent .obs columns across all datasets
  5. QC pipelines — scRNA (doublets, mito%, HVG), scATAC (TSS, LSI), bulk
  6. Ortholog mapping — BioMart + Compara FTP with cross-validation
  7. Cell-state integration — Marker scoring, Harmony/scVI joint embedding
  8. Decidualization scoring — 8 gene-set modules via scanpy
  9. DecidualAtlas — Streamlit app with UMAP, gene explorer, species comparison
  10. Reports — Methods, coverage, QC, ortholog, manifest with checksums

Phase 2: Execution against real data (E1–E5)

When the code hit real APIs and data, several fixes were needed:

  • BioMart mirrors down → added Ensembl Compara FTP fallback for ortholog mapping
  • GEO URL prefix bug → fixed accession slicing for 6+ digit accessions
  • DGE format detection → broadened file pattern matching for tab-separated count matrices
  • HVG numpy incompatibility → added cell_ranger fallback when seurat_v3 fails
  • Sparse matrix h5ad write → convert to dense for small bulk datasets
  • Harmony shape mismatch → auto-detect Z_corr transposition
  • Python 3.9 compat → datetime.UTC → timezone.utc, HVG flavor fallback

Each fix was committed individually and the pipeline re-run to validate.

Development

# Lint
ruff check .

# Test
pytest

# Validate configs
wombat validate-config

CI runs lint, test, and config validation on every push via GitHub Actions.

License

MIT

References

Releases

Packages

Contributors

Languages