This repository contains complementary analysis code to the Every Cure KG publication.
DOI to the datasets published: doi.org/10.5281/zenodo.20815441
Paper: to be added
If you are interested in using EC-KG, please see our hugging face org - this is where you can find information on EC-KG nodes, edges as well as supplementary drug list, disease list and indications list.
You can load EC-KG nodes and edges with a few lines of code with help of HF Datasets library:
from datasets import load_dataset
edges_ds = load_dataset("everycure/kg-edges")
nodes_ds = load_dataset("everycure/kg-nodes")
If you want to reproduce EC-KG construction, see code constructing EC-KG, see MATRIX repo. MATRIX pipeline can be used out of the box to re-construct KG by running a series of make commands:
uv venv
uv sync
make local_test # to see if pipeline works as expected
kedro run --pipeline data_engineering # to reconstructDetailed documentation how to use MATRIX pipeline can be found at https://docs.dev.everycure.org/getting_started/
Below code (wrapped in Makefile for ease of use) can be used to reproduce the analysis used in the technical validation section of the paper. Note that the computations performed are computationally intensive and require at least 32 GB of RAM, with 4-hop SOP calculation requiring up to 512 GB for extracting paths from EC-KG.
To reproduce the analysis for Emergence of Novel Context section, run:
make emergence_novel_context
To reproduce Knowledge Source Aggregation Analysis, run:
make knowledge_source_aggregation
To reproduce the ML-based validation of computational drug repurposing use case, one needs to reproduce the KG re-construction and modelling run using MATRIX modelling pipeline.
0. Re-construct EC-KG using MATRIX pipeline using git sha 60369715eff704384189ce11a22c4a0cbc0b9f24. This step is essential to ensure all datasets are aligned with the kedro data catalog, enabling MATRIX modelling pipeline running out of the box.
git checkout 60369715eff704384189ce11a22c4a0cbc0b9f24
kedro run --pipeline data_engineering - Using MATRIX pipeline, run
feature_modellingpipeline for RTX-KG2 using git sha5654eec23ca9662e168432614a7b40499a3cfb9f:
git checkout 5654eec23ca9662e168432614a7b40499a3cfb9f
kedro run --pipeline feature_and_modelling- Using MATRIX pipeline, run
feature_modellingpipeline for PrimeKG using git sha2d93853a2703b33c75ca535a8ec2c60407ce56e0:
git checkout 2d93853a2703b33c75ca535a8ec2c60407ce56e0
kedro run --pipeline feature_and_modelling- Using MATRIX pipeline, run
feature_modellingpipeline for ROBOKOP using git sha88d1c75ef534a02b2ac1ee3a808c057de601f8d9
git checkout 88d1c75ef534a02b2ac1ee3a808c057de601f8d9
kedro run --pipeline feature_and_modelling- Using MATRIX pipeline, run
feature_modellingpipeline for EC-KG using git sha58ff055483c1fe1fb6f3ee57379729e4ac753e4d
git checkout 58ff055483c1fe1fb6f3ee57379729e4ac753e4d
kedro run --pipeline feature_and_modellingOnce the modelling pipelines run to completion, the following makefile command can be executed to reproduce analysis present in the Technical Validation section:
make ml_validation
Scripts for the paper's figures. All outputs are PDF, produced with matplotlib.
| Directory | Script | Output |
|---|---|---|
sankey/ |
sankey.py |
sankey.pdf — full EC-KG edge-flow diagram |
distribution/ |
distribution.py |
distribution_nodes.pdf, distribution_edges.pdf |
distribution/ |
upstream_distribution.py |
upstream_nodes.pdf, upstream_edges.pdf, upstream_pks.pdf — grouped bar charts |
distribution/ |
upstream_stacked_bar.py |
upstream_stacked_bar.pdf — 3-panel version of the above; one stacked bar per item (top 15), segmented by each upstream source's contribution |
distribution/ |
predicate_distribution.py |
predicate_distribution.pdf — biolink hierarchy tree |
CoreEntities_sankey/ |
drug_sankey.py |
drug_sankey_full.pdf, drug_sankey_grouped.pdf |
CoreEntities_sankey/ |
disease_sankey.py |
disease_sankey_full.pdf, disease_sankey_grouped.pdf |
UpstreamData_Venn/ |
upstream_venn.py |
upstream_venn_nodes.pdf, upstream_venn_pks.pdf, upstream_venn_node_ids.pdf, upstream_venn_edge_triples.pdf |
UpstreamData_Venn/ |
upstream_upset.py |
upstream_upset.pdf — UpSet-style replacement for the node ID, edge triple, and PKS Venn diagrams above (3 panels, bar + membership matrix) |
Shared color constants live in eckg/colors.py, shared matplotlib styling in eckg/style.py, and shared category-grouping logic in eckg/grouping.py; all figure scripts import from these instead of redefining them locally.
The eckg package needs to be installed (editable) for these imports to resolve:
pip install -e .
gcloud auth application-default loginAll data is read from mtrx-hub-dev-3of.release_v0_15_19 on BigQuery; results are cached to local CSV/JSON files after the first query, so delete a cache file to force a re-query.
Final publication figures are assembled and polished in Adobe Illustrator from these PDF outputs.
