Skip to content

Repository files navigation

Tablassert

PyPI version CI Docs License

Extract knowledge assertions from tabular data into NCATS Translator-compliant KGX NDJSON, declaratively, with entity resolution built in and optional quality control.

Tablassert turns biomedical spreadsheets (Excel, CSV, TSV) into knowledge graphs ready for NCATS Translator. Declare how your columns map to subject-predicate-object statements in YAML; Tablassert resolves free text to standard CURIEs, attaches provenance and statistical annotations, and emits KGX-compliant nodes and edges.

Full Documentation: installation guides, tutorial, configuration reference, and API docs.

Statement of need

Biomedical knowledge lives in spreadsheets: association tables, assay results, curated gene-disease lists. Getting those rows into an NCATS Translator knowledge graph means writing an ingest that maps columns to Biolink statements, resolves free text ("TP53", "lung cancer") to standard CURIEs, and records provenance, terms of use, and statistical annotations. Today that ingest is Python code maintained per source, and entity resolution usually means calling a hosted name-resolution service at build time.

Tablassert replaces the per-source code with a declarative YAML mapping plus an embedded, offline entity-resolution database built from RENCI BABEL exports, so a build is reproducible, auditable, and network-free. It is aimed at Translator ingest authors and knowledge-graph data engineers who hold tabular biomedical sources, and at bioinformatics groups that want a validated KGX output without writing a transform pipeline. The configuration model and the Biolink/KGX contracts it emits follow current NCATSTranslator/translator-ingests practice, and each build can emit the Resource Ingest Guide (RIG) metadata those submissions require.

Getting help

Quick Start

pip install "tablassert[cli]"

Given a CSV of gene-disease associations with p-values and sample sizes, declare the mapping in a table config (table.yaml):

template:
  source:
    kind: text
    local: ./gene-disease.csv
    url: [https://example.com/data.csv]
    row_slice: [1, auto]
    delimiter: ","
  statement:
    subject: { method: column, encoding: A, prioritize: [Gene] }
    predicate: associated_with
    object: { method: column, encoding: B, prioritize: [Disease] }
  provenance: { repo: PMID, publication: "12345678" }
  annotations:
    - { annotation: p_value, method: column, encoding: C }
    - { annotation: study_size, method: column, encoding: D }

Wrap it in a graph config (graph.yaml) pointing at your fullmap entity-resolution database and carrying the required rig: metadata for the generated Resource Ingest Guide. Build or download that database once first (tablassert build-fullmap, a multi-GB download; see the Fullmap guide), which makes ./fullmap below the right path:

name: MY_KG
version: 1.0.0
tables:
  - ./table.yaml
fullmap: ./fullmap
rig:
  source_info:
    infores_id: infores:my-kg
    terms_of_use_info:
      terms_of_use_url: https://example.org/terms
    data_access_locations:
      - My source downloads - https://example.org/downloads
    source_status: maintained_regular_updates
  ingest_info:
    utility: Gene-disease associations support Translator disease-mechanism queries.
    scope: Gene-disease associations extracted from tabular sources.
  provenance_info:
    contributions:
      - "Author Name - code author, data modeling"
  artifact_base_url: https://example.org/my-kg
  artifact_base_path: ./published/my-kg

Build the knowledge graph:

tablassert build-kg graph.yaml

Output is one JSON object per line: nodes with Biolink categories, edges with annotations.

{"id":"HGNC:11998","name":"TP53","category":["biolink:Gene"],"taxon":"NCBITaxon:9606"}
{"id":"MONDO:0008903","name":"lung cancer","category":["biolink:Disease"]}
{"id":"2cfea591-0f8f-33af-a7df-03da531d3359","subject":"HGNC:11998","predicate":"biolink:associated_with","object":"MONDO:0008903","p_value":"1.0000e-03","statistical_significance_qualifier":"strongly_significant","has_supporting_studies":{"PMID:12345678":{"id":"PMID:12345678","name":"gene-disease.csv","study_size":450,"has_study_results":[{"id":"row:2"}]}},"publications":["PMID:12345678"]}

See the Tutorial for the full walkthrough.

Key Features

  • Declarative YAML configuration: define data transformations without writing code
  • Built-in entity resolution: map free text to genes, diseases, and chemicals with standard CURIEs, taxonomic filtering, and provenance, backed by an embedded redb database
  • Optional quality control: a four-stage audit (exact -> fuzzy -> abbreviation -> SapBERT embeddings) flags low-confidence mappings
  • KGX compliance: emits NCATS Translator-compatible node/edge NDJSON with Biolink categories and predicates
  • Autonomous agent (experimental): tablassert agent derives, builds, and refines configs for whole papers
  • Performance & reproducibility: lazy Polars pipelines and a deterministic UV-based development environment

Installation

pip install tablassert

Or with uv: uv tool install "tablassert[cli]". The base install provides the Python API; [cli] provides the tablassert command. Other extras are opt-in:

Extra Adds Install
cli tablassert command and rich terminal progress pip install "tablassert[cli]"
rt CPU-compatible Polars runtime pip install "tablassert[rt]"
aria2 bundled aria2c downloader, used automatically by build-fullmap when installed (Linux/Windows wheels only) pip install "tablassert[aria2]"
qc four-stage QC audit (exact -> fuzzy -> abbreviation -> SapBERT embeddings) pip install "tablassert[qc]"
agent autonomous agent (smolagents, litellm, article/table context) -- experimental, API may change pip install "tablassert[agent]"
optimize GEPA prompt optimization for agent --optimize (dspy) -- experimental, API may change pip install "tablassert[optimize]"
distill distillation dataset export (tablassert distill-export, HF datasets) -- experimental, API may change pip install "tablassert[distill]"
log loguru-backed file/progress logging (rotation, enqueue) pip install "tablassert[log]"

Reaching a feature whose extra is missing never produces a bare ModuleNotFoundError: the failure names the missing package and the install command that fixes it. QC is opt-in at build time (build-kg --qc). See the Installation guide for the full matrix, per-command preflight behavior, and the CLI Reference for every flag.

Entity Resolution API

from pathlib import Path
from tablassert.lib import resolve_many

results = resolve_many(
    col="gene",
    entities=["TP53", "BRCA1"],
    fullmap=Path("/path/to/fullmap"),
    taxon="9606",
)
# [{"original_gene": "TP53", "gene": "HGNC:11998", "gene_name": "TP53", ...}, ...]

Point resolve_many() at a fullmap database to resolve any iterable of entity strings to CURIEs, no LazyFrame setup or NLP preprocessing required. See the Batch Resolution API for the full reference.

Documentation

Developing

uv sync --group dev --extra cli --extra qc --extra log
uv run maturin develop --manifest-path rust/Cargo.toml
make check

See CONTRIBUTING.md for the full development loop, quality gates, and pull request guidelines.

Citation

If you use Tablassert, please cite it as described in CITATION.cff. The approach is described in:

Skye Lane Goetz, Amy K. Glen, and Gwênlyn Glusman. "MicrobiomeKG: bridging microbiome research and host health through knowledge graphs." Frontiers in Systems Biology 5 (2025). doi:10.3389/fsysb.2025.1544432

License

Apache License 2.0

Contributors

About

Tablassert turns biomedical spreadsheets (Excel, CSV, TSV) into knowledge graphs ready for NCATS Translator. Declare how your columns map to subject–predicate–object statements in YAML; Tablassert resolves free text to standard CURIEs, attaches provenance and statistical annotations, and emits KGX-compliant nodes and edges.

Topics

Resources

Code of conduct

Contributing

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages