SemCEB provides a benchmarking pipeline for evaluating cardinality estimation algorithms for semantic operators, specifically AI_FILTER and AI_JOIN, and for visualizing their results.
The benchmark includes 102 hand-curated queries spanning a wide range of selectivities and levels of difficulty. These queries cover single- and multi-column predicates, equality conditions, negations, spatial and temporal predicates, and more.
Each query intentionally contains exactly one semantic predicate and no relational predicates. This design isolates the core challenge of producing fast, inexpensive, and accurate cardinality estimates for semantic operators.
The dataset includes both (semi-structured) textual data and images, along with precomputed data and query embeddings.
Algorithms may consist of two phases:
- Setup phase: The algorithm receives access to the complete dataset and may compute any required statistics or auxiliary data structures.
- Estimation phase: The algorithm receives a single predicate and is expected to estimate the cardinality of the predicate’s output on the base table.
Clone the repository and install the project in editable mode from the project root.
The editable install means that local code changes under src/semceb/ are picked up immediately. This is useful when modifying the provided algorithm template in src/semceb/algorithms/custom_algorithm_template.py.
Linux
git clone --recurse-submodules https://github.com/utndatasystems/SemCEB.git
cd SemCEB
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .macOS
git clone --recurse-submodules https://github.com/utndatasystems/SemCEB.git
cd SemCEB
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .Windows
git clone --recurse-submodules https://github.com/utndatasystems/SemCEB.git
cd SemCEB
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e .Verify the installation by printing the available commands:
semcebNote:
The provided config.toml is configured to use an OpenAI model for LLM-based semantic operators. Therefore, an OpenAI API key is required when running the benchmark with the configuration included in this repository.
Create a local .env file from .env.example and configure the required OPENAI_API_KEY value there. If you configure or implement other LLM providers, add the corresponding API keys or credentials to the same .env file.
Runs the configured algorithms on the benchmark queries.
semceb runUses:
config.toml
benchmark_queries/queries.jsonl
Writes raw benchmark results to:
results/raw/result.jsonl
Creates result summaries, plots, and tables from the raw benchmark results.
semceb plotUses:
results/raw/result.jsonl
Writes plots and summary tables to:
results/plots
results/tables
The provided algorithm template is located at:
src/semceb/algorithms/custom_algorithm_template.py
This is the main file intended for users to modify. You can implement your own cardinality estimation logic there while keeping the rest of the benchmark pipeline unchanged.
After editing the algorithm file, run the benchmark with:
semceb runThen generate plots and summary tables with:
semceb plotNo reinstall is needed after changing files under src/semceb/, as long as the project was installed in editable mode; see Installation.
The benchmark is configured in config.toml.
Use this file to select which algorithms should run and to adjust benchmark or algorithm-specific settings. To exclude an algorithm from a run, comment out or remove its corresponding [[algorithms]] blocks.
The scale_factor setting defines how many rows are loaded from the main dataset table. Related tables are filtered to match the selected rows. Rows are shuffled deterministically before selection.
We warmly welcome contributions. If you do not receive timely feedback on your pull request (GitHub notifications can easily get lost) please feel free to contact us by email.
We particularly encourage contributions in the following areas:
-
New cardinality estimation algorithms for semantic operators. If you develop and benchmark a new approach using SemCEB, please consider contributing it to this repository so that others in the research community can build on your work. If you prefer to maintain your implementation separately, you can also add it as a Git submodule and provide a small integration layer.
-
Ground-truth computations for additional models and scale factors.
If you use this work, please cite the corresponding paper:
arXiv: https://arxiv.org/abs/2606.23081
@inproceedings{zimmerer2026semceb,
title = {{SemCEB}: A Cardinality Estimation Benchmark for Semantic Operators},
author = {Zimmerer, Andreas and Kühn, Claudius and Li, Yang and Stoian, Mihail and Borovica-Gajic, Renata and Kipf, Andreas},
year = {2026},
maintitle = {Proceedings of the {VLDB} Endowment},
booktitle = {2nd Workshop on Novel Optimizations for Visionary {AI} Systems ({NOVAS})},
publisher = {{VLDB} Endowment},
}