Skip to content

Repository files navigation

SemCEB: A Cardinality Estimation Benchmark for Semantic Operators

arXiv Python format check Python 3.10 Python 3.11 Python 3.12

SemCEB provides a benchmarking pipeline for evaluating cardinality estimation algorithms for semantic operators, specifically AI_FILTER and AI_JOIN, and for visualizing their results.

The benchmark includes 102 hand-curated queries spanning a wide range of selectivities and levels of difficulty. These queries cover single- and multi-column predicates, equality conditions, negations, spatial and temporal predicates, and more.

Each query intentionally contains exactly one semantic predicate and no relational predicates. This design isolates the core challenge of producing fast, inexpensive, and accurate cardinality estimates for semantic operators.

The dataset includes both (semi-structured) textual data and images, along with precomputed data and query embeddings.

Algorithms may consist of two phases:

  1. Setup phase: The algorithm receives access to the complete dataset and may compute any required statistics or auxiliary data structures.
  2. Estimation phase: The algorithm receives a single predicate and is expected to estimate the cardinality of the predicate’s output on the base table.

Installation

Clone the repository and install the project in editable mode from the project root.

The editable install means that local code changes under src/semceb/ are picked up immediately. This is useful when modifying the provided algorithm template in src/semceb/algorithms/custom_algorithm_template.py.

Linux
git clone --recurse-submodules https://github.com/utndatasystems/SemCEB.git
cd SemCEB

python3 -m venv .venv
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install -e .
macOS
git clone --recurse-submodules https://github.com/utndatasystems/SemCEB.git
cd SemCEB

python3 -m venv .venv
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install -e .
Windows
git clone --recurse-submodules https://github.com/utndatasystems/SemCEB.git
cd SemCEB

python -m venv .venv
.\.venv\Scripts\Activate.ps1

python -m pip install --upgrade pip
python -m pip install -e .

Verify the installation by printing the available commands:

semceb

Note:
The provided config.toml is configured to use an OpenAI model for LLM-based semantic operators. Therefore, an OpenAI API key is required when running the benchmark with the configuration included in this repository.
Create a local .env file from .env.example and configure the required OPENAI_API_KEY value there. If you configure or implement other LLM providers, add the corresponding API keys or credentials to the same .env file.

Modes

run

Runs the configured algorithms on the benchmark queries.

semceb run

Uses:

config.toml
benchmark_queries/queries.jsonl

Writes raw benchmark results to:

results/raw/result.jsonl

plot

Creates result summaries, plots, and tables from the raw benchmark results.

semceb plot

Uses:

results/raw/result.jsonl

Writes plots and summary tables to:

results/plots
results/tables

Implementing your own algorithm

The provided algorithm template is located at:

src/semceb/algorithms/custom_algorithm_template.py

This is the main file intended for users to modify. You can implement your own cardinality estimation logic there while keeping the rest of the benchmark pipeline unchanged.

After editing the algorithm file, run the benchmark with:

semceb run

Then generate plots and summary tables with:

semceb plot

No reinstall is needed after changing files under src/semceb/, as long as the project was installed in editable mode; see Installation.

Configuration

The benchmark is configured in config.toml.

Use this file to select which algorithms should run and to adjust benchmark or algorithm-specific settings. To exclude an algorithm from a run, comment out or remove its corresponding [[algorithms]] blocks.

The scale_factor setting defines how many rows are loaded from the main dataset table. Related tables are filtered to match the selected rows. Rows are shuffled deterministically before selection.

Contributions

We warmly welcome contributions. If you do not receive timely feedback on your pull request (GitHub notifications can easily get lost) please feel free to contact us by email.

We particularly encourage contributions in the following areas:

  1. New cardinality estimation algorithms for semantic operators. If you develop and benchmark a new approach using SemCEB, please consider contributing it to this repository so that others in the research community can build on your work. If you prefer to maintain your implementation separately, you can also add it as a Git submodule and provide a small integration layer.

  2. Ground-truth computations for additional models and scale factors.

Citation

If you use this work, please cite the corresponding paper:

arXiv: https://arxiv.org/abs/2606.23081

@inproceedings{zimmerer2026semceb,
  title = {{SemCEB}: A Cardinality Estimation Benchmark for Semantic Operators},
  author = {Zimmerer, Andreas and Kühn, Claudius and Li, Yang and Stoian, Mihail and Borovica-Gajic, Renata and Kipf, Andreas},
  year = {2026},
  maintitle = {Proceedings of the {VLDB} Endowment},
  booktitle = {2nd Workshop on Novel Optimizations for Visionary {AI} Systems ({NOVAS})},
  publisher = {{VLDB} Endowment},
}

About

SemCEB - A Cardinality Estimation Benchmark for Semantic Operators

Resources

Stars

22 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages