An AI coding benchmark task evaluating automated rebase conflict resolution, MLflow tracking security, secret redaction, and Hugging Face model evaluation safety.
-
Updated
Jul 25, 2026 - Python
An AI coding benchmark task evaluating automated rebase conflict resolution, MLflow tracking security, secret redaction, and Hugging Face model evaluation safety.
An automated media verification & reconciliation engine resolving conflicting SQLite attestations, Ed25519 signed release manifests, and lost Git verifier policies via Git reflog recovery.
Company-neutral workflow kit for creating, reviewing, calibrating, and packaging Harbor / Terminal-Bench tasks with LLM agents.
An offline benchmark task using Terraform, Python, and Graphviz to reconstruct corrupted machine learning model lineage metadata, query Hugging Face dataset split counts, and render color-coded lineage graphs.
Add a description, image, and links to the benchmark-tasks topic page so that developers can more easily learn about it.
To associate your repository with the benchmark-tasks topic, visit your repo's landing page and select "manage topics."