Internal lab for evaluating and improving TestBox's AI products. Leads: Lucas Wakigawa and Carlos Mattos.
| Folder | What it is | Start here |
|---|---|---|
monarch-benchmark/ |
Benchmark of Monarch against off-the-shelf models and agents (WorkflowBench) | monarch-benchmark/PLAN.md, then monarch-benchmark/docs/ai-labs-context.html |
- Every file in this repository is written in English.
- Each project keeps its own plan file (
PLAN.md) as the single tracking surface: methodology, deliverables, tasks, decisions. - Benchmark runs cost real money. CI runs only free checks and small smoke runs; full rounds are started by a person.