Summary
Add a built-in benchmarking surface to Potato OS so we can measure device, runtime, and model performance from inside the product.
This should eventually let Potato answer questions like:
- which model is actually usable on this device?
- how do
ik_llama and llama_cpp compare here?
- how much does prompt caching help on real multi-turn chat?
- what changes across Pi 4, Pi 5 8GB, Pi 5 16GB, SSD vs no SSD?
- what should Potato recommend or warn about based on measured performance?
Needs refinement
This is a parent discovery / direction-setting ticket, not a finalized implementation plan yet.
Open product and architecture questions still need to be narrowed down, including:
- should this ship as a separate built-in app, or as a platform/settings surface?
- which benchmark modes belong in v1?
- should Potato run only synthetic benchmarks, only realistic chat benchmarks, or both?
- how should results be stored, compared, and presented to the user?
- should benchmark outcomes later feed recommendations, warnings, or onboarding defaults?
Direction
The likely end state is some built-in Potato experience for benchmarking that can compare:
- devices
- runtimes
- models
- benchmark presets
But this ticket should stay generic enough to allow the implementation to land as:
- a standalone Benchmark app
- a platform tool
- or a hybrid approach
Scope
Explore and define a built-in benchmarking capability for Potato that can eventually cover:
- low-level runtime benchmarking:
- prompt processing
- token generation
- startup/load behavior
- realistic application-level benchmarking:
- multi-turn chat
- prompt caching effects
- long-context behavior
- practical usability
- comparison across:
- hardware profiles
- runtime families
- model families
- storage and deployment profiles where relevant
- result storage and comparison inside Potato
- a stable benchmark methodology Potato can reuse in spikes, release validation, and user-facing recommendations
Desired outcomes
This work should eventually make it possible to:
- benchmark a device from inside Potato
- compare one model against another on the same device
- compare
ik_llama against llama_cpp
- understand cached vs uncached behavior on multi-turn workloads
- capture measurements that are useful for release notes, docs, and future model recommendations
What this ticket is for right now
This ticket should first refine the problem into a concrete implementation plan.
Expected follow-up output from this epic:
- a clearer v1 product shape
- a benchmark methodology Potato wants to standardize on
- child issues for backend, UI/app, persistence, and validation work
- a decision on whether the first user-facing version is an app, a tool panel, or both
Acceptance criteria
- the benchmark feature area is defined clearly enough to split into concrete child tickets
- the intended v1 benchmark modes are identified
- the intended user surface is identified
- the intended result format and comparison model are identified
- related one-off benchmark spikes are linked into the plan
Non-goals
- building the entire benchmark experience in this ticket
- benchmarking every model family immediately
- creating a public leaderboard
- distributed or multi-device benchmarking in v1
Related
- Builds on
#274
- Builds on
#186
- May later inform model recommendations and runtime defaults
Summary
Add a built-in benchmarking surface to Potato OS so we can measure device, runtime, and model performance from inside the product.
This should eventually let Potato answer questions like:
ik_llamaandllama_cppcompare here?Needs refinement
This is a parent discovery / direction-setting ticket, not a finalized implementation plan yet.
Open product and architecture questions still need to be narrowed down, including:
Direction
The likely end state is some built-in Potato experience for benchmarking that can compare:
But this ticket should stay generic enough to allow the implementation to land as:
Scope
Explore and define a built-in benchmarking capability for Potato that can eventually cover:
Desired outcomes
This work should eventually make it possible to:
ik_llamaagainstllama_cppWhat this ticket is for right now
This ticket should first refine the problem into a concrete implementation plan.
Expected follow-up output from this epic:
Acceptance criteria
Non-goals
Related
#274#186