Skip to content

epic(platform): add built-in benchmarking and performance measurement to Potato OS #276

Description

@slomin

Summary

Add a built-in benchmarking surface to Potato OS so we can measure device, runtime, and model performance from inside the product.

This should eventually let Potato answer questions like:

  • which model is actually usable on this device?
  • how do ik_llama and llama_cpp compare here?
  • how much does prompt caching help on real multi-turn chat?
  • what changes across Pi 4, Pi 5 8GB, Pi 5 16GB, SSD vs no SSD?
  • what should Potato recommend or warn about based on measured performance?

Needs refinement

This is a parent discovery / direction-setting ticket, not a finalized implementation plan yet.

Open product and architecture questions still need to be narrowed down, including:

  • should this ship as a separate built-in app, or as a platform/settings surface?
  • which benchmark modes belong in v1?
  • should Potato run only synthetic benchmarks, only realistic chat benchmarks, or both?
  • how should results be stored, compared, and presented to the user?
  • should benchmark outcomes later feed recommendations, warnings, or onboarding defaults?

Direction

The likely end state is some built-in Potato experience for benchmarking that can compare:

  • devices
  • runtimes
  • models
  • benchmark presets

But this ticket should stay generic enough to allow the implementation to land as:

  • a standalone Benchmark app
  • a platform tool
  • or a hybrid approach

Scope

Explore and define a built-in benchmarking capability for Potato that can eventually cover:

  • low-level runtime benchmarking:
    • prompt processing
    • token generation
    • startup/load behavior
  • realistic application-level benchmarking:
    • multi-turn chat
    • prompt caching effects
    • long-context behavior
    • practical usability
  • comparison across:
    • hardware profiles
    • runtime families
    • model families
    • storage and deployment profiles where relevant
  • result storage and comparison inside Potato
  • a stable benchmark methodology Potato can reuse in spikes, release validation, and user-facing recommendations

Desired outcomes

This work should eventually make it possible to:

  • benchmark a device from inside Potato
  • compare one model against another on the same device
  • compare ik_llama against llama_cpp
  • understand cached vs uncached behavior on multi-turn workloads
  • capture measurements that are useful for release notes, docs, and future model recommendations

What this ticket is for right now

This ticket should first refine the problem into a concrete implementation plan.

Expected follow-up output from this epic:

  • a clearer v1 product shape
  • a benchmark methodology Potato wants to standardize on
  • child issues for backend, UI/app, persistence, and validation work
  • a decision on whether the first user-facing version is an app, a tool panel, or both

Acceptance criteria

  • the benchmark feature area is defined clearly enough to split into concrete child tickets
  • the intended v1 benchmark modes are identified
  • the intended user surface is identified
  • the intended result format and comparison model are identified
  • related one-off benchmark spikes are linked into the plan

Non-goals

  • building the entire benchmark experience in this ticket
  • benchmarking every model family immediately
  • creating a public leaderboard
  • distributed or multi-device benchmarking in v1

Related

  • Builds on #274
  • Builds on #186
  • May later inform model recommendations and runtime defaults

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:backendAPI and backendarea:opsDeploy, service, and runtime opsarea:uiFrontend and UXtype:epicParent issue grouping a larger body of work

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions