Skip to content

spike(runtime): benchmark Gemma 4 on supported Pi hardware with multi-turn cached-chat workloads #274

Description

@slomin

Summary

Benchmark Gemma 4 on the main Potato hardware profiles using a realistic multi-turn chat workload instead of a single short prompt.

The goal is to measure how Gemma 4 actually feels in Potato under longer conversations with prompt caching enabled, 16k context, and sustained generation, so we can make better runtime and product decisions for each device class.

Research question

How do the supported Gemma 4 variants perform on the main Potato hardware profiles during a realistic multi-turn chat session with prompt caching, and which model/device combinations are actually usable?

Scope

Benchmark Gemma 4 on this hardware matrix:

  • Raspberry Pi 5 16GB, no SSD
  • Raspberry Pi 5 8GB, with SSD
  • Raspberry Pi 4 8GB
    • E2B and E4B only on Pi 4

Benchmark these model paths where practical:

  • gemma-4-E2B-it-*
  • gemma-4-E4B-it-*
  • gemma-4-26B-A4B-it-*
    • only on hardware where it is realistic to attempt

Use a more realistic benchmark methodology instead of a single synthetic prompt:

  • run a fixed 4-5 message chat sequence
  • keep 16k context enabled
  • ensure prompt caching is enabled and actually exercised across turns
  • target roughly 1k-2k generated tokens across the run
  • capture both first-turn and later-turn behavior so cached vs non-cached performance is visible

For each tested model/device combination, capture:

  • whether the model starts successfully
  • whether the conversation completes successfully
  • prompt throughput / prompt-processing behavior
  • generation throughput
  • whether prompt caching is working as expected
  • memory pressure / RSS / swap behavior where available
  • any obvious instability, restart, or degraded behavior
  • a short usability judgment, not just raw numbers

Success thresholds

This spike is successful only if it leaves behind a clear recommendation per device profile.

At minimum, it must answer:

  • which Gemma 4 models are realistically usable on each target device
  • whether prompt caching materially improves the multi-turn experience on those devices
  • whether Pi 4 8GB can support meaningful Gemma 4 use at all for E2B / E4B
  • whether 26B-A4B is only viable on Pi 5 16GB without SSD

Evidence required

  • exact model filename and quant for each run
  • exact runtime used for each run
  • hardware profile for each run
  • the fixed multi-turn prompt set used
  • context size and generation settings used
  • total generated tokens for the run
  • prompt and generation timing/results by turn where possible
  • memory / RSS / swap notes
  • final recommendation by device class

Test expectations (required)

  • Manual:
    • run the benchmark on real hardware for each target profile
    • verify prompt caching is active and reused across the conversation
    • record whether the run remains usable across the full chat sequence
  • Automated:
    • only add lightweight helpers if they materially improve benchmark repeatability or result capture

Non-goals

  • changing runtime defaults in this ticket
  • adding new Gemma 4 support
  • full upstream Gemma 4 multimodal work
  • broad Inferno refactors
  • benchmarking non-Gemma model families

Related

  • Related to #265
  • Related to #268
  • Related to #270
  • Follow-up may inform #272

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:backendAPI and backendarea:opsDeploy, service, and runtime opstype:spikeExploratory spike / evaluation work

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions