Summary
Benchmark Gemma 4 on the main Potato hardware profiles using a realistic multi-turn chat workload instead of a single short prompt.
The goal is to measure how Gemma 4 actually feels in Potato under longer conversations with prompt caching enabled, 16k context, and sustained generation, so we can make better runtime and product decisions for each device class.
Research question
How do the supported Gemma 4 variants perform on the main Potato hardware profiles during a realistic multi-turn chat session with prompt caching, and which model/device combinations are actually usable?
Scope
Benchmark Gemma 4 on this hardware matrix:
- Raspberry Pi 5
16GB, no SSD
- Raspberry Pi 5
8GB, with SSD
- Raspberry Pi 4
8GB
Benchmark these model paths where practical:
gemma-4-E2B-it-*
gemma-4-E4B-it-*
gemma-4-26B-A4B-it-*
- only on hardware where it is realistic to attempt
Use a more realistic benchmark methodology instead of a single synthetic prompt:
- run a fixed
4-5 message chat sequence
- keep
16k context enabled
- ensure prompt caching is enabled and actually exercised across turns
- target roughly
1k-2k generated tokens across the run
- capture both first-turn and later-turn behavior so cached vs non-cached performance is visible
For each tested model/device combination, capture:
- whether the model starts successfully
- whether the conversation completes successfully
- prompt throughput / prompt-processing behavior
- generation throughput
- whether prompt caching is working as expected
- memory pressure / RSS / swap behavior where available
- any obvious instability, restart, or degraded behavior
- a short usability judgment, not just raw numbers
Success thresholds
This spike is successful only if it leaves behind a clear recommendation per device profile.
At minimum, it must answer:
- which Gemma 4 models are realistically usable on each target device
- whether prompt caching materially improves the multi-turn experience on those devices
- whether Pi 4
8GB can support meaningful Gemma 4 use at all for E2B / E4B
- whether
26B-A4B is only viable on Pi 5 16GB without SSD
Evidence required
- exact model filename and quant for each run
- exact runtime used for each run
- hardware profile for each run
- the fixed multi-turn prompt set used
- context size and generation settings used
- total generated tokens for the run
- prompt and generation timing/results by turn where possible
- memory / RSS / swap notes
- final recommendation by device class
Test expectations (required)
- Manual:
- run the benchmark on real hardware for each target profile
- verify prompt caching is active and reused across the conversation
- record whether the run remains usable across the full chat sequence
- Automated:
- only add lightweight helpers if they materially improve benchmark repeatability or result capture
Non-goals
- changing runtime defaults in this ticket
- adding new Gemma 4 support
- full upstream Gemma 4 multimodal work
- broad Inferno refactors
- benchmarking non-Gemma model families
Related
- Related to
#265
- Related to
#268
- Related to
#270
- Follow-up may inform
#272
Summary
Benchmark Gemma 4 on the main Potato hardware profiles using a realistic multi-turn chat workload instead of a single short prompt.
The goal is to measure how Gemma 4 actually feels in Potato under longer conversations with prompt caching enabled,
16kcontext, and sustained generation, so we can make better runtime and product decisions for each device class.Research question
How do the supported Gemma 4 variants perform on the main Potato hardware profiles during a realistic multi-turn chat session with prompt caching, and which model/device combinations are actually usable?
Scope
Benchmark Gemma 4 on this hardware matrix:
16GB, no SSD8GB, with SSD8GBE2BandE4Bonly on Pi 4Benchmark these model paths where practical:
gemma-4-E2B-it-*gemma-4-E4B-it-*gemma-4-26B-A4B-it-*Use a more realistic benchmark methodology instead of a single synthetic prompt:
4-5message chat sequence16kcontext enabled1k-2kgenerated tokens across the runFor each tested model/device combination, capture:
Success thresholds
This spike is successful only if it leaves behind a clear recommendation per device profile.
At minimum, it must answer:
8GBcan support meaningful Gemma 4 use at all forE2B/E4B26B-A4Bis only viable on Pi 516GBwithout SSDEvidence required
Test expectations (required)
Non-goals
Related
#265#268#270#272