Skip to content

feat(runtime): move Gemma 4 to merged upstream ik_llama with full multimodal support #272

Description

@slomin

Summary

Replace Potato's temporary Gemma 4 runtime path based on an unmerged upstream ik_llama branch with a proper merged upstream ik_llama release or commit, and restore full Gemma 4 multimodal support once upstream support is complete.

Today Potato uses Gemma 4 on ik_llama as a speed-first compromise:

  • built from an unmerged upstream Gemma 4 branch
  • text-only on ik_llama
  • vision intentionally suppressed because gemma4v projector support is not ready in that path

This follow-up is the “do it properly” ticket: move to merged upstream ik_llama, remove the temporary compatibility masking, and support Gemma 4 vision on the ik_llama path when upstream can do it correctly.

Upstream coordination

This work depends on upstream ik_llama landing and stabilizing complete Gemma 4 support in main, including the pieces Potato currently has to work around.

Before Potato switches over fully, confirm upstream provides:

  • merged Gemma 4 support in upstream ik_llama
  • stable q8_0 KV-cache behavior for the supported Gemma 4 sizes
  • working Gemma 4 projector / mmproj handling for multimodal use
  • any required tokenizer/chat-template support for normal Gemma 4 operation

Primary upstream reference:

  • https://github.com/ikawrakow/ik_llama.cpp/pull/1581

Implementation requirement

This ticket must be developed TDD-first for the Potato-side changes.

Scope

  • replace the temporary Gemma 4 runtime dependency on the unmerged upstream branch with a merged upstream ik_llama commit or release
  • update Potato's Pi runtime build and packaging flow to use that finalized upstream source
  • restore Gemma 4 multimodal behavior on ik_llama when upstream projector support is verified
  • remove the temporary Gemma 4 text-only masking currently used on ik_llama, including:
    • runtime-env projector suppression
    • status payload capabilities.vision=false masking
    • UI-side preservation workarounds that only exist because vision is being hidden
  • verify runtime routing still prefers the correct backend for Gemma 4 on supported hardware
  • keep safe fallback behavior where needed if a device or model variant is still not supported by the finalized upstream path
  • document the finalized upstream commit or release used for Potato validation

Acceptance criteria

  • Potato no longer depends on the temporary unmerged upstream Gemma 4 branch for ik_llama
  • Gemma 4 runs on a merged upstream ik_llama commit or release validated by Potato
  • Gemma 4 vision works on ik_llama for the supported model variants once a compatible projector is installed
  • /status and model payloads report the true Gemma 4 capabilities instead of temporary masked text-only behavior
  • image upload and multimodal inference work again for supported Gemma 4 models on ik_llama
  • the finalized upstream path does not regress the q8_0 KV-cache performance and memory behavior that justified #270
  • unsupported devices such as Pi 4 still avoid bad auto-switch behavior

Test expectations (required)

  • Unit:
    • update runtime-preference and capability tests for the final Gemma 4 ik_llama behavior
    • add or update projector and runtime-env coverage for restored Gemma 4 multimodal support
  • API:
    • add or update /status, activation, and projector-related coverage for the final non-masked behavior
  • UI/e2e:
    • add Gemma 4 coverage proving image upload is available again when the finalized ik_llama path supports vision
    • verify settings persistence still behaves correctly after the temporary masking logic is removed
  • Manual:
    • real Pi QA on the finalized upstream runtime build
    • validate E2B and E4B multimodal flows on supported hardware
    • validate 26B-A4B runtime behavior on Pi 5 16GB
    • capture a basic speed and memory comparison against the temporary path from #270

Non-goals

  • carrying another long-lived custom Gemma 4 fork if upstream still is not ready
  • broad Inferno or runtime abstraction work beyond what is needed for this migration
  • adding support for new Gemma families outside the existing Potato scope
  • changing Qwen runtime behavior in this ticket

Related

  • Follow-up to #270
  • Builds on #268
  • Builds on #265
  • Related to #264
  • Upstream dependency: https://github.com/ikawrakow/ik_llama.cpp/pull/1581

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:backendAPI and backendarea:opsDeploy, service, and runtime opsblockedBlocked by dependency or external constrainttype:featureFeature work

Type

No type

Projects

Relationships

None yet

Development

No branches or pull requests

Issue actions