Automatic image understanding for any model in Pi.
Important
This is a hard fork for my personal usage, and its design is highly opinionated.
Attach an image (or reference an image file path) while the active model cannot
accept image input, and this extension routes the image to a vision-capable
model. The description is returned as a tool result and stored as a session
entry, then substituted back into the conversation on later turns — so the
model "sees" it, now and in later turns. The agent can also re-query any image
on demand through the analyze_image tool, with an optional question passed
verbatim to the vision model.
A hard fork of pungggi/pi-multimodal-proxy
that keeps only the image path — video, audio, YouTube, commands, and the
configuration surface are gone. The design follows
nmdra/opencode-vision: a tiny
model chain, retries, and a single tool.
| Live widget while analyzing | Batch: 5 images in one call |
|---|---|
![]() |
![]() |
pi install git:github.com/nmdra/pi-vision-proxyNote
Requires Node 22+ and Pi with the opencode and openrouter providers configured (built-in providers; /login opencode, /login openrouter).
The flow, model chain, retry semantics, placeholders, persistence, and widget
behavior are documented in docs/ARCHITECTURE.md.
Video, audio, and YouTube handling from the upstream multimodal proxy are deliberately not part of this fork. If you need them, use the upstream package.
pungggi/pi-multimodal-proxy(MIT) — upstream: image/video/audio proxy; this fork keeps the image pipeline and its security hardening.nmdra/opencode-vision(MIT) — reference pattern: model chain, retry semantics, and prompt design.
MIT

