Skip to content
 
 

Repository files navigation

pi-vision-proxy

Automatic image understanding for any model in Pi.

Important

This is a hard fork for my personal usage, and its design is highly opinionated.

Attach an image (or reference an image file path) while the active model cannot accept image input, and this extension routes the image to a vision-capable model. The description is returned as a tool result and stored as a session entry, then substituted back into the conversation on later turns — so the model "sees" it, now and in later turns. The agent can also re-query any image on demand through the analyze_image tool, with an optional question passed verbatim to the vision model.

A hard fork of pungggi/pi-multimodal-proxy that keeps only the image path — video, audio, YouTube, commands, and the configuration surface are gone. The design follows nmdra/opencode-vision: a tiny model chain, retries, and a single tool.

Live widget while analyzing Batch: 5 images in one call
Live progress widget during an analysis Batch analysis of 5 images in one call

Install

pi install git:github.com/nmdra/pi-vision-proxy

Note

Requires Node 22+ and Pi with the opencode and openrouter providers configured (built-in providers; /login opencode, /login openrouter).

Architecture

The flow, model chain, retry semantics, placeholders, persistence, and widget behavior are documented in docs/ARCHITECTURE.md.

Out of scope

Video, audio, and YouTube handling from the upstream multimodal proxy are deliberately not part of this fork. If you need them, use the upstream package.

Credits

License

MIT

About

Automatic image understanding for any model in Pi.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Contributors

Languages