Skip to content

examples: Streamlit chat over the Gemini proxy demo - #401

Open
Ar9av wants to merge 1 commit into
mainfrom
demo/gemini-streamlit
Open

Ar9av wants to merge 1 commit into
mainfrom
demo/gemini-streamlit

Conversation

@Ar9av

@Ar9av Ar9av commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

examples/gemini-proxy-demo/demo.py (#388) shows the proxy governing Google Gen AI against a stub upstream, offline and keyless. This adds the other half: a real chat UI on the real API, where the only Prismor-aware line is the base_url on the client.

client = genai.Client(
    api_key=os.environ["GEMINI_API_KEY"],
    http_options=types.HttpOptions(base_url=os.getenv("PRISMOR_PROXY", "http://127.0.0.1:7080")),
)

22 lines total, including the chat loop.

Verified against prismor 1.50.0

check result
streams through the proxy in enforce mode PONG
prompt screening reaches the wire credential pasted into chat reached Gemini as @@SECRET:...@@
overhead vs direct call 2.0-2.5s direct, 2.6-4.5s proxied (n=4)

Two things worth knowing

Port 7080 collides. It is both the proxy default and demo.py's hardcoded port. A second listener binds alongside an existing one rather than failing, and requests then land on whichever, under the wrong workspace and policy. This cost real debugging time on a box already running a long-lived proxy under a LaunchAgent. The README says to check lsof and pass --port; making demo.py's port overridable is a reasonable follow-up, left out here to keep the diff to new files.

gemini-2.5-flash is retired — the API returns 404 pointing new users at gemini-3.6-flash, which is what this uses.

demo.py (#388) shows the proxy governing Google Gen AI against a stub
upstream, offline. This adds the other half: a real chat UI on the real
API, where the only Prismor-aware line is the base_url on the client.

Verified against prismor 1.50.0: streams through the proxy in enforce
mode, and a credential pasted into the chat reaches Gemini as a
@@secret:...@@ placeholder. Overhead measured at 0.5-2s per turn over a
direct call (2.0-2.5s direct, 2.6-4.5s proxied, n=4).

The README notes the 7080 collision: the proxy default is also demo.py's
hardcoded port, and a second listener binds alongside an existing one
rather than failing, sending requests to the wrong workspace and policy.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant