Skip to content

feat(api): serve a public demo with question limits, MCP over HTTP, and a Docker image - #18

Merged
cl0ver012 merged 5 commits into
mainfrom
feat/m9-hosted-demo
Sep 29, 2026
Merged

cl0ver012 merged 5 commits into
mainfrom
feat/m9-hosted-demo

Conversation

@cl0ver012

Copy link
Copy Markdown
Owner

Summary

Milestone 9, first part: everything the hosted demo needs except the hosting itself.

  • chatlore serve --public serves a library to the internet safely enough for a demo, with limits on the questions paid for with the host's key.
  • chatlore serve now also serves MCP over HTTP at /mcp, which is how ChatGPT connects.
  • A Docker image runs the demo library in public mode. CI builds it and checks it.

Deploying the image is left for a later PR, once a host is chosen.

Changes

Public mode (chatlore serve --public), for a server behind a proxy on the internet:

  • Question limits. Only questions that reach the language model count against them. The defaults are 10 an hour per visitor and 300 a day for everyone; --questions-per-hour and --questions-per-day change them. When a limit is reached, chat streams an error that says why and suggests pip install chatlore. Questions that find no passages never reach the model and are not counted. The counter (_Allowance) is shared by the request threads. It keeps a sliding hour per visitor and a UTC day for everyone, and forgets idle visitors once it tracks 10,000.
  • Visitors by proxy address. uvicorn runs with proxy_headers and trusts forwarded addresses, since behind a proxy every visitor would otherwise share the proxy's address.
  • Demo note. /health reports "public": true. The web interface then shows a short note in the sidebar and changes the ask screen's line to say the conversations are made up.
  • MCP from any host name. /mcp accepts requests for any host.

A private server is unchanged, except for /mcp.

MCP over HTTP. The API app now serves the same MCP server as chatlore mcp at /mcp. It uses the SDK's streamable HTTP transport, stateless with plain JSON replies, so any worker can answer any request. The session manager runs in the app's lifespan. A private server keeps the SDK's DNS rebinding protection, so /mcp only answers requests addressed to 127.0.0.1 or localhost.

Docker image (Dockerfile, .dockerignore):

  • Built on uv's Python 3.12 image, installing the locked dependencies without dev tools.
  • Loads the demo library and downloads the embedding model while building, so a container answers at once.
  • Runs as user 1000 and serves the public demo on $PORT or 7860.
  • .dockerignore sends only pyproject.toml, uv.lock, README.md, LICENSE, and src/ to the build, so a local .env or library cannot end up in the image.
  • The model key is passed at run time, e.g. -e OPENROUTER_API_KEY=....

CI. A new Docker image job builds the image, starts it, and checks four things: /health says public, the web interface loads, /stats has the 32 demo conversations, and a search tool call over /mcp finds the Postgres to SQLite conversation.

Docs.

  • New docs/hosting.md: the image, the model key and why to cap it, what --public changes, and what a host needs.
  • docs/mcp.md gains "Over HTTP", covering local use with Claude Code and the hosted demo for ChatGPT, Claude, and Cursor.
  • README docs links and roadmap updated: M9 has its server and image; deployment is next.
  • CHANGELOG updated.

How it was tested

  • 9 new tests, 312 in total:
    • the per-visitor hour and the daily limit that resets at midnight UTC, with a fake clock;
    • a public server refusing the third question with the reason;
    • questions that find nothing not being counted;
    • a private server without limits, and /health in both modes;
    • listing and calling tools over /mcp;
    • a private server refusing an MCP request for another host (421) and answering one for 127.0.0.1;
    • serve --public passing the proxy options and limits to uvicorn.
  • Ran chatlore serve --public on the demo library:
    • the SDK's own MCP client connected over HTTP, listed the six tools, and ask_context returned the cited passages;
    • the web interface showed the demo note, and the third question gave the limit message;
    • the process used about 400 MB after searches through both the API and MCP.
  • The Docker image could not be built locally (no Docker here), so the new CI job is the image's first build.
  • ruff, ruff format --check, strict mypy, and pytest pass locally on Python 3.14.

Checklist

  • uv run ruff check . and uv run ruff format --check . pass
  • uv run mypy passes
  • uv run pytest passes
  • No real exports, databases, or keys are included
  • CHANGELOG.md updated under Unreleased

cl0ver012 and others added 5 commits September 25, 2026 14:00
chatlore serve --public is for a server on the internet behind a
proxy. Every question that reaches the language model is paid for with
the host's key, so they are limited: 10 an hour for each visitor and
300 a day for everyone, by default, with the visitor told why and how
to install ChatLore instead. Questions that find no passages never
reach the model and are not counted. Visitors are told apart by the
address the proxy forwards. /health says the server is public, and the
web interface then says it is a demo.

chatlore serve now also serves the MCP server at /mcp over streamable
HTTP, stateless and with plain JSON replies, for assistants that
connect to a URL; ChatGPT only connects that way. A private server
keeps the SDK's protection against DNS rebinding, so /mcp only answers
requests addressed to this machine; a public one accepts any host name.
The image installs the locked dependencies with uv, loads the demo
library, and downloads the embedding model while building, so a
container answers as soon as it starts. It runs as user 1000, as
Hugging Face Spaces and other hosts expect, and serves the public demo
on $PORT or 7860. Only the files the image needs are sent to the build,
so a local .env or library can never end up in it.

CI builds the image, starts it, and checks the web interface, the API,
and a tool call over MCP.
docs/hosting.md covers the image, the model key and its limits, what
--public changes, and what a host needs. The MCP guide gains a section
on /mcp, locally and on the hosted demo, for ChatGPT, Claude, and
Cursor.
The server, MCP over HTTP, and the image are done; deploying the demo
is the part of M9 left, and ChatGPT connects once it is online.
@cl0ver012
cl0ver012 merged commit ec37b03 into main Sep 29, 2026
7 checks passed
@cl0ver012
cl0ver012 deleted the feat/m9-hosted-demo branch September 29, 2026 18:08
cl0ver012 added a commit that referenced this pull request Sep 30, 2026
#18 was merged after #19 and its conflicts were resolved in favour of
#19 in three files, so main lost the Docker image job from CI, the
changelog entries for serve --public, MCP over HTTP, and the
Dockerfile, the README's link to docs/hosting.md, and M9's roadmap
status. The code of both pull requests merged intact; these put the
rest back as #18 had it.
cl0ver012 added a commit that referenced this pull request Sep 30, 2026
…s a private library (#21)

* fix: restore what the hosted demo merge dropped

#18 was merged after #19 and its conflicts were resolved in favour of
#19 in three files, so main lost the Docker image job from CI, the
changelog entries for serve --public, MCP over HTTP, and the
Dockerfile, the README's link to docs/hosting.md, and M9's roadmap
status. The code of both pull requests merged intact; these put the
rest back as #18 had it.

* feat(archive): let archive import and export use a store already open

A visitor's library on a public server keeps its graph in its own
SQLite file whatever store the server uses, so importing into it or
exporting it cannot open the library's store by the global setting.
export_archive and import_archive take the store instead, and open the
library's own as before when given none.

* feat(api): upload exports, download the library, and give visitors their own

POST /library/import takes an export as the request body and imports
it in the background, through the same steps as chatlore import,
process, and extract, one upload at a time on a worker thread so the
server keeps answering. GET /library reports each step's progress;
GET /library/export downloads the library as an archive or Markdown.
Without a model key everything but the knowledge graph is built, and a
model that stops answering keeps what it read, as on the command line.
A zip of Markdown notes is unpacked and imported as a vault.

With serve --public --uploads, a visitor's first upload makes a library
of their own, tied to their browser by a random token in an HttpOnly,
SameSite=Lax cookie; each request is then served from it, and everyone
else still sees the server's library. Its folder is named after a hash
of the token and keeps its graph in SQLite, so deleting the folder
deletes everything. DELETE /library deletes it at once, stopping any
import first, and a sweep every ten minutes deletes those older than
--keep-hours, 24 by default.

Changes must carry an X-ChatLore header, which another website cannot
send through a visitor's browser. Uploads are limited to 200 MB
(--max-upload-mb), a zip to 2 GB of unpacked text, and --extract-limit
caps how many passages of an upload the model reads. On a private
server uploads go into the library itself, which cannot be deleted
from the web.

* feat(web): add Your data to import exports and download the library

A Your data button opens a dialog to drop or choose an export, shows
the upload and then each step of the import with its progress, and
downloads the library as an archive or Markdown. On a public server it
says where the upload goes, how long it stays, and that building the
knowledge graph sends it to the language model; a visitor with their
own library can delete it, and the ask screen then speaks of their own
conversations instead of the demo.

* build: let visitors of the demo image upload their own data

The hosted demo runs with --uploads. The Docker job uploads an export
as a visitor, waits for the import, checks that only that visitor sees
it, downloads it, and deletes it, which also proves the container's
user can write visitors' libraries.

* docs: describe Your data and visitors' own libraries

docs/web.md covers importing and downloading from the web interface;
docs/hosting.md covers what --uploads does for visitors: privacy, how
long their data stays, limits, and what building the graph costs.
Roadmap: M11.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant