feat(web): import your own data, export the library, and give visitors a private library - #21
Merged
Merged
Conversation
#18 was merged after #19 and its conflicts were resolved in favour of #19 in three files, so main lost the Docker image job from CI, the changelog entries for serve --public, MCP over HTTP, and the Dockerfile, the README's link to docs/hosting.md, and M9's roadmap status. The code of both pull requests merged intact; these put the rest back as #18 had it.
A visitor's library on a public server keeps its graph in its own SQLite file whatever store the server uses, so importing into it or exporting it cannot open the library's store by the global setting. export_archive and import_archive take the store instead, and open the library's own as before when given none.
…eir own POST /library/import takes an export as the request body and imports it in the background, through the same steps as chatlore import, process, and extract, one upload at a time on a worker thread so the server keeps answering. GET /library reports each step's progress; GET /library/export downloads the library as an archive or Markdown. Without a model key everything but the knowledge graph is built, and a model that stops answering keeps what it read, as on the command line. A zip of Markdown notes is unpacked and imported as a vault. With serve --public --uploads, a visitor's first upload makes a library of their own, tied to their browser by a random token in an HttpOnly, SameSite=Lax cookie; each request is then served from it, and everyone else still sees the server's library. Its folder is named after a hash of the token and keeps its graph in SQLite, so deleting the folder deletes everything. DELETE /library deletes it at once, stopping any import first, and a sweep every ten minutes deletes those older than --keep-hours, 24 by default. Changes must carry an X-ChatLore header, which another website cannot send through a visitor's browser. Uploads are limited to 200 MB (--max-upload-mb), a zip to 2 GB of unpacked text, and --extract-limit caps how many passages of an upload the model reads. On a private server uploads go into the library itself, which cannot be deleted from the web.
A Your data button opens a dialog to drop or choose an export, shows the upload and then each step of the import with its progress, and downloads the library as an archive or Markdown. On a public server it says where the upload goes, how long it stays, and that building the knowledge graph sends it to the language model; a visitor with their own library can delete it, and the ask screen then speaks of their own conversations instead of the demo.
The hosted demo runs with --uploads. The Docker job uploads an export as a visitor, waits for the import, checks that only that visitor sees it, downloads it, and deletes it, which also proves the container's user can write visitors' libraries.
docs/web.md covers importing and downloading from the web interface; docs/hosting.md covers what --uploads does for visitors: privacy, how long their data stays, limits, and what building the graph costs. Roadmap: M11.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Milestone 11: your own data in the web interface. Your data uploads a ChatGPT, Claude, or Gemini export, Markdown notes, or a ChatLore archive. It imports, embeds, and reads the upload into the knowledge graph in the background, with progress, and downloads the library as an archive or Markdown.
With
chatlore serve --public --uploads, which the hosted demo image now uses, each visitor who uploads gets a private library of their own. It is tied to their browser and deleted after 24 hours or when they choose. Everyone else keeps seeing the demo.Changes
Imports in the background (
chatlore.imports):Importsruns an upload through the same steps aschatlore import,process, andextract, in the same order, one upload at a time on a worker thread, so the server keeps answering.doneoftotal: importing, embedding, reading, summarising, linking, topics.Visitors' own libraries (
chatlore.spaces):--keep-hours(24). Unreadable leftovers count as expired.API:
GET /library: whose library, whether uploads are taken, limits, and the latest import's status.POST /library/import: the export as the request body, its name inX-Filename. Answers 202, or 409 while the library has an import running.GET /library/export?format=archive|markdown.DELETE /library: stops any import first, then deletes. The server's own library cannot be deleted from the web.X-ChatLore: 1header, which another website cannot send through a visitor's browser without this server allowing it./healthreportsuploads.--uploads.CLI:
servegains--uploads,--keep-hours,--spaces-dir(default~/.chatlore-spaces),--max-upload-mb(200), and--extract-limit. The graph is built in full by default, as locally.Archive:
export_archiveandimport_archivetake a store that is already open.Web interface:
Docker and CI:
--uploads.Docs: a "Your data" section in
docs/web.md, and "Visitors' own data" indocs/hosting.md: privacy, retention, limits, and what building the graph costs. README and CHANGELOG updated; roadmap row M11.How it was tested
--uploads, without the header, empty, or too large;--extract-limit;serve --public --uploads, and the downloaded archive importing again.serve --public --uploadson the demo library, uploading the made-up ChatGPT export through the dialog, with the real model (DeepSeek V4 Flash):ruff,ruff format --check, strictmypy, andpytestpass locally on Python 3.14. The Docker job's new steps run for the first time in this PR's CI (no Docker here).Checklist
uv run ruff check .anduv run ruff format --check .passuv run mypypassesuv run pytestpassesCHANGELOG.mdupdated under Unreleased