Skip to content

feat(web): import your own data, export the library, and give visitors a private library - #21

Merged
cl0ver012 merged 7 commits into
mainfrom
feat/m11-your-own-data
Sep 30, 2026
Merged

cl0ver012 merged 7 commits into
mainfrom
feat/m11-your-own-data

Conversation

@cl0ver012

@cl0ver012 cl0ver012 commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Summary

Milestone 11: your own data in the web interface. Your data uploads a ChatGPT, Claude, or Gemini export, Markdown notes, or a ChatLore archive. It imports, embeds, and reads the upload into the knowledge graph in the background, with progress, and downloads the library as an archive or Markdown.

With chatlore serve --public --uploads, which the hosted demo image now uses, each visitor who uploads gets a private library of their own. It is tied to their browser and deleted after 24 hours or when they choose. Everyone else keeps seeing the demo.

Changes

Imports in the background (chatlore.imports):

  • Imports runs an upload through the same steps as chatlore import, process, and extract, in the same order, one upload at a time on a worker thread, so the server keeps answering.
  • Each step reports done of total: importing, embedding, reading, summarising, linking, topics.
  • Without a model key, everything but the graph is built, with a note. A model that stops answering keeps what it read, as on the command line.
  • An import can be cancelled between steps; what finished is kept.
  • A zip of Markdown notes is unpacked and imported as a vault.
  • A zip may unpack to at most 2 GB of text.

Visitors' own libraries (chatlore.spaces):

  • A library is made on a visitor's first upload and reached by a random token in a cookie (HttpOnly, SameSite=Lax, Secure over HTTPS).
  • Its folder is named after a SHA-256 of the token, so the server's files do not give tokens away.
  • It keeps its graph in its own SQLite file even when the server's library is on FalkorDB, so deleting the folder deletes everything.
  • Libraries expire after --keep-hours (24). Unreadable leftovers count as expired.

API:

  • A small ASGI middleware serves each request from the visitor's library when their cookie reaches one, through a context variable, so the existing endpoints are unchanged.
  • New endpoints:
    • GET /library: whose library, whether uploads are taken, limits, and the latest import's status.
    • POST /library/import: the export as the request body, its name in X-Filename. Answers 202, or 409 while the library has an import running.
    • GET /library/export?format=archive|markdown.
    • DELETE /library: stops any import first, then deletes. The server's own library cannot be deleted from the web.
  • Changes need the X-ChatLore: 1 header, which another website cannot send through a visitor's browser without this server allowing it.
  • /health reports uploads.
  • A sweep in the app's lifespan deletes expired libraries every ten minutes.
  • One embedding model is shared by requests and imports, loaded once under a lock.
  • On a private server, uploads go into the library itself; on a public one, only with --uploads.

CLI: serve gains --uploads, --keep-hours, --spaces-dir (default ~/.chatlore-spaces), --max-upload-mb (200), and --extract-limit. The graph is built in full by default, as locally.

Archive: export_archive and import_archive take a store that is already open.

Web interface:

  • A Your data button in the navigation, visible on phones too, opens a dialog with a drop zone and upload progress. It then shows each import step with a progress bar and notes, followed by Show the library.
  • Download buttons for the archive and Markdown.
  • On a public server the dialog shows:
    • a notice of where the upload goes, how long it stays, and that building the graph sends it to the language model;
    • a Delete button once the visitor has a library;
    • on the ask screen, a line about the visitor's own conversations instead of the demo.

Docker and CI:

  • The image runs with --uploads.
  • The Docker job now also uploads the Claude fixture as a visitor and waits for the import. It checks that only that visitor sees the 3 conversations while others see the demo's 32, downloads the archive, and deletes the library.

Docs: a "Your data" section in docs/web.md, and "Visitors' own data" in docs/hosting.md: privacy, retention, limits, and what building the graph costs. README and CHANGELOG updated; roadmap row M11.

How it was tested

  • 18 new tests, 341 in total, each upload test run through the real API with the import finishing in the background:
    • an export imported, embedded, and read into the graph, with the upload file not kept;
    • uploads refused on a public server without --uploads, without the header, empty, or too large;
    • an unreadable zip failing with its reason;
    • a Markdown zip and an archive imported;
    • no model key, and --extract-limit;
    • downloads as an archive and as Markdown, and an empty library;
    • each visitor seeing only their own library, and the folder not revealing the token;
    • deleting, including while an import runs, and the server's library not deletable;
    • expiry with a fake clock, and the sweep at start-up;
    • serve --public --uploads, and the downloaded archive importing again.
  • The upload tests passed three runs in a row.
  • By hand, in a browser against serve --public --uploads on the demo library, uploading the made-up ChatGPT export through the dialog, with the real model (DeepSeek V4 Flash):
    • 11 conversations, 32 passages embedded, 58 entities, and 7 topics;
    • the graph view and search worked on them, and the 81 KB archive and 6 KB Markdown downloads worked;
    • a browser without the cookie still saw the demo's 32 conversations;
    • after a reload the page showed the visitor's library;
    • Delete removed the folder from the server, cleared the cookie, and brought back the demo.
  • ruff, ruff format --check, strict mypy, and pytest pass locally on Python 3.14. The Docker job's new steps run for the first time in this PR's CI (no Docker here).

Checklist

  • uv run ruff check . and uv run ruff format --check . pass
  • uv run mypy passes
  • uv run pytest passes
  • No real exports, databases, or keys are included
  • CHANGELOG.md updated under Unreleased

#18 was merged after #19 and its conflicts were resolved in favour of
#19 in three files, so main lost the Docker image job from CI, the
changelog entries for serve --public, MCP over HTTP, and the
Dockerfile, the README's link to docs/hosting.md, and M9's roadmap
status. The code of both pull requests merged intact; these put the
rest back as #18 had it.
A visitor's library on a public server keeps its graph in its own
SQLite file whatever store the server uses, so importing into it or
exporting it cannot open the library's store by the global setting.
export_archive and import_archive take the store instead, and open the
library's own as before when given none.
…eir own

POST /library/import takes an export as the request body and imports
it in the background, through the same steps as chatlore import,
process, and extract, one upload at a time on a worker thread so the
server keeps answering. GET /library reports each step's progress;
GET /library/export downloads the library as an archive or Markdown.
Without a model key everything but the knowledge graph is built, and a
model that stops answering keeps what it read, as on the command line.
A zip of Markdown notes is unpacked and imported as a vault.

With serve --public --uploads, a visitor's first upload makes a library
of their own, tied to their browser by a random token in an HttpOnly,
SameSite=Lax cookie; each request is then served from it, and everyone
else still sees the server's library. Its folder is named after a hash
of the token and keeps its graph in SQLite, so deleting the folder
deletes everything. DELETE /library deletes it at once, stopping any
import first, and a sweep every ten minutes deletes those older than
--keep-hours, 24 by default.

Changes must carry an X-ChatLore header, which another website cannot
send through a visitor's browser. Uploads are limited to 200 MB
(--max-upload-mb), a zip to 2 GB of unpacked text, and --extract-limit
caps how many passages of an upload the model reads. On a private
server uploads go into the library itself, which cannot be deleted
from the web.
A Your data button opens a dialog to drop or choose an export, shows
the upload and then each step of the import with its progress, and
downloads the library as an archive or Markdown. On a public server it
says where the upload goes, how long it stays, and that building the
knowledge graph sends it to the language model; a visitor with their
own library can delete it, and the ask screen then speaks of their own
conversations instead of the demo.
The hosted demo runs with --uploads. The Docker job uploads an export
as a visitor, waits for the import, checks that only that visitor sees
it, downloads it, and deletes it, which also proves the container's
user can write visitors' libraries.
docs/web.md covers importing and downloading from the web interface;
docs/hosting.md covers what --uploads does for visitors: privacy, how
long their data stays, limits, and what building the graph costs.
Roadmap: M11.
Base automatically changed from fix/restore-hosted-demo-pieces to main September 30, 2026 01:19
@cl0ver012
cl0ver012 merged commit 59dcd66 into main Sep 30, 2026
8 checks passed
@cl0ver012
cl0ver012 deleted the feat/m11-your-own-data branch September 30, 2026 01:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant