feat(store): keep the graph in FalkorDB as well as SQLite - #19
Merged
Merged
Conversation
The FalkorDB store needs the falkordb client, which most people do not, so it is an extra: pip install 'chatlore[falkordb]'. mypy skips its missing type information, as for the other untyped libraries.
CHATLORE_STORE=falkordb keeps the graph in a FalkorDB server instead of the SQLite file in the library folder, with CHATLORE_FALKORDB_URL and CHATLORE_FALKORDB_GRAPH saying where. Everything above the GraphStore interface is unchanged, and the store passes the same contract tests. Nodes carry their id and label as indexed properties and their props as JSON, since FalkorDB properties cannot hold maps or nulls, with simple values copied out for find_nodes and search filters. Where FalkorDB behaves differently from SQLite, the store makes up for it, each found by running against a real server: - full-text search does not fold accents, so indexed text and queries are folded here, and snippets are built here; stemming and stop words are turned off, so a word matches only itself, as in SQLite; - a vector of the wrong length and an edge to a missing node are accepted silently, so both are checked first; - a query returns at most 10,000 rows, so long reads come in pages; - within one query the first of two writes to the same node or edge wins, so the last one given is kept, as in SQLite; - a client made from a URL does not close its sockets, so they are closed here. chatlore index --rebuild drops either store, and chatlore doctor says which is in use. With CHATLORE_STORE=falkordb every test gets a graph of its own, so the whole suite runs against FalkorDB; the few tests that assumed the SQLite file now drop the store instead.
The new job starts FalkorDB 4.20.7 as a service, runs the contract tests on both stores, then runs every test with FalkorDB as the store.
docs/falkordb.md covers setting it up, moving a library into it from the caches without computing anything again, how the graph looks for Cypher queries, and how it differs from SQLite, with timings from a library of 12,000 notes. Roadmap: M10 is done.
The web interface and the MCP server open the store for every request, and each FalkorDB store connected anew: several round trips, 3.5 seconds to a FalkorDB Cloud server in another region, before the first query. Stores now share one client per server for the whole process, closed when it exits, which brought a stats request there from 5 seconds to 1. Networks drop idle connections, which then timed out, so a connection unused for half a minute is checked before it is used. chatlore doctor printed the server URL with its password; it now shows *** instead. The docs say to keep the server close, with the timings seen against a distant one.
A search against a FalkorDB server a continent away took 10 seconds: it sent 46 queries, 40 of them one lookup per passage found by meaning, and each waited 0.16 seconds on the network. On SQLite those lookups cost nothing, so the pattern went unnoticed; the graph view did the same with about 400 queries. GraphStore gains get_nodes and neighbors_many, with a default that asks one node at a time, so any backend keeps working. FalkorDB answers each in one query, in pages. Search, the passages gathered for a question, the graph view, and the entity and topic lists use them. The same search now takes 7 queries and 1 second; the graph view takes 5 queries, and the rest of its time is the download of what it shows.
Against a FalkorDB server a continent away, every query waits for the network, so pages still paid for many round trips and for data they did not use. Each change below returns exactly what it returned before, and a contract test holds each batch method to its one-at-a-time version on both stores. - The graph view fetched every neighbour's full record, and each entity's topic with the topic's whole report, to draw links it only needed the ends of. edges_many returns just the links; the outgoing links now come from the ones already fetched, and each topic is read once. About 500 KB became about 120 KB: 28 seconds to 2.4. - Finding the entities a question names read every entity. An entity's id is made from the same normal form of its name the question is matched in, so only the ids its words could name are fetched, and each is checked against its name again. 5.7 seconds to 0.9. - Stats counted each label in its own query; count_by_label counts them in one. - The FalkorDB store read the embedding dimension on every open; it now reads it the first time a request needs it. - Importing, indexing, and archive imports wrote each conversation in about five queries; a ConversationWriter hands them to upsert_conversations 200 at a time, and writes what is waiting even when the import stops with an error. Embeddings are written a batch at a time with set_embeddings. Loading the demo went from 96 seconds to 41. - Deleting a conversation's messages read the chunks without pages, so more than 10,000 rows would have been cut off; it pages now.
The plain CI jobs do not install the falkordb extra, so mypy could not find redis, which the FalkorDB store imports; it is skipped like the other untyped libraries. With FalkorDB as the store, the per-test graph was deleted using the URL a test had just pointed elsewhere; the URL is now read before the test runs.
5 tasks done
cl0ver012
added a commit
that referenced
this pull request
Sep 30, 2026
#18 was merged after #19 and its conflicts were resolved in favour of #19 in three files, so main lost the Docker image job from CI, the changelog entries for serve --public, MCP over HTTP, and the Dockerfile, the README's link to docs/hosting.md, and M9's roadmap status. The code of both pull requests merged intact; these put the rest back as #18 had it.
cl0ver012
added a commit
that referenced
this pull request
Sep 30, 2026
…s a private library (#21) * fix: restore what the hosted demo merge dropped #18 was merged after #19 and its conflicts were resolved in favour of #19 in three files, so main lost the Docker image job from CI, the changelog entries for serve --public, MCP over HTTP, and the Dockerfile, the README's link to docs/hosting.md, and M9's roadmap status. The code of both pull requests merged intact; these put the rest back as #18 had it. * feat(archive): let archive import and export use a store already open A visitor's library on a public server keeps its graph in its own SQLite file whatever store the server uses, so importing into it or exporting it cannot open the library's store by the global setting. export_archive and import_archive take the store instead, and open the library's own as before when given none. * feat(api): upload exports, download the library, and give visitors their own POST /library/import takes an export as the request body and imports it in the background, through the same steps as chatlore import, process, and extract, one upload at a time on a worker thread so the server keeps answering. GET /library reports each step's progress; GET /library/export downloads the library as an archive or Markdown. Without a model key everything but the knowledge graph is built, and a model that stops answering keeps what it read, as on the command line. A zip of Markdown notes is unpacked and imported as a vault. With serve --public --uploads, a visitor's first upload makes a library of their own, tied to their browser by a random token in an HttpOnly, SameSite=Lax cookie; each request is then served from it, and everyone else still sees the server's library. Its folder is named after a hash of the token and keeps its graph in SQLite, so deleting the folder deletes everything. DELETE /library deletes it at once, stopping any import first, and a sweep every ten minutes deletes those older than --keep-hours, 24 by default. Changes must carry an X-ChatLore header, which another website cannot send through a visitor's browser. Uploads are limited to 200 MB (--max-upload-mb), a zip to 2 GB of unpacked text, and --extract-limit caps how many passages of an upload the model reads. On a private server uploads go into the library itself, which cannot be deleted from the web. * feat(web): add Your data to import exports and download the library A Your data button opens a dialog to drop or choose an export, shows the upload and then each step of the import with its progress, and downloads the library as an archive or Markdown. On a public server it says where the upload goes, how long it stays, and that building the knowledge graph sends it to the language model; a visitor with their own library can delete it, and the ask screen then speaks of their own conversations instead of the demo. * build: let visitors of the demo image upload their own data The hosted demo runs with --uploads. The Docker job uploads an export as a visitor, waits for the import, checks that only that visitor sees it, downloads it, and deletes it, which also proves the container's user can write visitors' libraries. * docs: describe Your data and visitors' own libraries docs/web.md covers importing and downloading from the web interface; docs/hosting.md covers what --uploads does for visitors: privacy, how long their data stays, limits, and what building the graph costs. Roadmap: M11.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Milestone 10: a FalkorDB backend. With
CHATLORE_STORE=falkordbthe graph lives in a FalkorDB server instead of the SQLite file in the library folder. Every command, the web interface, the API, and the MCP server work on it unchanged. It passes the same contract tests, and CI runs the whole test suite with it as the store.Changes
Store (
chatlore.store.falkordb.FalkorDBStore):Node, with_idand_labelas indexed properties and all props as JSON in_props, since FalkorDB properties cannot hold maps or nulls. Strings, numbers, and booleans are copied intop_<name>properties forfind_nodesand the search filters. Edges keep their ChatLore type, with props as JSON. Store-wide state sits onMetanodes.clear_embeddings, so a new dimension works after clearing, as with SQLite.Where FalkorDB differs, and what the store does about it. Each of these was found by probing a real FalkorDB 4.20.7:
find_nodes,neighbors,nodes_without_embedding) are fetched in pages, and writes go in batches.transaction()only groups calls; each write is atomic.Choosing a store (
chatlore.store):open_storereadsCHATLORE_STORE(sqlite, the default, orfalkordb), plusCHATLORE_FALKORDB_URLandCHATLORE_FALKORDB_GRAPH. It says to installchatlore[falkordb]when the client is missing, and refuses unknown values.drop_storeanddescribe_store.chatlore index --rebuildnow drops either store, andchatlore doctorshows which is in use.Speed on a distant server. Running the demo against a free FalkorDB Cloud server a continent away (0.16 s per round trip, about 20 KB/s) showed pages sending one query per item and data they did not use. SQLite's cost per query is near zero, so this had gone unnoticed.
GraphStore. Each has a default that goes one item at a time, so any backend keeps working; FalkorDB answers each in one query, in pages:get_nodes,neighbors_many, andedges_manyread many nodes, neighbours, or links at once;count_by_labelcounts every label in one call;upsert_conversationsandset_embeddingswrite many at once.ConversationWriter200 at a time, which writes what is waiting even when an import stops with an error. Embedding sync writes a batch at a time.chatlore doctorhid nothing before; it now shows the server URL with the password as***.Dependencies.
falkordb>=1.7as the optional extrafalkordb. mypy skips its missing type information.Tests.
CHATLORE_TEST_FALKORDB_URLis set and is skipped otherwise, with a fresh graph per test.CHATLORE_STORE=falkordb, an autouse fixture gives every test its own graph, so the whole suite runs against FalkorDB.chatlore.dbor expecting a second folder to be a second database. They now calldrop_store, or pin SQLite where the test is about SQLite.CI. A new "FalkorDB backend" job starts
falkordb/falkordb:v4.20.7as a service. It runs the contract tests on both stores, then every test with FalkorDB as the store.Docs.
docs/falkordb.md: setup, the settings, and moving a library in from the caches without computing anything again. It also describes the graph's shape, with a Cypher example, and the differences from SQLite, with timings.How it was tested
chatlore demoloaded 32 conversations, 143 entities, and 17 topics;entity Tidewater,topics, andexportall gave the same results as on SQLite.ruff,ruff format --check, strictmypy, andpytestpass locally on Python 3.14 (FalkorDB tests skipped there).Checklist
uv run ruff check .anduv run ruff format --check .passuv run mypypassesuv run pytestpassesCHANGELOG.mdupdated under Unreleased