Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -68,3 +68,34 @@ jobs:
uvx --from "$wheel" chatlore --version
uvx --from "$wheel" chatlore demo --no-serve
uvx --from "$wheel" chatlore --home ~/.chatlore-demo search "litestream"

falkordb:
name: FalkorDB backend
runs-on: ubuntu-latest
services:
falkordb:
image: falkordb/falkordb:v4.20.7
ports: ["6379:6379"]
env:
CHATLORE_FALKORDB_URL: redis://localhost:6379
CHATLORE_TEST_FALKORDB_URL: redis://localhost:6379
steps:
- name: Check out
uses: actions/checkout@v4

- name: Set up uv
uses: astral-sh/setup-uv@v6
with:
python-version: "3.12"
enable-cache: true

- name: Install dependencies
run: uv sync --locked --extra falkordb

- name: Contract tests on both stores
run: uv run pytest tests/store

- name: Every test with FalkorDB as the store
run: uv run pytest
env:
CHATLORE_STORE: falkordb
20 changes: 20 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,26 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Added

- A FalkorDB store, chosen with `CHATLORE_STORE=falkordb`, with `CHATLORE_FALKORDB_URL` and
`CHATLORE_FALKORDB_GRAPH` saying where. Every command, the web interface, and the MCP server
work on it as on SQLite, and it passes the same contract tests. The client is an optional
dependency: `pip install 'chatlore[falkordb]'`. Guide in `docs/falkordb.md`.
- `chatlore doctor` shows which store is in use, and `chatlore index --rebuild` rebuilds either.
- Batch methods on `GraphStore`, each with a default that goes one item at a time:
`get_nodes`, `neighbors_many`, and `edges_many` read many nodes, neighbours, or links at once;
`count_by_label` counts every label in one call; `upsert_conversations` and `set_embeddings`
write many at once. Search, the passages gathered for a question, the graph view, stats,
entity and topic lists, importing, indexing, and embedding use them, so a store on a server
answers in a few queries instead of one per item. Against a distant FalkorDB server a search
went from 46 queries and 10.6 seconds to 7 queries and 1 second, the graph view from about 400
queries and 105 seconds to 4 queries and 2.4 seconds, and loading the demo from 96 to 41
seconds, with the same results.
- Finding the entities a question names looks up only the ones its words could name, by id,
instead of reading every entity.
- CI runs the contract tests on both stores, and every test with FalkorDB as the store.

## [0.1.0] - 2026-09-24

### Added
Expand Down
5 changes: 3 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,8 @@ see [docs/models.md](docs/models.md). The knowledge graph is described in
[docs/extraction.md](docs/extraction.md), chat and the API in
[docs/chat.md](docs/chat.md), the web interface in [docs/web.md](docs/web.md),
setting up assistants over MCP in [docs/mcp.md](docs/mcp.md), and archives and
Markdown export in [docs/export.md](docs/export.md).
Markdown export in [docs/export.md](docs/export.md), and keeping the graph in
FalkorDB instead of SQLite in [docs/falkordb.md](docs/falkordb.md).

## What ChatLore will do

Expand Down Expand Up @@ -95,7 +96,7 @@ extraction is an optional enrichment you can re-run with a better model later.
| M7 | MCP server for Claude Desktop, Claude Code, Cursor, ChatGPT | done for local assistants; ChatGPT with the hosted demo |
| M8 | Easy to try: PyPI package, demo library, export and import | done, v0.1.0 |
| M9 | Hosted demo | planned |
| M10 | FalkorDB backend | planned |
| M10 | FalkorDB backend | done |

## Development setup

Expand Down
84 changes: 84 additions & 0 deletions docs/falkordb.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
# Keeping the graph in FalkorDB

```bash
pip install 'chatlore[falkordb]'
docker run -d -p 6379:6379 --name falkordb falkordb/falkordb
export CHATLORE_STORE=falkordb
chatlore index # copy the library into FalkorDB
chatlore process # chunks and embeddings
chatlore extract # entities and topics, from the cached answers where possible
```

By default ChatLore keeps its graph in a SQLite file in the library folder,
which needs nothing else. [FalkorDB](https://www.falkordb.com) is a graph
database that runs as a server. Use it when you want to query the graph with
Cypher, look at it with FalkorDB's browser, or share one graph between several
machines on a network. Every command, the web interface, and the MCP server
work the same on both.

## Settings

| Variable | Meaning |
|---|---|
| `CHATLORE_STORE` | `sqlite` (default) or `falkordb`. |
| `CHATLORE_FALKORDB_URL` | The server, `redis://localhost:6379` by default. Add a password as `redis://:password@host:6379`. |
| `CHATLORE_FALKORDB_GRAPH` | The graph's name, `chatlore` by default. Give each library its own. |

`chatlore doctor` shows which store is in use.

## Moving a library into FalkorDB

The conversations stay in the library folder either way, and the caches of
embeddings and model answers too, so the graph can be rebuilt in FalkorDB
without computing anything again:

1. Set the variables above.
2. `chatlore index` copies every conversation into the graph.
3. `chatlore process` makes the chunks and takes their embeddings from the cache.
4. `chatlore extract` rebuilds entities, relationships, and topics from the
cached model answers. It needs a model key, but only asks the model about
text it has not read yet.

Going back to SQLite is the same with `CHATLORE_STORE` unset. An archive from
`chatlore export` imports into either store. `chatlore index --rebuild` deletes
the graph and builds it again from the library.

## What is in the graph

Every node has the label `Node`, with its id in `_id`, its ChatLore label
(`Conversation`, `Message`, `Chunk`, `Entity`, `Topic`) in `_label`, and all
of its properties as JSON in `_props`. Simple values are copied into `p_`
properties as well, so they can be queried: an entity's name is `p_name`, a
topic's title is `p_title`. Edges have their ChatLore type, such as `MENTIONS`,
`RELATED_TO`, or `IN_TOPIC`, with their properties as JSON in `_props`.

```cypher
MATCH (e:Node {_label: 'Entity'})-[r:RELATED_TO]-(o:Node)
WHERE e.p_name = 'Tidewater'
RETURN o.p_name, r._props
```

Everything ChatLore adds for search, the `_ft_` text copies and the
`_embedding` vectors, is indexed in FalkorDB. Treat the graph as ChatLore's:
changes made to it directly are lost on the next `chatlore index --rebuild`.

## Differences from SQLite

- FalkorDB has no transactions spanning several queries. Each write is atomic,
but an import interrupted halfway leaves the conversations written so far;
running it again finishes the job, as with SQLite.
- Word search matches the same words as SQLite, without stemming or stop words,
and ignores accents. Its scores are FalkorDB's own, so results with equal
matches can come back in a different order.
- The server has to be running. Commands fail with a connection error when it
is not.
- Keep the server close: on the same machine or network. ChatLore asks for
what a page needs in a few queries and writes in batches, but each query
still waits for the network. Against a free FalkorDB Cloud server a
continent away, with 0.16 seconds per round trip and about 20 KB per second,
a search took about a second, the graph view 2.4 seconds, gathering the
passages for a question under a second, and loading the demo 41 seconds.
- Each conversation takes a few round trips to the server, so importing is
slower. On a made-up library of 12,000 notes, importing took 31 seconds
against 8 with SQLite, chunking 26 seconds against 145, and a search the same
time on both.
7 changes: 6 additions & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,11 @@ Changelog = "https://github.com/cl0ver012/chatlore/blob/main/CHANGELOG.md"
[project.scripts]
chatlore = "chatlore.cli:run"

[project.optional-dependencies]
falkordb = [
"falkordb>=1.7",
]

[build-system]
requires = ["hatchling>=1.25"]
build-backend = "hatchling.build"
Expand Down Expand Up @@ -103,7 +108,7 @@ pretty = true
plugins = ["pydantic.mypy"]

[[tool.mypy.overrides]]
module = ["sqlite_vec", "fastembed", "networkx"]
module = ["sqlite_vec", "fastembed", "networkx", "falkordb", "falkordb.*", "redis", "redis.*"]
ignore_missing_imports = true

[tool.pytest.ini_options]
Expand Down
65 changes: 40 additions & 25 deletions src/chatlore/api.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@
from chatlore.llm import LLMError, make_llm
from chatlore.paths import default_home
from chatlore.search import hybrid_search
from chatlore.store import EdgeType, GraphStore, Label, Node, open_store
from chatlore.store import Edge, EdgeType, GraphStore, Label, Node, open_store

WEB = Path(__file__).parent / "web"
"""The web interface's files, served at /."""
Expand Down Expand Up @@ -121,13 +121,14 @@ def health() -> dict[str, str]:
@app.get("/stats")
def stats() -> dict[str, int]:
with store() as graph:
counts = graph.count_by_label()
return {
"conversations": graph.count_nodes(Label.CONVERSATION),
"messages": graph.count_nodes(Label.MESSAGE),
"chunks": graph.count_nodes(Label.CHUNK),
"conversations": counts[Label.CONVERSATION],
"messages": counts[Label.MESSAGE],
"chunks": counts[Label.CHUNK],
"embeddings": graph.count_embeddings(),
"entities": graph.count_nodes(Label.ENTITY),
"topics": graph.count_nodes(Label.TOPIC),
"entities": counts[Label.ENTITY],
"topics": counts[Label.TOPIC],
}

@app.get("/search")
Expand Down Expand Up @@ -224,7 +225,8 @@ def entities(
with store() as graph:
if q:
hits = graph.search_text(q, limit=limit, labels=[Label.ENTITY])
nodes = [node for hit in hits if (node := graph.get_node(hit.node_id))]
found = graph.get_nodes(hit.node_id for hit in hits)
nodes = [found[hit.node_id] for hit in hits if hit.node_id in found]
else:
nodes = sorted(
graph.find_nodes(Label.ENTITY), key=lambda node: -int(node.props["mentions"])
Expand Down Expand Up @@ -277,7 +279,8 @@ def topics(
with store() as graph:
if q:
hits = graph.search_text(q, limit=limit, labels=[Label.TOPIC])
nodes = [node for hit in hits if (node := graph.get_node(hit.node_id))]
found = graph.get_nodes(hit.node_id for hit in hits)
nodes = [found[hit.node_id] for hit in hits if hit.node_id in found]
else:
nodes = sorted(
graph.find_nodes(Label.TOPIC), key=lambda node: -int(node.props["size"])
Expand Down Expand Up @@ -335,30 +338,42 @@ def graph(
for _, node in sorted(members, key=lambda pair: -int(pair[1].props["mentions"]))
}
chosen = dict(list(chosen.items())[:limit])
else:
linked = [
node
for node in graph.find_nodes(Label.ENTITY)
if graph.neighbors(node.id, [EdgeType.RELATED_TO], "both", limit=1)
]
links: dict[str, list[Edge]] | None = None
if not entity and not topic:
every = graph.find_nodes(Label.ENTITY)
related = graph.edges_many(
(node.id for node in every), [EdgeType.RELATED_TO], "both"
)
linked = [node for node in every if related[node.id]]
linked.sort(key=lambda node: -int(node.props["mentions"]))
chosen = {node.id: node for node in linked[:limit]}
# The outgoing links are among the ones just fetched.
links = {
node_id: [edge for edge in related[node_id] if edge.src == node_id]
for node_id in chosen
}

nodes, edges, topics = [], [], {}
nodes, edges, topic_of = [], [], {}
if links is None:
links = graph.edges_many(chosen, [EdgeType.RELATED_TO], "out")
for node_id, in_topic in graph.edges_many(chosen, [EdgeType.IN_TOPIC]).items():
if in_topic:
topic_of[node_id] = in_topic[0].dst
topic_nodes = graph.get_nodes(dict.fromkeys(topic_of.values()))
topics = {
topic_id: topic_nodes[topic_id].props.get("title")
for topic_id in dict.fromkeys(topic_of.values())
if topic_id in topic_nodes
}
for node in chosen.values():
in_topic = graph.neighbors(node.id, [EdgeType.IN_TOPIC])
topic_node = in_topic[0][1] if in_topic else None
if topic_node is not None:
topics[topic_node.id] = topic_node.props.get("title")
nodes.append({**_entity(node), "topic": topic_node.id if topic_node else None})
for edge, other in graph.neighbors(
node.id, [EdgeType.RELATED_TO], "out", limit=1_000_000
):
if other.id in chosen:
topic_id = topic_of.get(node.id)
nodes.append({**_entity(node), "topic": topic_id if topic_id in topics else None})
for edge in links[node.id]:
if edge.dst in chosen:
edges.append(
{
"source": node.id,
"target": other.id,
"target": edge.dst,
"weight": edge.props.get("weight", 1),
"description": (edge.props.get("descriptions") or [None])[0],
}
Expand Down
15 changes: 8 additions & 7 deletions src/chatlore/archive.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@
from chatlore.extraction import ExtractionCache
from chatlore.library import AddOutcome, Library
from chatlore.models import Conversation, Role
from chatlore.store import Edge, GraphStore, Label, Node, open_store
from chatlore.store import ConversationWriter, Edge, GraphStore, Label, Node, open_store

FORMAT = "chatlore-archive"
FORMAT_VERSION = 1
Expand Down Expand Up @@ -216,12 +216,13 @@ def import_archive(path: Path, home: Path, dry_run: bool = False) -> ImportRepor
return report

with Library(home) as library, open_store(home) as store, store.transaction():
for conversation in read_conversations(path):
outcome = library.add(conversation)
report.outcomes[outcome.value] += 1
report.messages += len(conversation.messages)
if outcome is not AddOutcome.UNCHANGED:
store.upsert_conversation(conversation)
with ConversationWriter(store) as writer:
for conversation in read_conversations(path):
outcome = library.add(conversation)
report.outcomes[outcome.value] += 1
report.messages += len(conversation.messages)
if outcome is not AddOutcome.UNCHANGED:
writer.add(conversation)
report.nodes, report.edges = restore_graph(store, *read_graph(path))
restore_caches(path, home)
return report
Expand Down
Loading
Loading