Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,7 @@ jobs:

- name: Upload an export as a visitor, then download and delete it
run: |
docker exec demo bsdtar --version
curl -fs -c visitor.txt http://127.0.0.1:7860/library/import \
-H 'X-ChatLore: 1' -H 'X-Filename: conversations.json' \
--data-binary @tests/fixtures/claude/conversations.json
Expand Down
9 changes: 9 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- A `Dockerfile` for the hosted demo, with the demo library and embedding model built in. CI
builds it, starts it, and checks the web interface, the API, and a tool call over MCP. Guide
in `docs/hosting.md`.
- Import anything: `chatlore import` takes any number of files, folders, and archives, and
**Your data** in the web interface takes many files or a whole folder. Archives (zip, tar,
gz, and with bsdtar 7z and rar) are unpacked, nested ones too; chat exports are recognised by
their content wherever they sit; documents become notes: PDF, Word, PowerPoint, Excel,
OpenDocument, EPUB, RTF, web pages, CSV, JSON, XML, code, and any plain text; email becomes a
conversation per message. Skipped files are listed with the reason. New sources `document`
and `email`. Guide in `docs/importers.md`.
- `POST /library/files` adds a file, with its path, to a batch that `POST /library/import`
then imports.
- **Your data** in the web interface: upload a ChatGPT, Claude, or Gemini export, Markdown
notes, or a ChatLore archive, and watch it be imported, embedded, and read into the knowledge
graph in the background; download the library as a ChatLore archive or as Markdown.
Expand Down
3 changes: 3 additions & 0 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,9 @@ FROM ghcr.io/astral-sh/uv:0.12-python3.12-trixie-slim
ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never
WORKDIR /app

# bsdtar unpacks 7z and rar uploads; zip and tar need nothing extra.
RUN apt-get update && apt-get install -y --no-install-recommends libarchive-tools && rm -rf /var/lib/apt/lists/*

# Dependencies first, so changing the code does not reinstall them.
COPY pyproject.toml uv.lock ./
RUN uv sync --locked --no-dev --no-install-project
Expand Down
8 changes: 5 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,9 @@
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
[![Python 3.12+](https://img.shields.io/badge/python-3.12%2B-blue.svg)](pyproject.toml)

**Status: pre-alpha.** Importing and search work today: ChatGPT, Claude, Gemini,
and Markdown exports land in a local library and a SQLite graph store you can
**Status: pre-alpha.** Importing and search work today: ChatGPT, Claude, and
Gemini exports, notes, documents (PDF, Word, Excel, and more), email, and
archives of any of them land in a local library and a SQLite graph store you can
search from the terminal by words, by meaning, or both. A language model turns
them into a knowledge graph of entities, relationships, and topics that you can
browse and search, and you can ask questions and get answers with sources, in
Expand All @@ -34,7 +35,7 @@ Asking questions also needs a model key; see [docs/models.md](docs/models.md).

```bash
uv tool install chatlore # or: pipx install chatlore
chatlore import path/to/chatgpt-export.zip
chatlore import path/to/chatgpt-export.zip ~/Documents/notes # exports, documents, archives
chatlore search "postgres index"
chatlore process # chunk and embed, local model
chatlore search "why was my query slow" --semantic
Expand Down Expand Up @@ -102,6 +103,7 @@ extraction is an optional enrichment you can re-run with a better model later.
| M9 | Hosted demo | public mode, MCP over HTTP, and Docker image done; deployment next |
| M10 | FalkorDB backend | done |
| M11 | Your own data in the web interface: upload, export, private visitor libraries | done |
| M12 | Import anything: documents, email, data files, folders, and archives | done |

## Development setup

Expand Down
47 changes: 43 additions & 4 deletions docs/importers.md
Original file line number Diff line number Diff line change
@@ -1,15 +1,18 @@
# Importing your data

```bash
chatlore import <path> # detects the source from the content
chatlore import <path>... # anything: files, folders, archives
chatlore import <path> --source chatgpt
chatlore import <path> --dry-run # parse and report, write nothing
chatlore stats # what the library holds
```

`<path>` can be the zip exactly as you downloaded it, the extracted folder, or
the JSON file itself. Importing is idempotent: run it again after a fresh
export and only new or changed conversations are written.
Give it as many files, folders, and archives as you like, and it works out what
each one is from its content: the export zip exactly as you downloaded it, the
extracted folder, a folder of documents, or all of them at once. The web
interface does the same under **Your data**, for files or a whole folder.
Importing is idempotent: run it again after a fresh export and only new or
changed conversations are written.

Everything lands in `~/.chatlore/conversations/<source>/<id>.json` (override
the location with `--home` or `CHATLORE_HOME`). The files are plain JSON and stay readable
Expand All @@ -19,6 +22,42 @@ SQLite database holding the graph and the search indexes.
An archive written by `chatlore export` is recognised too, and brings its
knowledge graph and caches along; see [export.md](export.md).

## What can be imported

| What | Files | Becomes |
|---|---|---|
| Chat exports | ChatGPT and Claude `conversations.json`, Gemini `MyActivity.json`, wherever they sit | conversations |
| Notes | Markdown and text (`.md`, `.markdown`, `.txt`), Obsidian vaults | notes, source `markdown` |
| Documents | PDF, Word (`.docx`), PowerPoint (`.pptx`), Excel (`.xlsx`), OpenDocument (`.odt`, `.odp`, `.ods`), EPUB, RTF, web pages (`.html`) | notes, source `document` |
| Data | CSV and TSV, JSON and JSON Lines, XML | notes: a line per row, or `key.path: value` lines |
| Code and text | source code, `.log`, `.rst`, `.org`, subtitles, and any other file that is plain text | notes; code keeps its language |
| Email | `.eml`, and `.mbox` mailboxes | one conversation per email, source `email` |
| Archives | zip, tar, tar.gz, tgz, tar.bz2, tar.xz, gz, 7z, rar | whatever they hold |

**Archives** are unpacked, and archives inside them too, four levels deep, so a
zip of folders of zips works, and so does a Google Takeout split into several
zips. zip, tar, and gz need nothing extra. 7z and rar are unpacked with
`bsdtar`, which comes with Windows 10 and later and with macOS; on Linux,
install `libarchive-tools`. Together they may unpack to at most 2 GB, 100,000
files, and nothing is ever written outside the folder they are unpacked into.

**Chat exports** are found by their content, not their names or places. The
other files of an export, such as ChatGPT's `chat.html`, which holds the same
chats again, are skipped. In a Google Takeout, Gemini's activity is imported
and other products' activity is skipped, while other files, such as Drive
documents, are read like any others.

**Documents** become notes of one message each, titled from the document when
it names itself (a Word title, a PDF's metadata, a web page's `<title>`) and
otherwise by file name. Each is known by its path within what was imported, so
importing the same folder again updates the same notes. A scanned PDF with no
text layer has no text to read and is skipped.

**Skipped files** are listed at the end with the reason: pictures, audio,
video, programs, files that are not text, and documents that could not be read.
Folders such as `.git` and `node_modules` are left out, and so are system files
like `.DS_Store`.

## Searching

```bash
Expand Down
16 changes: 10 additions & 6 deletions docs/web.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,15 +32,19 @@ follows the system's light or dark setting.

**Your data**, in the navigation, imports an export and downloads the library:

- Drop a file on the dialog, or choose one: a ChatGPT or Claude export (.zip),
Gemini Takeout (.zip or MyActivity.json), Markdown notes (.zip or .md), or a
ChatLore archive. Uploads may be up to 200 MB, or what `--max-upload-mb` sets.
- Drop files or folders on the dialog, or choose files or a whole folder: chat
exports, documents, notes, email, data files, and archives of any of them, as
listed in [importers.md](importers.md#what-can-be-imported). Each file is sent
with its path in its folder, then all of them are imported together. Uploads
may be up to 200 MB in all, or what `--max-upload-mb` sets; folders such as
`.git` and `node_modules` are left out before anything is sent.
- The import runs in the background, like `chatlore import`, `chatlore
process`, and `chatlore extract` one after another, and the dialog shows each
step with its progress: importing, preparing search, reading with the language
model, summarising, linking names for the same thing, and topics. Without a
model key, everything but the knowledge graph is built. When it is done,
**Show the library** reloads the page on it.
model, summarising, linking names for the same thing, and topics. It then
says what came from where, and lists the files it skipped with the reason.
Without a model key, everything but the knowledge graph is built. When it is
done, **Show the library** reloads the page on it.
- **ChatLore archive** downloads the whole library, graph and caches included,
to import anywhere; **Markdown** downloads one readable file per conversation.
See [export.md](export.md).
Expand Down
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ dependencies = [
"networkx>=3.7",
"openai>=3.17",
"pydantic>=2.9",
"pypdf>=6.0",
"python-dotenv>=1.2",
"pyyaml>=6.0",
"rich>=13.7",
Expand Down
178 changes: 127 additions & 51 deletions src/chatlore/api.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@
from __future__ import annotations

import functools
import re
import shutil
import tempfile
import threading
Expand All @@ -35,7 +36,7 @@
from contextlib import asynccontextmanager
from contextvars import ContextVar
from http.cookies import SimpleCookie
from pathlib import Path
from pathlib import Path, PurePosixPath
from typing import Annotated, Any
from urllib.parse import unquote

Expand All @@ -54,7 +55,7 @@
from chatlore.archive import export_archive, export_markdown
from chatlore.chat import MAX_SOURCES, ChatLimits, Context, Source, answer, cited, retrieve
from chatlore.embeddings import Embedder, EmbeddingError, make_embedder, normalise
from chatlore.imports import Imports, upload_suffix
from chatlore.imports import Imports, safe_relative
from chatlore.library import Library
from chatlore.llm import LLMError, make_llm
from chatlore.mcp_server import create_server
Expand Down Expand Up @@ -176,6 +177,10 @@ def _topic(node: Node, entities: int | None = None) -> dict[str, Any]:
without the browser first asking this server, which does not allow it, so it
cannot upload or delete through a visitor's browser."""
_SWEEP_SECONDS = 600
MAX_BATCH_FILES = 20_000
"""Files one upload may hold."""
_BATCH = re.compile(r"[0-9a-f]{32}")
_STALE_BATCH_SECONDS = 24 * 3600

# The visitor's own library, when the request carries a token for one.
_visitor: ContextVar[Space | None] = ContextVar("chatlore_visitor", default=None)
Expand Down Expand Up @@ -212,6 +217,24 @@ def _cookie(scope: Scope, name: str) -> str | None:
return None


def _forget_stale_batches(uploads: Path) -> None:
"""Delete batches that were uploaded but never imported, a day on."""
if not uploads.exists():
return
now = time.time()
for folder in uploads.iterdir():
try:
stale = folder.is_dir() and now - folder.stat().st_mtime > _STALE_BATCH_SECONDS
except OSError:
continue
if stale:
shutil.rmtree(folder, ignore_errors=True)


def _only_file(folder: Path) -> str:
return next((path.name for path in folder.rglob("*") if path.is_file()), "1 file")


def _changes_allowed(request: Request) -> None:
if request.headers.get(REQUEST_HEADER) != "1":
raise HTTPException(403, f"changes need the {REQUEST_HEADER} header")
Expand Down Expand Up @@ -334,65 +357,118 @@ def library_info() -> dict[str, Any]:
"import": status.as_dict() if status is not None else None,
}

@app.post("/library/import", status_code=202)
async def import_upload(request: Request, response: Response) -> dict[str, Any]:
"""Take an export as the request body and import it in the background.
staged: dict[tuple[Path, str], list[int]] = {}
staging = threading.Lock()

Send the file's name in ``X-Filename``. On a public server the upload goes
into the visitor's own library, made on the first upload, and a cookie
remembers it. ``GET /library`` shows how far the import got.
"""
_changes_allowed(request)
def _target(request: Request, response: Response) -> tuple[Path, Callable[[], GraphStore]]:
"""The library uploads go into: the server's, or the visitor's, made on first use."""
if not uploads:
raise HTTPException(403, "this server does not take uploads")
declared = request.headers.get("content-length")
if declared is not None and declared.isdigit() and int(declared) > max_upload:
raise HTTPException(413, f"uploads may be up to {max_upload // 1_000_000} MB")
if not public:
return library, lambda: open_store(library)
assert spaces is not None
space = _visitor.get()
created: Space | None = None
if public:
assert spaces is not None
if space is None:
token, space = spaces.create()
created = space
response.set_cookie(
COOKIE,
token,
max_age=int(spaces.keep.total_seconds()),
httponly=True,
samesite="lax",
secure=request.url.scheme == "https",
)
target, opener = space.home, space.open_store
else:
target, opener = library, lambda: open_store(library)
if space is None:
token, space = spaces.create()
response.set_cookie(
COOKIE,
token,
max_age=int(spaces.keep.total_seconds()),
httponly=True,
samesite="lax",
secure=request.url.scheme == "https",
)
return space.home, space.open_store

async def _receive(request: Request, path: Path, room: int) -> int:
"""Write the request body to ``path``, refusing more than ``room`` bytes."""
declared = request.headers.get("content-length")
if declared is not None and declared.isdigit() and int(declared) > room:
raise HTTPException(413, f"uploads may be up to {max_upload // 1_000_000} MB in all")
path.parent.mkdir(parents=True, exist_ok=True)
size = 0
try:
if imports.running(target):
raise HTTPException(409, "an import is already running for this library")
name = unquote(request.headers.get("x-filename") or "upload")[:200]
folder = target / "uploads"
folder.mkdir(parents=True, exist_ok=True)
path = folder / f"{uuid.uuid4().hex}{upload_suffix(name)}"
size = 0
with path.open("wb") as handle:
async for piece in request.stream():
size += len(piece)
if size > room:
raise HTTPException(
413, f"uploads may be up to {max_upload // 1_000_000} MB in all"
)
handle.write(piece)
except BaseException:
path.unlink(missing_ok=True)
raise
return size

def _batch_folder(target: Path, batch: str) -> Path:
if not _BATCH.fullmatch(batch):
raise HTTPException(422, "a batch is 32 hexadecimal digits")
return target / "uploads" / batch

@app.post("/library/files")
async def upload_file(
request: Request, response: Response, batch: Annotated[str, Query()]
) -> dict[str, Any]:
"""Add one file to a batch of uploads; ``POST /library/import`` imports the batch.

Send the file as the request body and its path in ``X-Filename``, such as
``Export/conversations.json`` for a file chosen with its folder. ``batch``
is any 32 hexadecimal digits the client picks for the whole upload. The
files of a batch may be up to the upload limit in all.
"""
_changes_allowed(request)
target, _ = _target(request, response)
folder = _batch_folder(target, batch)
with staging:
used = staged.setdefault((target, batch), [0, 0])
if used[1] >= MAX_BATCH_FILES:
raise HTTPException(413, f"a batch may hold up to {MAX_BATCH_FILES:,} files")
used[1] += 1
if not folder.exists():
_forget_stale_batches(target / "uploads")
name = safe_relative(unquote(request.headers.get("x-filename") or "upload"))
size = await _receive(request, folder / name, max_upload - used[0])
with staging:
used[0] += size
return {"batch": batch, "files": used[1], "bytes": used[0]}

@app.post("/library/import", status_code=202)
async def import_upload(
request: Request, response: Response, batch: Annotated[str | None, Query()] = None
) -> dict[str, Any]:
"""Import a batch of files sent to ``POST /library/files``, or one file sent here.

Without ``batch``, the request body is the file and ``X-Filename`` its name.
Archives are unpacked, chat exports recognised, and other files read as
notes, in the background. On a public server the upload goes into the
visitor's own library, made on the first upload, and a cookie remembers it.
``GET /library`` shows how far the import got.
"""
_changes_allowed(request)
target, opener = _target(request, response)
if imports.running(target):
raise HTTPException(409, "an import is already running for this library")
if batch is not None:
folder = _batch_folder(target, batch)
with staging:
used = staged.pop((target, batch), [0, 0])
if not folder.exists() or not any(folder.iterdir()):
raise HTTPException(400, "the batch holds no files")
name = f"{used[1]:,} files" if used[1] != 1 else _only_file(folder)
else:
name = safe_relative(unquote(request.headers.get("x-filename") or "upload"))
folder = _batch_folder(target, uuid.uuid4().hex)
_forget_stale_batches(target / "uploads")
try:
with path.open("wb") as handle:
async for piece in request.stream():
size += len(piece)
if size > max_upload:
raise HTTPException(
413, f"uploads may be up to {max_upload // 1_000_000} MB"
)
handle.write(piece)
size = await _receive(request, folder / name, max_upload)
if size == 0:
raise HTTPException(400, "the upload is empty")
except BaseException:
path.unlink(missing_ok=True)
shutil.rmtree(folder, ignore_errors=True)
raise
except BaseException:
if created is not None and spaces is not None:
spaces.delete(created)
raise
return imports.start(target, path, name, opener).as_dict()
name = PurePosixPath(name).name
return imports.start(target, folder, name, opener).as_dict()

@app.get("/library/export")
def export_library(
Expand Down
Loading
Loading