Skip to content

feat(importers): import anything: documents, email, data files, folders, and archives - #22

Merged
cl0ver012 merged 7 commits into
mainfrom
feat/m12-import-anything
Sep 30, 2026
Merged

cl0ver012 merged 7 commits into
mainfrom
feat/m12-import-anything

Conversation

@cl0ver012

Copy link
Copy Markdown
Owner

Summary

Milestone 12: import anything. Users no longer need to think about formats:

  • chatlore import takes any number of files, folders, and archives.
  • Your data in the web interface takes many files or a whole folder.

ChatLore works out what each file is. Chat exports are recognised by content wherever they sit, documents, email, and data files become notes, archives are unpacked (nested ones too), and every file it skips is listed with the reason.

Changes

Documents as notes (chatlore.importers.documents). Each document becomes a note of one message, known by its path within the import, so it is chunked, searched, and read into the graph like a Markdown note, and re-importing changes nothing.

Kind Read with
Word, PowerPoint, Excel, OpenDocument (odt, odp, ods), EPUB the standard library, since they are zipped XML or HTML; titles and dates come from their own properties
PDF pypdf, the one new dependency (pure Python, loaded only for PDFs)
RTF, web pages (scripts and styles dropped) the standard library
CSV and TSV a line per row, column: value
JSON, JSON Lines, XML key.path: value lines
Code kept as code, with its language
.log, .rst, .org, subtitles, and any file that sniffs as plain text read as text, in whatever encoding
  • .eml becomes a conversation; .mbox becomes one per email.
  • New sources document and email. Markdown and .txt stay Markdown notes, as before.
  • A document with no text, such as a scanned PDF, is refused with that reason.

Intake (chatlore.importers.intake) walks everything it is given:

  • Archives are unpacked, four levels deep: zip, tar (with the standard library's data filter), tar.gz, tgz, tar.bz2, tar.xz, and gz. 7z and rar go through bsdtar, found on PATH or as Windows' own tar.exe or macOS's /usr/bin/tar, and checked to really be bsdtar. Unpacking is capped at 2 GB and 100,000 files, checked before writing, and never writes outside its folder; bsdtar's listing is read for sizes first.
  • Chat exports:
    • ChatGPT and Claude conversations.json and Gemini MyActivity.json go to their importers wherever they are.
    • An export's own .json and .html files, such as ChatGPT's chat.html with the same chats, are skipped as part of it.
    • In a Google Takeout, other products' activity is skipped, while Drive documents are read like any files.
    • ChatLore archives inside are handed back to the caller, which imports them with their graph.
  • Skipped without reading: .git, node_modules, __MACOSX, and similar folders, plus system files like .DS_Store.
  • Listed as skipped, with the reason: pictures, audio, video, and programs.
  • Paths: a folder's files keep the paths the Markdown importer gave them, so a vault imported before updates in place.

CLI.

  • chatlore import <path>... takes many paths.
  • The summary adds from <source> rows when there are several sources, a skipped files count, and the skipped files grouped by reason.
  • When nothing is readable, it says so before touching the library.
  • --source still imports one path at a time with that importer, and a single ChatLore archive imports as before.

API.

  • POST /library/files?batch=<32 hex digits> adds one file to a batch, with its path from X-Filename made safe: no absolute paths, no .., no reserved characters.
  • POST /library/import?batch= imports the batch through the intake.
  • The upload limit counts the whole batch; a batch holds up to 20,000 files, and batches never imported are deleted a day on.
  • The import status gains sources, skipped_files, and the first 50 skipped files with reasons.
  • One file sent to POST /library/import works as before.

Web interface.

  • The dialog takes several files, a folder through Choose a whole folder, or files and folders dropped together, whose folders it walks.
  • Each file is sent with its path under one progress bar, with .git, node_modules, and similar folders left out before sending.
  • The summary says what came from where, with a collapsible list of skipped files and reasons.

Docker and CI. The image installs libarchive-tools for 7z and rar, and the Docker job checks that bsdtar is there.

Docs.

  • docs/importers.md: a "What can be imported" table and how archives, exports inside them, documents, and skipped files are handled.
  • docs/web.md: dropping files and folders.
  • README, CHANGELOG, and roadmap (M12) updated.

How it was tested

  • 20 new tests, 361 in total:
    • every document reader on generated samples (Word, PowerPoint, Excel, OpenDocument, EPUB, PDF, email, mailbox, CSV, JSON, HTML, RTF, a cp1252 text file, code), with titles, dates, and sources;
    • sniffing unknown files, refusing empty or broken documents, and stable identities;
    • a mixed folder, with an export and its companions skipped, .git ignored, and a picture listed;
    • a zip holding a tar.gz holding a zip, and a lone .gz;
    • a Takeout with Gemini and other activity;
    • ChatLore archives set aside;
    • zip-slip names staying inside, and the size and file-count limits;
    • a 7z made and unpacked with bsdtar (skipped where bsdtar is missing);
    • chatlore import with several paths of different kinds, and with nothing readable;
    • a batch of four files with folder paths through the API, a .. path kept inside, and batch checks (bad id, empty, over the limit, missing header).
  • By hand, in a browser against a local server with the real model (DeepSeek V4 Flash), uploading seven files at once (the made-up ChatGPT export zip in a folder, a PDF and a Word file in another, an EPUB, an email, a tar.gz of notes, and a photo) plus a file under node_modules:
    • the node_modules file was left out before sending;
    • 17 conversations came in: 11 from ChatGPT, 3 documents, 2 notes from inside the tar.gz, and 1 email;
    • the photo was listed as skipped with its reason;
    • the graph got 57 entities and 10 topics;
    • word searches found the Word file, the email, the note inside the tar.gz, the PDF, and the EPUB by their contents.
  • ruff, ruff format --check, strict mypy, and pytest pass locally on Python 3.14.

Checklist

  • uv run ruff check . and uv run ruff format --check . pass
  • uv run mypy passes
  • uv run pytest passes
  • No real exports, databases, or keys are included
  • CHANGELOG.md updated under Unreleased

PDF is the one document format the standard library cannot read. pypdf
is pure Python with no dependencies of its own, and is loaded only when
a PDF is imported.
Each document becomes a note of one message, known by its path within
what was imported, so it is chunked, searched, and read into the graph
like a Markdown note, and importing it again changes nothing. Word,
PowerPoint, Excel, OpenDocument, and EPUB are zipped XML or HTML and
are read with the standard library, titled from their own properties;
PDF uses pypdf; web pages drop scripts and styles; CSV rows and JSON
values become lines of text; code keeps its language. An email is a
conversation of its own, and a mailbox one per email, under the new
source email; other documents use the new source document. Markdown and
text files stay Markdown notes, as before, and a text file in another
encoding is still read. A document with no text, such as a scanned PDF,
is refused with that reason.
chatlore import takes any number of files, folders, and archives, and
works out what each file is. Archives are unpacked, and archives inside
them, four levels deep: zip, tar, and gz with the standard library, 7z
and rar with the bsdtar that ships with Windows and macOS. Chat exports
are recognised by content wherever they sit, and the rest of an export,
such as ChatGPT's chat.html, is skipped; in a Google Takeout only
Gemini's activity is taken, while other files are read like any others.
Everything else ChatLore can read becomes a note, and every file left
out is listed with the reason.

Unpacking is limited to 2 GB and 100,000 files, checked before writing,
and never writes outside its folder. A folder's files keep the paths
the Markdown importer gave them, so a vault imported before updates in
place. --source still imports one path with one importer, as before.
POST /library/files adds one file to a batch, keeping the path the
browser gives it within the folder it was chosen from, made safe;
POST /library/import?batch= then imports the batch through the same
intake as the command line. The upload limit counts the whole batch,
and a batch holds up to 20,000 files. Batches never imported are
deleted a day on. The import's status says what came from which source
and lists the files it skipped. One file sent to POST /library/import
works as before.
The dialog takes several files, a folder chosen with the folder button,
or files and folders dropped together, whose folders it walks. Each
file is sent with its path, with one progress bar for all of them, and
folders like .git and node_modules are left out before anything is
sent. The summary says what came from where and lists skipped files.
zip and tar need nothing extra, but 7z and rar are unpacked with
bsdtar, which Debian ships in libarchive-tools. The Docker job checks
that it is there.
docs/importers.md lists what can be imported and how archives, chat
exports inside them, documents, and skipped files are handled; the web
guide covers dropping files and folders. Roadmap: M12.
@cl0ver012
cl0ver012 merged commit c336cf0 into main Sep 30, 2026
8 checks passed
@cl0ver012
cl0ver012 deleted the feat/m12-import-anything branch September 30, 2026 02:09
@cl0ver012 cl0ver012 mentioned this pull request Sep 30, 2026
5 tasks done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant