feat(importers): import anything: documents, email, data files, folders, and archives - #22
Merged
Merged
Conversation
PDF is the one document format the standard library cannot read. pypdf is pure Python with no dependencies of its own, and is loaded only when a PDF is imported.
Each document becomes a note of one message, known by its path within what was imported, so it is chunked, searched, and read into the graph like a Markdown note, and importing it again changes nothing. Word, PowerPoint, Excel, OpenDocument, and EPUB are zipped XML or HTML and are read with the standard library, titled from their own properties; PDF uses pypdf; web pages drop scripts and styles; CSV rows and JSON values become lines of text; code keeps its language. An email is a conversation of its own, and a mailbox one per email, under the new source email; other documents use the new source document. Markdown and text files stay Markdown notes, as before, and a text file in another encoding is still read. A document with no text, such as a scanned PDF, is refused with that reason.
chatlore import takes any number of files, folders, and archives, and works out what each file is. Archives are unpacked, and archives inside them, four levels deep: zip, tar, and gz with the standard library, 7z and rar with the bsdtar that ships with Windows and macOS. Chat exports are recognised by content wherever they sit, and the rest of an export, such as ChatGPT's chat.html, is skipped; in a Google Takeout only Gemini's activity is taken, while other files are read like any others. Everything else ChatLore can read becomes a note, and every file left out is listed with the reason. Unpacking is limited to 2 GB and 100,000 files, checked before writing, and never writes outside its folder. A folder's files keep the paths the Markdown importer gave them, so a vault imported before updates in place. --source still imports one path with one importer, as before.
POST /library/files adds one file to a batch, keeping the path the browser gives it within the folder it was chosen from, made safe; POST /library/import?batch= then imports the batch through the same intake as the command line. The upload limit counts the whole batch, and a batch holds up to 20,000 files. Batches never imported are deleted a day on. The import's status says what came from which source and lists the files it skipped. One file sent to POST /library/import works as before.
The dialog takes several files, a folder chosen with the folder button, or files and folders dropped together, whose folders it walks. Each file is sent with its path, with one progress bar for all of them, and folders like .git and node_modules are left out before anything is sent. The summary says what came from where and lists skipped files.
zip and tar need nothing extra, but 7z and rar are unpacked with bsdtar, which Debian ships in libarchive-tools. The Docker job checks that it is there.
docs/importers.md lists what can be imported and how archives, chat exports inside them, documents, and skipped files are handled; the web guide covers dropping files and folders. Roadmap: M12.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Milestone 12: import anything. Users no longer need to think about formats:
chatlore importtakes any number of files, folders, and archives.ChatLore works out what each file is. Chat exports are recognised by content wherever they sit, documents, email, and data files become notes, archives are unpacked (nested ones too), and every file it skips is listed with the reason.
Changes
Documents as notes (
chatlore.importers.documents). Each document becomes a note of one message, known by its path within the import, so it is chunked, searched, and read into the graph like a Markdown note, and re-importing changes nothing.pypdf, the one new dependency (pure Python, loaded only for PDFs)column: valuekey.path: valuelines.log,.rst,.org, subtitles, and any file that sniffs as plain text.emlbecomes a conversation;.mboxbecomes one per email.documentandemail. Markdown and.txtstay Markdown notes, as before.Intake (
chatlore.importers.intake) walks everything it is given:datafilter), tar.gz, tgz, tar.bz2, tar.xz, and gz. 7z and rar go throughbsdtar, found on PATH or as Windows' owntar.exeor macOS's/usr/bin/tar, and checked to really be bsdtar. Unpacking is capped at 2 GB and 100,000 files, checked before writing, and never writes outside its folder; bsdtar's listing is read for sizes first.conversations.jsonand GeminiMyActivity.jsongo to their importers wherever they are..jsonand.htmlfiles, such as ChatGPT'schat.htmlwith the same chats, are skipped as part of it..git,node_modules,__MACOSX, and similar folders, plus system files like.DS_Store.CLI.
chatlore import <path>...takes many paths.from <source>rows when there are several sources, askipped filescount, and the skipped files grouped by reason.--sourcestill imports one path at a time with that importer, and a single ChatLore archive imports as before.API.
POST /library/files?batch=<32 hex digits>adds one file to a batch, with its path fromX-Filenamemade safe: no absolute paths, no.., no reserved characters.POST /library/import?batch=imports the batch through the intake.sources,skipped_files, and the first 50 skipped files with reasons.POST /library/importworks as before.Web interface.
.git,node_modules, and similar folders left out before sending.Docker and CI. The image installs
libarchive-toolsfor 7z and rar, and the Docker job checks thatbsdtaris there.Docs.
docs/importers.md: a "What can be imported" table and how archives, exports inside them, documents, and skipped files are handled.docs/web.md: dropping files and folders.How it was tested
.gitignored, and a picture listed;.gz;chatlore importwith several paths of different kinds, and with nothing readable;..path kept inside, and batch checks (bad id, empty, over the limit, missing header).node_modules:node_modulesfile was left out before sending;ruff,ruff format --check, strictmypy, andpytestpass locally on Python 3.14.Checklist
uv run ruff check .anduv run ruff format --check .passuv run mypypassesuv run pytestpassesCHANGELOG.mdupdated under Unreleased