Description
Add DirectoryLoader that walks a directory (optionally recursive, with glob filtering) and delegates each file to the right loader by extension, returning the concatenated list[Document].
loader = DirectoryLoader(
loaders={".txt": TextFileLoader(), ".md": MarkdownLoader(), ".pdf": PDFLoader()},
glob="**/*",
recursive=True,
on_error="skip", # or "raise"
)
pipeline = RAGPipeline(loader=loader, ...)
pipeline.ingest("./knowledge_base")
Motivation
Today RAGPipeline.ingest() takes a single file path, so indexing a corpus means writing a loop by hand and choosing loaders manually. This is the single most requested convenience in every RAG library.
Acceptance criteria
Files to touch
ragframework/document/loaders.py
ragframework/document/__init__.py
tests/test_document/test_loaders.py
Resources
Estimated effort: Medium (half day)
Description
Add
DirectoryLoaderthat walks a directory (optionally recursive, with glob filtering) and delegates each file to the right loader by extension, returning the concatenatedlist[Document].Motivation
Today
RAGPipeline.ingest()takes a single file path, so indexing a corpus means writing a loop by hand and choosing loaders manually. This is the single most requested convenience in every RAG library.Acceptance criteria
DirectoryLoader(DocumentLoader)inragframework/document/loaders.pyloadersmapping (.txt,.md,.markdown,.pdfwhenpypdfis importable) built lazily so importing the module never requires optional depson_error="raise"re-raises asLoaderErrorwith the offending pathsourcemust be an existing directory, otherwiseLoaderErrorDocument.metadataincludes"relative_path"from the root directorytmp_pathwith mixed extensions, nested folders, an unreadable file with bothon_errormodesragframework/document/__init__.py;CHANGELOG.mdupdatedFiles to touch
ragframework/document/loaders.pyragframework/document/__init__.pytests/test_document/test_loaders.pyResources
Path.rglobEstimated effort: Medium (half day)