Description
Only ChromaRetriever survives a process restart. FAISSRetriever and InMemoryRetriever must re-embed the whole corpus every run. Add:
retriever.save("index_dir/")
retriever = FAISSRetriever.load("index_dir/")
Motivation
Embedding is the slow, expensive step (API cost or GPU time). Losing the index on exit makes the two fastest retrievers impractical beyond demos.
Acceptance criteria
Files to touch
ragframework/retriever/faiss.py
ragframework/retriever/in_memory.py
tests/test_retriever/test_faiss.py, tests/test_retriever/test_in_memory.py
examples/README.md
Resources
Estimated effort: Medium (half day)
Description
Only
ChromaRetrieversurvives a process restart.FAISSRetrieverandInMemoryRetrievermust re-embed the whole corpus every run. Add:Motivation
Embedding is the slow, expensive step (API cost or GPU time). Losing the index on exit makes the two fastest retrievers impractical beyond demos.
Acceptance criteria
save(path: str | Path) -> Noneand@classmethod load(path) -> Selfon both classesFAISSRetriever.savewritesindex.faiss(faiss.write_index) +chunks.json(ids, content, metadata) +meta.json(dimension,m,ef_construction,ef_search, format version)InMemoryRetriever.savewritesmatrix.npy+chunks.jsonnp.load(..., allow_pickle=False)load()on a missing/corrupt directory raisesRetrieverError; format-version mismatch raisesRetrieverErrorwith guidancetmp_path: save → load → identicallen()and identicalretrieve()resultsexamples/README.mdsnippet;CHANGELOG.mdupdatedFiles to touch
ragframework/retriever/faiss.pyragframework/retriever/in_memory.pytests/test_retriever/test_faiss.py,tests/test_retriever/test_in_memory.pyexamples/README.mdResources
faiss.write_index/read_indexnumpy.save/numpy.loadEstimated effort: Medium (half day)