Description
RAGPipeline.ingest(path) called twice on the same file doubles the index:
pipe.ingest("doc.txt") # 12 chunks
pipe.ingest("doc.txt") # 12 chunks
len(pipe.retriever) # 24 — same 12 chunk ids stored twice
ChromaRetriever.add() uses upsert, so it is idempotent on chunk id. InMemoryRetriever and FAISSRetriever blindly append. The Retriever.add docstring in base.py does not say which behaviour is correct.
Motivation
Re-running an ingestion script (the normal workflow while iterating) silently pollutes results: the same chunk appears multiple times in top_k, crowding out genuinely different context. Chunk ids are already deterministic (<doc_id>:<index>, where doc_id is a hash of the source path), so idempotency is achievable.
Acceptance criteria
Files to touch
ragframework/base.py — Retriever.add docstring
ragframework/retriever/in_memory.py
ragframework/retriever/faiss.py
tests/test_retriever/test_in_memory.py, tests/test_retriever/test_faiss.py, tests/test_retriever/test_chroma.py
Resources
Estimated effort: Medium (half day)
Description
RAGPipeline.ingest(path)called twice on the same file doubles the index:ChromaRetriever.add()usesupsert, so it is idempotent on chunk id.InMemoryRetrieverandFAISSRetrieverblindly append. TheRetriever.adddocstring inbase.pydoes not say which behaviour is correct.Motivation
Re-running an ingestion script (the normal workflow while iterating) silently pollutes results: the same chunk appears multiple times in
top_k, crowding out genuinely different context. Chunk ids are already deterministic (<doc_id>:<index>, wheredoc_idis a hash of the source path), so idempotency is achievable.Acceptance criteria
Retriever.adddocstring inragframework/base.pydefines the contract: adding a chunk whoseidalready exists replaces it (upsert semantics)InMemoryRetriever.add()replaces existing rows by id (maintain anid -> row indexdict)FAISSRetriever.add()handles duplicates.IndexHNSWFlatdoes not support removal, so pick and document one approach: (a) keep anid -> labelmap and tombstone superseded labels (filtered out inretrieve, over-fetching to compensate), or (b) wrap the index infaiss.IndexIDMap2if it supports the chosen index type. Explain the trade-off in the docstring.len()unchanged, retrieval returns it once, and the new content/embedding winsCHANGELOG.mdupdated under[Unreleased]→### FixedFiles to touch
ragframework/base.py—Retriever.adddocstringragframework/retriever/in_memory.pyragframework/retriever/faiss.pytests/test_retriever/test_in_memory.py,tests/test_retriever/test_faiss.py,tests/test_retriever/test_chroma.pyResources
ChromaRetriever.addfor the upsert reference behaviourEstimated effort: Medium (half day)