Description
Dense retrieval alone misses exact matches (error codes, product SKUs, names). Add a lexical BM25Retriever and a HybridRetriever that fuses dense + lexical results with Reciprocal Rank Fusion.
The API problem this exposes
Retriever.retrieve(query_embedding, top_k) receives only the vector. A lexical retriever needs the text. Proposed, backward-compatible change to the ABC:
def retrieve(self, query_embedding: list[float], top_k: int = 5, *, query_text: str | None = None) -> list[Chunk]:
RAGPipeline.query() always passes query_text=query. Existing retrievers ignore it. BM25Retriever raises RetrieverError if it is None.
Acceptance criteria
Files to touch
ragframework/base.py
ragframework/pipeline/rag.py
ragframework/retriever/bm25.py, ragframework/retriever/hybrid.py — new
ragframework/retriever/__init__.py, pyproject.toml
tests/test_retriever/test_bm25.py, tests/test_retriever/test_hybrid.py — new
Resources
Estimated effort: Large (1-2 days)
Description
Dense retrieval alone misses exact matches (error codes, product SKUs, names). Add a lexical
BM25Retrieverand aHybridRetrieverthat fuses dense + lexical results with Reciprocal Rank Fusion.The API problem this exposes
Retriever.retrieve(query_embedding, top_k)receives only the vector. A lexical retriever needs the text. Proposed, backward-compatible change to the ABC:RAGPipeline.query()always passesquery_text=query. Existing retrievers ignore it.BM25RetrieverraisesRetrieverErrorif it isNone.Acceptance criteria
Retriever.retrievesignature extended with keyword-onlyquery_text; all built-in retrievers andRAGPipeline.query()updated; existing tests passBM25Retriever(Retriever)inragframework/retriever/bm25.pyusingrank_bm25(new[bm25]extra, guarded import). Simple whitespace + lowercase tokeniser by default, injectabletokenizer: Callable[[str], list[str]]BM25Retriever.add()accepts chunks without embeddings (document this deviation in the docstring — it is the one retriever where that is legitimate)HybridRetriever(dense: Retriever, sparse: Retriever, k: int = 60)inragframework/retriever/hybrid.pyimplementing RRF; fetchestop_k * 2from each and returns the fused top_kCHANGELOG.mdupdated under[Unreleased]→### Added/### ChangedFiles to touch
ragframework/base.pyragframework/pipeline/rag.pyragframework/retriever/bm25.py,ragframework/retriever/hybrid.py— newragframework/retriever/__init__.py,pyproject.tomltests/test_retriever/test_bm25.py,tests/test_retriever/test_hybrid.py— newResources
Estimated effort: Large (1-2 days)