Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
48 changes: 30 additions & 18 deletions docs-website/reference/haystack-api/retrievers_api.md
Original file line number Diff line number Diff line change
Expand Up @@ -324,7 +324,7 @@ Create the InMemoryBM25Retriever component.

- **document_store** (<code>InMemoryDocumentStore</code>) – An instance of InMemoryDocumentStore where the retriever should search for relevant documents.
- **filters** (<code>dict\[str, Any\] | None</code>) – A dictionary with filters to narrow down the retriever's search space in the document store.
- **top_k** (<code>int</code>) – The maximum number of documents to retrieve.
- **top_k** (<code>int</code>) – The maximum number of documents to retrieve. Must be greater than 0.
- **scale_score** (<code>bool</code>) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant.
When `False`, uses raw similarity scores.
- **filter_policy** (<code>FilterPolicy</code>) – The filter policy to apply during retrieval.
Expand All @@ -336,7 +336,7 @@ Create the InMemoryBM25Retriever component.
**Raises:**

- <code>TypeError</code> – If the document_store is not an instance of InMemoryDocumentStore.
- <code>ValueError</code> – If the specified `top_k` is not > 0.
- <code>ValueError</code> – If `top_k` is not greater than 0.

#### to_dict

Expand Down Expand Up @@ -383,7 +383,7 @@ Run the InMemoryBM25Retriever on the given input data.

- **query** (<code>str</code>) – The query string for the Retriever.
- **filters** (<code>dict\[str, Any\] | None</code>) – A dictionary with filters to narrow down the search space when retrieving documents.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return. If 0, no documents are returned.
- **scale_score** (<code>bool | None</code>) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant.
When `False`, uses raw similarity scores.

Expand All @@ -393,7 +393,8 @@ Run the InMemoryBM25Retriever on the given input data.

**Raises:**

- <code>ValueError</code> – If the specified DocumentStore is not found or is not a InMemoryDocumentStore instance.
- <code>ValueError</code> – If the specified DocumentStore is not found or is not a InMemoryDocumentStore instance,
or if `top_k` is negative.

#### run_async

Expand All @@ -412,7 +413,7 @@ Run the InMemoryBM25Retriever on the given input data.

- **query** (<code>str</code>) – The query string for the Retriever.
- **filters** (<code>dict\[str, Any\] | None</code>) – A dictionary with filters to narrow down the search space when retrieving documents.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return. If 0, no documents are returned.
- **scale_score** (<code>bool | None</code>) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant.
When `False`, uses raw similarity scores.

Expand All @@ -422,7 +423,8 @@ Run the InMemoryBM25Retriever on the given input data.

**Raises:**

- <code>ValueError</code> – If the specified DocumentStore is not found or is not a InMemoryDocumentStore instance.
- <code>ValueError</code> – If the specified DocumentStore is not found or is not a InMemoryDocumentStore instance,
or if `top_k` is negative.

## in_memory/embedding_retriever

Expand Down Expand Up @@ -483,7 +485,7 @@ Create the InMemoryEmbeddingRetriever component.

- **document_store** (<code>InMemoryDocumentStore</code>) – An instance of InMemoryDocumentStore where the retriever should search for relevant documents.
- **filters** (<code>dict\[str, Any\] | None</code>) – A dictionary with filters to narrow down the retriever's search space in the document store.
- **top_k** (<code>int</code>) – The maximum number of documents to retrieve.
- **top_k** (<code>int</code>) – The maximum number of documents to retrieve. Must be greater than 0.
- **scale_score** (<code>bool</code>) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant.
When `False`, uses raw similarity scores.
- **return_embedding** (<code>bool</code>) – When `True`, returns the embedding of the retrieved documents.
Expand All @@ -497,7 +499,7 @@ Create the InMemoryEmbeddingRetriever component.
**Raises:**

- <code>TypeError</code> – If the document_store is not an instance of InMemoryDocumentStore.
- <code>ValueError</code> – If the specified top_k is not > 0.
- <code>ValueError</code> – If `top_k` is not greater than 0.

#### to_dict

Expand Down Expand Up @@ -545,7 +547,7 @@ Run the InMemoryEmbeddingRetriever on the given input data.

- **query_embedding** (<code>list\[float\]</code>) – Embedding of the query.
- **filters** (<code>dict\[str, Any\] | None</code>) – A dictionary with filters to narrow down the search space when retrieving documents.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return. If 0, no documents are returned.
- **scale_score** (<code>bool | None</code>) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant.
When `False`, uses raw similarity scores.
- **return_embedding** (<code>bool | None</code>) – When `True`, returns the embedding of the retrieved documents.
Expand All @@ -557,7 +559,8 @@ Run the InMemoryEmbeddingRetriever on the given input data.

**Raises:**

- <code>ValueError</code> – If the specified DocumentStore is not found or is not an InMemoryDocumentStore instance.
- <code>ValueError</code> – If the specified DocumentStore is not found or is not an InMemoryDocumentStore instance,
or if `top_k` is negative.

#### run_async

Expand All @@ -577,7 +580,7 @@ Run the InMemoryEmbeddingRetriever on the given input data.

- **query_embedding** (<code>list\[float\]</code>) – Embedding of the query.
- **filters** (<code>dict\[str, Any\] | None</code>) – A dictionary with filters to narrow down the search space when retrieving documents.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return. If 0, no documents are returned.
- **scale_score** (<code>bool | None</code>) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant.
When `False`, uses raw similarity scores.
- **return_embedding** (<code>bool | None</code>) – When `True`, returns the embedding of the retrieved documents.
Expand All @@ -589,7 +592,8 @@ Run the InMemoryEmbeddingRetriever on the given input data.

**Raises:**

- <code>ValueError</code> – If the specified DocumentStore is not found or is not an InMemoryDocumentStore instance.
- <code>ValueError</code> – If the specified DocumentStore is not found or is not an InMemoryDocumentStore instance,
or if `top_k` is negative.

## multi_query_embedding_retriever

Expand Down Expand Up @@ -1031,6 +1035,10 @@ Create the MultiRetriever component.
- `concatenate`: Combines all results into a single list and deduplicates.
- `reciprocal_rank_fusion`: Deduplicates and assigns scores based on reciprocal rank fusion.

**Raises:**

- <code>ValueError</code> – If `top_k` or `top_k_per_retriever` is set and is not greater than 0.

#### warm_up

```python
Expand Down Expand Up @@ -1084,11 +1092,12 @@ Runs retrievers in parallel on the given query and returns deduplicated results.
- **filters** (<code>dict\[str, Any\] | None</code>) – Filters to apply. Defaults to the value set at initialization.
- **top_k_per_retriever** (<code>int | None</code>) – The maximum number of documents to return per retriever. When set, this will override the `top_k`
parameter for each retriever. If None, the `top_k` parameter set for retrievers will be used.
Defaults to the value set at initialization.
If 0, no documents are returned. Defaults to the value set at initialization.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return overall, extracted from the combined results of all
retrievers. When set, the results are always merged using reciprocal rank fusion (regardless of
`join_mode`) so that the combined list has a consistent global ranking before it is truncated to
`top_k`. If None, all results are returned. Defaults to the value set at initialization.
`top_k`. If None, all results are returned. If 0, no documents are returned.
Defaults to the value set at initialization.
- **active_retrievers** (<code>list\[str\] | None</code>) – Names of retrievers to run. Defaults to all. Must match keys in the `retrievers` dictionary.

**Returns:**
Expand All @@ -1098,7 +1107,8 @@ Runs retrievers in parallel on the given query and returns deduplicated results.

**Raises:**

- <code>ValueError</code> – If any name in `active_retrievers` does not match a retriever name.
- <code>ValueError</code> – If any name in `active_retrievers` does not match a retriever name,
or if the resolved `top_k` or `top_k_per_retriever` is negative.

#### run_async

Expand All @@ -1123,11 +1133,12 @@ Uses each retriever's `run_async` method if available, otherwise runs `run` in a
- **filters** (<code>dict\[str, Any\] | None</code>) – Filters to apply. Defaults to the value set at initialization.
- **top_k_per_retriever** (<code>int | None</code>) – The maximum number of documents to return per retriever. When set, this will override the `top_k`
parameter for each retriever. If None, the `top_k` parameter set for retrievers will be used.
Defaults to the value set at initialization.
If 0, no documents are returned. Defaults to the value set at initialization.
- **top_k** (<code>int | None</code>) – The maximum number of documents to return overall, extracted from the combined results of all
retrievers. When set, the results are always merged using reciprocal rank fusion (regardless of
`join_mode`) so that the combined list has a consistent global ranking before it is truncated to
`top_k`. If None, all results are returned. Defaults to the value set at initialization.
`top_k`. If None, all results are returned. If 0, no documents are returned.
Defaults to the value set at initialization.
- **active_retrievers** (<code>list\[str\] | None</code>) – Names of retrievers to run. Defaults to all. Must match keys in the `retrievers` dictionary.

**Returns:**
Expand All @@ -1137,7 +1148,8 @@ Uses each retriever's `run_async` method if available, otherwise runs `run` in a

**Raises:**

- <code>ValueError</code> – If any name in `active_retrievers` does not match a retriever name.
- <code>ValueError</code> – If any name in `active_retrievers` does not match a retriever name,
or if the resolved `top_k` or `top_k_per_retriever` is negative.

#### to_dict

Expand Down