From 5e12cc2b8a959abc5ca465802bd915b1cdda86b3 Mon Sep 17 00:00:00 2001 From: davidsbatista <7937824+davidsbatista@users.noreply.github.com> Date: Thu, 24 Sep 2026 16:53:05 +0000 Subject: [PATCH] Sync Haystack API reference on Docusaurus --- .../reference/haystack-api/retrievers_api.md | 48 ++++++++++++------- 1 file changed, 30 insertions(+), 18 deletions(-) diff --git a/docs-website/reference/haystack-api/retrievers_api.md b/docs-website/reference/haystack-api/retrievers_api.md index ab4a9ecd0c..f908c47e66 100644 --- a/docs-website/reference/haystack-api/retrievers_api.md +++ b/docs-website/reference/haystack-api/retrievers_api.md @@ -324,7 +324,7 @@ Create the InMemoryBM25Retriever component. - **document_store** (InMemoryDocumentStore) – An instance of InMemoryDocumentStore where the retriever should search for relevant documents. - **filters** (dict\[str, Any\] | None) – A dictionary with filters to narrow down the retriever's search space in the document store. -- **top_k** (int) – The maximum number of documents to retrieve. +- **top_k** (int) – The maximum number of documents to retrieve. Must be greater than 0. - **scale_score** (bool) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant. When `False`, uses raw similarity scores. - **filter_policy** (FilterPolicy) – The filter policy to apply during retrieval. @@ -336,7 +336,7 @@ Create the InMemoryBM25Retriever component. **Raises:** - TypeError – If the document_store is not an instance of InMemoryDocumentStore. -- ValueError – If the specified `top_k` is not > 0. +- ValueError – If `top_k` is not greater than 0. #### to_dict @@ -383,7 +383,7 @@ Run the InMemoryBM25Retriever on the given input data. - **query** (str) – The query string for the Retriever. - **filters** (dict\[str, Any\] | None) – A dictionary with filters to narrow down the search space when retrieving documents. -- **top_k** (int | None) – The maximum number of documents to return. +- **top_k** (int | None) – The maximum number of documents to return. If 0, no documents are returned. - **scale_score** (bool | None) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant. When `False`, uses raw similarity scores. @@ -393,7 +393,8 @@ Run the InMemoryBM25Retriever on the given input data. **Raises:** -- ValueError – If the specified DocumentStore is not found or is not a InMemoryDocumentStore instance. +- ValueError – If the specified DocumentStore is not found or is not a InMemoryDocumentStore instance, + or if `top_k` is negative. #### run_async @@ -412,7 +413,7 @@ Run the InMemoryBM25Retriever on the given input data. - **query** (str) – The query string for the Retriever. - **filters** (dict\[str, Any\] | None) – A dictionary with filters to narrow down the search space when retrieving documents. -- **top_k** (int | None) – The maximum number of documents to return. +- **top_k** (int | None) – The maximum number of documents to return. If 0, no documents are returned. - **scale_score** (bool | None) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant. When `False`, uses raw similarity scores. @@ -422,7 +423,8 @@ Run the InMemoryBM25Retriever on the given input data. **Raises:** -- ValueError – If the specified DocumentStore is not found or is not a InMemoryDocumentStore instance. +- ValueError – If the specified DocumentStore is not found or is not a InMemoryDocumentStore instance, + or if `top_k` is negative. ## in_memory/embedding_retriever @@ -483,7 +485,7 @@ Create the InMemoryEmbeddingRetriever component. - **document_store** (InMemoryDocumentStore) – An instance of InMemoryDocumentStore where the retriever should search for relevant documents. - **filters** (dict\[str, Any\] | None) – A dictionary with filters to narrow down the retriever's search space in the document store. -- **top_k** (int) – The maximum number of documents to retrieve. +- **top_k** (int) – The maximum number of documents to retrieve. Must be greater than 0. - **scale_score** (bool) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant. When `False`, uses raw similarity scores. - **return_embedding** (bool) – When `True`, returns the embedding of the retrieved documents. @@ -497,7 +499,7 @@ Create the InMemoryEmbeddingRetriever component. **Raises:** - TypeError – If the document_store is not an instance of InMemoryDocumentStore. -- ValueError – If the specified top_k is not > 0. +- ValueError – If `top_k` is not greater than 0. #### to_dict @@ -545,7 +547,7 @@ Run the InMemoryEmbeddingRetriever on the given input data. - **query_embedding** (list\[float\]) – Embedding of the query. - **filters** (dict\[str, Any\] | None) – A dictionary with filters to narrow down the search space when retrieving documents. -- **top_k** (int | None) – The maximum number of documents to return. +- **top_k** (int | None) – The maximum number of documents to return. If 0, no documents are returned. - **scale_score** (bool | None) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant. When `False`, uses raw similarity scores. - **return_embedding** (bool | None) – When `True`, returns the embedding of the retrieved documents. @@ -557,7 +559,8 @@ Run the InMemoryEmbeddingRetriever on the given input data. **Raises:** -- ValueError – If the specified DocumentStore is not found or is not an InMemoryDocumentStore instance. +- ValueError – If the specified DocumentStore is not found or is not an InMemoryDocumentStore instance, + or if `top_k` is negative. #### run_async @@ -577,7 +580,7 @@ Run the InMemoryEmbeddingRetriever on the given input data. - **query_embedding** (list\[float\]) – Embedding of the query. - **filters** (dict\[str, Any\] | None) – A dictionary with filters to narrow down the search space when retrieving documents. -- **top_k** (int | None) – The maximum number of documents to return. +- **top_k** (int | None) – The maximum number of documents to return. If 0, no documents are returned. - **scale_score** (bool | None) – When `True`, scales the score of retrieved documents to a range of 0 to 1, where 1 means extremely relevant. When `False`, uses raw similarity scores. - **return_embedding** (bool | None) – When `True`, returns the embedding of the retrieved documents. @@ -589,7 +592,8 @@ Run the InMemoryEmbeddingRetriever on the given input data. **Raises:** -- ValueError – If the specified DocumentStore is not found or is not an InMemoryDocumentStore instance. +- ValueError – If the specified DocumentStore is not found or is not an InMemoryDocumentStore instance, + or if `top_k` is negative. ## multi_query_embedding_retriever @@ -1031,6 +1035,10 @@ Create the MultiRetriever component. - `concatenate`: Combines all results into a single list and deduplicates. - `reciprocal_rank_fusion`: Deduplicates and assigns scores based on reciprocal rank fusion. +**Raises:** + +- ValueError – If `top_k` or `top_k_per_retriever` is set and is not greater than 0. + #### warm_up ```python @@ -1084,11 +1092,12 @@ Runs retrievers in parallel on the given query and returns deduplicated results. - **filters** (dict\[str, Any\] | None) – Filters to apply. Defaults to the value set at initialization. - **top_k_per_retriever** (int | None) – The maximum number of documents to return per retriever. When set, this will override the `top_k` parameter for each retriever. If None, the `top_k` parameter set for retrievers will be used. - Defaults to the value set at initialization. + If 0, no documents are returned. Defaults to the value set at initialization. - **top_k** (int | None) – The maximum number of documents to return overall, extracted from the combined results of all retrievers. When set, the results are always merged using reciprocal rank fusion (regardless of `join_mode`) so that the combined list has a consistent global ranking before it is truncated to - `top_k`. If None, all results are returned. Defaults to the value set at initialization. + `top_k`. If None, all results are returned. If 0, no documents are returned. + Defaults to the value set at initialization. - **active_retrievers** (list\[str\] | None) – Names of retrievers to run. Defaults to all. Must match keys in the `retrievers` dictionary. **Returns:** @@ -1098,7 +1107,8 @@ Runs retrievers in parallel on the given query and returns deduplicated results. **Raises:** -- ValueError – If any name in `active_retrievers` does not match a retriever name. +- ValueError – If any name in `active_retrievers` does not match a retriever name, + or if the resolved `top_k` or `top_k_per_retriever` is negative. #### run_async @@ -1123,11 +1133,12 @@ Uses each retriever's `run_async` method if available, otherwise runs `run` in a - **filters** (dict\[str, Any\] | None) – Filters to apply. Defaults to the value set at initialization. - **top_k_per_retriever** (int | None) – The maximum number of documents to return per retriever. When set, this will override the `top_k` parameter for each retriever. If None, the `top_k` parameter set for retrievers will be used. - Defaults to the value set at initialization. + If 0, no documents are returned. Defaults to the value set at initialization. - **top_k** (int | None) – The maximum number of documents to return overall, extracted from the combined results of all retrievers. When set, the results are always merged using reciprocal rank fusion (regardless of `join_mode`) so that the combined list has a consistent global ranking before it is truncated to - `top_k`. If None, all results are returned. Defaults to the value set at initialization. + `top_k`. If None, all results are returned. If 0, no documents are returned. + Defaults to the value set at initialization. - **active_retrievers** (list\[str\] | None) – Names of retrievers to run. Defaults to all. Must match keys in the `retrievers` dictionary. **Returns:** @@ -1137,7 +1148,8 @@ Uses each retriever's `run_async` method if available, otherwise runs `run` in a **Raises:** -- ValueError – If any name in `active_retrievers` does not match a retriever name. +- ValueError – If any name in `active_retrievers` does not match a retriever name, + or if the resolved `top_k` or `top_k_per_retriever` is negative. #### to_dict