From b6f4b48cce409db2c52509c978791d8346e88f55 Mon Sep 17 00:00:00 2001 From: davidsbatista <7937824+davidsbatista@users.noreply.github.com> Date: Fri, 25 Sep 2026 14:30:19 +0000 Subject: [PATCH] Sync Core Integrations API reference (astra) on Docusaurus --- .../reference/integrations-api/astra.md | 361 ++++++- .../version-2.18/integrations-api/astra.md | 361 ++++++- .../version-2.19/integrations-api/astra.md | 361 ++++++- .../version-2.20/integrations-api/astra.md | 361 ++++++- .../version-2.21/integrations-api/astra.md | 361 ++++++- .../version-2.22/integrations-api/astra.md | 361 ++++++- .../version-2.23/integrations-api/astra.md | 361 ++++++- .../version-2.24/integrations-api/astra.md | 361 ++++++- .../version-2.25/integrations-api/astra.md | 361 ++++++- .../version-2.26/integrations-api/astra.md | 361 ++++++- .../version-2.27/integrations-api/astra.md | 361 ++++++- .../version-2.28/integrations-api/astra.md | 361 ++++++- .../version-2.29/integrations-api/astra.md | 361 ++++++- .../version-2.30/integrations-api/astra.md | 361 ++++++- .../version-2.31/integrations-api/astra.md | 361 ++++++- .../version-3.0/integrations-api/astra.md | 361 ++++++- .../version-3.1/integrations-api/astra.md | 361 ++++++- .../integrations-api/astra.md | 880 ++++++++++++++++++ .../version-3.2/integrations-api/astra.md | 361 ++++++- 19 files changed, 7270 insertions(+), 108 deletions(-) create mode 100644 docs-website/reference_versioned_docs/version-3.2-unstable/integrations-api/astra.md diff --git a/docs-website/reference/integrations-api/astra.md b/docs-website/reference/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference/integrations-api/astra.md +++ b/docs-website/reference/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.18/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.18/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.18/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.18/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.19/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.19/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.19/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.19/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.20/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.20/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.20/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.20/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.21/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.21/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.21/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.21/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.22/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.22/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.22/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.22/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.23/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.23/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.23/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.23/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.24/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.24/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.24/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.24/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.25/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.25/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.25/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.25/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.26/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.26/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.26/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.26/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.27/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.27/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.27/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.27/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.28/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.28/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.28/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.28/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.29/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.29/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.29/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.29/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.30/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.30/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.30/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.30/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-2.31/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-2.31/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-2.31/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-2.31/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-3.0/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-3.0/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-3.0/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-3.0/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-3.1/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-3.1/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-3.1/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-3.1/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError diff --git a/docs-website/reference_versioned_docs/version-3.2-unstable/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-3.2-unstable/integrations-api/astra.md new file mode 100644 index 00000000000..ef41ce88690 --- /dev/null +++ b/docs-website/reference_versioned_docs/version-3.2-unstable/integrations-api/astra.md @@ -0,0 +1,880 @@ +--- +title: "Astra" +id: integrations-astra +description: "Astra integration for Haystack" +slug: "/integrations-astra" +--- + + +## haystack_integrations.components.retrievers.astra.retriever + +### AstraEmbeddingRetriever + +A component for retrieving documents from an AstraDocumentStore. + +Usage example: + +```python +from haystack_integrations.document_stores.astra import AstraDocumentStore +from haystack_integrations.components.retrievers.astra import AstraEmbeddingRetriever + +document_store = AstraDocumentStore( + api_endpoint=api_endpoint, + token=token, + collection_name=collection_name, + duplicates_policy=DuplicatePolicy.SKIP, + embedding_dim=384, +) + +retriever = AstraEmbeddingRetriever(document_store=document_store) +``` + +#### __init__ + +```python +__init__( + document_store: AstraDocumentStore, + filters: dict[str, Any] | None = None, + top_k: int = 10, + filter_policy: str | FilterPolicy = FilterPolicy.REPLACE, +) -> None +``` + +Initialize the AstraEmbeddingRetriever. + +**Parameters:** + +- **document_store** (AstraDocumentStore) – An instance of AstraDocumentStore. +- **filters** (dict\[str, Any\] | None) – a dictionary with filters to narrow down the search space. +- **top_k** (int) – the maximum number of documents to retrieve. +- **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. + +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + +#### run + +```python +run( + query_embedding: list[float], + filters: dict[str, Any] | None = None, + top_k: int | None = None, +) -> dict[str, list[Document]] +``` + +Retrieve documents from the AstraDocumentStore. + +**Parameters:** + +- **query_embedding** (list\[float\]) – floats representing the query embedding +- **filters** (dict\[str, Any\] | None) – Filters applied to the retrieved Documents. The way runtime filters are applied depends on + the `filter_policy` chosen at retriever initialization. See init method docstring for more + details. +- **top_k** (int | None) – the maximum number of documents to retrieve. + +**Returns:** + +- dict\[str, list\[Document\]\] – a dictionary with the following keys: +- `documents`: A list of documents retrieved from the AstraDocumentStore. + +#### run_async + +```python +run_async( + query_embedding: list[float], + filters: dict[str, Any] | None = None, + top_k: int | None = None, +) -> dict[str, list[Document]] +``` + +Retrieve documents from the AstraDocumentStore asynchronously. + +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – floats representing the query embedding +- **filters** (dict\[str, Any\] | None) – Filters applied to the retrieved Documents. The way runtime filters are applied depends on + the `filter_policy` chosen at retriever initialization. See init method docstring for more + details. +- **top_k** (int | None) – the maximum number of documents to retrieve. + +**Returns:** + +- dict\[str, list\[Document\]\] – a dictionary with the following keys: +- `documents`: A list of documents retrieved from the AstraDocumentStore. + +#### to_dict + +```python +to_dict() -> dict[str, Any] +``` + +Serializes the component to a dictionary. + +**Returns:** + +- dict\[str, Any\] – Dictionary with serialized data. + +#### from_dict + +```python +from_dict(data: dict[str, Any]) -> AstraEmbeddingRetriever +``` + +Deserializes the component from a dictionary. + +**Parameters:** + +- **data** (dict\[str, Any\]) – Dictionary to deserialize from. + +**Returns:** + +- AstraEmbeddingRetriever – Deserialized component. + +## haystack_integrations.document_stores.astra.document_store + +### AstraDocumentStore + +An AstraDocumentStore document store for Haystack. + +Example Usage: + +```python +from haystack_integrations.document_stores.astra import AstraDocumentStore + +document_store = AstraDocumentStore( + api_endpoint=api_endpoint, + token=token, + collection_name=collection_name, + duplicates_policy=DuplicatePolicy.SKIP, + embedding_dim=384, +) +``` + +#### __init__ + +```python +__init__( + api_endpoint: Secret = Secret.from_env_var("ASTRA_DB_API_ENDPOINT"), + token: Secret = Secret.from_env_var("ASTRA_DB_APPLICATION_TOKEN"), + collection_name: str = "documents", + embedding_dimension: int = 768, + duplicates_policy: DuplicatePolicy = DuplicatePolicy.NONE, + similarity: str = "cosine", + namespace: str | None = None, +) -> None +``` + +The connection to Astra DB is established and managed through the JSON API. + +The required credentials (api endpoint and application token) can be generated +through the UI by clicking and the connect tab, and then selecting JSON API and +Generate Configuration. + +**Parameters:** + +- **api_endpoint** (Secret) – the Astra DB API endpoint. +- **token** (Secret) – the Astra DB application token. +- **collection_name** (str) – the current collection in the keyspace in the current Astra DB. +- **embedding_dimension** (int) – dimension of embedding vector. +- **duplicates_policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. + +**Raises:** + +- ValueError – if the API endpoint or token is not set. + +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async + +```python +close_async() -> None +``` + +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. + +#### from_dict + +```python +from_dict(data: dict[str, Any]) -> AstraDocumentStore +``` + +Deserializes the component from a dictionary. + +**Parameters:** + +- **data** (dict\[str, Any\]) – Dictionary to deserialize from. + +**Returns:** + +- AstraDocumentStore – Deserialized component. + +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + +#### to_dict + +```python +to_dict() -> dict[str, Any] +``` + +Serializes the component to a dictionary. + +**Returns:** + +- dict\[str, Any\] – Dictionary with serialized data. + +#### write_documents + +```python +write_documents( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + +#### count_documents + +```python +count_documents() -> int +``` + +Counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + +#### filter_documents + +```python +filter_documents(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + +#### get_documents_by_id + +```python +get_documents_by_id(ids: list[str]) -> list[Document] +``` + +Gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + +#### get_document_by_id + +```python +get_document_by_id(document_id: str) -> Document +``` + +Gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + +#### search + +```python +search( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Perform a search for a list of queries. + +**Parameters:** + +- **query_embedding** (list\[float\]) – a list of query embeddings. +- **top_k** (int) – the number of results to return. +- **filters** (dict\[str, Any\] | None) – filters to apply during search. + +**Returns:** + +- list\[Document\] – matching documents. + +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + +#### delete_documents + +```python +delete_documents(document_ids: list[str]) -> None +``` + +Deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + +#### delete_all_documents + +```python +delete_all_documents(*, recreate_index: bool = False) -> None +``` + +Deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_by_filter + +```python +delete_by_filter(filters: dict[str, Any]) -> int +``` + +Deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + +#### update_by_filter + +```python +update_by_filter(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + +#### count_documents_by_filter + +```python +count_documents_by_filter(filters: dict[str, Any]) -> int +``` + +Applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + +#### count_unique_metadata_by_filter + +```python +count_unique_metadata_by_filter( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Applies a filter selecting documents and counts the unique values for each meta field of the matched documents. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### get_metadata_fields_info + +```python +get_metadata_fields_info() -> dict[str, dict[str, str]] +``` + +Returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + +#### get_metadata_field_min_max + +```python +get_metadata_field_min_max(metadata_field: str) -> dict[str, Any] +``` + +For a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + +#### get_metadata_field_unique_values + +```python +get_metadata_field_unique_values( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Retrieves unique values for a field matching a search term or all possible values if no search term is given. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + +## haystack_integrations.document_stores.astra.errors + +### AstraDocumentStoreError + +Bases: DocumentStoreError + +Parent class for all AstraDocumentStore errors. + +### AstraDocumentStoreFilterError + +Bases: FilterError + +Raised when an invalid filter is passed to AstraDocumentStore. + +### AstraDocumentStoreConfigError + +Bases: AstraDocumentStoreError + +Raised when an invalid configuration is passed to AstraDocumentStore. diff --git a/docs-website/reference_versioned_docs/version-3.2/integrations-api/astra.md b/docs-website/reference_versioned_docs/version-3.2/integrations-api/astra.md index 8c4f892ec57..ef41ce88690 100644 --- a/docs-website/reference_versioned_docs/version-3.2/integrations-api/astra.md +++ b/docs-website/reference_versioned_docs/version-3.2/integrations-api/astra.md @@ -49,6 +49,22 @@ Initialize the AstraEmbeddingRetriever. - **top_k** (int) – the maximum number of documents to retrieve. - **filter_policy** (str | FilterPolicy) – Policy to determine how filters are applied. +#### close + +```python +close() -> None +``` + +Release the synchronous resources of the underlying Document Store. + +#### close_async + +```python +close_async() -> None +``` + +Release the asynchronous resources of the underlying Document Store. + #### run ```python @@ -86,7 +102,8 @@ run_async( Retrieve documents from the AstraDocumentStore asynchronously. -Runs the sync search in a thread pool to avoid blocking the event loop. +Uses the native async Astra DB API with a reusable connection. Call `close_async()` (or +`Pipeline.close_async()`) when finished, before closing the event loop. **Parameters:** @@ -182,19 +199,36 @@ Generate Configuration. - `DuplicatePolicy.SKIP`: if a Document with the same ID already exists, it is skipped and not written. - `DuplicatePolicy.OVERWRITE`: if a Document with the same ID already exists, it is overwritten. - `DuplicatePolicy.FAIL`: if a Document with the same ID already exists, an error is raised. -- **similarity** (str) – the similarity function used to compare document vectors. +- **similarity** (str) – Similarity metric for new collections: `cosine`, `dot_product`, or `euclidean`. + Existing collections retain their configured metric. +- **namespace** (str | None) – The keyspace containing the collection, or the SDK default when omitted. **Raises:** - ValueError – if the API endpoint or token is not set. -#### index +#### close + +```python +close() -> None +``` + +Drop the cached synchronous collection without deleting documents. + +AstraPy 2 exposes no way to release synchronous connections, so this only discards the collection; the next +synchronous operation creates a new one. + +#### close_async ```python -index: AstraClient +close_async() -> None ``` -Return the AstraClient index, initializing it if necessary. +Release the cached async collection connection without deleting documents. + +Call this on the same event loop as the async methods, after they have finished and before +closing the loop. Repeated calls are safe; a later call opens a new connection. Calls on a new +event loop open a new connection automatically. #### from_dict @@ -212,6 +246,10 @@ Deserializes the component from a dictionary. - AstraDocumentStore – Deserialized component. +**Raises:** + +- ValueError – The serialized `duplicates_policy` is not a valid policy name. + #### to_dict ```python @@ -256,6 +294,38 @@ Indexes documents for later queries. - DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. - Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. +#### write_documents_async + +```python +write_documents_async( + documents: list[Document], policy: DuplicatePolicy = DuplicatePolicy.NONE +) -> int +``` + +Asynchronously indexes documents for later queries. + +**Parameters:** + +- **documents** (list\[Document\]) – a list of Haystack Document objects. +- **policy** (DuplicatePolicy) – handle duplicate documents based on DuplicatePolicy parameter options. + Parameter options : (`SKIP`, `OVERWRITE`, `FAIL`, `NONE`) +- `DuplicatePolicy.NONE`: Default policy, If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.SKIP`: If a Document with the same ID already exists, + it is skipped and not written. +- `DuplicatePolicy.OVERWRITE`: If a Document with the same ID already exists, it is overwritten. +- `DuplicatePolicy.FAIL`: If a Document with the same ID already exists, an error is raised. + +**Returns:** + +- int – number of documents written. + +**Raises:** + +- ValueError – if the documents are not of type Document or dict. +- DuplicateDocumentError – if a document with the same ID already exists and policy is set to FAIL. +- Exception – if the document ID is not a string or if `id` and `_id` are both present in the document. + #### count_documents ```python @@ -268,6 +338,18 @@ Counts the number of documents in the document store. - int – the number of documents in the document store. +#### count_documents_async + +```python +count_documents_async() -> int +``` + +Asynchronously counts the number of documents in the document store. + +**Returns:** + +- int – the number of documents in the document store. + #### filter_documents ```python @@ -288,6 +370,26 @@ Returns at most 1000 documents that match the filter. - AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. +#### filter_documents_async + +```python +filter_documents_async(filters: dict[str, Any] | None = None) -> list[Document] +``` + +Asynchronously returns at most 1000 documents that match the filter. + +**Parameters:** + +- **filters** (dict\[str, Any\] | None) – filters to apply. + +**Returns:** + +- list\[Document\] – matching documents. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported by this class. + #### get_documents_by_id ```python @@ -304,6 +406,22 @@ Gets documents by their IDs. - list\[Document\] – the matching documents. +#### get_documents_by_id_async + +```python +get_documents_by_id_async(ids: list[str]) -> list[Document] +``` + +Asynchronously gets documents by their IDs. + +**Parameters:** + +- **ids** (list\[str\]) – the IDs of the documents to retrieve. + +**Returns:** + +- list\[Document\] – the matching documents. + #### get_document_by_id ```python @@ -324,6 +442,26 @@ Gets a document by its ID. - MissingDocumentError – if the document is not found +#### get_document_by_id_async + +```python +get_document_by_id_async(document_id: str) -> Document +``` + +Asynchronously gets a document by its ID. + +**Parameters:** + +- **document_id** (str) – the ID to filter by + +**Returns:** + +- Document – the found document + +**Raises:** + +- MissingDocumentError – if the document is not found + #### search ```python @@ -346,6 +484,31 @@ Perform a search for a list of queries. - list\[Document\] – matching documents. +#### search_async + +```python +search_async( + query_embedding: list[float], + top_k: int, + filters: dict[str, Any] | None = None, +) -> list[Document] +``` + +Search using AstraPy's native async API. + +The collection connection is initialized lazily and reused across searches on the same event loop. +Call `close_async()` when finished, including after failures or cancellation, before closing the loop. + +**Parameters:** + +- **query_embedding** (list\[float\]) – A list of query embeddings. +- **top_k** (int) – The number of results to return. +- **filters** (dict\[str, Any\] | None) – Filters to apply during search. + +**Returns:** + +- list\[Document\] – Matching documents, including embeddings, metadata and similarity scores. + #### delete_documents ```python @@ -362,14 +525,60 @@ Deletes documents from the document store. - MissingDocumentError – if no document was deleted but document IDs were provided. +#### delete_documents_async + +```python +delete_documents_async(document_ids: list[str]) -> None +``` + +Asynchronously deletes documents from the document store. + +**Parameters:** + +- **document_ids** (list\[str\]) – IDs of the documents to delete. + +**Raises:** + +- MissingDocumentError – if no document was deleted but document IDs were provided. + #### delete_all_documents ```python -delete_all_documents() -> None +delete_all_documents(*, recreate_index: bool = False) -> None ``` Deletes all documents from the document store. +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + +#### delete_all_documents_async + +```python +delete_all_documents_async(*, recreate_index: bool = False) -> None +``` + +Asynchronously deletes all documents from the document store. + +**Parameters:** + +- **recreate_index** (bool) – If `True`, drops the collection and recreates it with its current definition (vector + dimension, metric and indexing settings) instead of deleting its documents. Dropping and creating a + collection takes several seconds and is not atomic: if the creation fails, the next operation creates the + collection from this store's settings. + +**Raises:** + +- DocumentStoreError – if the documents could not be deleted or the collection could not be recreated. + #### delete_by_filter ```python @@ -390,6 +599,26 @@ Deletes documents that match the provided filters. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### delete_by_filter_async + +```python +delete_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously deletes documents that match the provided filters. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to delete. + +**Returns:** + +- int – The number of documents deleted. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### update_by_filter ```python @@ -411,6 +640,27 @@ Updates documents that match the provided filters with the given metadata. - AstraDocumentStoreFilterError – if the filter is invalid or not supported. +#### update_by_filter_async + +```python +update_by_filter_async(filters: dict[str, Any], meta: dict[str, Any]) -> int +``` + +Asynchronously updates documents that match the provided filters with the given metadata. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to find documents to update. +- **meta** (dict\[str, Any\]) – The metadata fields to update. This will be merged with existing metadata. + +**Returns:** + +- int – The number of documents updated. + +**Raises:** + +- AstraDocumentStoreFilterError – if the filter is invalid or not supported. + #### count_documents_by_filter ```python @@ -427,6 +677,22 @@ Applies a filter and counts the documents that matched it. - int – The number of documents that match the filter. +#### count_documents_by_filter_async + +```python +count_documents_by_filter_async(filters: dict[str, Any]) -> int +``` + +Asynchronously applies a filter and counts the documents that matched it. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. + +**Returns:** + +- int – The number of documents that match the filter. + #### count_unique_metadata_by_filter ```python @@ -444,6 +710,26 @@ Applies a filter selecting documents and counts the unique values for each meta **Returns:** +- dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique + values. + +#### count_unique_metadata_by_filter_async + +```python +count_unique_metadata_by_filter_async( + filters: dict[str, Any], metadata_fields: list[str] +) -> dict[str, int] +``` + +Asynchronously counts the unique values of each meta field across the documents matching a filter. + +**Parameters:** + +- **filters** (dict\[str, Any\]) – The filters to apply to the document list. +- **metadata_fields** (list\[str\]) – The metadata fields to count unique values for. + +**Returns:** + - dict\[str, int\] – A dictionary where the keys are the metadata field names and the values are the count of unique values. @@ -459,6 +745,18 @@ Returns the metadata fields and the corresponding types. - dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. +#### get_metadata_fields_info_async + +```python +get_metadata_fields_info_async() -> dict[str, dict[str, str]] +``` + +Asynchronously returns the metadata fields and the corresponding types. + +**Returns:** + +- dict\[str, dict\[str, str\]\] – A dictionary mapping field names to dictionaries with a `type` key. + #### get_metadata_field_min_max ```python @@ -475,6 +773,22 @@ For a given metadata field, find its max and min value. - dict\[str, Any\] – A dictionary with `min` and `max`. +#### get_metadata_field_min_max_async + +```python +get_metadata_field_min_max_async(metadata_field: str) -> dict[str, Any] +``` + +Asynchronously, for a given metadata field, find its max and min value. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. + +**Returns:** + +- dict\[str, Any\] – A dictionary with `min` and `max`. + #### get_metadata_field_unique_values ```python @@ -510,6 +824,41 @@ Exception are floats with a fractional part (e.g. `1.5`) are unaffected and roun - tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. +#### get_metadata_field_unique_values_async + +```python +get_metadata_field_unique_values_async( + metadata_field: str, + search_term: str | None = None, + from_: int = 0, + size: int = 10, + filters: dict[str, Any] | None = None, +) -> tuple[list[Any], int] +``` + +Asynchronously retrieves unique values for a field, optionally matching a search term. + +**Note**: values of different types are kept distinct even when they compare equal in Python +(e.g. the int `1`, the bool `True` and the str `"1"` are returned as three separate values), with +one exception. +AstraDB's Data API canonicalizes any whole-number float (e.g. `1.0`) to an int on storage, unconditionally so a +whole-number float is always returned back as an int, never as a float. +Example: 1.0 (float) is sent to storage and comes back a 1 (int) + +Exception are floats with a fractional part (e.g. `1.5`) are unaffected and round-trip normally. + +**Parameters:** + +- **metadata_field** (str) – The metadata field to inspect. +- **search_term** (str | None) – Optional case-insensitive substring search term. +- **from\_** (int) – The starting index for pagination. +- **size** (int) – The number of values to return. +- **filters** (dict\[str, Any\] | None) – Optional filters to restrict the documents considered. + +**Returns:** + +- tuple\[list\[Any\], int\] – A tuple containing the paginated values (in their original type) and the total count. + ## haystack_integrations.document_stores.astra.errors ### AstraDocumentStoreError