Skip to content

feat: add client-side rate limiting - #262

Closed
GenZ-CODER-X wants to merge 3 commits into
codeforstartups:developmentfrom
GenZ-CODER-X:feat/client-rate-limiting
Closed

GenZ-CODER-X wants to merge 3 commits into
codeforstartups:developmentfrom
GenZ-CODER-X:feat/client-rate-limiting

Conversation

@GenZ-CODER-X

Copy link
Copy Markdown
Contributor

Description

Adds client-side token-bucket rate limiting to S3 Vectors requests to help avoid exceeding AWS request-rate limits.

The limiter is configurable independently for write and query operations and is disabled by default for backward compatibility.

Related issue

Fixes #51

Changes

  • Added thread-safe TokenBucket rate limiter with configurable RPS and capacity.
  • Added put_rps and query_rps configuration with validation.
  • Applied rate limiting to PutVectors batches and paginated QueryVectors requests.
  • Added unit tests for token consumption, refill, waiting behavior, capacity limits, request wiring, and paginator exhaustion.

Testing

  • Tests pass locally
  • Ruff checks pass
  • Documentation updated, if applicable

Validation performed:

  • uv run --no-sync pytest -q --ignore=tests/test_langchain.py --ignore=tests/test_llamaindex.py
  • Ruff checks passed for all changed files.
  • git diff --check passed.

Checklist

  • My changes are focused and relevant to this pull request.
  • I have added or updated tests where appropriate.
  • I have reviewed my changes for unrelated modifications.
  • I have updated documentation where necessary.

@GenZ-CODER-X

Copy link
Copy Markdown
Contributor Author

Design assumptions / implementation notes

While implementing #51, I made the following design decisions based on the issue requirements:

  • Rate-limit at the S3 Vectors request boundary rather than at Dynavec.upsert() / Dynavec.search(), so every underlying AWS request is controlled consistently.

  • Separate limiters for writes and queries using put_rps and query_rps. This prevents write traffic from consuming the entire request budget and delaying queries.

  • Disabled by default for backward compatibility. When put_rps / query_rps is None, no rate limiting is applied.

  • One token per AWS request/batch. A PutVectors batch consumes one token, while each paginated QueryVectors request consumes one token.

  • Thread-safe token bucket using threading.Lock and time.monotonic(). The lock is released before sleeping so other threads are not blocked while waiting for tokens.

  • Bucket capacity defaults to the configured RPS. This allows a short initial burst while maintaining the configured long-term request rate.

  • Query pagination is rate-limited per page/request while avoiding an extra limiter acquisition after the paginator has reached its final page.

  • Retries remain outside the limiter call so each actual AWS request is rate-limited while preserving the existing retry behavior.

I also added tests covering token consumption, refill, waiting, capacity limits, configuration validation, write request limiting, query-page limiting, and paginator exhaustion.

These are the main implementation assumptions I made for #51. Please let me know if any of these differ from the intended behavior.

@GenZ-CODER-X

Copy link
Copy Markdown
Contributor Author

CI status update

The current CI / typecheck failure appears to originate from the existing
development branch configuration rather than from the changes in this PR.

src/dynavec/integrations/semantic_kernel.py is included in the mypy check,
but semantic-kernel is not included in the typecheck dependency extra.
As a result, mypy cannot resolve semantic_kernel.data.vector and reports
cascading type errors in that integration.

I previously added semantic-kernel>=1.44 to the typecheck extra locally to
verify this, and the typecheck then passed successfully. I have reverted that
change because it is outside the scope of this PR.

From my side, the rate-limiting implementation and tests are passing, and the
Python test and Ruff checks are clear.

Please let me know whether the existing development typecheck configuration
should be fixed separately.

codeforstartups pushed a commit that referenced this pull request Sep 28, 2026
Opt-in TokenBucket limiters (put_rps / query_rps in config) throttle S3
Vectors put/query calls to stay under account/service limits. Rebased onto
current development by the maintainer (its CI typecheck failure was a stale
base, not a code issue); mypy + full suite green.

Co-authored-by: GenZ-CODER-X <GenZ-CODER-X@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@codeforstartups

Copy link
Copy Markdown
Owner

Landed in development as 48428ae — thanks @GenZ-CODER-X! 🙌 Clean opt-in TokenBucket rate limiting (put_rps / query_rps) to keep S3 Vectors calls under service/account limits.\n\nHeads-up on the red typecheck: it was a stale base, not your code — your branch predated some mypy config on development. I rebased onto current development (0 conflicts) and it is green: mypy Success, ruff clean, full suite (758) passing. Landed with your authorship; closing as merged-manually. (Tip: your commits have an unset git email — git config user.email — so I attributed via your GitHub noreply address.)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Client-side rate limiting / backpressure

2 participants