Skip to content

Best Practices for Load Shedding in Envoy Gateway #9653

Description

@am36166

Hello,

I'm looking for guidance on implementing load shedding in Envoy Gateway to protect backend services under high load.

In networking, mechanisms such as Random Early Detection (RED) and traffic policing proactively prevent congestion by dropping or limiting traffic before the network becomes saturated. I'm looking for the equivalent approach for HTTP/API traffic using Envoy Gateway.

Specifically, I'd like to know:

Does Envoy Gateway support proactive load shedding based on resource pressure (e.g., request latency, concurrency, queue depth, CPU utilization, or other overload signals)?
Is Envoy's adaptive concurrency filter currently supported and configurable through Envoy Gateway?
What is the recommended way to reject excess requests before backend services become overloaded?
Are there best practices for combining:
Local or global rate limiting
Circuit breakers
Adaptive concurrency
Overload Manager
Kubernetes HPA
If some of these capabilities are not yet exposed by Envoy Gateway, what is the recommended production approach today?

Our goal is not only to enforce rate limits, but to gracefully shed load when the platform approaches its safe operating capacity, ensuring that the system remains responsive instead of allowing latency to grow until services become unavailable.

If there are existing examples, documentation, or recommended configuration patterns for this use case, I would greatly appreciate being pointed to them.

Thank you!

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions