Why Rate Limiting Fails in Production
Rate limiting is one of those controls that is trivially implemented in a single-instance local environment and genuinely difficult to get right in production. The failure modes are consistent: engineers pick the cheapest algorithm without understanding its boundary burst behaviour, choose the wrong identity (IP address instead of authenticated principal), implement counting locally in each gateway node instead of in a shared store, and then discover the actual protection is a fraction of what the configuration claims.
The distributed problem is the root cause. A Kong cluster with 6 pods, each counting requests locally, provides each consumer with 6 times the configured limit — not the limit. This is not a Kong bug; it is the correct behaviour of the local policy. To provide a genuine cluster-wide limit, all gateway nodes must consult a shared counter store — Redis in production, with the associated operational overhead that implies.
The four canonical algorithms each make a different trade-off between accuracy, memory, and burst behaviour. Understanding the trade-off before choosing one is the difference between a rate limiter that protects the downstream service and one that provides a false sense of control.
Algorithm Comparison
Four algorithms cover the vast majority of production rate-limiting use cases. The choice is not primarily about accuracy — all four can be made accurate enough — but about burst behaviour and operational cost.
| Algorithm | Accuracy | Burst handling | Redis cost | SAMA suitability |
|---|---|---|---|---|
| Fixed window | Low — up to 2× burst at window boundary | Allows 2× limit in burst at reset boundary (e.g. 200 req in 2 s when limit is 100/min) | Single INCR + EXPIRE per request; lowest cost | Unsuitable for payment initiation APIs — SAMA expects consistent throttling |
| Sliding window | High — consistent at all points in the window | No boundary burst; rate is smoothly enforced | ZADD + ZREMRANGEBYSCORE + ZCARD per request; moderate cost | Preferred for PIS and AIS APIs; consistent with SAMA quota audit expectations |
| Token bucket | High for steady state; allows burst up to bucket capacity | Intentional burst support: bucket fills during idle periods and drains under load | Two HSET operations + GET per request; moderate cost | Suitable for AIS read APIs where burst is expected; not for PIS without burst cap |
| Leaky bucket | High for egress smoothing; doesn’t count inbound rate | Queues excess requests; smooths egress rate regardless of inbound spikes | Queue data structure in Redis; highest cost; requires timeout management | Useful for smoothing outbound calls to core banking; rarely correct for inbound API rate limiting |
The practical recommendation for a regulated bank running SAMA Open Banking APIs: sliding window for payment initiation (PIS) and account information (AIS) APIs, token bucket for internal BFF-to-gateway APIs where developers expect burst tolerance, and fixed window only for non-regulated internal monitoring endpoints where accuracy does not matter.
Fixed window rate limiting is cheap but has a well-known boundary burst vulnerability: a consumer can send 100 requests in the last second of window N and 100 requests in the first second of window N+1, effectively sending 200 requests in two seconds against a 100/minute limit. For PIS APIs where each request may initiate a payment, this doubles the maximum throughput to SARIE or SWIFT in a burst, potentially causing downstream queue saturation. Sliding window eliminates this by counting requests over a true rolling window: the consumer sees exactly 100 allowed requests per 60-second period, regardless of when in the period they fall.
Redis-Backed Distributed Counters
A Redis-backed sliding window counter uses a sorted set where each member is a unique request ID (UUID or timestamp with nanosecond precision to avoid collisions) and the score is the Unix timestamp in milliseconds. To count requests in the current window, remove all members older than now - window_ms and count the remaining members. The entire sequence must be atomic to avoid race conditions between the ZREMRANGEBYSCORE (clean) and ZCARD (count) operations.
Atomicity is achieved through a Redis Lua script executed with EVAL or EVALSHA. Lua scripts execute as a single Redis command and cannot be interrupted. Kong’s rate-limiting-advanced plugin implements this pattern natively when strategy: redis is configured; the script below is the equivalent for direct Redis clients or custom gateway integrations.
-- KEYS[1]: rate limit key, e.g. "rl:pis:{jwt_sub}:60000"
-- ARGV[1]: current timestamp in milliseconds
-- ARGV[2]: window size in milliseconds (e.g. 60000 for 1 minute)
-- ARGV[3]: request limit (e.g. 100)
-- ARGV[4]: unique request ID (UUID)
-- Returns: {allowed (0/1), current_count, ttl_ms}
local key = KEYS[1]
local now = tonumber(ARGV[1])
local window_ms = tonumber(ARGV[2])
local limit = tonumber(ARGV[3])
local req_id = ARGV[4]
local cutoff = now - window_ms
-- Remove expired entries (older than the window)
redis.call('ZREMRANGEBYSCORE', key, '-inf', cutoff)
-- Count remaining entries in the window
local count = redis.call('ZCARD', key)
if count < limit then
-- Add this request to the window
redis.call('ZADD', key, now, req_id)
-- Set key TTL to window duration (auto-cleanup)
redis.call('PEXPIRE', key, window_ms)
return {1, count + 1, window_ms} -- allowed
else
-- Get TTL to inform Retry-After header
local oldest = redis.call('ZRANGE', key, 0, 0, 'WITHSCORES')
local reset_in = window_ms - (now - tonumber(oldest[2]))
return {0, count, reset_in} -- rejected, ms until oldest entry expires
end
Redis Cluster key sharding for rate limiting requires careful key design. The default Redis Cluster shards by hash of the entire key, which means two rate limit keys for the same consumer (rl:{sub}:pis and rl:{sub}:ais) will land on different slots and cannot be accessed in the same Lua transaction. Use hash tags to force co-location: rl:{sub}:pis and rl:{sub}:ais both hash on sub, ensuring all rate limit keys for one consumer land on the same slot and can be atomically evaluated together if a multi-key script is needed.
Kong Rate Limiting Configuration
Kong’s rate-limiting-advanced plugin (Kong Enterprise) provides sliding window counting in Redis out of the box, eliminating the need to implement the Lua script above directly. The key configuration decisions are: which limit_by identity to use, which window_type (sliding vs fixed), which strategy (redis vs local vs cluster), and how to structure per-consumer and per-route policies.
Policy hierarchy in Kong: a plugin configured at the route level overrides one at the service level, which overrides one at the global level. A consumer-specific plugin configuration overrides route-level. This hierarchy allows you to set a base rate limit on the service, tighten it on specific high-sensitivity routes (payment initiation), and give premium consumer groups a higher quota.
# Base rate limit on the payment initiation route (PIS)
name: rate-limiting-advanced
route: pis-initiate-payment
config:
limit: [100] # 100 requests per window
window_size: [60] # 60-second window (1 minute)
window_type: sliding
limit_by: consumer # JWT sub resolved to Kong consumer via JWT plugin
strategy: redis # cluster-wide distributed counter
redis:
cluster_addresses:
- "redis-0.redis-headless.infra.svc:6379"
- "redis-1.redis-headless.infra.svc:6379"
- "redis-2.redis-headless.infra.svc:6379"
database: 1
timeout: 2000 # 2 s Redis timeout; fall back to local on timeout
ssl: true
ssl_verify: true
hide_client_headers: false
# Headers injected on every response:
# X-RateLimit-Limit-60: 100
# X-RateLimit-Remaining-60: N
# X-RateLimit-Reset-60: {epoch}
# Retry-After: {seconds} (on 429 only)
error_code: 429
error_message: "Rate limit exceeded. See Retry-After header."
---
# Account information route (AIS) — separate quota class per SAMA
name: rate-limiting-advanced
route: ais-account-info
config:
limit: [60, 500] # 60/min AND 500/hour — multi-window enforcement
window_size: [60, 3600]
window_type: sliding
limit_by: consumer
strategy: redis
---
# Premium PSP consumer group override — higher quota approved by SAMA sandbox
name: rate-limiting-advanced
route: pis-initiate-payment
consumer_group: premium-psp
config:
limit: [300]
window_size: [60]
window_type: sliding
strategy: redis
IBM API Connect Rate Plans
IBM API Connect enforces rate limiting through Rate Plans attached to Products. A Product groups one or more APIs and defines the terms of access (who can subscribe, at what quota). Rate plans set per-plan limits that apply to all consumers subscribed to that plan, with per-application and per-operation overrides possible at the plan level.
API Connect’s rate plan model maps naturally to SAMA Open Banking’s subscription tier model: a PSP subscribes to the “Open Banking Standard” plan, which enforces the SAMA-mandated quotas. Premium PSPs with a separate agreement can be subscribed to a “Premium PSP” plan with higher limits, without any code change to the API itself.
product: 1.0.0
info:
title: SAIB Open Banking APIs
name: saib-open-banking
version: 1.3.0
apis:
payment-initiation:
$ref: "./apis/pis-api.yaml"
account-information:
$ref: "./apis/ais-api.yaml"
plans:
open-banking-standard:
title: Open Banking Standard
description: SAMA-mandated quotas for licensed PSPs and TPs
approval: true # subscriptions require manual approval
rate-limits:
default:
value: 100
unit: minute
hard-limit: true # reject at limit; no grace
burst-limits:
default:
value: 20
unit: second
apis:
payment-initiation:
rate-limits:
pis-initiate:
value: 100
unit: minute # SAMA PIS quota: 100 req/min per PSP
hard-limit: true
operations:
createPayment:
rate-limits:
per-operation:
value: 50
unit: minute # tighter limit on the mutation endpoint
account-information:
rate-limits:
ais-query:
value: 60
unit: minute # SAMA AIS quota: 60 req/min per app
hard-limit: true
premium-psp:
title: Premium PSP
description: Higher quotas for PSPs with separate SAMA approval
approval: true
rate-limits:
default:
value: 300
unit: minute
hard-limit: true
SAMA Regulatory Context
SAMA’s Open Banking Framework defines quota classes per API type and per participant tier. These are not suggested defaults — they are regulatory limits that licensed PSPs and Third-Party Providers (TPPs) are bound by, and that the bank must enforce as the Account Servicing Payment Service Provider (ASPSP).
The applicable quotas as of the current SAMA Open Banking Technical Standards (v3.0.2):
- Payment Initiation Service (PIS) APIs — 100 requests per minute per PSP. “Per PSP” means per PSP organisation, identified by the
x-fapi-financial-idheader resolved to a SAMA-registered PSP identifier — not per application or per end-user session. A PSP with 10 applications gets 100 req/min total, not 1 000. - Account Information Service (AIS) APIs — 60 requests per minute per TPP-application combination. AIS quotas are per-application (identified by the
client_idin the OAuth token) rather than per-organisation, because different AIS applications serve different customer consent contexts. - Confirmation of Funds (CoF) — not separately quota-controlled in the current SAMA standard; treated as PIS quota class.
- SAMA sandbox environment — quotas are relaxed: 500 req/min for PIS, 300 req/min for AIS, with no hard-limit enforcement (soft-limit only; 429 responses are informational and requests are still processed). This is intentional to allow PSPs to load-test integrations without hitting production limits. Do not carry sandbox quota values into the production configuration.
Audit requirements: every 429 response on a SAMA Open Banking API must be logged with the PSP identifier, the API endpoint, the count at rejection time, and the window reset timestamp. SAMA examination teams have requested 90-day retention of throttle event logs in previous reviews. Configure Kong’s file-log or http-log plugin to emit these events to the SIEM separately from the standard access log.
Adaptive Rate Limiting
Static rate limits protect the gateway but not necessarily the backend. A backend service with a p99 latency of 800 ms and a rate limit of 100 req/min has a maximum concurrent request depth of approximately 1.3 — under any burst, the backend will accumulate queued requests that exceed its processing capacity. Rate limiting must be combined with concurrency limits and latency-based backpressure to protect the full request path.
- Concurrency limits (semaphore). Kong’s
response-ratelimitingplugin can count concurrently active requests (requests initiated but not yet responded to) in addition to request rate. For synchronous payment execution backends, a concurrency limit of 50 is a practical ceiling that prevents the backend thread pool from exhausting while the rate limiter still allows traffic. - Latency-based backpressure. When the upstream service’s p99 response time exceeds a threshold (e.g. 2× the SLO target), the rate limiter should reduce the allowed rate by a factor. This is not native to Kong; implement it as a Lua plugin that reads the upstream latency metric from Prometheus (via an API call) and dynamically adjusts the rate limit key’s effective ceiling.
- Priority lanes for RTGS vs SWIFT. Not all payment traffic has equal urgency. SARIE RTGS payments (same-day settlement) have a higher regulatory priority than SWIFT correspondent payments. Implement separate rate limit keys and Redis namespaces per payment rail; in a capacity constraint scenario, apply stricter limits to the lower-priority SWIFT lane first.
- Circuit breaker integration. When the circuit breaker opens on the payments backend (Kong’s upstream health check detects consecutive failures), reduce the rate limit to near-zero immediately — not to let more traffic hammer a degraded backend. A half-open circuit that still allows 100 req/min will fail those 100 requests and potentially worsen the recovery time.
Rate Limit Bypass Prevention
Rate limiting is only as strong as the identity it counts against. The most common bypass patterns in banking API contexts all exploit identity ambiguity — the rate limiter is counting against something other than the true consumer.
- IP rotation. A consumer uses a pool of IP addresses (NAT gateway rotation, residential proxies) to spread requests across multiple IP-based rate limit keys. Mitigation: never rate-limit by source IP alone for authenticated APIs. Always rate-limit by authenticated principal (JWT
suborclient_id). IP-based limits are a supplementary layer to prevent unauthenticated abuse, not the primary control. - API key rotation. A consumer with multiple API keys rotates through them to multiply their effective quota. Mitigation: in APIC, rate plans apply at the subscription level, not the credential level. A PSP with 5 API key credentials on one subscription still hits the plan quota shared across all credentials.
- JWT
subas canonical identity. For OAuth 2.0 flows where the consumer can rotateclient_idby registering multiple applications, rate-limit onsub(the end-user identity from the IdP) combined withclient_id, notclient_idalone. For PIS APIs, SAMA’s PSP identifier (x-fapi-financial-idheader) is the authoritative identity for quota enforcement; validate this header against the SAMA PSP registry on every request. - Distributed shadow counters. For high-value API endpoints, implement a secondary rate counter in the application layer (not just the gateway) using the same Redis Lua pattern. This catches bypass attempts that route around the gateway through internal service-to-service paths — a known attack vector when internal APIs are not properly protected by NetworkPolicy.
Corporate banking clients and SAMA-licensed TPPs almost always connect through a NAT gateway or a shared enterprise internet egress. A single IP address may represent thousands of legitimate users from the same organisation. Rate-limiting by IP on a shared NAT will throttle an entire corporate client when any one of their sessions is busy, without reducing load from the actual problem consumer. Use the authenticated JWT sub or SAMA PSP identifier as the rate limit key for all authenticated endpoints; reserve IP-based limits for pre-authentication DDoS mitigation only.
Common Pitfalls
- Sliding window cost at high cardinality. A sorted set per consumer per API class per window accumulates quickly. At 50 000 active consumers, each with 3 API classes and a 60-second window holding up to 100 entries, the Redis memory footprint is approximately 50 000 × 3 × 100 × ~80 bytes = 1.2 GB. This is manageable but must be planned; unexpected cardinality growth (e.g. per-session rather than per-consumer rate limit keys) can exhaust Redis memory quickly.
- Redis split-brain behaviour. In a Redis Cluster, a network partition can cause nodes to accept writes independently. Rate limit keys on both sides of the partition accumulate counts separately; when the partition heals, the merged count may be below the actual request count, temporarily allowing more traffic than configured. This is a known CAP trade-off of Redis Cluster in AP mode. For strictly regulated limits (PIS quotas), evaluate Redis Sentinel with a primary-only write policy to sacrifice availability for consistency during partitions.
- Retry storms amplifying load. When rate limiting returns 429, consumers retry. If retry is not governed by an exponential backoff with jitter, the retry burst hits at the rate limit reset boundary, exceeding the limit again immediately. The
Retry-Afterheader tells the consumer when to retry; the consumer must respect it. Include retry policy guidance in your API documentation and reject consumers at the gateway whose User-Agent identifies known misconfigured retry clients.
Redis key expiry is not guaranteed to fire at exactly the configured TTL — it fires when the key is next accessed, or during the background expiry cycle. Under high load, expired rate limit keys may persist briefly past their expiry time, and the ZREMRANGEBYSCORE cleanup inside the Lua script may have a few milliseconds of lag. In practice this gap is under 50 ms and is not operationally significant. Where it does matter: if you set a rate limit of 100/min and the consumer hits exactly 100 requests in the last 10 ms of a window, the next window’s first request may incorrectly see a count of 101 before cleanup runs. Accept this as a design trade-off of the sliding window approach; it under-protects by at most one request per window boundary.
Production Checklist
Steps to roll out rate limiting without breaking existing consumers
- Measure before you limit. Deploy Kong with logging-only mode first (set rate limit to an unrealistically high value like 100 000 req/min). Collect actual per-consumer request rates from Kong’s access log for 2–4 weeks. The p99 actual rate per consumer across the collection period becomes your baseline for choosing the production limit.
- Set limits above the measured baseline. Start at 2× the p99 observed rate for non-SAMA APIs, and at the SAMA-mandated limit for regulated APIs. This ensures the limit does not break any currently working consumer while still providing protection against runaway or malicious traffic.
- Enable response headers. Set
hide_client_headers: falseand ensure consumers receiveX-RateLimit-RemainingandX-RateLimit-Resetheaders on every response. Give consumers at least 4 weeks to instrument their clients to read and respect these headers before enforcing hard limits. - Alert on 429 rate before enforcing. Create a Grafana alert on Kong’s 429 response rate per consumer. Notify affected consumer teams at 10% of their limit, then 50%, then enforce at 100%. This gives teams time to adapt their retry logic before they experience hard throttling.
- Switch to Redis sliding window. Migrate from
strategy: local(or no rate limiting) tostrategy: rediswithwindow_type: slidingin a maintenance window, with Redis Cluster deployed and tested. Validate that the cluster-wide limit is equivalent to the per-node limit you measured during baseline collection. - Enforce hard limits for SAMA APIs. Only after steps 1–5 are complete, set
hard-limit: truefor PIS and AIS routes. Inform all registered PSPs and TPPs via the SAMA Open Banking developer portal before the enforcement date. - Establish a quota breach review process. After hard limits are in production, any PSP hitting their limit repeatedly should trigger a review — are they misconfigured, have they grown beyond their plan, or are they attempting bypass? Build a Grafana dashboard that shows per-PSP 429 frequency over time and route it to the API governance team for monthly review.
Production checklist
- Sliding window algorithm configured for all SAMA PIS and AIS routes; fixed window only for non-regulated endpoints.
- Redis Cluster deployed with at least 3 nodes; TLS enabled on all cluster connections from Kong.
- Rate limit identity uses JWT
subor SAMAx-fapi-financial-id; IP-based limits applied as supplementary DDoS layer only. - Multi-window enforcement for AIS: 60/min AND 500/hour configured as separate window entries in the rate-limiting plugin.
- APIC rate plans reflect SAMA quotas: 100/min for PIS, 60/min for AIS;
hard-limit: trueon both. - SAMA sandbox environment configured with soft limits only; hard limits never carried over to production config.
Retry-AfterandX-RateLimit-Resetheaders present on every 429 response.- Throttle event log (PSP ID, endpoint, count, timestamp) retained for 90 days in SIEM per SAMA audit requirement.
- Concurrency limits set on payment execution backend; latency-based backpressure evaluated for high-volume routes.
- Redis Cluster key design uses hash tags to co-locate all rate limit keys for a single consumer on one slot.
- Retry storm mitigation documented in consumer API guide; per-consumer 429 rate alerting configured in Grafana.