Why Rate Limiting Fails in Production

Rate limiting is one of those controls that is trivially implemented in a single-instance local environment and genuinely difficult to get right in production. The failure modes are consistent: engineers pick the cheapest algorithm without understanding its boundary burst behaviour, choose the wrong identity (IP address instead of authenticated principal), implement counting locally in each gateway node instead of in a shared store, and then discover the actual protection is a fraction of what the configuration claims.

The distributed problem is the root cause. A Kong cluster with 6 pods, each counting requests locally, provides each consumer with 6 times the configured limit — not the limit. This is not a Kong bug; it is the correct behaviour of the local policy. To provide a genuine cluster-wide limit, all gateway nodes must consult a shared counter store — Redis in production, with the associated operational overhead that implies.

The four canonical algorithms each make a different trade-off between accuracy, memory, and burst behaviour. Understanding the trade-off before choosing one is the difference between a rate limiter that protects the downstream service and one that provides a false sense of control.

Algorithm Comparison

Four algorithms cover the vast majority of production rate-limiting use cases. The choice is not primarily about accuracy — all four can be made accurate enough — but about burst behaviour and operational cost.

Algorithm Accuracy Burst handling Redis cost SAMA suitability
Fixed window Low — up to 2× burst at window boundary Allows 2× limit in burst at reset boundary (e.g. 200 req in 2 s when limit is 100/min) Single INCR + EXPIRE per request; lowest cost Unsuitable for payment initiation APIs — SAMA expects consistent throttling
Sliding window High — consistent at all points in the window No boundary burst; rate is smoothly enforced ZADD + ZREMRANGEBYSCORE + ZCARD per request; moderate cost Preferred for PIS and AIS APIs; consistent with SAMA quota audit expectations
Token bucket High for steady state; allows burst up to bucket capacity Intentional burst support: bucket fills during idle periods and drains under load Two HSET operations + GET per request; moderate cost Suitable for AIS read APIs where burst is expected; not for PIS without burst cap
Leaky bucket High for egress smoothing; doesn’t count inbound rate Queues excess requests; smooths egress rate regardless of inbound spikes Queue data structure in Redis; highest cost; requires timeout management Useful for smoothing outbound calls to core banking; rarely correct for inbound API rate limiting

The practical recommendation for a regulated bank running SAMA Open Banking APIs: sliding window for payment initiation (PIS) and account information (AIS) APIs, token bucket for internal BFF-to-gateway APIs where developers expect burst tolerance, and fixed window only for non-regulated internal monitoring endpoints where accuracy does not matter.

Use sliding window for financial APIs

Fixed window rate limiting is cheap but has a well-known boundary burst vulnerability: a consumer can send 100 requests in the last second of window N and 100 requests in the first second of window N+1, effectively sending 200 requests in two seconds against a 100/minute limit. For PIS APIs where each request may initiate a payment, this doubles the maximum throughput to SARIE or SWIFT in a burst, potentially causing downstream queue saturation. Sliding window eliminates this by counting requests over a true rolling window: the consumer sees exactly 100 allowed requests per 60-second period, regardless of when in the period they fall.

Redis-Backed Distributed Counters

A Redis-backed sliding window counter uses a sorted set where each member is a unique request ID (UUID or timestamp with nanosecond precision to avoid collisions) and the score is the Unix timestamp in milliseconds. To count requests in the current window, remove all members older than now - window_ms and count the remaining members. The entire sequence must be atomic to avoid race conditions between the ZREMRANGEBYSCORE (clean) and ZCARD (count) operations.

Atomicity is achieved through a Redis Lua script executed with EVAL or EVALSHA. Lua scripts execute as a single Redis command and cannot be interrupted. Kong’s rate-limiting-advanced plugin implements this pattern natively when strategy: redis is configured; the script below is the equivalent for direct Redis clients or custom gateway integrations.

sliding_window_ratelimit.lualua
-- KEYS[1]: rate limit key, e.g. "rl:pis:{jwt_sub}:60000"
-- ARGV[1]: current timestamp in milliseconds
-- ARGV[2]: window size in milliseconds (e.g. 60000 for 1 minute)
-- ARGV[3]: request limit (e.g. 100)
-- ARGV[4]: unique request ID (UUID)
-- Returns: {allowed (0/1), current_count, ttl_ms}

local key        = KEYS[1]
local now        = tonumber(ARGV[1])
local window_ms  = tonumber(ARGV[2])
local limit      = tonumber(ARGV[3])
local req_id     = ARGV[4]
local cutoff     = now - window_ms

-- Remove expired entries (older than the window)
redis.call('ZREMRANGEBYSCORE', key, '-inf', cutoff)

-- Count remaining entries in the window
local count = redis.call('ZCARD', key)

if count < limit then
  -- Add this request to the window
  redis.call('ZADD', key, now, req_id)
  -- Set key TTL to window duration (auto-cleanup)
  redis.call('PEXPIRE', key, window_ms)
  return {1, count + 1, window_ms}   -- allowed
else
  -- Get TTL to inform Retry-After header
  local oldest = redis.call('ZRANGE', key, 0, 0, 'WITHSCORES')
  local reset_in = window_ms - (now - tonumber(oldest[2]))
  return {0, count, reset_in}          -- rejected, ms until oldest entry expires
end

Redis Cluster key sharding for rate limiting requires careful key design. The default Redis Cluster shards by hash of the entire key, which means two rate limit keys for the same consumer (rl:{sub}:pis and rl:{sub}:ais) will land on different slots and cannot be accessed in the same Lua transaction. Use hash tags to force co-location: rl:{sub}:pis and rl:{sub}:ais both hash on sub, ensuring all rate limit keys for one consumer land on the same slot and can be atomically evaluated together if a multi-key script is needed.

Kong Rate Limiting Configuration

Kong’s rate-limiting-advanced plugin (Kong Enterprise) provides sliding window counting in Redis out of the box, eliminating the need to implement the Lua script above directly. The key configuration decisions are: which limit_by identity to use, which window_type (sliding vs fixed), which strategy (redis vs local vs cluster), and how to structure per-consumer and per-route policies.

Policy hierarchy in Kong: a plugin configured at the route level overrides one at the service level, which overrides one at the global level. A consumer-specific plugin configuration overrides route-level. This hierarchy allows you to set a base rate limit on the service, tighten it on specific high-sensitivity routes (payment initiation), and give premium consumer groups a higher quota.

kong-rate-limiting-plugin.yamlyaml
# Base rate limit on the payment initiation route (PIS)
name: rate-limiting-advanced
route: pis-initiate-payment
config:
  limit: [100]          # 100 requests per window
  window_size: [60]     # 60-second window (1 minute)
  window_type: sliding
  limit_by: consumer   # JWT sub resolved to Kong consumer via JWT plugin
  strategy: redis       # cluster-wide distributed counter
  redis:
    cluster_addresses:
      - "redis-0.redis-headless.infra.svc:6379"
      - "redis-1.redis-headless.infra.svc:6379"
      - "redis-2.redis-headless.infra.svc:6379"
    database: 1
    timeout: 2000        # 2 s Redis timeout; fall back to local on timeout
    ssl: true
    ssl_verify: true
  hide_client_headers: false
  # Headers injected on every response:
  # X-RateLimit-Limit-60: 100
  # X-RateLimit-Remaining-60: N
  # X-RateLimit-Reset-60: {epoch}
  # Retry-After: {seconds} (on 429 only)
  error_code: 429
  error_message: "Rate limit exceeded. See Retry-After header."

---

# Account information route (AIS) — separate quota class per SAMA
name: rate-limiting-advanced
route: ais-account-info
config:
  limit: [60, 500]      # 60/min AND 500/hour — multi-window enforcement
  window_size: [60, 3600]
  window_type: sliding
  limit_by: consumer
  strategy: redis

---

# Premium PSP consumer group override — higher quota approved by SAMA sandbox
name: rate-limiting-advanced
route: pis-initiate-payment
consumer_group: premium-psp
config:
  limit: [300]
  window_size: [60]
  window_type: sliding
  strategy: redis

IBM API Connect Rate Plans

IBM API Connect enforces rate limiting through Rate Plans attached to Products. A Product groups one or more APIs and defines the terms of access (who can subscribe, at what quota). Rate plans set per-plan limits that apply to all consumers subscribed to that plan, with per-application and per-operation overrides possible at the plan level.

API Connect’s rate plan model maps naturally to SAMA Open Banking’s subscription tier model: a PSP subscribes to the “Open Banking Standard” plan, which enforces the SAMA-mandated quotas. Premium PSPs with a separate agreement can be subscribed to a “Premium PSP” plan with higher limits, without any code change to the API itself.

open-banking-product.yamlyaml
product: 1.0.0
info:
  title: SAIB Open Banking APIs
  name: saib-open-banking
  version: 1.3.0

apis:
  payment-initiation:
    $ref: "./apis/pis-api.yaml"
  account-information:
    $ref: "./apis/ais-api.yaml"

plans:
  open-banking-standard:
    title: Open Banking Standard
    description: SAMA-mandated quotas for licensed PSPs and TPs
    approval: true    # subscriptions require manual approval
    rate-limits:
      default:
        value: 100
        unit: minute
        hard-limit: true   # reject at limit; no grace
    burst-limits:
      default:
        value: 20
        unit: second
    apis:
      payment-initiation:
        rate-limits:
          pis-initiate:
            value: 100
            unit: minute    # SAMA PIS quota: 100 req/min per PSP
            hard-limit: true
        operations:
          createPayment:
            rate-limits:
              per-operation:
                value: 50
                unit: minute   # tighter limit on the mutation endpoint
      account-information:
        rate-limits:
          ais-query:
            value: 60
            unit: minute    # SAMA AIS quota: 60 req/min per app
            hard-limit: true

  premium-psp:
    title: Premium PSP
    description: Higher quotas for PSPs with separate SAMA approval
    approval: true
    rate-limits:
      default:
        value: 300
        unit: minute
        hard-limit: true

SAMA Regulatory Context

SAMA’s Open Banking Framework defines quota classes per API type and per participant tier. These are not suggested defaults — they are regulatory limits that licensed PSPs and Third-Party Providers (TPPs) are bound by, and that the bank must enforce as the Account Servicing Payment Service Provider (ASPSP).

The applicable quotas as of the current SAMA Open Banking Technical Standards (v3.0.2):

  • Payment Initiation Service (PIS) APIs — 100 requests per minute per PSP. “Per PSP” means per PSP organisation, identified by the x-fapi-financial-id header resolved to a SAMA-registered PSP identifier — not per application or per end-user session. A PSP with 10 applications gets 100 req/min total, not 1 000.
  • Account Information Service (AIS) APIs — 60 requests per minute per TPP-application combination. AIS quotas are per-application (identified by the client_id in the OAuth token) rather than per-organisation, because different AIS applications serve different customer consent contexts.
  • Confirmation of Funds (CoF) — not separately quota-controlled in the current SAMA standard; treated as PIS quota class.
  • SAMA sandbox environment — quotas are relaxed: 500 req/min for PIS, 300 req/min for AIS, with no hard-limit enforcement (soft-limit only; 429 responses are informational and requests are still processed). This is intentional to allow PSPs to load-test integrations without hitting production limits. Do not carry sandbox quota values into the production configuration.

Audit requirements: every 429 response on a SAMA Open Banking API must be logged with the PSP identifier, the API endpoint, the count at rejection time, and the window reset timestamp. SAMA examination teams have requested 90-day retention of throttle event logs in previous reviews. Configure Kong’s file-log or http-log plugin to emit these events to the SIEM separately from the standard access log.

Adaptive Rate Limiting

Static rate limits protect the gateway but not necessarily the backend. A backend service with a p99 latency of 800 ms and a rate limit of 100 req/min has a maximum concurrent request depth of approximately 1.3 — under any burst, the backend will accumulate queued requests that exceed its processing capacity. Rate limiting must be combined with concurrency limits and latency-based backpressure to protect the full request path.

  • Concurrency limits (semaphore). Kong’s response-ratelimiting plugin can count concurrently active requests (requests initiated but not yet responded to) in addition to request rate. For synchronous payment execution backends, a concurrency limit of 50 is a practical ceiling that prevents the backend thread pool from exhausting while the rate limiter still allows traffic.
  • Latency-based backpressure. When the upstream service’s p99 response time exceeds a threshold (e.g. 2× the SLO target), the rate limiter should reduce the allowed rate by a factor. This is not native to Kong; implement it as a Lua plugin that reads the upstream latency metric from Prometheus (via an API call) and dynamically adjusts the rate limit key’s effective ceiling.
  • Priority lanes for RTGS vs SWIFT. Not all payment traffic has equal urgency. SARIE RTGS payments (same-day settlement) have a higher regulatory priority than SWIFT correspondent payments. Implement separate rate limit keys and Redis namespaces per payment rail; in a capacity constraint scenario, apply stricter limits to the lower-priority SWIFT lane first.
  • Circuit breaker integration. When the circuit breaker opens on the payments backend (Kong’s upstream health check detects consecutive failures), reduce the rate limit to near-zero immediately — not to let more traffic hammer a degraded backend. A half-open circuit that still allows 100 req/min will fail those 100 requests and potentially worsen the recovery time.

Rate Limit Bypass Prevention

Rate limiting is only as strong as the identity it counts against. The most common bypass patterns in banking API contexts all exploit identity ambiguity — the rate limiter is counting against something other than the true consumer.

  • IP rotation. A consumer uses a pool of IP addresses (NAT gateway rotation, residential proxies) to spread requests across multiple IP-based rate limit keys. Mitigation: never rate-limit by source IP alone for authenticated APIs. Always rate-limit by authenticated principal (JWT sub or client_id). IP-based limits are a supplementary layer to prevent unauthenticated abuse, not the primary control.
  • API key rotation. A consumer with multiple API keys rotates through them to multiply their effective quota. Mitigation: in APIC, rate plans apply at the subscription level, not the credential level. A PSP with 5 API key credentials on one subscription still hits the plan quota shared across all credentials.
  • JWT sub as canonical identity. For OAuth 2.0 flows where the consumer can rotate client_id by registering multiple applications, rate-limit on sub (the end-user identity from the IdP) combined with client_id, not client_id alone. For PIS APIs, SAMA’s PSP identifier (x-fapi-financial-id header) is the authoritative identity for quota enforcement; validate this header against the SAMA PSP registry on every request.
  • Distributed shadow counters. For high-value API endpoints, implement a secondary rate counter in the application layer (not just the gateway) using the same Redis Lua pattern. This catches bypass attempts that route around the gateway through internal service-to-service paths — a known attack vector when internal APIs are not properly protected by NetworkPolicy.
Never rate-limit by IP alone in bank APIs — shared NAT

Corporate banking clients and SAMA-licensed TPPs almost always connect through a NAT gateway or a shared enterprise internet egress. A single IP address may represent thousands of legitimate users from the same organisation. Rate-limiting by IP on a shared NAT will throttle an entire corporate client when any one of their sessions is busy, without reducing load from the actual problem consumer. Use the authenticated JWT sub or SAMA PSP identifier as the rate limit key for all authenticated endpoints; reserve IP-based limits for pre-authentication DDoS mitigation only.

Common Pitfalls

  • Sliding window cost at high cardinality. A sorted set per consumer per API class per window accumulates quickly. At 50 000 active consumers, each with 3 API classes and a 60-second window holding up to 100 entries, the Redis memory footprint is approximately 50 000 × 3 × 100 × ~80 bytes = 1.2 GB. This is manageable but must be planned; unexpected cardinality growth (e.g. per-session rather than per-consumer rate limit keys) can exhaust Redis memory quickly.
  • Redis split-brain behaviour. In a Redis Cluster, a network partition can cause nodes to accept writes independently. Rate limit keys on both sides of the partition accumulate counts separately; when the partition heals, the merged count may be below the actual request count, temporarily allowing more traffic than configured. This is a known CAP trade-off of Redis Cluster in AP mode. For strictly regulated limits (PIS quotas), evaluate Redis Sentinel with a primary-only write policy to sacrifice availability for consistency during partitions.
  • Retry storms amplifying load. When rate limiting returns 429, consumers retry. If retry is not governed by an exponential backoff with jitter, the retry burst hits at the rate limit reset boundary, exceeding the limit again immediately. The Retry-After header tells the consumer when to retry; the consumer must respect it. Include retry policy guidance in your API documentation and reject consumers at the gateway whose User-Agent identifies known misconfigured retry clients.
Redis key expiry gaps can allow over-quota traffic

Redis key expiry is not guaranteed to fire at exactly the configured TTL — it fires when the key is next accessed, or during the background expiry cycle. Under high load, expired rate limit keys may persist briefly past their expiry time, and the ZREMRANGEBYSCORE cleanup inside the Lua script may have a few milliseconds of lag. In practice this gap is under 50 ms and is not operationally significant. Where it does matter: if you set a rate limit of 100/min and the consumer hits exactly 100 requests in the last 10 ms of a window, the next window’s first request may incorrectly see a count of 101 before cleanup runs. Accept this as a design trade-off of the sliding window approach; it under-protects by at most one request per window boundary.

Production Checklist

Steps to roll out rate limiting without breaking existing consumers

  1. Measure before you limit. Deploy Kong with logging-only mode first (set rate limit to an unrealistically high value like 100 000 req/min). Collect actual per-consumer request rates from Kong’s access log for 2–4 weeks. The p99 actual rate per consumer across the collection period becomes your baseline for choosing the production limit.
  2. Set limits above the measured baseline. Start at 2× the p99 observed rate for non-SAMA APIs, and at the SAMA-mandated limit for regulated APIs. This ensures the limit does not break any currently working consumer while still providing protection against runaway or malicious traffic.
  3. Enable response headers. Set hide_client_headers: false and ensure consumers receive X-RateLimit-Remaining and X-RateLimit-Reset headers on every response. Give consumers at least 4 weeks to instrument their clients to read and respect these headers before enforcing hard limits.
  4. Alert on 429 rate before enforcing. Create a Grafana alert on Kong’s 429 response rate per consumer. Notify affected consumer teams at 10% of their limit, then 50%, then enforce at 100%. This gives teams time to adapt their retry logic before they experience hard throttling.
  5. Switch to Redis sliding window. Migrate from strategy: local (or no rate limiting) to strategy: redis with window_type: sliding in a maintenance window, with Redis Cluster deployed and tested. Validate that the cluster-wide limit is equivalent to the per-node limit you measured during baseline collection.
  6. Enforce hard limits for SAMA APIs. Only after steps 1–5 are complete, set hard-limit: true for PIS and AIS routes. Inform all registered PSPs and TPPs via the SAMA Open Banking developer portal before the enforcement date.
  7. Establish a quota breach review process. After hard limits are in production, any PSP hitting their limit repeatedly should trigger a review — are they misconfigured, have they grown beyond their plan, or are they attempting bypass? Build a Grafana dashboard that shows per-PSP 429 frequency over time and route it to the API governance team for monthly review.

Production checklist

  1. Sliding window algorithm configured for all SAMA PIS and AIS routes; fixed window only for non-regulated endpoints.
  2. Redis Cluster deployed with at least 3 nodes; TLS enabled on all cluster connections from Kong.
  3. Rate limit identity uses JWT sub or SAMA x-fapi-financial-id; IP-based limits applied as supplementary DDoS layer only.
  4. Multi-window enforcement for AIS: 60/min AND 500/hour configured as separate window entries in the rate-limiting plugin.
  5. APIC rate plans reflect SAMA quotas: 100/min for PIS, 60/min for AIS; hard-limit: true on both.
  6. SAMA sandbox environment configured with soft limits only; hard limits never carried over to production config.
  7. Retry-After and X-RateLimit-Reset headers present on every 429 response.
  8. Throttle event log (PSP ID, endpoint, count, timestamp) retained for 90 days in SIEM per SAMA audit requirement.
  9. Concurrency limits set on payment execution backend; latency-based backpressure evaluated for high-volume routes.
  10. Redis Cluster key design uses hash tags to co-locate all rate limit keys for a single consumer on one slot.
  11. Retry storm mitigation documented in consumer API guide; per-consumer 429 rate alerting configured in Grafana.