Every API gateway product ships with a rate limiting plugin. The configuration is typically three fields: a limit number, a time window, and a storage backend. Engineers configure it in an afternoon, commit it to the platform repository, and mark the story done. Six months later, a PSP partner sends a complaint that their payment initiation requests are being rejected at peak hours. SAMA’s examination team asks for a log showing which API calls were throttled, when, and why, and the team realises the throttle events are written to an access log that is rotated after 7 days. The “done” story was not done. The rate limiting was configured; it was not decided.

Why rate limiting is a business decision, not a technical one

Rate limiting defines your service tiers. The limit you set on payment initiation APIs is not a technical ceiling — it is a commercial term that determines how many transactions per minute a PSP can initiate through your infrastructure. A PSP with a 100 req/min limit and a settlement window that requires 250 initiations in a 3-minute burst is a PSP that cannot meet its contractual obligations to its own customers through your bank’s API. The commercial relationship, the SLA, and the API limit must be aligned. That alignment is a business decision, not an engineering one.

Rate limits also define your service quality promise to internal teams. An internal BFF serving the mobile application that is throttled during a promotional campaign has a user-facing impact on the bank’s digital channel — login failures, balance queries that time out, transaction status pages that return errors. The engineering team did not budget for that impact when they set the limit. The product team did not know the limit existed. Rate limits that are invisible to the business until they fire are limits that will create operational incidents at the worst time.

A rate limit is a contract. Set it without consulting the parties who will be bound by it, and you will renegotiate it under pressure in a production incident, which is the worst possible time to make infrastructure policy.

The practical framework: before any rate limit goes into production enforcement, three questions must be answered with named owners. Who approved this limit value, and on what basis? Which consumers will be affected, and have they been notified? What is the process for a consumer to request a quota increase, and who has the authority to approve it? Without answers to all three, the rate limit is a technical artefact with no business legitimacy, and it will be bypassed or overridden at the first sign of friction.

What SAMA requires

SAMA’s Open Banking Technical Standards define quota classes for each API category. These are not suggestions — they are regulatory limits that the Account Servicing Payment Service Provider (ASPSP) — the bank — is required to enforce. Payment Initiation Service APIs allow 100 requests per minute per PSP organisation; Account Information Service APIs allow 60 requests per minute per TPP application. These limits are tested during SAMA sandbox certification and are expected to be in production enforcement before go-live.

Three SAMA compliance requirements consistently catch banks off-guard during examination. The first is the distinction between sandbox and production quotas. SAMA’s sandbox environment uses relaxed limits (typically 5× the production limit, with soft enforcement) to allow PSPs to load-test their integrations. Banks that carry the sandbox configuration into production — or that configure production with sandbox-level limits “temporarily” while they finalise the production values — are not compliant. SAMA examinations have specifically checked the effective rate limit on the production environment as a separate step from reviewing the configuration file.

The second is the audit log for quota breaches. SAMA requires that every throttle event on a regulated API endpoint is logged with sufficient detail to reconstruct the event: the PSP identifier (not just the IP address or API key), the API endpoint, the count at the time of rejection, and the window reset time. This log must be retained for a minimum period consistent with SAMA’s data retention requirements (90 days minimum for transaction-adjacent audit data). An access log that records HTTP 429 responses without the PSP identifier, or a log that is rotated after 7 days, does not satisfy this requirement. The engineering configuration of the rate limiter and the engineering configuration of the audit log are separate stories, and both must be done before examination.

The third is per-organisation vs per-application quota enforcement. SAMA’s PIS quota is per-PSP organisation, not per-application or per-API-key. A PSP with multiple registered applications still gets 100 req/min total, shared across all applications. This requires the rate limit identity to be the SAMA-registered PSP identifier (x-fapi-financial-id header, validated against the SAMA PSP registry), not the OAuth client_id or the API key. A rate limiter keyed on client_id when a PSP has 5 applications effectively gives them 500 req/min, which is non-compliant. This identity mapping must be built into the gateway configuration before production go-live, not as a future improvement.

The trade-off that leadership must own

Rate limiting presents a genuine tension between two legitimate objectives: protecting the infrastructure and the regulatory compliance posture on one side, and minimising friction for commercial partners and internal teams on the other. Both sides of this tension have a legitimate claim, and neither side resolves the trade-off alone.

Set limits too aggressively and you create customer friction. A PSP whose payment initiation requests are rejected at the legal limit during peak settlement windows will escalate. An internal mobile BFF team that hits a rate limit during a marketing campaign will bypass the gateway or demand an exception. Exceptions granted under pressure, without process, undermine the rate limiting policy for every other consumer. The commercial cost of excessive throttling — in partner relationship damage, in lost payment volume, in internal delivery friction — is real and must be weighed.

Set limits too permissively and you expose the infrastructure to load that the downstream systems — the core banking platform, the SARIE connection, the SWIFT gateway — cannot absorb. A DDoS attack against an Open Banking payment initiation endpoint that is not rate-limited will not just take down the API; it will exhaust the bank’s SARIE transaction capacity for legitimate customers. The operational cost of under-throttling is not just a 429 response — it is a payment system outage during trading hours, with regulatory, reputational, and financial consequences that dwarf any partner complaint about a throttled API.

The leadership decision is to set limits that reflect actual backend capacity, not aspirational throughput figures from architectural presentations. This requires knowing the p99 throughput of the core banking integration layer under sustained load, not the theoretical maximum. It requires a conversation between the VP overseeing digital channels and the VP overseeing core banking operations, not just an engineering estimate from the team building the gateway.

What leadership must decide before enforcement goes live

Four decisions belong to leadership rather than to the engineering team configuring the rate limiter.

Which endpoints get protected first. Not every API endpoint has the same risk profile. Payment initiation endpoints carry regulatory and financial risk; they must be rate-limited before any other endpoint. Account balance query endpoints carry customer-experience risk if over-throttled; their limits should be set more conservatively. Internal service-to-service APIs carry operational risk if under-throttled when a runaway process starts looping. The prioritisation of which endpoints get hard rate limit enforcement first — as opposed to monitoring-only soft limits — is a risk decision that belongs at VP level, not at sprint planning.

The retry policy as a first-class customer communication. When a consumer receives a 429, what should they do? The technical answer is “read the Retry-After header and wait.” The business answer is: do they have an SLA guarantee on eventual processing? Do they receive a push notification when the window resets? Is there a priority queue for payment initiations that failed due to throttling but must be processed before end-of-business? These questions are customer experience and commercial decisions, not technical ones. The Retry-After header is infrastructure; the retry policy is a product feature that must be designed, documented, and communicated to affected partners before hard limits go live.

Customer communication when throttled. The bank’s relationship with PSPs and TPPs operating under its Open Banking infrastructure is a regulated relationship with disclosure requirements. A PSP that discovers their rate limit through a 429 response in production — rather than through a documented quota published in the developer portal, confirmed during sandbox certification, and notified in advance of any change — has a legitimate grievance. The communication plan for rate limits — how new limits are announced, how changes are notified, what the appeal and exception process is — is a governance decision, not a configuration decision. It belongs in the API product governance framework before enforcement begins.

The escalation path for limit breaches. When a SAMA-regulated PSP hits its PIS rate limit repeatedly, two things must happen simultaneously: the 429 must fire at the gateway (the technical enforcement), and someone at the bank must be notified that a PSP is operating at or above their quota limit (the relationship management). The monitoring alert that fires when a PSP hits 80% of their quota, and the process for that alert to reach an account manager who can open a conversation about a quota increase or a usage pattern review — that is a cross-functional process that must be designed before the first production incident creates it under pressure.

Rate limiting done right is invisible to well-behaved consumers and impassable to poorly-behaved ones. It requires correct technical implementation and it requires business decisions that give the technical implementation legitimacy and a governance framework to operate within. The engineering team can configure Redis and Kong in a day. The business decisions that make the configuration durable take longer and require different people in the room.

For the engineering depth behind these decisions — sliding window vs token bucket algorithms, Redis Lua scripts for atomic counters, Kong rate-limiting-advanced configuration, IBM API Connect rate plans, SAMA quota enforcement patterns, and bypass prevention — see the companion Lab article: API Gateway: Rate Limiting & Throttling at Scale.