Every API has a latency floor — the minimum response time it can achieve regardless of how much infrastructure you throw at it. For a banking API that reads from an Oracle or DB2 database on every request, that floor is 20–80 ms of database latency, irreducible. For a banking API with a Redis L2 cache, the floor drops to 1–3 ms on a cache hit. For an API with a Caffeine L1 in-process cache, it drops to sub-microsecond on the hottest keys. The multi-tier architecture is the mechanism that pushes the floor down. The decision of whether to build it, and how, determines what your customers experience at the p99 percentile under peak load.

What looks like a performance engineering decision is actually an architecture commitment. L1 and L2 caching together create a coherence problem: when a value changes in Redis, every pod’s in-process Caffeine cache may hold a stale copy. Managing that staleness — deciding what data types can tolerate it, implementing Redis keyspace notifications to push invalidations across pods, tuning TTLs to match acceptable staleness budgets — is not a sprint task. It is a platform-level capability that must be designed correctly before it is deployed to a system that carries payment traffic.

The business case in latency terms

A Saudi retail bank with 2 million mobile customers and 10% of them active simultaneously at peak hour has 200,000 concurrent sessions. Each session load involves a balance read, a recent-transactions read, and several product or limit checks. At 5 requests per session load, that is 1 million API calls in the peak minute — roughly 17,000 requests per second. If those 17,000 requests per second all reach Redis, the Redis cluster must handle 17,000 round trips per second at 0.5–1 ms each. That is feasible. If those requests all reach the database, the database must handle 17,000 queries per second at 20–50 ms each — which almost certainly exceeds the database connection pool and produces latency degradation under exactly the conditions you cannot afford it.

An L1 Caffeine cache with a 90% hit rate reduces Redis load to 1,700 requests per second for exchange rates and limit matrices — a 10x reduction. The 10% of requests that miss L1 go to Redis at L2 latency; the 1% that miss both go to the database at L3 latency. The p99 for exchange-rate reads drops from Redis RTT (1 ms) to Caffeine heap read (sub-microsecond). The p99 for the overall API drops because the system is no longer serialising on Redis connection pool slots at peak throughput.

The tier you miss on under load determines your p99. Not your average latency — your p99. And p99 is what your customers experience during the moments when they most need the system to be fast: salary day, end-of-month settlement, Ramadan retail peak.

The trade-offs that matter at scale

Three production risks in multi-tier caching deserve leadership attention, because each one requires a mitigation that has operational cost and cannot be deferred until the first incident.

L1 divergence across pods. An in-process Caffeine cache in a 10-pod deployment means 10 independent copies of the same data, each expiring on its own schedule. Without Redis keyspace notifications pushing invalidations across pods, those 10 copies can diverge from each other and from the database for up to the L1 TTL. For exchange rates with a 60-second TTL, this means a customer might see a rate that is up to 60 seconds old on one pod and current on another. For limit matrices, it means a limit change may not be reflected uniformly across the fleet for up to 30 seconds — a window during which transactions might slip through under the old limit. The keyspace notification mechanism closes this window; implementing it correctly is not complex, but deciding which data types require it and which can tolerate natural TTL expiry requires a business answer.

Redis as a single point of failure. In a multi-tier cache architecture, Redis (L2) is in the read path for every L1 miss. If Redis is unavailable, every L1 miss falls through to the database. Under typical conditions, that might be 10% of requests. Under cold-cache conditions (post-deployment, post-Redis-restart), it could be 100% of requests reaching the database simultaneously. Redis Sentinel or Cluster provides automatic failover, typically within 5–15 seconds. During those 15 seconds, every cache miss reaches the database. Whether the database connection pool can absorb that spike — and whether the application has a circuit breaker that can shed load rather than queuing waiting database connections — is a capacity planning question, not just a Redis HA question.

Cache warming cost on startup. A rolling deployment brings up new pods with cold L1 caches. Each new pod must warm its L1 from L2 Redis. If Redis is also cold (after a Redis failover or a cluster restart during the deployment window), each pod’s warming reads go to the database. At 10 pods each pre-fetching 5,000 limit matrix entries, that is 50,000 database reads in the first 30 seconds of the deployment. In a high-concurrency environment where the deployment happens during peak hours, this startup stampede can degrade latency for live traffic. Staggering the warming startup (random delay per pod), warming from Redis before touching the database, and gating the readiness probe behind warming completion are all engineering mitigations — but the decision to allow deployment during peak hours without a maintenance window, knowing there is a warming cost, is a business and operations policy decision.

What leadership must own

The engineering team can implement multi-tier caching. What it cannot do is decide which data types belong in which tier, because that requires a business answer about acceptable staleness. An exchange rate that is 60 seconds old is acceptable for a currency display; it may not be acceptable as the rate used in a confirmed trade. An account limit that is 30 seconds stale is acceptable for a display; it must not be used as the authoritative check for a payment that is about to post to core banking. These are not implementation questions; they are risk and compliance questions. The tier assignment table for each entity must be owned and signed off by the business and risk owners, not determined by which TTL made the latency numbers look best in the load test.

The second ownership question is who is on-call when the Redis cache cluster fails at peak hours. In a multi-tier architecture, a Redis L2 failure during peak is not a “cache layer problem” — it is a database load problem that may exceed the database’s capacity to handle the sudden traffic increase. The incident response requires someone who can simultaneously manage Redis failover, monitor database connection pool utilisation, and make a call about whether to continue serving (accepting degraded latency) or invoke the platform’s circuit breakers and shed non-critical load. That decision authority needs to be named in the runbook before the incident, not established during it.

The third is the staleness tolerance policy for banking-specific entities. Which entities can use L1 caching with natural TTL expiry? Which require L1 coherence via keyspace notifications? Which must not use L1 at all because any staleness is unacceptable? The answer to these questions determines how much of the system’s complexity is justified and where the keyspace notification overhead is spent. A blanket policy of “all data uses L1 with 30-second TTL” will serve stale account balances; a blanket policy of “no L1 caching for any financial data” will give up the latency benefits that make the architecture worthwhile. The correct policy is entity-specific, and it requires someone with business knowledge and authority to set it.

For the engineering depth behind this decision — Caffeine configuration, Redis keyspace notification coherence, write policies for financial data, cache warming strategies, and banking-specific tiering patterns — see the companion Lab article: Distributed Caching: Multi-Tier Caching Strategies.