Overview
A single-tier cache forces an architectural choice: pay the network round trip to Redis on every read (microseconds per request, tens of milliseconds aggregate overhead per second at 100k RPS), or pay the JVM heap and GC cost of an in-process store. The multi-tier answer refuses the choice. L1 in the JVM absorbs hot reads at nanosecond latency, no network. L2 in Redis provides cross-pod consistency at microsecond latency, one network hop. L3 in the database is the source of truth, accessed only on a full cache miss.
The cost of multi-tier caching is coherence: when a value in Redis (L2) changes, every pod’s in-process L1 cache may hold a stale copy. Without a coherence mechanism, pods serve stale data for up to the L1 TTL — which is fine for exchange rates (60-second staleness acceptable) and wrong for account balances (any staleness unacceptable). The coherence mechanism adds complexity; understanding when it is necessary and when it is not is the core judgment call in multi-tier cache design.
This article assumes a Spring Boot application running on OpenShift with 3+ replicas. L1 is Caffeine (in-JVM, per-pod). L2 is a Redis Cluster. The coherence signal is Redis keyspace notifications. The three-tier pattern applies equally to Hazelcast Near Cache (where Hazelcast provides the coherence protocol internally) — the principles are the same; the wiring differs.
A Redis round trip at p50 is 0.5–1 ms. At 100,000 requests per second, that is 50–100 seconds of cumulative Redis RTT per second of wall-clock time — absorbed by connection pooling and pipelining, but never eliminated. An L1 cache with a 95% hit rate reduces Redis load by 20x and eliminates the RTT for 95% of reads. At high-throughput banking APIs (mobile dashboard loads, balance pre-checks for payment initiation), the L1 hit rate directly determines whether Redis is a performance multiplier or a bottleneck.
Write Policies
The write policy determines when the cache is updated relative to the database. Choosing the wrong policy for financial data is not a performance issue; it is a correctness issue.
| Policy | Write path | Read path | Safe for banking | Notes |
|---|---|---|---|---|
| Cache-aside | Application writes to DB only; cache populated on read miss | Miss → DB → populate cache | Yes — default choice | Cache is always consistent with DB on miss; may serve stale until TTL/invalidation |
| Read-through | Same as cache-aside | Cache fetches from DB on miss (behind cache abstraction) | Yes | Application code simpler; cache library handles miss logic |
| Write-through | Application writes to cache AND DB together | Cache always populated after every write | Yes — good for read-heavy entities | Two writes per update; transactionality is complex (Lua + DB transaction) |
| Write-behind | Application writes to cache; cache asynchronously writes to DB | Cache always has latest value | No — pod crash before async write = data loss | Acceptable for non-financial data (analytics counters); never for account state |
Cache-aside is the correct default for financial workloads. It makes the application code responsible for fetching data on a miss — which means the DB is the system of record, always. Write-through adds value for reference data that must be in the cache immediately after an update (e.g. a limit matrix change that must be reflected in the next transaction). Write-behind must never be used for account balances, limit structures, or any state that a SAMA examination would treat as authoritative.
L1 with Caffeine
Caffeine is the successor to Guava Cache and the default in-process cache for Spring Boot 3. Its W-TinyLFU eviction algorithm (Window Tiny Least Frequently Used) provides near-optimal hit rates for most access patterns by combining a frequency-biased main cache with a small recency window. For banking reference data (product configurations, exchange rates, limit matrices), W-TinyLFU outperforms simple LRU because access is frequency-biased: the exchange rate for SAR/USD is read far more often than the exchange rate for SAR/PKR, and W-TinyLFU learns this and protects the hot entry from eviction.
Size L1 conservatively. Every entry in the Caffeine cache is on the JVM heap, competing with application working set, thread stacks, and code cache. A Caffeine cache of 10,000 entries where each entry is 500 bytes of serialised object consumes approximately 5 MB per pod. At 10 pods, that is 50 MB of aggregate in-JVM cache data. At 10,000 entries averaging 50 KB (large objects), it is 500 MB per pod — which will trigger GC pressure in a 2 GB container. Set maximumSize based on the average entry size, not the entry count.
@Configuration
@EnableCaching
public class MultiTierCacheConfig {
// L1: Caffeine in-JVM cache with per-cache TTL and size limits.
// Sizes and TTLs tuned for KSA retail banking access patterns.
@Bean
public CaffeineSpec exchangeRateCaffeineSpec() {
return CaffeineSpec.parse("maximumSize=200,expireAfterWrite=60s,recordStats");
}
@Bean
public CaffeineSpec limitMatrixCaffeineSpec() {
return CaffeineSpec.parse("maximumSize=5000,expireAfterWrite=30s,recordStats");
}
// L2: Redis cache with RedisCacheManager and per-cache TTL.
// Keys serialized as Strings; values serialized as JSON (forward-compatible).
@Bean
public RedisCacheManager redisCacheManager(RedisConnectionFactory factory) {
RedisCacheConfiguration defaults = RedisCacheConfiguration.defaultCacheConfig()
.serializeKeysWith(RedisSerializationContext.SerializationPair.fromSerializer(
new StringRedisSerializer()))
.serializeValuesWith(RedisSerializationContext.SerializationPair.fromSerializer(
new GenericJackson2JsonRedisSerializer()))
.disableCachingNullValues();
return RedisCacheManager.builder(factory)
.cacheDefaults(defaults)
.withCacheConfiguration("exchange-rates",
defaults.entryTtl(Duration.ofMinutes(5)))
.withCacheConfiguration("limit-matrix",
defaults.entryTtl(Duration.ofMinutes(10)))
.withCacheConfiguration("bic-routing",
defaults.entryTtl(Duration.ofHours(24)))
.build();
}
// CompositeCacheManager: tries L1 Caffeine first, falls through to L2 Redis.
// Spring @Cacheable uses the first CacheManager that returns a non-null cache.
@Bean
@Primary
public CacheManager cacheManager(CaffeineCacheManager l1, RedisCacheManager l2) {
CompositeCacheManager composite = new CompositeCacheManager(l1, l2);
composite.setFallbackToNoOpCache(false); // error if no cache found
return composite;
}
}
L2 with Redis
Redis serves as the shared distributed L2 cache, visible to all pods simultaneously. Its role in a multi-tier architecture is narrower than in a single-tier setup: it is the authoritative cache tier that L1 populates from on a miss. When a pod’s L1 entry expires or is evicted, it fetches from L2 — only falling through to the database if L2 also misses.
Redis keyspace notifications are the coherence signal. When a Redis key is deleted or updated (by the application, by a CDC invalidation consumer, or by TTL expiry), Redis can publish a notification to a channel. Applications subscribed to that channel receive the notification and evict the corresponding key from their L1 Caffeine cache. This is the pull coherence model: L1 invalidation is triggered by L2 events, not by the application write path.
Enable keyspace notifications with the minimal necessary event types. KEA (keyevent + expired + all) publishes on every key operation and produces significant Redis CPU overhead for high-write workloads. For cache coherence, you need only deletion events (KEg$lz — keyevent + generic + string + list + set commands that delete keys, plus expired). Benchmark notification overhead in your load test environment before enabling on the production cache Redis.
L1–L2 Coherence
Without coherence, each pod’s L1 Caffeine cache is an independent island of potentially stale data. Pod A updates a value in Redis (L2); Pod B’s L1 still holds the old value for up to the L1 TTL. If your L1 TTL is 60 seconds and you have 10 pods, at any point in time you may have 10 different values for the same key spread across your fleet. For exchange rates (60-second staleness budget), this is acceptable. For limit matrices (30-second staleness budget), it is borderline. For anything with a zero-staleness requirement, L1 must not be used at all.
// Subscribes to Redis keyspace notifications for cache key deletions.
// On notification: evicts the corresponding key from L1 Caffeine.
@Component
public class L1InvalidationListener implements MessageListener {
private static final Logger log = LoggerFactory.getLogger(L1InvalidationListener.class);
private final Cache l1ExchangeRates;
private final Cache l1LimitMatrix;
public L1InvalidationListener(CaffeineCacheManager caffeine) {
this.l1ExchangeRates = caffeine.getCache("exchange-rates");
this.l1LimitMatrix = caffeine.getCache("limit-matrix");
}
// Called by RedisMessageListenerContainer on every keyspace notification.
// message.getBody() = the Redis key that was deleted/expired.
@Override
public void onMessage(Message message, byte[] pattern) {
String key = new String(message.getBody());
log.debug("L2 keyspace event: key='{}'", key);
// Route to the correct L1 cache based on key prefix.
if (key.startsWith("exchange-rates::")) {
l1ExchangeRates.evict(key.substring("exchange-rates::".length()));
} else if (key.startsWith("limit-matrix::")) {
l1LimitMatrix.evict(key.substring("limit-matrix::".length()));
}
// Unknown keys are ignored — no-op; avoids coupling to every cache name.
}
// Register the listener on application startup.
@Bean
public RedisMessageListenerContainer keyspaceNotificationContainer(
RedisConnectionFactory factory, L1InvalidationListener listener) {
RedisMessageListenerContainer container = new RedisMessageListenerContainer();
container.setConnectionFactory(factory);
// Subscribe to key-event notifications for the cache database (db 0).
// Adjust pattern for your Redis keyspace notification configuration.
container.addMessageListener(listener,
new PatternTopic("__keyevent@0__:del"));
container.addMessageListener(listener,
new PatternTopic("__keyevent@0__:expired"));
return container;
}
}
If you deploy L1 Caffeine without Redis keyspace notifications (because enabling notifications seems complex or adds Redis CPU overhead), you have deployed a stale-data machine. Each pod’s L1 cache expires on its own schedule. Pod A’s Caffeine entry for the exchange rate was written at 14:00:00 and expires at 14:01:00. Pod B’s entry was written at 14:00:30 and expires at 14:01:30. The exchange rate changes at 14:00:45. Between 14:00:45 and 14:01:00, Pod A serves the old rate; between 14:00:45 and 14:01:30, Pod B serves the old rate. For exchange rates, the 45-second divergence is typically acceptable. For limit matrices, it is not. Know your staleness budget per data type before deciding whether coherence is required.
Eviction Policies
Caffeine’s default eviction algorithm, W-TinyLFU (Window Tiny LFU), is superior to LRU for most financial workloads. It divides the cache into three segments: a small window for recently accessed entries (protecting against burst access), a protected segment for frequently accessed entries (the hot set), and a probationary segment for newly admitted entries. Entries graduate from probationary to protected when accessed again; they are evicted when a new entry needs space and both the candidate and the victim’s frequency are compared.
| Algorithm | Behaviour | Best for | Worst for |
|---|---|---|---|
| LRU | Evicts least recently used entry | Recency-biased access (session data, user contexts) | Frequency-biased access with occasional scans (scan evicts the hot set) |
| LFU | Evicts least frequently used entry | Frequency-biased access (exchange rates, product configs) | Bursty access (new entries never get high frequency before eviction) |
| W-TinyLFU | Frequency + recency hybrid; Caffeine default | General purpose; best hit rate across mixed access patterns | Workloads so skewed that the admission filter causes freshness issues (rare) |
| Size-based | Evicts by weight (byte count) rather than entry count | Caches with highly variable entry sizes | Adds complexity; only needed when entry size variance is >10× |
For Redis L2 eviction, use volatile-lru — evict only keys that have a TTL set, using LRU ordering. All cache entries should have TTLs (even if long); session data must not be evicted. volatile-lru gives you the separation you need without risking eviction of keys that must survive memory pressure.
Cache Warming
A cold cache on pod startup causes a thundering herd: all pods start simultaneously (during a rolling deploy or after a cluster restart), all have empty L1 caches, all send concurrent requests to Redis for the same reference data, and if Redis also misses (after a Redis restart), all concurrent requests go to the database simultaneously. At 10 pods each making 100 concurrent requests, that is 1,000 simultaneous DB queries for the same reference data that a single DB read per cache key could have handled.
// Warms reference data caches on application startup.
// Loads from Redis (L2) first; falls back to DB only if L2 is also cold.
// Staggers DB reads across pods using a randomized delay to prevent stampede.
@Component
public class CacheWarmingService implements ApplicationListener<ApplicationReadyEvent> {
private final ExchangeRateRepository exchangeRateRepo;
private final LimitMatrixRepository limitMatrixRepo;
private final CacheManager cacheManager;
@Override
public void onApplicationEvent(ApplicationReadyEvent event) {
// Stagger pod startup warming: random delay 0-10 seconds.
// Reduces simultaneous DB reads from N pods to ~N/10 concurrent reads.
long delay = (long)(Math.random() * 10_000);
try { Thread.sleep(delay); } catch (InterruptedException e) { Thread.currentThread().interrupt(); return; }
// warmCache() calls @Cacheable methods which check L2 before DB.
// If L2 is warm from another pod, DB is not reached at all.
warmExchangeRates();
warmLimitMatrix();
log.info("Cache warming complete. Exchange rates: {}, Limit matrix: {}",
exchangeRateCount, limitMatrixCount);
}
private void warmExchangeRates() {
// Load all active exchange rate pairs; @Cacheable populates L2 on DB read.
exchangeRateRepo.findAllActive().forEach(rate -> {
Cache l1 = cacheManager.getCache("exchange-rates");
if (l1 != null) l1.put(rate.getPair(), rate);
});
}
private void warmLimitMatrix() {
// Limit matrix: load all product/segment combinations.
// TTL-aware: only warm L1 if L2 TTL still has > 60% of its lifetime remaining.
limitMatrixRepo.findAllWithValidCache().forEach(entry -> {
Cache l1 = cacheManager.getCache("limit-matrix");
if (l1 != null) l1.put(entry.getKey(), entry.getValue());
});
}
}
Banking-Specific Patterns
Different financial data types have fundamentally different staleness budgets, update frequencies, and consistency requirements. Assigning every entity to the same tier with the same TTL is the most common caching design error in banking systems. The correct approach is to assign each data type to the appropriate tier explicitly, with a documented rationale.
| Data type | L1 (Caffeine) | L2 (Redis) | Invalidation | Rationale |
|---|---|---|---|---|
| Account balance | No L1 — too risky | 30 s safety-net TTL | CDC on every debit/credit | Zero staleness requirement; any L1 staleness unacceptable under SAMA |
| Exchange rates | 60 s TTL, 200 entries | 5 min TTL | Scheduled refresh at SAMA window | 60-second staleness acceptable; high read frequency makes L1 essential |
| BIC/IBAN routing table | 1 h TTL, 10,000 entries | 24 h TTL | CDC on SWIFT table update | Changes only during settlement windows; very high read-to-write ratio |
| Limit matrix | 30 s TTL, 5,000 entries | 10 min TTL | Application-level on config change | Updated by ops, not by transactions; L1 coherence required across pods |
| Product catalogue | 5 min TTL, 1,000 entries | 1 h TTL | Application-level on publish | Low update frequency; L1 reduces Redis load from product listing APIs |
Account balance is the canonical exception to multi-tier caching. The SAMA balance display requirement eliminates L1 as an option: any pod-local L1 entry for a balance will diverge from the DB state within milliseconds of a transaction. Store balances only in L2 (Redis) with a short safety-net TTL and CDC-driven invalidation. The Redis RTT for a balance read is typically 0.5–1 ms — acceptable for a once-per-page-load dashboard read.
Pitfalls
Write-behind (or write-back) caching writes to the cache synchronously but writes to the database asynchronously. If the pod crashes before the async write completes, the data in the cache is lost and the database never received the update. For a payment limit counter or an account balance, this means real money or authorised payments may be lost. Write-behind is appropriate only for non-financial data where eventual persistence is acceptable (analytics counters, recommendation history). It must never be used for account state, limit structures, or any data that is an authoritative financial record.
When a popular cache entry expires simultaneously across multiple pods, all pods see a cache miss and all send concurrent requests to the database for the same key. At 10 pods with 100 concurrent users per pod, a single cache expiry can produce 1,000 simultaneous DB queries in the space of a few milliseconds. Mitigation: probabilistic early expiration (Caffeine’s refreshAfterWrite triggers a background refresh before expiry); per-key mutex (one pod fetches and repopulates while others wait on a Redis lock); or jitter on TTL (add ±10% random variation to TTLs so entries across pods expire at different times).
All pods in a rolling deployment start within a narrow time window. All have cold L1 caches. All send warming requests to Redis simultaneously. If Redis is also cold (after a Redis restart or a cluster failover), all warming requests go to the database simultaneously. At 10 pods warming 5,000 limit matrix entries each, that is 50,000 DB queries in the first 30 seconds after startup. Add a randomised delay to the warming service (0–10 seconds), warm from L2 before going to L3, and set the readiness probe to only pass after warming completes — so traffic is not routed to a pod with a cold cache.
Steps to Implement a Production Multi-Tier Cache with L1 Coherence
-
Classify all cached entities by staleness budget
Produce a data classification table: entity name, update frequency, acceptable staleness window, L1 eligible (yes/no), L1 TTL, L2 TTL, invalidation mechanism. This table drives every subsequent decision. Get sign-off from the business owner on the staleness window before implementing — discovering disagreement during a SAMA examination is expensive.
-
Deploy Caffeine L1 for eligible entities
Add
spring-boot-starter-cacheandcaffeinedependencies. ConfigureCaffeineCacheManagerwith per-cache specs (maximumSize, expireAfterWrite). Annotate service methods with@Cacheable(cacheManager="caffeineCacheManager"). Expose Caffeine stats via Micrometer to measure hit rate immediately. -
Add Redis L2 with JSON serialization
Configure
RedisCacheManagerwithGenericJackson2JsonRedisSerializerand per-cache TTLs. Set up aCompositeCacheManager(L1 first, L2 fallback) so@Cacheablechecks both tiers transparently. Verify that an L1 miss populates L1 from L2 correctly. -
Enable Redis keyspace notifications
Run
CONFIG SET notify-keyspace-events KEAon the cache Redis instance (or add it to the Redis configuration). ImplementRedisMessageListenerContainersubscribed to__keyevent@0__:deland__keyevent@0__:expiredchannels. Test: write a key to Redis, delete it, verify that the corresponding L1 entry is evicted in the same pod and in a separate pod instance. -
Implement cache warming with startup jitter
Implement
ApplicationListener<ApplicationReadyEvent>that pre-populates L1 from L2 (or L3 if L2 is cold) on startup. Add a random 0–10 second jitter to prevent stampede. Gate the readiness probe behind a warming-complete flag so traffic is not routed before the cache is warm. -
Add hit-rate and staleness metrics
Expose Caffeine stats via Micrometer (
cache.gets,cache.evictions,cache.size). Add a custom metric for L2 keyspace notification latency (time between a Redis key deletion and the L1 eviction). Alert: L1 hit rate below 80% for exchange-rates is a tuning problem; L1 hit rate below 50% means L1 adds cost without benefit. -
Load-test and tune before production
Run a load test at 1.5x expected peak RPS with all pods at steady state. Measure: L1 hit rate, L2 hit rate, DB query rate, Redis CPU, JVM heap usage, p99 latency. Tune maximumSize, TTLs, and warming scope based on observed hit rates. A L1 hit rate below 90% on exchange rates indicates the TTL is too short or maximumSize too small.
| Tier | Latency | Scope | TTL range | Consistency | Eviction control | Cost per lookup |
|---|---|---|---|---|---|---|
| L1 Caffeine | <1 µs (heap) | Pod-local | 10 s – 10 min | Eventual (TTL + keyspace notification) | W-TinyLFU, size-bounded | Heap memory allocation |
| L2 Redis | 0.5–2 ms (network) | All pods | 30 s – 24 h | Eventual (TTL + explicit invalidation) | volatile-lru (recommended) | Network RTT + Redis CPU |
| L3 Database | 10–100 ms (network + disk) | Global | N/A (source of truth) | Strong (ACID) | N/A | DB connection, query plan, I/O |