Every bank that has gone through a SAMA technology risk examination has produced a business continuity plan. Most of those plans describe an active-active deployment: two data centers, both accepting writes, both serving customers. Very few of them have tested what happens to an account balance when both data centers accept a debit simultaneously during a network partition, and the partition heals before anyone notices.
Active-active is an availability architecture. It means both data centers are live and serving traffic. It does not mean both data centers will produce consistent results under all network conditions. Those are different properties, and the architecture determines which you get. An active-active deployment where both masters accept writes and replicate asynchronously will produce inconsistent account balances under network partition. An active-active deployment where writes require a synchronous cross-datacenter quorum will not — but it will add latency to every write, and it requires a different database technology or a carefully engineered synchronous replication configuration. The decision about which architecture to adopt is not a database team decision. It is a business decision with a regulatory dimension, and it must be owned by the VP or CTO who signs the BCP attestation.
What the regulator will test
SAMA’s Business Continuity Planning framework sets explicit requirements for Tier 1 systems: those that process payments, hold customer deposits, or support core banking operations. The RTO is less than 4 hours; the RPO for financial data is effectively zero for institutions with a synchronous replication requirement. The “effectively zero” language in SAMA guidance is deliberate: for a payment processing system, an RPO of even 1 minute means that up to 1 minute of payment records could be lost in a failover. SAMA does not accept “we recovered most of the data” as a BCP outcome for a Tier 1 system.
The data centre separation requirement of at least 30 kilometres exists to ensure that a localised disaster — a power grid failure, a building fire, a flood — does not take both data centers simultaneously. This separation also introduces a physical reality that the BCP architecture must account for: light travels through fibre at approximately 200,000 km per second. A 30km separation adds a minimum of 0.15 milliseconds of round-trip latency on the replication path, before switching overhead, before TCP overhead, before the database protocol overhead. A synchronous write that requires acknowledgement from both data centers cannot complete in less than that floor. This is not a network engineering problem to be optimised away; it is physics, and the application and database layer must be designed to function correctly within it.
The examination team does not validate a BCP by reading the DR plan document. They test it. They ask you to execute the failover procedure in their presence, or they review documented test results with specific timestamps and data integrity validation evidence. A BCP that has never been tested against a real workload — with a real transaction volume, with a real conflict scenario, with the actual database configuration that will be in place on the day of a real outage — is not a BCP that will survive an examination. It is a document that describes what the team hoped would happen.
The examination tests your actual recovery capability. The DR plan document is your assertion about what that capability is. The gap between the two is the examination finding.
The business trade-off
Synchronous replication between two data centers separated by 30 or more kilometres adds latency to every write. At SAIB’s operational scale and the typical fibre distances between Saudi data centers, the measured addition is 1–3 milliseconds at the p99 percentile for the replication round-trip. On top of a baseline payment processing latency of 50–200 milliseconds end-to-end, 2 milliseconds of replication overhead is not a customer-perceptible difference. It is, however, a genuine cost: a payment processing system that handles 2,000 transactions per second and adds 2ms of latency per transaction is consuming 4,000 thread-milliseconds per second in additional blocking wait. That translates to database connection pool sizing, thread pool sizing, and timeout configuration — all of which must be designed to accommodate the synchronous replication latency from the start.
The alternative trade-off is not “save 2ms per transaction.” The alternative is the possibility of a negative balance incident under network partition. A bank that deploys active-active for account balance writes with asynchronous replication is accepting the scenario where two concurrent debits from the same account, one at each data center during a partition, both pass the balance check and both commit locally. When the partition heals, the account has a negative balance that neither debit should have produced. The reconciliation and remediation of that incident — identifying affected accounts, reversing incorrect transactions, notifying customers, reporting to SAMA — costs orders of magnitude more than the 2ms of replication latency would have cost across the lifetime of the system. The cost of the latency is always less than the cost of the incident. This is not a close call.
What makes this decision difficult is not the arithmetic. The arithmetic is straightforward. What makes it difficult is that synchronous replication for active-active requires either a fundamentally different database technology (CockroachDB or YugabyteDB, which implement Raft-based synchronous replication by design) or a significant engineering investment in the application layer to handle the serialization errors and retry logic that distributed transactions produce. That engineering investment requires budget, time, and engineers with distributed systems expertise. The path of least resistance is to deploy two PostgreSQL instances with asynchronous streaming replication, call it active-active in the BCP document, and leave the conflict scenario as a theoretical risk. The examination that exposes that theoretical risk as a real one is not theoretical.
What leadership must own
Four decisions cannot be delegated to the database team, the architecture team, or the platform team. All four require executive commitment because all four create constraints that other teams must work within, and constraints without executive backing are suggestions.
The first is Tier 1 system classification. The BCP framework applies to Tier 1 systems. The classification of which systems are Tier 1 is a business risk decision, not a technical one. Every payment processing system, every account balance system, and every customer data store that constitutes a financial liability if lost is a Tier 1 system. Classifying a system as Tier 2 to avoid the synchronous replication cost is a risk acceptance decision that must be made explicitly, by the person who is accountable for regulatory compliance, not implicitly by the database team choosing an architecture that makes the problem go away on paper.
The second is replication topology sign-off. The replication topology — synchronous or asynchronous, quorum-based or master-standby, which databases are in which data center, what the failover mechanism is — must be signed off at the VP level before it is deployed in production. Not because the VP needs to understand the technology in detail, but because the topology determines the actual RPO and RTO that the bank can deliver, and the VP who signs the BCP attestation is attesting to those numbers. If the topology cannot deliver the attested numbers, the attestation is incorrect. That is a regulatory exposure, not just a technical gap.
The third is conflict resolution policy per data type. Not all data has the same conflict sensitivity. Account balances require distributed transactions or synchronous replication. Notification preferences can tolerate Last Write Wins. Reference data (product rates, fee schedules) can tolerate a short window of eventual consistency. The policy that specifies which conflict resolution strategy applies to which data type must be formally adopted and documented, because it is the answer to the examination question: “How do you ensure that concurrent writes to this system do not produce incorrect account balances?” The answer must be specific, verifiable, and correct. “We use best practices” is not an answer.
The fourth is BCP test schedule and accountability. The BCP test must happen on a defined schedule — SAMA expects annual tests for Tier 1 systems, more frequent for higher-criticality classifications — and someone must be personally accountable for executing the test, documenting the results, and presenting them to governance. Not accountable for the team executing the test. Personally accountable for the outcome: did the failover succeed, did the RTO target get met, and was data integrity verified by running the reconciliation pipeline over the failover window? If no one is personally accountable for the BCP test outcome, the test will be deferred until the examination forces it, at which point it will be done under pressure with incomplete documentation and an examiner watching.
Multi-master replication with conflict-free design is achievable at a Saudi bank. The technology exists, the patterns are proven, and the engineering investment is bounded. What is not achievable is a conflict-free active-active deployment that does not require executive commitment to the architecture, the test schedule, and the accountability model. The technology is the easy part. Leadership owning the decision is what makes it real.
For the engineering depth behind these decisions — conflict types, LWW risks, CRDT patterns for banking, vector clocks, CockroachDB geo-partitioning SQL, PostgreSQL conflict resolution triggers, Redis PN-Counter implementation, and the full production checklist — see the companion Lab article: Data Synchronisation: Conflict Resolution in Multi-Master Replication.