Every database modernisation programme at a bank eventually reaches the same moment: someone has to move the accounts data. Not the product catalogue, not the transaction history archive, not the reference data tables — the accounts data. The ledger. The thing that says how much money each customer has, what limits apply, and what state their account is in. Moving that data without losing any of it, without serving incorrect balances during the transition, and without creating a SAMA-notifiable incident is the hardest migration problem in banking IT.
Most banks approach it by scheduling a maintenance window. The window gets shortened in review because it is too long. It gets approved anyway. The team stays up through the night, the migration runs, the reconciliation looks clean — and somewhere in the next 48 hours, something starts failing. An overnight batch job was writing to the old schema. A payment authorisation service cached account limits that are now stale. The EBCDIC-to-UTF-8 conversion handled Arabic names with hamza differently in the migration script than the application does in its regular reads. None of these failures were visible in staging because staging does not have production data, production concurrency, or production batch jobs.
This is not bad execution. It is a category error: treating database migration as a one-time event rather than a state synchronisation problem.
The decision that precedes the technology
The dual-write approach — writing every account modification to both the old database and the new one simultaneously, running shadow reads to validate the new database without serving from it, and shifting traffic gradually over weeks rather than in a single maintenance window — is not primarily a technology decision. The technology is well understood: Spring’s repository layer, a Kafka bus for backfill and compensation, Debezium for CDC. The decision is whether the organisation will accept the cost of running two databases in parallel for two to three months in exchange for removing the risk that a botched cutover corrupts the accounts ledger.
That cost is real. Write latency increases because every write touches two stores. Operational complexity doubles during the migration window. The reconciliation infrastructure has to be built, monitored, and maintained. Engineers who would otherwise be building features are building the migration scaffolding. And the programme timeline extends by the weeks needed to reach a credible zero-divergence signal before cutover.
Against that cost: a failed big-bang cutover on the accounts database of a Saudi bank is a SAMA notification event. It is an operational resilience failure under the CSAML and the Cybersecurity Framework. It is recoverable in principle — you can roll back and replay — but in practice, recovery from a corrupted accounts ledger takes longer than the window of acceptable downtime for SARIE payment channels, and the customer and regulatory damage happens in that gap.
The dual-write programme is expensive because it is designed to make failure impossible, not merely unlikely. That is the right trade-off for a payment accounts database.
What the VP has to decide
The dual-write migration programme requires four explicit decisions from the VP of Enterprise Integration or equivalent. Engineers cannot make these decisions because they involve authority over production systems and tolerance for programme risk that sits at the leadership level.
1. Accept the timeline extension. A dual-write migration to a new database for the accounts domain takes 10–14 weeks from Debezium deployment to cutover. A big-bang migration can be done in a single weekend. If the modernisation programme has a hard board-level deadline that does not accommodate a 14-week migration runway, that deadline needs to be renegotiated before engineering begins, not after the first failed cutover attempt. The VP owns that renegotiation.
2. Define the cutover gate criteria. The decision to proceed from shadow mode to traffic shifting, and from traffic shifting to write cutover, must be governed by quantitative criteria — not by schedule, not by stakeholder pressure, not by “it looks good enough.” The VP must define those criteria in writing before the migration starts: what shadow divergence rate is acceptable, how many consecutive hours of zero divergence constitute confidence, what the rollback trigger conditions are. Engineers will make optimistic calls under schedule pressure unless the gate criteria are pre-committed and auditable.
3. Authorise rollback without approval. The on-call engineer at 2am during cutover must have standing authority to roll back to the old database without escalating to a VP or a change advisory board. If rollback requires a CAB approval process, it will not happen fast enough to prevent customer impact. The VP authorises rollback as a pre-approved action in the change management system before the cutover window opens. The authority is time-bounded (the cutover window) and condition-bounded (specific rollback triggers), but within those bounds it is unilateral.
4. Commit the resourcing for the migration period. The dual-write infrastructure is not a side task. The reconciliation dashboard, the shadow comparator, the compensation event consumer, the batch job migration — these require engineers who are focused on them, not split across three other priorities. A migration done in the gaps of a busy delivery team is a migration done slowly with recurring gaps in the validation signal. The VP secures dedicated resourcing for the migration period as a separate workstream, not an add-on to existing delivery capacity.
What SAMA expects to see
SAMA examination teams increasingly ask for migration documentation as part of IT risk reviews, particularly for changes to payment-critical systems. The dual-write approach produces the documentation naturally: shadow validation reports, gate check records, compensation event metrics, and the post-cutover reconciliation query output are all artefacts that show a controlled, evidence-based migration. A big-bang cutover with a maintenance window produces a change ticket and a post-implementation review — a much thinner evidence trail for a much higher-risk operation.
The PDPL dimension is also significant. The Debezium CDC pipeline and Kafka migration bus carry personal data — customer names, IBANs, national IDs — across a new data path that did not exist before the migration programme. Under PDPL, this is a new data processing activity that requires appropriate controls: data residency within KSA, access controls on the migration topics, retention limits on the backfill data, and deletion of migration artefacts after the programme closes. A VP who approves the migration programme without approving a data protection review of the migration infrastructure has created a PDPL compliance gap alongside the technical migration.
What the team needs from leadership to ship it
The engineers running the dual-write migration need one thing more than tooling or architecture guidance: a VP who will not override the gate criteria when the programme is behind schedule. The migration will fall behind schedule — it always does, because reconciliation reveals problems that require time to diagnose and fix. The moment the programme is behind schedule, stakeholder pressure to cut corners on the validation gates begins. “The divergence rate is only 0.02%, that’s fine.” “We’ve been in shadow mode long enough.” “The business cannot wait another two weeks.”
The VP’s job at that moment is to hold the line on the gate criteria that were committed to before the programme started. Not because the criteria are sacred, but because the alternative — cutting over with a known divergence rate and hoping it does not affect payment authorisation — is the exact failure mode the 14-week migration runway was designed to avoid.
The technical companion to this post is the Dual-Write & Shadow Mode lab article, which covers the Spring Boot implementation of the dual-write repository, the Debezium DB2 connector configuration, the shadow comparator pattern, and the exact cutover sequence with pre-cutover gate checks.