Most banks that have a mainframe also have a polling problem. Every team that needs mainframe data has written a job that queries the DB2 table every five minutes, looking for rows where the update timestamp is newer than the last run. There are twenty of these jobs. They all run at different intervals. Three of them are owned by teams that have left the bank. None of them can see deletes. They collectively consume a share of MIPS that appears on the software bill every month as an undifferentiated infrastructure cost, because no one ever mapped the query load back to its origin.

Change Data Capture from the DB2 transaction log is the answer. It is also, in the context of a SAMA-regulated bank running z/OS and IBM i, a decision that carries more organisational and regulatory weight than the tool choice itself. The tool is IBM IIDR on z/OS or Debezium on IBM i. The decision is whether the bank is willing to change who owns the mainframe data contract — and what that ownership shift requires from leadership before the first event hits Kafka.

What the shift actually means

Today, if a new service needs account balance data, it goes to the mainframe team, negotiates a CICS interface or a JDBC query window, and takes what it can get. The data shape is a direct reflection of the DB2 schema — ACCT_NBR as CHAR(16), BAL_AVAIL as packed decimal, date fields in Julian format. The consuming team builds workarounds for all of it and calls it “integration overhead.”

With CDC, the integration team publishes a canonical event schema — accountNumber as a string, availableBalance as a normalised decimal, transactionDate as ISO 8601. The mainframe team no longer owns how consuming teams see the data; the integration team does. This is the shift. It is an organisational decision about who controls the data contract, not a technical decision about which log reader to use.

In practice this means two things need to happen before any tooling is deployed. First, the integration team needs explicit authority to define the canonical schema and to hold that schema stable even when the mainframe team changes a column name or type — which they will, because mainframe schema changes have historically been internal decisions invisible to application teams. Second, the mainframe team needs a channel for communicating schema changes before they happen, not after. Without this, CDC becomes a source of incidents: IIDR or Debezium surfaces a schema change as a broken event, consumers fail, and the integration team discovers the mainframe change in production.

The trade-off the business needs to own

CDC from a log reader has zero write overhead on the source database. It does not add triggers, does not run table scans, and does not consume database connections during peak load. This is the right engineering answer and it is true.

What the business needs to understand is the other side of that trade: log-based CDC reads the database transaction log at the point of commit, which means it is observing production data as it changes in real time, with minimal delay. The IIDR agent runs as a started task on z/OS with DB2 trusted client authority. The Debezium connector connects to IBM i via a service account with journal read access. These are not elevated permissions in the sense of being able to change data, but they are elevated in the sense of being able to see all data changes continuously, including changes that are sensitive under the bank’s data classification policy.

CDC is a continuous wire tap on the production database. The data governance team needs to classify what flows over it, and the security team needs to approve the service account scope, before the first connector goes live.

This is the conversation that gets skipped in most projects. The engineering team deploys CDC because it is technically superior to polling. The data governance team discovers six months later that the Kafka topic carrying account balances is accessible to the analytics team without the same access controls as the source DB2 table. Not because anyone was careless, but because no one mapped the CDC pipeline through the bank’s data access control framework before deploying it.

What SAMA will look for

SAMA’s Technology Risk Management framework requires that banks maintain an unbroken audit trail for core banking data. A CDC pipeline that publishes DB2 changes to Kafka introduces a new node in the data lineage chain. SAMA’s examination team will want to know two things: can you prove that every change to a regulated table was captured without gaps, and does the captured data stay within KSA?

The first question is answered by a reconciliation process — a scheduled job that counts and checksums source DB2 state against consumer state and alerts on any divergence. This is not a technical afterthought; it is a required control for a SAMA-regulated CDC implementation, and it needs to be running and monitored from Day 1, not added when the examination team asks about it.

The second question is about data residency. Kafka topics carrying CDC data from DB2 are core banking data in transit. They must be hosted in KSA. If the bank uses IBM Event Streams on a cloud provider, the cluster must be in a KSA-certified region or on-premises. Generic region labels from cloud provider documentation are not sufficient assurance for a SAMA examination; you need a written confirmation from the hosting provider that the data does not leave KSA borders, including for replication, backup, and disaster recovery.

What the team needs from leadership to ship it

The CDC project will move at the speed of three things: schema mapping decisions, access control approvals, and mainframe team cooperation. All three are leadership problems, not engineering problems.

Schema mapping decisions cannot be delegated to the engineer building the Avro schema. Every canonical field name and type choice is a contract that downstream teams will build on for years. The VP or integration architect needs to chair the schema review process, hold the boundary on stability commitments (“this field name will not change without a deprecation period”), and make the call when the mainframe team and the integration team disagree on the canonical shape of a field. Without this authority sitting above both teams, schema review meetings stall and the project waits.

Access control approvals for service accounts on z/OS and IBM i go through the bank’s identity and access management process. On a typical timeline, that process takes four to six weeks from request to approval for a new service account with non-standard access (and “DB2 trusted client” on z/OS counts as non-standard). Start this process eight weeks before the intended CDC go-live date. Discovering at Week 3 of a four-week sprint that the IIDR service account has not been approved is the most common schedule killer in mainframe CDC projects, and it is entirely avoidable.

Mainframe team cooperation is the dependency most often underweighted in project plans. The IBM i journal receiver retention period needs to increase from whatever the current default is to at least 48 hours; that is a mainframe DBA change with a change management cycle attached. IIDR needs a DB2 PTF level validated before installation; the mainframe team owns the PTF schedule. Schema changes need to be communicated to the integration team before they are applied; that requires a new step in the mainframe change process. None of these are large asks, but all of them require someone with organisational authority to make them happen on a schedule that the CDC project depends on.

For the engineering depth behind these decisions — IBM IIDR configuration, Debezium connector setup, schema mapping patterns, reconciliation SQL, SAMA audit trail continuity, and the tool comparison across IBM IIDR, Debezium, and custom polling — see the companion Lab article: Legacy Integration: CDC from Legacy DB2 to Modern Services.