Overview
The strangler fig gives you a modern interface in front of the mainframe. CDC gives you the data. These are two separate problems, and confusing them is the source of most long-running legacy integration headaches: teams build the façade, expose a REST API backed by CICS, and then discover that every downstream analytics system, every data warehouse, every new microservice still needs a feed of what changed in the DB2 tables. Without CDC, that feed is a polling job — and polling jobs compound their costs quietly until they consume a noticeable fraction of your mainframe MIPS.
Change Data Capture reads the database transaction log. It sees every insert, update, and delete as it commits, converts it to a structured event, and publishes it to a topic. Consumers subscribe to the topic. The mainframe does not know they exist. This is the only architecture that decouples consumer growth from mainframe overhead.
Two tools dominate in practice for IBM environments: IBM IIDR (InfoSphere Data Replication, now part of IBM DataStage) for DB2 on z/OS, and Debezium for DB2 on IBM i (AS/400). They solve the same problem against different log formats, and they are not interchangeable — choose based on your platform, not on open-source preference.
This article covers log-based CDC exclusively. Trigger-based CDC — where a database trigger writes to a shadow table and a job polls that table — works but adds write overhead to the production DB2 instance on z/OS and introduces a polling lag. On a mainframe running under MLC pricing, trigger overhead is directly visible in your software bill. Log-based CDC has zero write overhead on the source.
Why CDC over polling
The immediate objection to legacy polling is latency: a 5-minute batch job sees data that is 5 minutes old at best. But the real problem is not latency — it is resource and fragility:
- MIPS consumption. A SELECT * WHERE last_updated > :last_run against a 500M-row account table on z/OS consumes MIPS proportional to the rows scanned, not the rows changed. If 100 rows change per minute and the table has 500M rows, you scan 500M rows to find 100. With CDC you process exactly the 100 rows that changed.
- No tombstones. A poll-based job that reads last_updated cannot detect hard deletes. If a record is deleted, it disappears from the result set. CDC surfaces deletes as explicit tombstone events with the before-image of the deleted row.
- No delta for binary columns. BLOB and CLOB columns do not have an updatable timestamp you can poll on. CDC reads the full column value from the log on every change.
- Consumer proliferation. Every new consumer that needs a data feed writes another polling job. On the mainframe, each job consumes MIPS. CDC amortises the capture cost across all consumers via Kafka: capture once, consume many times.
IBM IIDR on z/OS
IBM IIDR reads the DB2 for z/OS active and archive logs via the DB2 IFCID 306 interface (the instrumentation facility component). It requires no table-level changes, no triggers, and no additional indexes. The IIDR agent runs as a started task on z/OS and must be authorised as a DB2 trusted client.
IIDR publishes to Kafka via its Kafka stage — a built-in producer that converts the internal IIDR row-change format to Avro or JSON. The Avro path is strongly preferred in production: it enforces schema at publish time and integrates with Confluent Schema Registry or IBM Event Streams’ built-in registry.
{
"subscriptionName": "SAIB_ACCOUNTS_TO_KAFKA",
"source": {
"type": "DB2ZOS",
"database": "SAIBPROD",
"schema": "ACCTDB",
"tables": ["ACCOUNTS", "LEDGER_ENTRIES", "HOLD_ORDERS"],
"captureMode": "LOG_BASED", // never TRIGGER
"beforeImage": true, // required for audit trail
"lobHandling": "INLINE_UP_TO_32KB"
},
"target": {
"type": "KAFKA",
"bootstrapServers": "kafka-1.saib.internal:9093,kafka-2.saib.internal:9093",
"schemaRegistry": "https://schema-registry.saib.internal",
"serializationFormat": "AVRO",
"topicMapping": {
"ACCOUNTS": "saib.legacy.accounts.cdc",
"LEDGER_ENTRIES": "saib.legacy.ledger.cdc",
"HOLD_ORDERS": "saib.legacy.holds.cdc"
},
"partitionKey": "ACCOUNT_NBR", // ordering guaranteed per account
"security": {
"protocol": "SSL",
"keystorePath": "/var/iidr/keystores/kafka-client.jks"
}
},
"latencyTargetMs": 500 // IIDR SLA: 500ms capture-to-publish
}
Key configuration decisions worth calling out:
beforeImage: true. IIDR captures both the before and after images of every changed row. This is more expensive (doubles the payload for updates) but is non-negotiable if SAMA audit trail continuity applies — the before-image is the audit record.partitionKey: ACCOUNT_NBR. Partitioning on the business key guarantees that all changes for one account land on the same partition in order. Without this, a consumer can see an UPDATE before the INSERT that created the row.- Latency SLA. 500ms is a realistic IIDR target for a well-sized z/OS agent. The log reader runs in a tight loop; the bottleneck is usually Kafka producer acknowledgement, not log read speed.
IBM IIDR requires specific DB2 for z/OS subsystem versions and PTF levels. Running IIDR 11.4 against a DB2 V12 subsystem that has not applied the required PTFs produces silent data loss — not an error. Validate the compatibility matrix before installing and revalidate after every DB2 maintenance window. IIDR subscription metadata is stored in a separate DB2 subsystem — not in the source — so a DB2 maintenance window that does not include the IIDR control database still requires an IIDR restart.
Debezium on IBM i
DB2 for IBM i uses journal receivers rather than the DB2 z/OS log format. Debezium’s IBM i connector reads from the journal via the QSYS2.DISPLAY_JOURNAL table function or, in newer versions, the Change Data Capture journal exit point. Unlike IIDR, Debezium runs outside the IBM i system — it runs as a Kafka Connect connector on your integration platform and connects to IBM i via JDBC and Toolbox for Java.
{
"name": "saib-ibmi-customers",
"config": {
"connector.class": "io.debezium.connector.db2i.Db2iConnector",
"database.hostname": "as400-prod.saib.internal",
"database.port": "8471",
"database.user": "CDCUSER",
"database.password": "${file:/opt/kafka/secrets/ibmi.properties:db.password}",
"database.dbname": "*LOCAL",
"schema.include.list": "CUSTLIB",
"table.include.list": "CUSTLIB.CUSTOMERS,CUSTLIB.LOANS,CUSTLIB.COLLATERAL",
"topic.prefix": "saib.legacy.ibmi",
"key.converter": "io.confluent.kafka.serializers.KafkaAvroSerializer",
"value.converter": "io.confluent.kafka.serializers.KafkaAvroSerializer",
"value.converter.schema.registry.url": "https://schema-registry.saib.internal",
"journal.name": "CUSTLIB/CUSTJRN",
"snapshot.mode": "initial", // full snapshot on first run
"snapshot.isolation.mode": "read_committed",
"decimal.handling.mode": "string", // packed decimal → string first
"time.precision.mode": "adaptive_time_microseconds",
"heartbeat.interval.ms": "10000"
}
}
The decimal.handling.mode: "string" setting is important: DB2 for i stores monetary amounts as packed decimal (COMP-3 equivalent). Debezium can emit them as Java BigDecimal through Kafka Connect’s logical types, but this requires the consuming schema to understand org.apache.kafka.connect.data.Decimal. Emitting as string and converting at the consumer is more portable and avoids precision surprises in intermediate frameworks like Flink that have their own decimal handling.
Debezium reads journal receivers sequentially. If Debezium is down for longer than the IBM i journal receiver retention period, it loses its position. The connector then falls back to a snapshot of the current table state — which is correct but expensive on large tables. Set journal receiver retention to at least 48 hours on any table that Debezium is watching, and alert if the connector lag approaches that threshold.
Schema mapping
Legacy DB2 schemas were designed for 1980s storage constraints and programmer conventions, not for API consumers. Mapping them is not optional — it is the job. The choices you make here determine the maintainability of every downstream consumer for the next decade.
Three categories of mapping work appear on every project:
- Type normalisation. Convert packed decimal to
BigDecimal, Julian dates to ISO 8601, EBCDIC character fields to UTF-8, and CHAR(3) currency codes to ISO 4217 strings. Do this at the CDC layer, not in each consumer. - Semantic rename.
ACCT_BAL_AVAILbecomesavailableBalance.TXN_DT_YYMMDDbecomestransactionDate. This is a one-time investment that pays back in every consumer that no longer needs to know thatCUST_TP_CDmeans customer type. - Derived fields. Sometimes a column does not map cleanly to a single modern concept because it encodes multiple meanings (a status field where values 1–9 mean “active” and values 10–19 mean “suspended with different reasons”). Add a derived
statusenum field. Keep the original column too, at least during a transition period.
{
"type": "record",
"name": "AccountCdcEvent",
"namespace": "info.wbadawi.saib.legacy.accounts",
"fields": [
{ "name": "eventType", "type": { "type": "enum", "name": "CdcOp", "symbols": ["INSERT","UPDATE","DELETE"] }},
{ "name": "capturedAt", "type": "string" }, // ISO 8601 from DB2 log timestamp
{ "name": "accountNumber", "type": "string" }, // was ACCT_NBR CHAR(16)
{ "name": "customerId", "type": "string" },
{ "name": "availableBalance", "type": "string" }, // string BigDecimal; scale=2
{ "name": "ledgerBalance", "type": "string" },
{ "name": "currency", "type": "string" }, // ISO 4217, was CHAR(3)
{ "name": "status", "type": { "type": "enum", "name": "AcctStatus", "symbols": ["ACTIVE","DORMANT","SUSPENDED","CLOSED"] }},
{ "name": "before", "type": ["null", "AccountCdcPayload"], "default": null },
{ "name": "after", "type": ["null", "AccountCdcPayload"], "default": null }
]
}
Event envelope design
A CDC event is not an API response. It is a fact: something changed in the source system at a point in time. The envelope must carry enough metadata that any consumer can answer three questions independently: what changed, when did it change, and was this change an insert, update, or delete?
Beyond the Debezium/IIDR defaults, a financial-services CDC envelope needs two additional fields:
- Source transaction ID. The DB2 commit unit identifier. Multiple row changes in the same DB2 transaction produce multiple CDC events. Grouping them by transaction ID is how a consumer reconstructs transactional consistency. IIDR provides this natively; Debezium provides it as
source.txId. - Change sequence number. A monotonically increasing number within the source DB2 log. This is the tie-breaker when two CDC events for the same partition key have the same millisecond timestamp. Both IIDR and Debezium expose this; make it a top-level field in your canonical envelope rather than burying it in the source metadata.
Kafka partition offsets are not DB2 log sequence numbers. A Kafka offset tells you where an event sits in a Kafka partition. A DB2 log sequence number tells you where a change sits in the source database transaction log. These are different numbering spaces. If consumers use Kafka offsets to deduplicate or order database changes, they will get it wrong on Kafka topic compaction, partition reassignment, or topic recreation. Always carry the source log sequence number in the event payload.
Deployment steps
-
Map the DB2 schema to a canonical event shape
For each table being captured, produce a canonical Avro schema. Map every column type to its modern equivalent. Identify derived fields. Register the schema in Schema Registry with
BACKWARD_TRANSITIVEcompatibility (consumers can read any older version; new versions may add optional fields only). -
Configure IIDR replication rules (z/OS path)
Install the IIDR agent as a z/OS started task. Grant it DB2 trusted client authority. Create one subscription per table group. Set
beforeImage: trueandcaptureMode: LOG_BASED. Validate the DB2 PTF compatibility matrix before starting. -
Deploy Debezium Kafka Connect (IBM i path)
Deploy a Kafka Connect cluster (or use your existing one) accessible to the IBM i network segment. Configure the DB2 for i connector with the journal receiver name, schema/table include lists, and Avro serialisation. Run the initial snapshot during a low-traffic window — it does a full table read via JDBC.
-
Apply SMTs for EBCDIC normalisation and field rename
Use Kafka Connect Single Message Transforms to rename columns, convert EBCDIC character data, and extract the source transaction ID into a top-level field. Keep SMTs stateless — any stateful transformation belongs in a Flink or Kafka Streams job downstream, not inside the connector.
-
Wire consumer services with schema evolution rules
Each consumer registers its expected schema version with Schema Registry. Set consumer compatibility to
FORWARD: the consumer can read messages written by any schema version it has registered, plus any newer version that only adds optional fields. Configure a dead-letter queue for any event that fails schema validation — never discard unreadable events silently. -
Enable reconciliation and alert on drift
Within 24 hours of enabling CDC, run the first reconciliation query (see next section). Set a pipeline health alert: if the lag between last captured event timestamp and wall-clock time exceeds 10 minutes, page the on-call engineer. This catches IIDR agent restarts, network partitions, and Kafka producer backpressure before consumers notice.
Reconciliation & drift detection
CDC pipelines drift silently. The common failure modes — an IIDR subscription that paused and resumed, a Debezium connector that fell behind and did a partial re-snapshot, an SMT that mishandled a null value in a column it had never seen null before — all produce a state where consumers believe they have accurate data but do not. The only reliable counter-measure is a reconciliation query that counts and checksums source vs. consumer state on a schedule.
-- Run on z/OS DB2 (source of truth)
SELECT
COUNT(*) AS src_count,
SUM(DECIMAL(BAL_AVAIL, 20, 2)) AS src_balance_sum,
MAX(LAST_UPD_TS) AS src_max_ts
FROM ACCTDB.ACCOUNTS
WHERE ACCT_STATUS != 'C' -- exclude closed accounts
AND LAST_UPD_TS < CURRENT TIMESTAMP - 5 MINUTES; -- allow CDC lag
-- Run on consumer DB (Postgres in the Payments service)
SELECT
COUNT(*) AS cdc_count,
SUM(available_balance) AS cdc_balance_sum,
MAX(captured_at) AS cdc_max_ts
FROM accounts_replica
WHERE status != 'CLOSED';
-- Alert if: ABS(src_count - cdc_count) > 0
-- OR ABS(src_balance_sum - cdc_balance_sum) > 0.01
-- OR (src_max_ts - cdc_max_ts) > INTERVAL '5 minutes'
Run this query every 4 hours in normal operation and after every IIDR or Debezium restart. Track the drift count as a metric in your observability stack. A drift of zero for 30 consecutive days before a production cutover is a meaningful confidence bar — not a guarantee, but the closest proxy available.
SAMA compliance
SAMA’s Technology Risk Management framework requires that core banking systems maintain a complete and tamper-evident audit trail through any change in how data is processed or replicated. A CDC pipeline that reads the DB2 transaction log and publishes to Kafka introduces two new points where the audit chain can break:
- Before-image retention. The before-image of every changed row must be retained, not just the after-image. If your Kafka topic compaction removes tombstones before the required retention period expires (SAMA expects 10 years for core banking records), you have a gap. Use Kafka topic retention policy
deletenotcompactfor CDC topics that carry SAMA-regulated data, and setretention.msaccordingly. - Data residency. All CDC data in transit must stay within KSA. This means your Kafka cluster, Schema Registry, and all consumer services must be hosted in KSA regions or on-premises. Specifically: if you use IBM Event Streams on IBM Cloud, the KSA region is
au-syd-equivalent for SAMA purposes only if the tenant is provisioned in your on-premises or co-location facility. Verify with IBM, not with generic cloud region documentation.
Kafka log compaction retains only the latest record per key in a topic. For a CDC topic, this means only the current state of each row survives compaction — the history of how it got there is gone. For a SAMA-regulated core banking table, you need the full history. Set cleanup.policy=delete (not compact) and set retention to match your regulatory requirement. Use compacted topics only for reference data materialisation, never for the primary CDC stream.
Tool comparison
| Capability | IBM IIDR 11.4 | Debezium 2.7 | Custom polling |
|---|---|---|---|
| DB2 z/OS | Native, IFCID 306 | Not supported | JDBC, polling only |
| DB2 for IBM i | Supported (IIDR for i) | Native, journal CDC | JDBC, polling only |
| Write overhead on source | Zero (log read) | Zero (journal read) | High (table scans) |
| Delete events | Yes, with before-image | Yes, with before-image | No (row disappears) |
| Schema Registry integration | Native Avro/JSON | Native via Kafka Connect | Manual, per-consumer |
| Open source | No (IBM licence) | Yes (Apache 2.0) | Yes (build it) |
| SAMA audit trail | Before-image native | Before-image with config | Manual instrumentation |
| Typical latency | <500ms | <1s (journal lag) | Minutes (poll interval) |
The operative recommendation: IBM IIDR for z/OS, Debezium for IBM i. IIDR is not open source and carries an IBM licence cost, but it is the only tool with native IFCID 306 access and IBM-supported SLAs for the z/OS environment. Debezium on IBM i is mature (the DB2 for i connector has been production-grade since 2.3) and avoids an additional IBM licence for that environment. Do not use Debezium on z/OS with an ODBC/JDBC bridge as a cost-saving measure — JDBC-based CDC on z/OS is still polling, and it costs MIPS proportional to the rows scanned.
Common pitfalls
A DBA adds a NOT NULL column to ACCOUNTS on z/OS. IIDR picks it up immediately. The Avro schema in Schema Registry has no field for it. Debezium consumers configured with BACKWARD compatibility reject the message and route it to the dead-letter queue silently if no alert is configured. You discover the break when a business analyst asks why account counts have dropped. Set up an alert on dead-letter queue depth — anything above zero in production requires immediate investigation.
When Debezium runs its initial snapshot, it reads the current table state via JDBC while live transactions are committing. The snapshot is not transaction-consistent unless you use snapshot.isolation.mode: serializable, which holds a read lock for the duration of the snapshot — on a large table, this means a multi-hour lock on a live production table. Use read_committed isolation and accept that the snapshot will contain a mix of states; the subsequent journal-based CDC stream fills in the gaps. Never run the initial snapshot during business hours on a high-write table.
IIDR handles EBCDIC-to-UTF-8 conversion internally for character fields. Debezium relies on the JDBC driver’s character set handling. If the JDBC connection does not specify the correct CCSID (Coded Character Set Identifier) for the IBM i job, you get question marks for non-ASCII characters in Arabic or French text stored in customer name fields. Set translate binary=true in the JDBC URL and specify ccsid=1208 (UTF-8) explicitly.