Every Saudi bank that has been operating for more than two decades has core banking data in COBOL-defined formats on a mainframe. Account balances, transaction histories, customer records, and payment references — the data that every digital channel, analytics pipeline, and new microservice needs — are stored in packed decimal, EBCDIC character sets, and fixed-length records whose layout is described in a copybook file that may not have changed since the 1990s. Making that data readable to a JSON-consuming REST service is not a trivial technical task, and the decision about how to approach it has consequences that persist for years.

What the decision is

The COBOL data mapping decision is a choice between three approaches, each with a different risk and longevity profile. The first is DFDL with IBM ACE: a schema-governed approach where the COBOL copybook is modelled in a Data Format Description Language schema, and the ACE integration platform applies that schema to parse binary records into a typed logical tree. The schema is the single source of truth; when the copybook changes, the schema is updated and regression tests run. The second is a bespoke Java parser: hand-coded parsing logic that reads the byte offsets, converts packed decimal fields, handles EBCDIC conversion, and emits JSON. Faster to ship for a single copybook, but every new copybook requires new code, and every change to an existing copybook requires a code change, a review, and a deployment. The third is mainframe retirement: replacing the mainframe system with a modern one and eliminating the COBOL data problem entirely. This is the right answer for some institutions on a long enough timeline. It is not a data mapping strategy; it is a multi-year programme that costs an order of magnitude more than either of the first two options and carries programme risk that organisations routinely underestimate.

The choice between DFDL and a bespoke parser is not a technology preference. It is a decision about how many copybooks the organisation expects to integrate over the next five years, and who will own the integration layer when the engineers who wrote the original parser have moved on.

The regulatory angle

SAMA supervises core banking systems and the data they produce. Account balances reported to customers through digital channels, payment amounts processed through SARIE and SWIFT, and transaction records retained for regulatory purposes all originate from mainframe data. If the COBOL-to-JSON mapping is wrong — if the packed decimal parser misreads the sign nibble, if the implicit decimal point is applied at the wrong scale, if the EBCDIC conversion uses the wrong code page — the downstream systems receive incorrect data. An incorrect balance displayed in a mobile banking app is a customer complaint. An incorrect amount in a payment instruction is a potential financial loss and a SAMA reportable incident. An incorrect amount in a regulatory report is a compliance failure.

None of these outcomes are hypothetical. Every team that has built a COBOL parser from scratch has at some point shipped one of these bugs. The sign nibble that treats D as positive because the developer misread the specification. The scale that is off by one because the PIC V clause was overlooked. The EBCDIC dollar sign that becomes a different character because the wrong code page was assumed. The regulatory significance of these bugs is that they are silent: the system continues to run, the data continues to flow, and the error accumulates in downstream records until an auditor or a customer finds a discrepancy.

DFDL does not eliminate mapping errors, but it makes the mapping specification explicit, testable, and independently reviewable. A DFDL schema can be reviewed by an integration architect who did not write it. A bespoke parser is often intelligible only to the person who wrote it and anyone willing to work through hundreds of lines of byte-offset arithmetic.

The business trade-off

DFDL is the right long-term investment for any organisation that expects to integrate more than a handful of COBOL copybooks or expects those copybooks to evolve over the integration’s lifetime. The investment pays back in two ways. First, each new copybook costs less to onboard: the DFDL tooling, the testing framework, and the governance process are already in place. Second, copybook changes are handled through schema versioning rather than code changes: the mainframe team delivers a new copybook version, the integration team updates the DFDL schema, runs the regression suite against the existing hex dump test fixtures, and deploys. The deployment cycle is measurably shorter and the audit trail is cleaner.

Bespoke parsers are faster to ship for the first copybook. That advantage is real — a motivated Java developer can write a working packed decimal parser and an EBCDIC converter in a day. The technical debt accrues on the second copybook, when the first parser’s conventions are not followed consistently. It compounds on the fifth, when the team is maintaining five separate parsing implementations with five different null-handling conventions, five different date format assumptions, and five different test coverage levels. By the tenth copybook, the bespoke parser estate is a maintenance burden that consumes more engineering capacity than a DFDL implementation would have required for all ten.

What leadership must own

The COBOL data mapping problem is fundamentally a governance problem. The technical implementation — whether DFDL or Java — is manageable by the integration engineering team. What the engineering team cannot manage without executive backing is the governance of the copybooks themselves: who approves changes to a copybook that affects a live integration, who is responsible for notifying the integration team when a mainframe maintenance release changes a field length, and what the rollback plan is when a copybook change breaks a mapping and incorrect data has already flowed to downstream systems for some period of time.

These questions must be answered before the first COBOL integration goes to production, not after the first incident. The mainframe team and the integration team operate on different release cycles, different change management processes, and different organisational hierarchies. A copybook change that is routine from the mainframe team’s perspective can silently break an integration that has been running correctly for three years. The governance structure that prevents this — a shared change advisory process, mandatory regression testing of the integration layer before any mainframe change goes live, and a clear escalation path when a conflict arises — requires VP-level sponsorship on both sides of the boundary to be effective. Without it, the COBOL data mapping implementation, however technically sound, will eventually produce an incident that could have been prevented.

For the engineering depth behind these decisions — PIC clause types and byte lengths, EBCDIC code page selection, packed decimal parsing in Java, DFDL schema construction in ACE 12, round-trip testing, and the production checklist — see the companion Lab article: Legacy Integration: COBOL Data Mapping & Type Conversion.