Overview

COBOL copybooks have defined data layouts in core banking systems for longer than most of the engineers working on them have been alive. That stability is a feature: the PAYMENT-RECORD that worked in 1985 still works today because the mainframe has never been forced to care about REST or JSON. The problem is that every modern integration — a new microservice, an analytics pipeline, an Open Banking API — needs to read that data, and the types do not map cleanly.

The specific mismatch categories are: packed decimal (COMP-3), where two decimal digits are encoded per byte with a sign nibble that Java has no native type for; EBCDIC character encoding, where the character set is entirely different from UTF-8 and several printable characters live at different code points; implicit decimal points (PIC V), where the decimal position is defined in the copybook rather than stored in the data; and group-level items with REDEFINES, where the same bytes have two different interpretations depending on a discriminator field that the parser must read first.

This article covers the full conversion path: what each PIC clause means in bytes, how to convert EBCDIC and packed decimal, how DFDL in IBM ACE models a copybook, and when a Java-based parser is the right alternative.

DFDL is the right long-term answer

DFDL (Data Format Description Language) separates data format description from transformation logic. A DFDL schema captures the copybook structure once; the ACE ibmDFDL:parseUnparse transform node applies it to any message. When a copybook changes, you update the DFDL schema — not the transformation code. This is the architectural equivalent of keeping schema and logic separate, and it matters when you have dozens of copybooks and the DBA team changes them quarterly.

PIC Clause Types

Every field in a COBOL copybook is described by a PIC (PICTURE) clause. The clause encodes both the logical type and the physical storage format. Understanding what bytes each type occupies is the prerequisite for writing any parser, DFDL schema, or ESQL mapping.

PIC Type Bytes on disk Java type JSON mapping Gotcha
PIC 9(n) n bytes (display numeric) Long / String integer or string Leading zeros present in raw bytes; stripping changes meaning for account numbers
PIC X(n) n bytes (EBCDIC) String string Requires EBCDIC→UTF-8 conversion; code page must match z/OS job CCSID
PIC A(n) n bytes (EBCDIC alpha) String string Same physical storage as X; only valid values are alphabetic — validate after conversion
PIC S9(n) COMP-3 ceil((n+1)/2) bytes BigDecimal string (decimal) Sign nibble C=positive, D=negative, F=unsigned; never use float or double
PIC S9(n) COMP 2 bytes (n≤4), 4 bytes (n≤9), 8 bytes (n≤18) Long integer Big-endian on z/OS; must byte-swap on little-endian JVMs before casting
PIC S9(n) COMP-5 Same as COMP Long integer Native endian (platform-dependent); dangerous if the copybook migrates between z/OS and x86
PIC V (implicit decimal) 0 bytes (logical only) N/A (scale modifier) N/A Shifts the decimal point in an adjacent PIC 9 or COMP-3; the scale is in the copybook, not the data
FILLER Varies (padding) N/A Omit from JSON Still occupies bytes; must be skipped in the offset calculation or the entire record shifts

The most consequential detail in the table: PIC V occupies zero bytes but changes the meaning of the digits before it. A field declared as PIC 9(9)V9(2) is 9 bytes on disk (nine display digits), but the value it represents has two decimal places — a stored value of 000001234500 is actually 12345.00. If you parse without knowing the scale from the copybook, every monetary amount in the system will be off by a factor of 100.

EBCDIC-to-UTF-8 Conversion

IBM mainframes use EBCDIC (Extended Binary Coded Decimal Interchange Code) as their native character set. It is not ASCII. Most printable characters have different numeric values, and several characters that developers use routinely — curly braces, square brackets, backslash, the pipe symbol — live at different code points in different EBCDIC variants.

The two EBCDIC code pages relevant to Saudi banking integrations are IBM Code Page 37 (US English, the historical default for most mainframe installations) and IBM Code Page 1047 (IBM Open systems, common in more recent z/OS configurations and preferred for UNIX-accessible data). Code page 37 and 1047 agree on most characters but diverge on a subset of special characters: the dollar sign ($), curly braces ({ }), square brackets ([ ]), and the exclamation mark (!) are at different positions. Identifying the wrong code page silently corrupts these characters in string fields.

In IBM ACE 12, EBCDIC conversion for MQ messages is handled by the CCSID property on the MQ channel definition. If the channel CCSID is set correctly (37 or 1047 matching the z/OS job), ACE converts PIC X fields transparently during parsing. If the CCSID is wrong or set to 0 (no conversion), you receive raw EBCDIC bytes in what ACE treats as a string — the fields look populated but contain garbage when printed.

For direct file-based integration (z/OS datasets transferred via FTP or SFTP), conversion does not happen automatically. Use IBM ICU4J in your ACE Java compute node or in a pre-processing step:

EbcdicConverter.javajava
import com.ibm.icu.text.CharsetDetector;
import java.nio.ByteBuffer;
import java.nio.charset.Charset;

public class EbcdicConverter {

    // Code page 37 = US EBCDIC (most z/OS SAIB mainframe jobs)
    private static final Charset CP37  = Charset.forName("IBM037");
    private static final Charset CP1047 = Charset.forName("IBM1047");

    public static String convertToUtf8(byte[] ebcdicBytes, String ccsid) {
        Charset source = "1047".equals(ccsid) ? CP1047 : CP37;
        // decode EBCDIC bytes to Java String (Unicode internally)
        return new String(ebcdicBytes, source).trim(); // trim EBCDIC spaces (X'40')
    }

    // Validate that a converted field contains only expected characters
    public static boolean isValidAlphanumeric(String value) {
        return value.chars().allMatch(c -> c >= 0x20 && c < 0x7F);
    }
}

One particularly important detail for Saudi Arabia deployments: Arabic characters stored in COBOL PIC X fields require a hybrid encoding arrangement. Core banking systems at Saudi banks sometimes store Arabic customer names in UTF-8 or Windows-1256 within what the copybook declares as a PIC X field, because the mainframe was configured with a dual code-page arrangement. If character set negotiation is wrong, Arabic names appear as question marks or Latin garbage. Validate every Arabic-capable string field against real mainframe test data before certifying the conversion.

Packed Decimal (COMP-3) Parsing

COMP-3 is the most common storage format for monetary amounts and large counters in COBOL. It packs two decimal digits into each byte, with the last nibble reserved for the sign. A field declared as PIC S9(11) COMP-3 stores 11 digits in 6 bytes (two digits per byte for the first 5 bytes, one digit plus sign in the last byte).

The sign nibble values: C (hex C, binary 1100) = positive, D (hex D, binary 1101) = negative, F (hex F, binary 1111) = unsigned. A value of positive 12345 in a PIC S9(5) COMP-3 field (3 bytes) looks like: 0x01 0x23 0x4C. Negative 12345 looks like: 0x01 0x23 0x4D.

The length calculation trap: a PIC S9(11) COMP-3 has 11 digits plus 1 sign nibble = 12 nibbles = 6 bytes. A PIC S9(10) COMP-3 has 10 digits plus 1 sign nibble = 11 nibbles = 5.5 bytes, rounded up to 6. Both take 6 bytes. The general formula: ceil((n+1)/2) where n is the digit count in the PIC clause.

PackedDecimalParser.javajava
import java.math.BigDecimal;

public class PackedDecimalParser {

    /**
     * Parse a COMP-3 (packed decimal) field.
     *
     * @param bytes     Raw bytes from the COBOL record
     * @param offset    Start offset within bytes[]
     * @param picDigits Number of digits declared in PIC clause (e.g. 11 for S9(11))
     * @param scale     Implicit decimal places from PIC V clause (0 if no V)
     * @return          BigDecimal representation with correct sign and scale
     */
    public static BigDecimal parse(byte[] bytes, int offset, int picDigits, int scale) {
        int byteLen = (picDigits + 2) / 2; // ceil((n+1)/2)
        StringBuilder digits = new StringBuilder(picDigits + 1);

        for (int i = 0; i < byteLen - 1; i++) {
            int b = bytes[offset + i] & 0xFF;
            digits.append((char) ('0' + (b >>> 4)));   // high nibble
            digits.append((char) ('0' + (b & 0x0F)));  // low nibble
        }

        // Last byte: high nibble = last digit, low nibble = sign
        int lastByte = bytes[offset + byteLen - 1] & 0xFF;
        digits.append((char) ('0' + (lastByte >>> 4)));
        int signNibble = lastByte & 0x0F;

        // D = negative; C and F = positive (unsigned treated as positive)
        boolean negative = (signNibble == 0x0D);

        BigDecimal unscaled = new BigDecimal(digits.toString());
        BigDecimal result   = unscaled.movePointLeft(scale);
        return negative ? result.negate() : result;
    }
}
Never use Java float or double for COMP-3 amounts

IEEE 754 floating-point cannot represent many decimal fractions exactly. A COMP-3 field holding 12345678901234567.89 (18 digits, scale 2 — plausible for a notional on a SAMA RTGS settlement) will lose the last several digits silently when stored in a Java double. There will be no exception. The wrong value will propagate to downstream systems, settlement instructions, and ledger entries. Always use java.math.BigDecimal for every COMP-3 field that represents a monetary amount, rate, or quantity. This is non-negotiable.

DFDL in IBM ACE 12

DFDL (Data Format Description Language, IBM standard) allows you to describe the structure of a binary or text data format in an XML schema, then parse or serialise messages of that format inside an IBM ACE message flow without writing parsing code. For COBOL copybooks, DFDL is the structured alternative to hand-coding an ESQL parser: you describe the copybook once in DFDL, and the ibmDFDL:parseUnparse transform node handles the byte-level parsing.

ACE 12.0.10 ships with a DFDL test tool (accessible via the Toolkit: Run → Run DFDL Schema Editor) that lets you paste raw hex bytes, select a DFDL schema, and see the parsed logical tree. Use this during development to verify your schema before wiring it into a message flow. A DFDL schema that parses a 20-field copybook correctly in the test tool will parse it correctly in production — the parser is the same.

The recommended workflow for generating a DFDL schema from a copybook is to first convert the copybook to XML using cb2xml (an open-source COBOL copybook parser), then hand-edit the cb2xml output into DFDL format. IBM also provides a DFDL schema wizard in ACE Toolkit that accepts COBOL copybook syntax directly, but it handles only a subset of COBOL syntax — REDEFINES, OCCURS DEPENDING ON, and nested group items often require manual correction.

payment-record.dfdl.xsdxml
<?xml version="1.0" encoding="UTF-8"?>
<xs:schema
  xmlns:xs="http://www.w3.org/2001/XMLSchema"
  xmlns:dfdl="http://www.ogf.org/dfdl/dfdl-1.0/"
  xmlns:ibmDFDL="http://www.ibm.com/dfdl/EDI">

  <!-- Global format: EBCDIC binary record, no line terminator -->
  <xs:annotation><xs:appinfo source="http://www.ogf.org/dfdl/dfdl-1.0/">
    <dfdl:format
      encoding="IBM037"
      byteOrder="bigEndian"
      bitOrder="mostSignificantBitFirst"
      lengthKind="implicit"
      representation="binary"
      occursCountKind="fixed"
      textTrimKind="none"
      initiator="" terminator="" separator=""/>
  </xs:appinfo></xs:annotation>

  <!-- Root: PAYMENT-RECORD -->
  <xs:element name="PAYMENT-RECORD">
    <xs:complexType><xs:sequence>

      <!-- PIC X(16): alphanumeric account number, EBCDIC -->
      <xs:element name="ACCOUNT-NBR" type="xs:string">
        <xs:annotation><xs:appinfo source="http://www.ogf.org/dfdl/dfdl-1.0/">
          <dfdl:element representation="text" encoding="IBM037"
            lengthKind="explicit" length="16" lengthUnits="bytes"/>
        </xs:appinfo></xs:annotation>
      </xs:element>

      <!-- PIC S9(11) COMP-3: payment amount, scale 2, 6 bytes -->
      <xs:element name="PAYMENT-AMT" type="xs:decimal">
        <xs:annotation><xs:appinfo source="http://www.ogf.org/dfdl/dfdl-1.0/">
          <dfdl:element
            representation="binary"
            binaryNumberRep="packed"
            lengthKind="explicit" length="6" lengthUnits="bytes"
            totalDigits="11" fractionDigits="2"/>
        </xs:appinfo></xs:annotation>
      </xs:element>

      <!-- PIC 9(7): Julian date YYYYDDD, display numeric -->
      <xs:element name="VALUE-DATE" type="xs:string">
        <xs:annotation><xs:appinfo source="http://www.ogf.org/dfdl/dfdl-1.0/">
          <dfdl:element representation="text" encoding="IBM037"
            lengthKind="explicit" length="7" lengthUnits="bytes"/>
        </xs:appinfo></xs:annotation>
      </xs:element>

      <!-- FILLER 3 bytes: must be accounted for even though omitted from output -->
      <xs:element name="FILLER-1" type="xs:hexBinary">
        <xs:annotation><xs:appinfo source="http://www.ogf.org/dfdl/dfdl-1.0/">
          <dfdl:element representation="binary"
            lengthKind="explicit" length="3" lengthUnits="bytes"/>
        </xs:appinfo></xs:annotation>
      </xs:element>

    </xs:sequence></xs:complexType>
  </xs:element>

</xs:schema>

ACE Mapping Flow

Once DFDL has parsed the raw binary record into an ACE logical tree, ESQL transforms individual fields from the COBOL representation to JSON-ready values. The DFDL parser handles byte layout; ESQL handles semantic conversion — sign propagation, decimal placement, null equivalents, and field renaming.

COBOL null equivalents differ by field type. A PIC X field containing all LOW-VALUE bytes (X'00...') typically represents “not present.” A PIC X field containing all SPACE bytes (X'40...' in EBCDIC) typically represents “initialised but empty.” A COMP-3 field containing all zero bytes with a positive sign is genuinely zero. These distinctions matter for downstream services: mapping LOW-VALUE to null and SPACE to "" is different from mapping both to null.

PaymentRecordMap.esqlsql
CREATE COMPUTE MODULE PaymentRecordToJSON_Map

  CREATE FUNCTION Main() RETURNS BOOLEAN
  BEGIN

    -- DFDL has already parsed InputBody; fields are now typed
    DECLARE srcAmt  DECIMAL;
    DECLARE srcDate CHARACTER;
    DECLARE scale   INTEGER 2;  -- PIC V9(2): 2 implicit decimal places

    -- COMP-3 PAYMENT-AMT: DFDL gives us xs:decimal with fractionDigits=2
    -- but we cast explicitly to ensure scale and sign are preserved
    SET srcAmt = CAST(InputBody.PAYMENT-RECORD.PAYMENT-AMT AS DECIMAL);

    -- Validate: COMP-3 field must not be negative for a debit-only record type
    IF InputBody.PAYMENT-RECORD.TXN-TYPE = 'D' AND srcAmt < 0 THEN
      THROW USER EXCEPTION VALUES('Negative amount on debit record',
        CAST(srcAmt AS CHARACTER));
    END IF;

    -- Julian date YYYYDDD → ISO 8601 YYYY-MM-DD
    SET srcDate = InputBody.PAYMENT-RECORD.VALUE-DATE;
    DECLARE isoDate CHARACTER JulianToISO(srcDate);

    -- Build JSON output tree
    SET OutputRoot.JSON.Data.accountNumber =
        TRIM(InputBody.PAYMENT-RECORD.ACCOUNT-NBR);  -- remove EBCDIC trailing spaces

    -- Emit amount as string to preserve scale (never as float)
    SET OutputRoot.JSON.Data.amount =
        CAST(srcAmt AS CHARACTER);                  -- e.g. "12345.67"

    SET OutputRoot.JSON.Data.currency = 'SAR';    -- ISO 4217
    SET OutputRoot.JSON.Data.valueDate = isoDate;

    -- Null equivalent: LOW-VALUE account number → omit field
    IF BITAND(CCSID(InputBody.PAYMENT-RECORD.ACCOUNT-NBR), 0xFF) = 0 THEN
      SET OutputRoot.JSON.Data.accountNumber = NULL;
    END IF;

    RETURN TRUE;
  END;

END MODULE;

Java-Based Parsing

Not every COBOL mapping runs inside ACE. Kafka Streams processors, Flink jobs, and standalone microservices that receive raw mainframe files need a Java-based parsing path. Three libraries are relevant in practice:

  • cb2xml. Parses a COBOL copybook file into an XML DOM that describes each field’s PIC type, length, and position. Use it to generate DFDL schemas or as the input to a code generator. cb2xml does not parse data — it parses copybook definitions.
  • net.sf.cb2java. A higher-level library that combines copybook parsing with record parsing. Given a copybook and a byte array, it returns a Java object tree with field values. Handles COMP-3 and EBCDIC natively. The library is mature but not actively maintained; pin a specific version and test against your copybooks.
  • JRecord. Designed for fixed-format COBOL records (and CSV/delimited files). Supports COMP-3, COMP, display numeric, and EBCDIC. Better maintained than cb2java. Useful for file-based integration where records arrive as byte arrays from z/OS datasets.

Whichever library you use, the integration test strategy is the same: capture a hex dump from a live mainframe test session, keep it as a test fixture, and run the parser against it in every build. This catches library updates that change EBCDIC or COMP-3 handling silently, and it catches copybook changes that the mainframe team forgot to communicate.

REDEFINES shares storage — choose the right interpretation

A REDEFINES clause means that two (or more) field definitions occupy the same bytes. The classic pattern in core banking is a group item that represents either a domestic or international payment address — same 40 bytes, interpreted as a domestic sort-code structure or as a BIC/IBAN structure depending on a discriminator field elsewhere in the record. If you parse without reading the discriminator first, you will always get the first definition, which is correct for roughly half the records and wrong for the other half. The DFDL dfdl:discriminator element handles this; in hand-coded Java you must read the discriminator before dispatching to the right field group.

Testing COBOL Mappings

COBOL mapping bugs are silent. A COMP-3 field that is 1 byte shorter than expected does not throw an exception — it shifts every subsequent field’s offset by one byte, producing plausible but wrong values for the rest of the record. The only reliable test strategy is round-trip testing against real mainframe data.

Testing approach:

  • Hex dump capture. Have the mainframe team produce a fixed set of test records (covering positive amounts, negative amounts, zero, maximum value, all-space fields, and LOW-VALUE fields) and export them as hex dumps. These become your canonical test fixtures.
  • ACE DFDL debug output. Enable DFDL trace in ACE (set dfdlDebugMode: true in server.conf.yaml for a test server). The trace log shows every field parsed, its byte offset, its value, and the DFDL schema element that matched it. Use this to validate the DFDL schema against the hex dump fixtures before connecting to a live source.
  • Round-trip testing. Parse the byte record to a logical tree, then serialise the logical tree back to bytes using DFDL. Compare the output bytes to the input bytes. Any difference is a schema or parser bug. This catches FILLER fields that are silently dropped during parsing but cause a length mismatch on re-serialisation.
  • Boundary cases. Test the maximum representable value for each COMP-3 field (all nines). Test negative zero (sign nibble D, all digits zero). Test SPACE vs LOW-VALUE for PIC X fields. These are the values that expose off-by-one errors in length calculations and null-handling decisions.

Pitfalls

COMP-3 display length vs storage length confusion

The most common COBOL mapping bug: a developer reads that a field is PIC S9(11) and allocates 11 bytes for it. COMP-3 storage is ceil((11+1)/2) = 6 bytes. The remaining 5 bytes belong to the next field. Every field after that one is misaligned. The parser produces no error — it just reads the wrong bytes for everything that follows. Calculate COMP-3 byte lengths from the formula, not from the digit count.

OCCURS DEPENDING ON creates variable-length records

Standard COBOL processing assumes fixed-length records. OCCURS DEPENDING ON creates records whose length depends on the value of another field in the same record — a count field that determines how many times a sub-group repeats. DFDL handles this with dfdl:occursCount="{../COUNT-FIELD}". Hand-coded Java parsers must read the count before allocating the array. Either way, the record length is no longer a fixed constant — you cannot assume a record is N bytes long without reading it first.

Hybrid COBOL and UTF-8 fields corrupt Arabic characters

Some Saudi bank COBOL systems were extended to hold Arabic customer names by storing UTF-8 bytes in PIC X fields (declared in the copybook with a wider length to account for multi-byte characters). A PIC X(100) field in EBCDIC is 100 bytes; the same field holding a UTF-8 Arabic name may need up to 200 bytes for the same number of visible characters. If the EBCDIC-to-UTF-8 conversion is applied to a field that is already UTF-8, the result is double-converted garbage. Establish which fields are EBCDIC and which are already UTF-8 with the mainframe team before applying any conversion, and document it in the DFDL schema as a comment.

Steps to Onboard a New COBOL Copybook

  1. Obtain the copybook and confirm the code page

    Get the .cpy or .cob copybook file from the mainframe team. Confirm the z/OS job CCSID (37 or 1047 are most common). If the mainframe team is uncertain, request a hex dump of a known test record and check the value of a known alphabetic field character (e.g. the letter A is 0xC1 in CP37 and CP1047, but the dollar sign is 0x5B in CP37 and 0x4A in CP1047).

  2. Calculate field byte lengths and build an offset table

    Work through every field in the copybook in order. Record the PIC type, the storage length (applying the COMP-3 formula where applicable), and the byte offset from the start of the record. Mark FILLER fields explicitly — they consume bytes even though they are not mapped. REDEFINES fields share the byte range of the item they redefine; document which discriminator field selects the active definition.

  3. Generate and validate a DFDL schema

    Run cb2xml against the copybook to produce an XML description. Use the ACE DFDL schema wizard or hand-edit the cb2xml output into a DFDL schema. Test the schema in the ACE DFDL test tool with the hex dump test fixtures. Verify that every field’s parsed value matches the expected value from the mainframe team.

  4. Design the canonical JSON / Avro target schema

    Map each COBOL field to a JSON field with a semantically meaningful name. Convert types: COMP-3 → JSON string (decimal notation), PIC X → JSON string (UTF-8, trimmed), Julian dates → ISO 8601, status codes → enum strings. Document the mapping in a table that both the mainframe team and the consuming service teams can read.

  5. Build the ACE ESQL mapping and unit test it

    Write the ESQL compute module that transforms the DFDL-parsed tree to the JSON output. Unit test every field conversion: amount sign handling, LOW-VALUE null treatment, REDEFINES discriminator logic. Use DFDL debug mode in a development ACE server to validate the parsed tree before the ESQL runs.

  6. Run round-trip and boundary case tests

    Parse every test fixture record. Verify the JSON output against expected values. Serialise the parsed tree back to bytes and compare with the original hex dump. Test the maximum value, negative zero, all-space fields, and at least one REDEFINES alternative. All cases must pass before the mapping is considered production-ready.

  7. Register the copybook and DFDL schema in source control with a governance tag

    Store the copybook, DFDL schema, offset table, and mapping documentation in the integration platform repository under a path that mirrors the mainframe library name. Tag the version with the mainframe SYSMOD level or the date it was extracted. When the mainframe team applies a maintenance release that changes a field length, the regression test suite must run automatically and alert if any test fixture fails.