Every Kafka adoption at scale eventually encounters the same crisis. A payments team renames a field in their event schema to align with a new internal naming convention. The change looks minor — a field called amount becomes instrAmt to match ISO 20022 terminology. The CI pipeline is green. The deployment succeeds. And then, silently, every consumer of that topic starts returning null for the amount field. The fraud scoring model receives a zero-value feature. The reconciliation job writes a zero to the ledger. The regulatory reporting job produces totals that are wrong by the entire day’s transaction volume. None of these consumers threw an exception; they all handled the null gracefully and continued. The failure mode is invisible until the business notices.
This is not a Kafka problem and it is not a technical problem. It is a team coordination problem that looks like a technical problem because teams that should be coordinating are instead operating independently on a shared interface — the event schema — without an enforced contract. Schema Registry is the technical mechanism that enforces that contract. But before the technology can help, leadership has to decide that the contract matters and that enforcement has teeth.
What this decision actually is
Schema governance is a decision about who owns the contract between producer and consumer teams, and what the consequence is when the contract is broken. In a traditional API world, this is resolved by the API owner publishing a versioned contract and consuming teams subscribing to the version they depend on. In an event-driven architecture, the same discipline is required but the enforcement mechanism is different: the schema registry is the versioned contract store, and the compatibility mode configured on each subject is the policy that determines what changes are and are not allowed.
The decision a VP or CIO needs to make is not “should we use Confluent Schema Registry?” — that is an implementation choice the platform team can evaluate. The decision is: what is the schema governance model for event streams in this organisation, and who enforces it? This means designating a schema governance board or function (which teams or roles review and approve compatibility mode changes and breaking schema changes), establishing a compatibility policy per topic class (payment event topics use FULL compatibility; internal monitoring topics may use BACKWARD; experimental topics use NONE), and allocating a migration budget when breaking changes are necessary (because sometimes they are, and the cost of a planned migration is always lower than the cost of a production incident).
The compatibility mode on a Kafka subject is an organisational decision disguised as a configuration flag. Setting it without a governance process is as meaningful as publishing an API versioning policy that no one is authorised to enforce.
What the regulator cares about
SAMA’s Technology Risk Management framework, and Saudi Arabia’s Personal Data Protection Law, impose specific requirements on event-driven architectures that go beyond what most platform teams consider when they deploy a schema registry. The regulator cares about three things in this space.
Immutable audit schemas. A transaction event published to a Kafka topic must be replayable with the exact schema that was in use at the time it was produced. Schema Registry’s versioned schema archive is the mechanism that makes this possible: each message on the wire carries a schema ID, and that schema ID resolves permanently to the schema used at production time. But this only holds if the registry is configured to retain all schema versions indefinitely for regulated topics — not to compact or delete them. The platform team must configure the _schemas internal topic with infinite retention for regulatory subjects, and must verify that schema deletion is disabled or gated behind a governance approval process.
PDPL data classification embedded in events. Events that carry personal data — customer IBANs, national IDs, transaction counterparty names — must be identifiable as such at the event level, not only at the database level. The correct pattern is to embed a pdplClassification field in every Avro schema that carries personal data, with a controlled vocabulary: PUBLIC, INTERNAL, CONFIDENTIAL, or RESTRICTED. This makes data classification machine-readable: downstream consumers, data governance pipelines, and DLP scanning tools can act on the classification without out-of-band configuration. A schema that lacks this field on a topic carrying customer payment data is a PDPL compliance gap.
Schema archive for regulatory replay. When SAMA requests a replay of transactions from a specific time window — for examination, investigation, or reconciliation — the replay must produce exactly the data that existed at that time, using the schema that was in use at that time. This is only possible if the schema registry has retained every historical schema version. An organisation that compacted or deleted old schema versions has made regulatory replay materially harder and potentially impossible for events produced before the deletion. Treat schema version retention as a data retention policy subject to the same governance controls as the event data itself.
The business trade-off
Strict compatibility enforcement slows delivery. This is not a theoretical concern — it is what product teams report when they encounter FULL compatibility for the first time. A developer who wants to add a new required field to a payment event schema discovers that FULL compatibility forbids it (adding a required field without a default breaks backward compatibility). They must instead add the field as optional with a default, negotiate a migration plan with all consumer teams, and follow a multi-phase rollout. What feels like a one-day schema change becomes a two-sprint migration. That is the correct outcome — but it requires the team to understand why and to have a governance process that helps them navigate it efficiently, not just a registry that rejects their change with a cryptic error.
The alternative — no schema governance — has a different and worse cost structure. Breaking schema changes are free to make and free to deploy right up until the moment they reach production, at which point the cost is a production incident, potentially data corruption in downstream systems, and a reconciliation effort that can take days. In a regulated environment, a payment amount incorrectly calculated because a field was renamed is not an operational incident that can be silently resolved. It is a material data integrity event that must be reported, investigated, and remediated under SAMA’s incident management framework.
The business trade-off is between the predictable friction of planned migrations under a governance regime and the unpredictable cost of unmanaged schema drift at production scale. For a bank with dozens of Kafka topics, multiple producer teams, and consuming systems that include real-time fraud engines, regulatory reporting pipelines, and core banking integrations, the calculus is unambiguous: the governance regime wins on expected value, even before the regulatory dimension is considered.
What leadership must own
Platform teams cannot own schema governance. They can operate the Schema Registry, enforce technical compatibility checks in CI, and advise on migration patterns. But the governance decisions — who can relax a compatibility mode, what constitutes an approved breaking change, how long schema versions are retained — require cross-team authority that platform engineers do not have and should not be asked to exercise.
The VP of Integration or the Chief Architect needs to own three specific commitments. First: a compatibility policy per topic class, published and communicated to all producer and consumer teams before they onboard to the event platform. Payment topics: FULL. Internal infrastructure topics: BACKWARD. Experimental topics: NONE, with a mandatory promotion to BACKWARD or FULL before any consumer outside the producing team onboards. Second: a governance process for compatibility exceptions. When a team genuinely needs to make a breaking schema change, there must be a defined path: a migration plan reviewed by all affected consumer teams, a migration budget allocated, and an approval that names the person who accepted the risk. Without this path, teams will route around the registry — setting subjects to NONE mode to bypass validation — and the governance regime will fail silently. Third: schema retention aligned with data retention policy. Every regulated event topic’s schema history must be retained for the same period as the event data. If payments data is retained for seven years for regulatory purposes, schema versions for payment topics must be retained for seven years. This is a data governance policy, not a platform configuration detail.
A Schema Registry that is deployed, that runs CI compatibility checks, and that is backed by a cross-team governance process is a genuine data contract enforcement mechanism. It makes the event-driven architecture more reliable, more auditable, and more defensible in a regulatory examination. It is worth the delivery friction it introduces. The leadership decision is to commit to the governance model that makes it work, not just to deploy the technology.
For the engineering depth behind these decisions — Avro vs Protobuf format selection, compatibility mode semantics, ISO 20022 Avro mapping, CI integration patterns, and the production checklist — see the companion Lab article: Data Streaming: Schema Registry & Data Contracts.