Every messaging platform will eventually encounter a message it cannot process. A payment initiation that fails schema validation because the originating channel deployed a new field the consumer had not yet adopted. An account update that references a customer ID that was merged in the core banking system an hour ago. A notification event that triggers a downstream call to a service in the middle of a deployment. The question is not whether this will happen. It will. The question leadership must answer in advance is: when it does, what happens next?
The naive answer — the default in most messaging frameworks without explicit configuration — is retry. The consumer tries again. And again. And again. The message stays at the head of the queue, blocking every other message behind it. Consumer throughput collapses. The downstream system that depends on this queue falls behind its SLA. A single malformed payment message from a channel deployment has now blocked a queue that normally processes 50,000 transactions per hour. The incident that follows is not a messaging failure; it is a design failure.
The leadership question
The dead letter queue (DLQ) is the mechanism that answers the question “what happens next?” correctly. When a consumer exhausts its retry budget on a message it cannot process, the message is routed to a separate dead letter queue rather than blocking the main pipeline. The main queue continues processing. The failed message is preserved for analysis, repair, and replay. The operations team is alerted. Normal processing continues.
This sounds simple. The decision that leadership must make is not whether to have a DLQ — every production messaging platform in a financial institution should have one — but whether to invest in making the DLQ operational rather than merely configured. A DLQ with no consumer process, no retention policy, no alert, and no replay procedure is functionally equivalent to dropping the message silently. The message is preserved, but it is invisible to the team and unreachable by the customer whose transaction it represents. The DLQ becomes a message graveyard rather than a safety valve.
The DLQ is not a solution to message failure. It is a holding pen that gives you time to find the solution without blocking the pipeline. Its value is only realised if someone is watching it and has the tools and authority to act on what lands there.
The leadership question, then, is: who owns the DLQ? Who monitors it? What is the SLA for a message to be classified and either replayed or archived? Who has the authority to approve a bulk replay operation after a consumer bug is fixed? These are organisational decisions, not engineering decisions, and they must be made before the first message lands in the DLQ, not during the first incident where it matters.
What the regulator cares about
SAMA’s Technology Risk Management framework and the operational requirements for IPS participation create a specific set of obligations around payment message failures. A payment message that fails processing — a pacs.008 credit transfer that is rejected, a payment status update that cannot be applied — is not simply a failed transaction. It is a financial event that must be recoverable, auditable, and resolved within a defined timeframe.
The SAMA examination question on messaging failures is not “do you have retry logic?” It is “can you demonstrate that no payment message was lost, that every failure was logged with the reason and timestamp, and that failed messages were either successfully replayed or formally written off within your operational SLA?” A DLQ that is properly monitored, with a handler that classifies each failure and routes it to a replay queue or a compliance archive, produces exactly the audit trail that satisfies this question.
The compliance dimension is more acute for payment messages than for other event types. A pacs.008 that fails SAMA IPS processing and lands in a DLQ may be subject to a regulatory hold if the reason for failure involves a sanctions check or an AML alert. Such messages must never be auto-deleted or auto-expired. The DLQ retention policy for payment topics must be set with compliance requirements, not operational convenience, as the primary constraint. 90 days is a reasonable minimum; permanently archiving flagged messages to immutable cold storage before the DLQ retention window expires is the correct disposition for any message with a compliance flag.
The business trade-off
The argument against a properly operationalised DLQ is the same argument made against any resilience infrastructure: it adds complexity and it costs engineering time to build and operate. A consumer that retries forever is simpler to write than a consumer that has explicit retry budgets, configures a dead letter topic, and publishes structured metadata to the DLQ handler. A DLQ handler that classifies failures and routes them to different downstream queues is more complex than one that simply alerts and waits for manual intervention.
The question is what the alternative costs. A consumer without a DLQ that encounters a poison message blocks its queue. In a Kafka partition, that means every message behind the poison message is also blocked, for as long as it takes to diagnose and resolve the issue. In IBM MQ, a message that causes a consumer to roll back reaches its backout threshold and lands on the MQ dead letter queue — but if the DLQ itself has no consumer, the queue manager’s DLQ eventually fills, at which point new messages that would ordinarily land there instead cause the message put to fail, propagating the error upward to the sending application.
The customer-visible consequence of a blocked payment queue is a payment that neither completes nor fails explicitly — it is simply pending, indefinitely. The customer’s mobile app shows “payment in progress” while the message that represents their transaction sits behind a poison message in a blocked consumer. The operations team receives an alert about consumer lag growing, investigates, and manually resolves the blockage. Total elapsed time: 30 minutes to 2 hours, depending on how quickly the on-call engineer is available and how well the runbook covers this scenario. For a payment that a customer needed to complete before the end of business, 2 hours is a customer-visible outage with reputational and potentially regulatory consequences.
The DLQ, properly built, reduces this scenario to an alert, a classification, and a replay. The failing message is routed out of the pipeline in milliseconds. Normal processing continues. The fix is developed and deployed. The message is replayed. Total customer impact: the payment is delayed, not blocked. The distinction matters to the customer. It also matters to the SAMA examiner.
What the team needs
The team that operates a messaging platform with a proper DLQ needs three things that engineering cannot provide for itself.
Authority to define message failure SLAs. How long can a payment message sit in the DLQ before the operations team must act? This is a business decision, not a technical one. For a payment message during business hours, 15 minutes is a defensible SLA. For a payment message received after banking hours, 4 hours before the next business day’s opening may be acceptable. These SLAs must be agreed with the business and the compliance team before they can be operationalised in an alerting policy. Engineering can implement whatever SLA is agreed; it cannot unilaterally decide what the SLA should be for a message that represents a customer’s financial transaction.
Budget for DLQ monitoring tooling. A DLQ handler, a Grafana dashboard showing per-topic DLQ depth and message age, and the alert routing that connects DLQ events to the on-call rotation are not free. They require engineering time to build and ongoing time to maintain. The operations case for this investment is straightforward — the cost of a two-hour payment queue outage, measured in customer impact and potential regulatory notification obligations, exceeds the cost of the tooling many times over — but the budget must be explicitly allocated. Treating DLQ tooling as “nice to have” in sprint planning means it never gets built before the first incident that demonstrates its necessity.
Process for replay sign-off. Replaying messages from a DLQ is not a purely technical operation, especially for payment messages. Before a bulk replay is executed, someone must verify that the consumer bug that caused the failures has been fixed in the deployed version, that the messages being replayed are safe to reprocess (idempotency has been confirmed), and that no regulatory hold applies to any message in the replay batch. This sign-off process requires a named approver with both technical context (the engineer who deployed the fix) and business authority (the payments operations lead). Without a defined process, replay becomes a high-risk ad hoc operation that engineers either avoid for fear of double-processing or execute without adequate verification.
For the engineering depth behind these decisions — Kafka @RetryableTopic configuration, IBM MQ AMQMDLQ handler, poison message classification, replay with idempotency guarantees, and the production checklist — see the companion Lab article: Event-Driven: Dead Letter Queue Design & Event Replay Patterns.