Every modern payment platform built on microservices faces the same structural problem: a payment does not complete in a single service. Fraud screening lives in one domain, liquidity management in another, external network submission in a third. When all three must succeed for a payment to be valid — and any one of them can fail after the others have already succeeded — you have a distributed transaction problem.
The engineering answer to this problem has a name: the saga pattern. Instead of a single atomic transaction spanning all three services (which requires locking resources across network calls and is incompatible with asynchronous network protocols like SAMA’s IPS), you decompose the workflow into a sequence of local transactions, each of which publishes an event or sends a command that triggers the next step. Failures are handled by compensating transactions — explicit, designed rollback actions for each step that completed before the failure.
The saga pattern is not new. It was first described by Hector Garcia-Molina and Kenneth Salem in 1987. What is new is that it is now the standard approach at every KSA bank modernising its payments stack. And the decisions that determine whether sagas work in production are not technical decisions. They are leadership decisions.
The decision that looks like an architecture choice but isn’t
The first question the architecture team will bring to leadership is: orchestration saga or choreography saga? In an orchestration saga, a central coordinator service drives the workflow — it sends commands to participant services, waits for responses, and decides what to do if a step fails. In a choreography saga, there is no coordinator; each service reacts to events from the previous step autonomously, and the workflow emerges from the sequence of reactions.
Both are technically valid. The choice is architectural — but the trade-off is organisational. Choreography feels cleaner on a diagram: each service is autonomous, there is no central point of failure, teams can develop and deploy independently. In practice, choreography distributes the saga logic across every service that participates in it. When a payment gets stuck — fraud cleared, liquidity reserved, IPS submission timed out, no acknowledgement arrived — the payments operations team needs to know whose responsibility it is to diagnose and resolve the stuck payment. In a choreography saga, the answer is “it depends which service processed it last,” which means three different on-call engineers are involved before the customer’s payment is resolved.
For regulated payment workflows, orchestration is the right answer. A central coordinator means a single saga state table that records every step of every payment saga in one place, queryable by payment ID. When SAMA asks for the audit trail on a specific transaction — and they will — the answer is one SQL query, not a distributed trace across five services.
This is a decision leadership should make explicitly, not leave to the architecture team to discover through the implementation. The organisational implication of choreography — multiple teams responsible for the audit trail of a single payment — is not visible in the sequence diagram.
Who has authority to compensate?
A compensating transaction for a reserved liquidity position releases funds that have been held against a customer account. This is a balance adjustment. In most banks, balance adjustments have an approval workflow, a posting authority limit, and an audit requirement. The engineering team cannot give itself posting authority for compensation transactions; that authority flows from the CFO or the Head of Finance, and it needs to be formally delegated to the payments service before the system goes live.
This sounds procedural, but the practical consequence is significant. If the payments saga cannot automatically release a liquidity reservation that became stranded by an IPS rejection — because the automatic release has not been given posting authority — the operations team must manually intervene on every IPS rejection. At low payment volumes, that is manageable. At SAIB’s IPS transaction volumes, where rejections are a small but non-zero percentage of submissions, it is an operational bottleneck that defeats the purpose of the automation.
The posting authority question must be resolved before the compensation logic is coded. The engineering team needs to know exactly which compensation actions can be executed automatically, which require a dual-control approval from an operations supervisor, and which require escalation to the finance function. Those three tiers map directly to code paths in the saga coordinator.
The operations team is part of the saga design
Every saga implementation must plan for a state it cannot escape from automatically: the payment whose IPS submission status is genuinely unknown because the IPS network is unreachable and has not responded. Compensating an unknown submission risks abandoning funds in transit. Not compensating it locks the liquidity reservation until the network recovers. The saga coordinator must escalate these cases to a human operations team, and that team must have a resolution process before the system goes live.
In practice, this means the payments operations team needs a dashboard that shows stuck sagas by state, age, and payment amount. They need a runbook for each stuck state: what to check in the IPS participant portal, what constitutes confirmation that the submission was received versus rejected, and what approval is needed to manually trigger compensation or settlement. That runbook is not an engineering deliverable — it is a process deliverable owned by the Head of Payments Operations, and it needs the same sign-off as the system’s go-live checklist.
The SLA for resolving stuck sagas during business hours is a leadership decision. At SAIB, the target is 15 minutes from detection to resolution for stuck IPS payments during business hours, with a 30-minute SLA outside business hours backed by an on-call arrangement. Those numbers come from the Head of Retail Banking and the Head of Payments, not from the engineering team. The engineering team builds the tooling to make those SLAs achievable; they cannot set the SLAs themselves.
What SAMA expects to see
SAMA’s Operational Risk Management Guidelines require financial institutions to have documented, tested incident response procedures for payment processing failures. A stuck saga is a payment processing failure. SAMA’s payment examination process will ask: what is your procedure when an IPS credit transfer is in a partially-completed state? Who is responsible? What is the resolution time? How is the customer notified?
A bank that can answer those questions with a documented runbook, an SLA, and evidence of a quarterly test is in a defensible position. A bank that has a sophisticated orchestration saga implementation but no documented resolution process for stuck sagas has built something technically impressive that it cannot defend to a regulator.
Three decisions that cannot be delegated
Which compensation actions are fully automated, and which require human approval? The automatic compensation tier needs formal posting authority delegated to the payments service. The manual tier needs a defined approval workflow and an SLA for each action type. The escalation tier needs clear ownership — which individual or committee has authority to trigger a compensation above the manual tier threshold.
Who owns the saga state table for audit purposes? In an orchestration saga, the coordinator’s saga state table is the authoritative record of every step in every payment workflow. It must be governed with the same retention and access controls as the payment ledger — which means a named data owner, an access control review process, and a retention policy that meets SAMA’s 10-year record-keeping requirement. The engineering team cannot self-appoint as the data owner of a financial record.
What is the escalation path for a saga stuck in an UNKNOWN state? The UNKNOWN state — where the IPS network has not confirmed whether it received a submission — cannot be resolved by software alone. The resolution process involves a human querying the SAMA IPS participant portal and making a judgement call about whether to compensate or await settlement. That person, their backup, and their authority to act must be identified and documented before the system goes live.
For the engineering depth behind this topic — Axon Framework saga implementation, choreography vs orchestration state machines, compensating transaction patterns, idempotency in saga steps, deadline handling for IPS timeouts, and integration testing with Testcontainers — see the companion Lab article: Event-Driven Architecture: Saga Patterns for Distributed Transactions.