Why sagas, not 2PC

A payment initiation at SAIB crosses at least four bounded contexts before it settles: fraud screening, liquidity reservation, IPS network submission, and notification delivery. Each context has its own database, its own team, and its own deployment cadence. They cannot share a transaction boundary.

Two-phase commit (2PC) was the industry answer to this problem for two decades. The XA protocol locks resources across participants during the prepare phase and releases them on commit or rollback. At the scale and latency profile of a modern payment system, 2PC has two fatal problems. First, it holds locks across network calls — if the IPS network submission times out at 800ms, every row touched in the fraud and liquidity contexts is locked for those 800ms, blocking concurrent payments. Second, XA requires all participants to implement the XA protocol, which eliminates Kafka, most REST services, and any external party. In practice, IPS is a SAMA-operated network with a SWIFT-style asynchronous API. It cannot participate in a synchronous XA commit.

The saga pattern replaces the distributed transaction with a sequence of local transactions, each of which publishes an event or sends a command that triggers the next step. Failures are handled by compensating transactions — explicit, designed rollback actions for each step. The trade-off is clear: you give up atomicity and get availability and decoupling in return. For a payment system, that is the right trade-off. A payment that partially fails and rolls back cleanly is better than a payment system that deadlocks under load.

BASE, not ACID

Sagas operate on BASE semantics: Basically Available, Soft state, Eventual consistency. The intermediate states of a saga — where fraud is cleared but liquidity is not yet reserved — are visible to other parts of the system. Design your read models to handle these transient states gracefully rather than pretending they don’t exist. A payment in FRAUD_CLEARED status is a valid business state that the customer-facing status API must represent correctly.

Choreography sagas

In a choreography saga, each service listens for events and reacts autonomously. There is no central coordinator. The workflow emerges from the sequence of events and reactions.

The choreography model has significant appeal at the start of a project: no new infrastructure, no coordinator service to build and operate, no single point of failure in the workflow. Each team owns their slice of the workflow and their event contract. The complexity is distributed, not centralised.

The catch surfaces as the saga grows. When a payment fails at the IPS submission step and requires a compensating liquidity release, which service is responsible for emitting the compensation trigger? In a choreography model, the answer is the ips-svc emits an IPSRejected event, and the liquidity-svc listens for it and releases the reservation. This works as long as the logic is simple. When you add a third failure path — a SAMA regulatory hold that can arrive after liquidity is reserved but before IPS submission — the event graph gains a new edge, and every service that might need to react must be updated. In a choreography saga, the implicit workflow is encoded in the event subscriptions of dozens of services. It is untraceable without distributed tracing, and it cannot be changed without coordinating changes across multiple teams simultaneously.

Orchestration sagas

In an orchestration saga, a dedicated coordinator service drives the workflow by issuing commands to participant services and reacting to their responses. The workflow state lives in the coordinator, not distributed across the event stream.

PaymentSaga.java — Axon Framework orchestration sagajava
@Saga
public class PaymentSaga {

    @Autowired
    private transient CommandGateway commandGateway;

    private String paymentId;
    private String debtorIban;
    private BigDecimal amount;
    private String reservationRef;       // stored for compensation

    @StartSaga
    @SagaEventHandler(associationProperty = "paymentId")
    public void on(PaymentInitiatedEvent e) {
        paymentId  = e.getPaymentId();
        debtorIban = e.getDebtorIban();
        amount     = e.getAmount();
        // Step 1: request fraud check
        commandGateway.send(new RequestFraudCheckCommand(paymentId, amount, debtorIban));
    }

    @SagaEventHandler(associationProperty = "paymentId")
    public void on(FraudCheckPassedEvent e) {
        // Step 2: reserve liquidity
        commandGateway.send(new ReserveLiquidityCommand(paymentId, amount, debtorIban));
    }

    @SagaEventHandler(associationProperty = "paymentId")
    public void on(FraudCheckFailedEvent e) {
        // Terminal failure — no compensation needed (nothing reserved yet)
        commandGateway.send(new FailPaymentCommand(paymentId, "FRAUD_DECLINED", e.getReason()));
        SagaLifecycle.end();
    }

    @SagaEventHandler(associationProperty = "paymentId")
    public void on(LiquidityReservedEvent e) {
        reservationRef = e.getReservationRef();
        // Step 3: submit to IPS
        commandGateway.send(new SubmitToIPSCommand(paymentId, amount, debtorIban, e.getCreditorIban()));
    }

    @SagaEventHandler(associationProperty = "paymentId")
    public void on(LiquidityInsufficientEvent e) {
        // Terminal failure — no liquidity reserved, so no compensation
        commandGateway.send(new FailPaymentCommand(paymentId, "INSUFFICIENT_FUNDS", e.getReason()));
        SagaLifecycle.end();
    }

    @SagaEventHandler(associationProperty = "paymentId")
    public void on(IPSRejectedEvent e) {
        // IPS rejected: must release the liquidity reservation
        commandGateway.send(new ReleaseLiquidityCommand(paymentId, reservationRef, "IPS_REJECTED"));
    }

    @SagaEventHandler(associationProperty = "paymentId")
    public void on(LiquidityReleasedEvent e) {
        // Compensation complete — mark payment failed
        commandGateway.send(new FailPaymentCommand(paymentId, "IPS_REJECTED", e.getReason()));
        SagaLifecycle.end();
    }

    @SagaEventHandler(associationProperty = "paymentId")
    public void on(PaymentSettledEvent e) {
        // Happy path complete
        SagaLifecycle.end();
    }
}

The saga state — the reservationRef field — is serialised by Axon and stored in the saga store (a saga_entry table in PostgreSQL). If the coordinator restarts between steps, Axon reloads the saga state and continues from where it left off. The workflow is durable: a JVM crash between the LiquidityReservedEvent and the SubmitToIPSCommand does not orphan a reserved liquidity position — the saga resumes and reissues the command on restart.

Choosing between the two

DimensionChoreographyOrchestration
Workflow visibilityDistributed across event subscriptions — hard to traceCentralised in the coordinator — one class, one place
CouplingServices coupled only to event contractsCoordinator knows participant command interfaces
Single point of failureNone — each service is independentCoordinator is a critical service (mitigated by HA)
Compensation logicEncoded in event subscriptions — hard to reason about under partial failureEncoded in the coordinator — explicit, testable
Adding a new stepNew service subscribes; existing services unawareCoordinator updated; all steps in one place
Regulatory auditabilityDistributed trace required to reconstruct the workflowSaga state table is a complete, queryable workflow log
Best fitSimple, stable workflows with 2–3 steps; event fan-out patternsComplex multi-step workflows with compensation; regulated processes

For payment workflows in a regulated environment, the recommendation is orchestration. The SAMA audit requirement alone tips the balance: a single saga_entry table that records every state transition of every payment saga — with timestamps — satisfies the “reconstruct any transaction at time T” requirement without stitching together distributed traces from five different services. The coordinator is also easier to explain to an examiner than “the workflow is implicit in the event subscriptions across twelve microservices.”

Compensating transactions

A compensating transaction is not a rollback. A database rollback undoes a local transaction atomically — as if it never happened. A saga compensation acknowledges that the prior step did happen and applies an explicit corrective action. The prior step’s effects may have been visible to other parts of the system during the window between the step completing and the compensation completing. Design compensations accordingly.

  1. Design compensations before you design the forward steps

    Every step in a saga must have a defined compensation before the step is implemented. If you cannot design a compensation, the step should not be asynchronous — either make it synchronous (so it participates in the local transaction boundary) or restructure the domain to make compensation unnecessary (e.g., soft-reserve before hard-reserve).

  2. Compensations must be idempotent

    A compensation command may be delivered multiple times if the network times out between the coordinator sending it and the participant acknowledging it. The participant must handle a duplicate ReleaseLiquidityCommand for an already-released reservation without error or double-release. The simplest implementation: check the reservation status before acting; if already released, return success.

  3. Compensations must be retriable

    A compensation that fails must be retried until it succeeds. There is no “fail the compensation” path — if the compensation fails, the system is in an inconsistent state. The saga coordinator must be configured with a retry policy and a dead-letter queue for compensation commands that cannot be delivered after N retries. A human operations process must exist to resolve dead-lettered compensations.

  4. Order compensations in reverse saga order

    If steps A → B → C all succeeded and step D failed, compensate in the order C′ → B′ → A′. Do not compensate A before compensating C — earlier steps may have created dependencies that the later steps rely on. At SAIB, the IPS submission acknowledgement contains the IPS reference number that the liquidity compensation uses; releasing liquidity before the IPS rejection is confirmed loses that reference.

  5. Compensations are business events, not technical housekeeping

    A LiquidityReleasedEvent is a business event that fraud analytics, the customer notification service, and the SAMA audit trail all care about. Emit it as a first-class domain event through the same Kafka topic as the forward events. Do not treat compensations as internal plumbing that bypasses the event log.

  6. Some steps are non-compensable

    Once a pacs.008 is submitted to the IPS network, it cannot be withdrawn — only returned via a pacs.004 payment return, which is a new forward transaction, not a compensation. Model this correctly: the saga compensates by initiating a payment return, not by “undoing” the IPS submission. The IPS submission is a pivot point — steps before it can be compensated cleanly; steps after it require a new forward flow.

Pivot points change the compensation design

Identify the pivot point in your saga — the step after which compensations become forward transactions — and document it explicitly. For a KSA IPS payment, the pivot is the moment the pacs.008 is accepted by the IPS network. Before the pivot: cancel and release. After the pivot: initiate a return. The PaymentSaga needs different code paths for pre-pivot and post-pivot failures.

IPS payment saga end-to-end

A concrete walkthrough of the happy path and the two most common failure paths for an outbound IPS credit transfer at SAIB.

The IPS rejection path is the one that matters most operationally. An IPS rejection arrives asynchronously — the pacs.002 rejection response from SAMA can arrive seconds to minutes after the pacs.008 submission. The saga must hold state (the reservationRef) across that window and act correctly when the rejection arrives. This is the case that synchronous, stateless approaches cannot handle without a polling loop against the IPS network, which introduces its own failure modes.

Failure & recovery patterns

Three failure scenarios have distinct handling requirements. Get them wrong and you either orphan resources (reserved funds that are never released) or create phantom payments (payments recorded as settled that were actually rejected).

Saga step timeout is not the same as saga failure

A ReserveLiquidityCommand that times out at 500ms does not mean liquidity reservation failed. It means the coordinator did not receive an acknowledgement. The liquidity service may have processed the command successfully and the response was lost. Before compensating, the coordinator must query the liquidity service’s idempotency record for the given paymentId. If it confirms reservation, continue the saga. If it confirms no reservation, retry the command once. Only compensate after a confirmed failure, never after a timeout alone.

saga-config.yaml — Axon deadline for IPS acknowledgement timeoutyaml
axon:
  deadline:
    enabled: true
  serializer:
    general: xstream

# In PaymentSaga.java — schedule a deadline after IPS submission
# If IPSAcknowledgedEvent not received in 90s, fire the deadline handler
deadlineManager.schedule(
    Duration.ofSeconds(90),
    "ipsSubmissionTimeout",
    new IPSSubmissionDeadline(paymentId)
);

# Handler: query IPS status, compensate or extend depending on result
@DeadlineHandler(deadlineName = "ipsSubmissionTimeout")
public void onIPSTimeout(IPSSubmissionDeadline deadline) {
    IPSStatus status = ipsQueryService.checkStatus(deadline.getPaymentId());
    if (status == IPSStatus.PENDING) {
        // Extend: IPS is slow, reschedule
        deadlineManager.schedule(Duration.ofSeconds(30), "ipsSubmissionTimeout", deadline);
    } else if (status == IPSStatus.REJECTED) {
        AggregateLifecycle.apply(new IPSRejectedEvent(paymentId, "TIMEOUT_CONFIRMED_REJECTED"));
    } else if (status == IPSStatus.UNKNOWN) {
        // Escalate to ops dead-letter queue
        deadLetterQueue.park(paymentId, "IPS_STATUS_UNKNOWN");
    }
}

The UNKNOWN status branch is the one that requires an operational process. When the IPS network is unreachable and the coordinator cannot determine whether the submission was received, the payment cannot be compensated automatically — compensating an IPS-received submission would abandon the funds in transit. These cases go to a dead-letter queue monitored by the payments operations team, who reconcile directly with SAMA’s IPS participant portal before acting.

The dead-letter queue is a design requirement, not a failure mode

Every saga implementation that touches an external network must design for the UNKNOWN state. Build the dead-letter queue, the operational dashboard, and the resolution workflow before the system goes live. At SAIB, the IPS operations team has a reconciliation SLA of 15 minutes for unknown-state payments during business hours. That SLA is achievable only if the tooling was built ahead of go-live.

Idempotency in saga steps

Every saga step must be idempotent at the participant service level. The coordinator may resend a command if it does not receive an acknowledgement. The participant must handle the duplicate without creating a duplicate effect.

LiquidityService.java — idempotent reservationjava
@CommandHandler
public ReservationResult handle(ReserveLiquidityCommand cmd) {
    // Idempotency: check whether reservation already exists for this paymentId
    return reservationRepo
        .findByPaymentId(cmd.getPaymentId())
        .map(existing -> {
            // Duplicate command: return the existing reservation
            log.info("Duplicate ReserveLiquidity for {}: returning existing ref {}",
                cmd.getPaymentId(), existing.getReservationRef());
            return new ReservationResult(existing.getReservationRef(), existing.getReservedAt());
        })
        .orElseGet(() -> {
            // First attempt: create the reservation
            Account account = accountRepo.findByIban(cmd.getDebtorIban())
                .orElseThrow(() -> new AccountNotFoundException(cmd.getDebtorIban()));
            if (account.getAvailableBalance().compareTo(cmd.getAmount()) < 0) {
                throw new InsufficientFundsException(cmd.getPaymentId(), cmd.getAmount());
            }
            Reservation reservation = new Reservation(
                UUID.randomUUID().toString(),
                cmd.getPaymentId(), cmd.getAmount(), Instant.now()
            );
            account.holdAmount(cmd.getAmount());
            reservationRepo.save(reservation);
            accountRepo.save(account);
            return new ReservationResult(reservation.getRef(), reservation.getCreatedAt());
        });
}

The idempotency check uses the paymentId as the natural key, not the CommandMessage identifier. Using the Axon message ID would require the coordinator to always resend with the same message ID, which Axon does not guarantee across restarts. The business key (paymentId) is stable across retries because the coordinator always sends the same command for the same payment.

Testing sagas

Sagas are integration-heavy components. Unit testing the saga class with mocked events is necessary but not sufficient — the interesting failure modes require a real Kafka cluster, a real PostgreSQL saga store, and real participant services (or stub implementations with realistic failure injection).

PaymentSagaIntegrationTest.java — Testcontainersjava
@SpringBootTest
@Testcontainers
class PaymentSagaIntegrationTest {

    @Container
    static PostgreSQLContainer<?> postgres =
        new PostgreSQLContainer<>("postgres:15")
            .withDatabaseName("payments_test");

    @Container
    static KafkaContainer kafka =
        new KafkaContainer(DockerImageName.parse("confluentinc/cp-kafka:7.7.0"));

    @Autowired
    private CommandGateway commandGateway;

    @Autowired
    private PaymentStatusRepository statusRepo;

    @Test
    void ipsRejection_shouldReleaseLiquidityAndFailPayment() throws Exception {
        String paymentId = UUID.randomUUID().toString();

        // Configure stub IPS service to return rejection
        ipsStub.willReturn("REJECTED", "AC04");  // AC04 = closed creditor account

        CompletableFuture<?> result = commandGateway.send(new InitiatePaymentCommand(
            paymentId, "SA4420000001234567891234", "SA7710000000012345678901",
            new BigDecimal("5000.00"), "SAR",
            "E2E-20260808-001", LocalDate.now()
        ));
        result.get();

        // Wait for saga to complete (async: IPS rejection → compensation → terminal)
        Awaitility.await().atMost(Duration.ofSeconds(10))
            .until(() -> statusRepo.findByPaymentId(paymentId)
                .map(s -> PaymentStatus.FAILED == s.getStatus())
                .orElse(false));

        PaymentStatus status = statusRepo.findByPaymentId(paymentId).orElseThrow();
        assertThat(status.getFailureReason()).isEqualTo("IPS_REJECTED");

        // Verify liquidity was released
        assertThat(reservationRepo.findByPaymentId(paymentId).get().getStatus())
            .isEqualTo(ReservationStatus.RELEASED);
    }
}

Integration tests for sagas must cover: the happy path, each individual step failure, the timeout/deadline path, and the duplicate command path (idempotency). The timeout path requires configuring short deadlines in the test context — a 90-second deadline is untestable in a normal test suite. Use a Spring profile that overrides the deadline to 500ms for integration tests.

Anti-patterns

Saga as a substitute for domain design

A saga that spans ten steps and compensates eight of them is a sign that the domain model has not been designed correctly. If almost every step needs a compensation, the workflow boundary is probably wrong — either some steps should be in the same local transaction (same database, same service), or some steps should be decoupled from the main flow and treated as eventual consistency without compensation. Sagas are for the genuinely distributed coordination that cannot be avoided; they are not a way to avoid thinking about bounded context boundaries.

Storing saga state in the event log

The saga state table (saga_entry) and the event store serve different purposes. The event store records what happened; the saga table records the coordinator’s current position in the workflow. Do not try to derive saga state by replaying events — the saga coordinator needs to look up its state quickly on restart, not replay a potentially large event history. Axon’s saga store serialises and stores the saga state directly; use it.

Choreography for workflows that need auditability

A choreography saga that spans five services produces an audit trail distributed across five event streams and five databases. Reconstructing the complete history of a single payment requires joining those streams on paymentId across five different data stores — which may have different retention policies, different access controls, and different log formats. For SAMA audit purposes, an orchestration saga with a centralised saga state table is significantly easier to present to an examiner. If you are already using choreography, invest in a correlation-ID-based event correlation service that aggregates the per-payment audit trail into a single queryable view.