Real-time fraud detection sounds like a technology project. Deploy Apache Flink, wire it to your Kafka topics, configure a velocity check, and ship it. In practice, the technology is the easy part. The hard part is answering a set of architectural questions whose answers are not reversible once you are in production with live payment traffic running through them: How many milliseconds of latency is acceptable between a transaction arriving and a fraud decision being made? What does it mean, at the business layer rather than the Kafka layer, for processing to happen “exactly once”? Who owns the audit trail when a declined transaction is later found to have been a false positive? What happens to the fraud engine’s state during a recovery after a system failure?
These are leadership decisions. Engineers can implement the answer once the answer is given, but they cannot set the answer alone, because the answer involves trade-offs between latency, cost, regulatory obligation, and acceptable business risk that only the business can authorise.
What this decision actually is
The fundamental decision in stateful stream processing is not which framework to use. It is the consistency model the fraud infrastructure operates under, and what the consequences of that model are when things go wrong.
A stream processing job that uses at-least-once delivery semantics will process every payment event at least once, and potentially more than once during recovery from a failure. For most streaming use cases, this is acceptable — count the extra processing, absorb the duplicate, move on. For a fraud velocity check on a payment system, it is not. If the velocity check counts a transaction twice because the Flink job recovered from a checkpoint and re-processed the last batch of events, the customer’s transaction count for the window is wrong. The fraud signal is wrong. The decision — approve, flag, or block — is made on incorrect state. In the worst case, a legitimate transaction is blocked because a double-counted window exceeded the threshold.
Exactly-once processing — the guarantee that every event updates operator state exactly once, even under failure and recovery — addresses this. But it is not a configuration flag. It is an architecture commitment that changes how the job is deployed (checkpointing interval, state backend, Kafka transactional producer), how the downstream systems must behave (idempotent writes keyed on transaction ID), and what latency the end-to-end system delivers (checkpointing adds measurable overhead; Kafka transactions add commit latency). The commitment must be made at the design stage, not retrofitted after the first production incident.
Exactly-once at the stream processor layer is a necessary condition for correct fraud detection, but it is not sufficient. The downstream authorisation system must also be idempotent. A Flink job that guarantees exactly-once state updates but writes to an authorisation gateway that does not deduplicate decisions can still produce duplicate BLOCK outcomes for a single transaction.
The regulatory angle
SAMA’s real-time fraud requirements are not advisory. The Technology Risk Management framework expects that velocity-based controls — limits on transaction count and amount within defined time windows — operate continuously and correctly on all regulated payment channels. For Mada card authorisations, the screening decision must be made within the network’s SLA window. For IPS credit transfers, SAMA expects a status response within seconds. These are not aspirational targets; they are operational requirements that a SAMA examination will test.
The velocity limits themselves are regulatory controls. When SAMA defines that a card must not exceed N transactions in a rolling window without step-up authentication, that window computation must be correct. An at-least-once pipeline that double-counts transactions during recovery is not implementing the regulatory control correctly — it is implementing a control that looks correct in normal operation and fails silently in edge cases. The examination question is not “do you have a velocity check?” but “can you demonstrate that the velocity check is applied consistently and that its results are auditable?”
The audit trail for declined transactions is a regulatory obligation that most teams design as an afterthought. Every transaction that is flagged or blocked by the fraud engine must generate an immutable audit event: the transaction ID, the decision, the features that drove the decision (velocity count at time of scoring, risk score, threshold), and the model version in use. This event must be stored in a way that cannot be altered — an append-only Kafka topic with infinite retention backed by object storage, not a database table with an update path. When a customer disputes a blocked transaction and the case reaches SAMA’s consumer protection function, the bank must produce this audit record. Without it, the bank cannot defend the decision and risks the regulatory finding that the control is not operating correctly.
The business trade-off
Exactly-once processing adds latency. The checkpoint mechanism that provides the exactly-once guarantee requires the stream processor to periodically snapshot its state to durable storage and to use Kafka transactions for output commits. In production, with RocksDB as the state backend and a 60-second checkpoint interval, the overhead is typically 5–8% additional end-to-end latency compared to at-least-once. For most payment channels, this is acceptable — the latency budget for IPS screening is measured in seconds, not the tens of milliseconds that exactly-once overhead adds.
For Mada card authorisation, where the Mada network SLA demands a decision in under 80–100 milliseconds end-to-end, the latency budget is tighter. The fraud check contribution to that budget is typically 10–20 milliseconds — the time to look up velocity state in RocksDB, score the transaction, and produce the decision to the authorisation topic. Exactly-once overhead within that window is manageable if the state backend and checkpoint interval are tuned correctly. But it requires explicit measurement and validation, not assumption.
At-least-once processing with carefully designed idempotency downstream is a legitimate design choice for high-throughput, low-latency pipelines where the exactly-once Kafka transaction overhead is genuinely prohibitive. The condition is that the downstream system must deduplicate on transaction ID, the state update logic must be idempotent (counting a transaction twice must produce the same result as counting it once), and the consequence of occasional double-processing must be understood and accepted by the business and the compliance function. This is a business and compliance decision, not an engineering one.
What leadership must decide
Four decisions belong at the leadership level and cannot be delegated to the engineering team to resolve independently.
The latency budget for fraud checks per payment channel. Mada, IPS, SADAD, and SARIE each have different network SLA constraints. The total latency budget for a fraud decision must be partitioned across the components that contribute to it: Kafka consumer fetch latency, state lookup time, model scoring time, Kafka producer commit time. The engineering team can measure each component’s contribution, but only the business can decide how much of the network SLA window is available for fraud screening versus core transaction processing. This budget directly determines whether exactly-once processing is feasible or whether a carefully managed at-least-once approach is required for the highest-throughput channels.
What “exactly-once” means at the business layer, not the Kafka layer. Engineers can configure Flink for exactly-once state updates and Kafka transactional sinks. What they cannot define is whether “exactly-once” at the business level means that a customer is charged exactly once, that a fraud alert is created exactly once, or that a regulatory event is recorded exactly once. Each of these may require additional idempotency controls beyond what the stream processor provides — deduplication at the authorisation gateway, idempotent writes at the alert management system, exactly-once event publishing to the audit log. The business must define what the end-to-end guarantee means before the architecture can be designed to satisfy it.
The state retention and recovery policy. A Flink fraud pipeline accumulates velocity state for every active customer key. If the job fails and cannot recover from its last checkpoint, some state is lost. The question “how much state loss is acceptable?” is a business question. A 60-second checkpoint interval means at most 60 seconds of velocity state is lost on an unrecoverable failure; transactions in that window may be incorrectly scored until the state warms up again. For a bank processing tens of thousands of transactions per minute, this may mean hundreds of transactions scored without accurate velocity context. Whether that is acceptable depends on the risk appetite and the regulatory exposure, not on the engineering team’s judgment alone.
The operational capability to run a stateful stream processor. Apache Flink is a distributed system with a learning curve. Diagnosing a checkpoint timeout, managing a parallelism change on a running job, performing a state migration during a job upgrade — these require engineers with specific Flink expertise. The decision to deploy stateful stream processing for a regulatory control like fraud velocity checking is also a decision to invest in that expertise. “We will learn it as we go” is not an acceptable posture for infrastructure that carries a SAMA compliance obligation. The headcount, training, and vendor support decisions must be made before the pipeline goes live on a production payment channel, not after the first 3 a.m. incident.
For the engineering depth behind these decisions — window types, Flink and Kafka Streams state backends, exactly-once checkpoint configuration, velocity check implementation, stream enrichment patterns, and the production checklist — see the companion Lab article: Data Streaming: Stateful Stream Processing Patterns.