At most KSA banks I have observed, a production deployment still looks like this: the change advisory board approves a maintenance window, the deployment engineer promotes a build in the ITSM ticket, someone runs through a verification checklist, and the service goes from 0% to 100% of traffic in a single step. If it works, the ticket is closed. If it does not work, the rollback is manual, it takes another change window, and the incident report attributes the failure to “insufficient testing in the lower environment.”
The failure mode is not that the testing was insufficient. The failure mode is that the risk model is wrong. A lower environment that matches production in infrastructure but not in traffic patterns, not in data volume, and not in the combination of concurrent user behaviours that only manifests at scale will always miss some failures. Production is the only environment that looks like production. The question is not how to make the lower environment a better proxy for production — it never will be fully — but how to expose production failures before they affect the entire user population.
Progressive delivery — canary deployments, feature flags, metric-gated promotion — is the answer to that question. But adopting it requires a leadership decision, not a tooling decision. The tooling (Argo Rollouts, Flagger, Unleash) is straightforward. The decision is about the organisation: who owns the promotion gate, what happens when the automated rollback fires at 2am, and whether the change management process can adapt to a deployment model where rollback is measured in seconds rather than change windows.
What the risk model actually looks like
In a standard canary deployment for a payment API, the new version begins serving 5% of traffic. Automated analysis measures the payment success rate, the HTTP error rate, and the p99 latency every minute. If any of those metrics breach a threshold, the canary is rolled back automatically — the traffic weight resets to zero in under 5 seconds, before the operations team has finished reading the alert. The maximum impact of a bad deployment is whatever damage can be done at 5% canary weight before the first analysis interval catches it: typically 2–5 minutes of partial exposure, not the full production blast radius of the traditional deployment model.
This is not theoretical risk reduction. It is a structural change to the economics of a failed deployment. Under the traditional model, a bad deployment that is caught in post-deployment verification has already affected 100% of users. Under progressive delivery, the same failure is caught at 5% exposure, with an automatic rollback that does not require the on-call engineer to be awake and logged in. The blast radius shrinks by an order of magnitude.
Progressive delivery does not prevent failures. It constrains their blast radius and automates the response. The decision leadership has to make is whether to invest in the observability infrastructure that makes metric-gated analysis possible, and whether the change management process is flexible enough to accommodate a promotion gate that is partly automated and partly human.
The trade-off leadership has to make explicit
Progressive delivery introduces a new failure mode that the traditional deployment model does not have: a stuck canary. A canary that is neither promoted nor rolled back — because the metric analysis is inconclusive, because the traffic volume is too low to collect statistically significant data, or because the manual gate at 50% is waiting on a change manager who is in a meeting — leaves the production environment in a mixed-version state indefinitely. Two versions of the service are serving traffic simultaneously. Support tickets about inconsistent behaviour are difficult to diagnose when the routing decision that determined which version a user hit is not logged.
This is a real operational risk, and it requires explicit mitigations: a maximum canary window after which the promotion either succeeds or is automatically aborted, logging of the pod template hash alongside every payment transaction for post-hoc version attribution, and a clear ownership model for the promotion gate. Without those mitigations, the stuck canary is worse than the traditional binary deployment.
The ownership model is the leadership decision. Someone must be responsible for monitoring the canary during its analysis window, responding to alert escalations when the automated analysis is inconclusive, and making the promotion decision at the manual gate. In a project-delivery organisation, that responsibility falls to whoever is on call — which means it falls to no one in particular, and the gate is either rushed through (cancelling the risk reduction benefit) or left open indefinitely (creating the stuck-canary risk). In a platform organisation, that responsibility belongs to a platform team with a defined SLA for gate decisions and a runbook for inconclusive analysis.
What SAMA expects
SAMA’s Technology Risk Management framework does not prescribe deployment methodology. It requires that changes be risk-assessed, tested, and reversible. Progressive delivery satisfies all three requirements — more rigorously than the traditional maintenance-window model. The risk assessment is live: the canary analysis is watching the actual production risk surface in real time, not a pre-deployment estimate. The testing evidence is production data: the AnalysisRun record shows exactly which metrics were measured and what values were observed during the canary window. The rollback capability is verified: the automated rollback was tested in the lower environment before the production canary, and the rollback time is documented as a concrete number, not a change-window estimate.
The specific alignment that matters for a SAMA examination is the manual gate at 50% canary weight. SAMA TRM requires segregation of duties for significant changes. The manual gate — wired to an ITSM approval action — ensures that the person who deployed the change is not the same person who authorised it to reach the majority of production traffic. The change manager reviews actual canary metric data, not a hypothetical risk assessment. When an auditor asks “how do you verify that a change is behaving correctly before full rollout?”, the answer is a Grafana dashboard showing the canary metric time series during the analysis window, not a verbal assurance.
Feature flags are a separate decision
Feature flags and canary deployments address different problems and are often conflated. A canary deployment manages the risk of deploying a new version of an application. A feature flag manages the risk of releasing a new capability within an already-deployed version. They are complementary, not equivalent.
The leadership decision on feature flags is about ownership and discipline. Flags are cheap to create and expensive to maintain. Every flag that is not removed after its feature is fully rolled out is a code path that exists but is not tested, a toggle that could be accidentally flipped, and an audit question that needs an answer (“what is this flag controlling in production?”). Banks that adopt feature flags without an expiry policy and a flag hygiene process within six months have a flag inventory that is larger and more opaque than their deployment inventory was before. The tool solves a problem and creates a different one.
The governance that makes feature flags safe in a regulated environment: a flag registry in Unleash that is reviewed quarterly, an expiry date attached to every release flag at creation, a named owner for every flag that controls a regulatory-adjacent feature (payment routing, consent management, SAMA Open Banking flows), and a process for removing flags from code within 30 days of their being fully enabled in production. Without that governance, feature flags are a liability. With it, they are one of the highest-leverage tools available for reducing the cost and risk of releasing in a regulated environment.
What the team needs from leadership to ship it
Three things, in order of importance:
Permission to instrument. Progressive delivery requires that every service on the canary path expose the metrics the analysis needs: HTTP error rate, p99 latency, and a business-level KPI. For a legacy payments service that was built before observability was a team norm, adding those metrics requires a code change and a deployment. That code change will not happen unless the team has explicit permission to treat observability instrumentation as a first-class delivery item, not something to be deferred until “we have capacity.” The VP has to name it as a priority.
A change management policy that accommodates automated rollback. If the current change management policy requires that every production rollback be approved by a change advisory board, an automated rollback that fires at 2am when no one is available to convene a board is non-compliant. The change management team needs to agree in advance that automated rollback triggered by a Prometheus metric threshold is an approved rollback mechanism for services operating under a canary strategy. This is a policy change, not a technical one, and it requires VP-level sponsorship to get through the governance process.
An explicit on-call model for the promotion gate. The manual gate at 50% canary weight will fire during business hours, after hours, and on weekends. Someone has to be reachable to make the promotion decision. If the change manager who approved the original deployment is not on call, the gate will either be left open (stuck canary) or bypassed (missing the segregation of duties control). Define who owns the gate, what the response SLA is (30 minutes during business hours, 2 hours after hours), and what the automated fallback is if the gate is not acted on within the SLA window.
For the engineering depth behind this topic — Argo Rollouts manifests, Flagger Canary CRDs, AnalysisTemplate Prometheus queries, OpenFeature SDK wiring, SAMA TRM change ticket mapping, and the rollback automation steps that satisfy a SAMA audit — see the companion Lab article: Progressive Delivery: Flagger, Argo Rollouts & Feature Flags for Regulated Financial Services.