Overview

The conventional deployment risk model in financial services is binary: a change is either in production or it is not. Every new version goes from zero traffic to one hundred percent in a single step, gated by a change advisory board window and a post-deployment verification checklist that takes thirty minutes to run. When the change fails in production — and some percentage of changes always fail in production, regardless of the quality of pre-production testing — the blast radius is the entire user population, the rollback is manual and takes another change window, and the incident ticket attributes the failure to “insufficient testing.”

Progressive delivery is the architectural inversion of that model. Instead of treating production exposure as binary, progressive delivery treats it as a variable: a new version begins serving a small fraction of traffic, automated analysis measures whether that version is performing within defined bounds, and traffic shifts progressively toward the new version as confidence accumulates. If analysis detects a regression — p99 latency above threshold, error rate spike, a business-level KPI like payment success rate dropping below baseline — traffic is automatically shifted back to the stable version. The blast radius of a bad deployment is the percentage of traffic the canary was serving when the regression was detected, typically less than ten percent.

The critical framing for a regulated KSA bank: progressive delivery is not a circumvention of SAMA’s Technology Risk Management (TRM) change management requirements. It is a more rigorous implementation of them. SAMA’s TRM framework requires that changes be tested, that risk be assessed, and that rollback capability be verified before a change goes live. A canary deployment that is automatically rolled back when a Prometheus metric breaches a threshold provides stronger rollback capability guarantees than a manual rollback procedure documented in a runbook. The engineering challenge is wiring the automated analysis to the right metrics and aligning the rollout strategy to the change management approval your board already granted.

Progressive delivery requires observability as a prerequisite

You cannot run metric-gated canary analysis without the metrics. Before implementing Flagger or Argo Rollouts canary analysis, your application must expose the signals the analysis will query: HTTP error rate, p99 request latency, and at least one business-level KPI (payment success rate, message throughput, queue depth). If those metrics do not exist in Prometheus at the time the canary starts, the analysis template will fail to collect enough data points and the promotion will time out. Wire observability first; the OTel pipeline for integration services is covered in the Observability Pipeline Architecture article in this series.

Progressive Delivery Landscape

Three strategies are in common use on Kubernetes, with significantly different trade-off profiles.

StrategyTraffic split mechanismRollback speedBest fit
Rolling updatePod replacement; no explicit traffic splitMinutes (pod restarts)Stateless services with low blast-radius tolerance; fastest to configure
Blue/greenKubernetes Service selector flip; 0% or 100%Seconds (selector change)Batch jobs, schema migrations, services that cannot tolerate mixed-version traffic
CanaryService mesh weight; Ingress header routing; Argo Rollouts TrafficRoutingSeconds (weight reset to 0%)High-traffic APIs where partial exposure reduces risk; metric-gated promotion

For integration services at a KSA bank the recommendation is canary for live payment API paths and blue/green for batch processors and schema migrations. Rolling updates are appropriate for supporting services (config servers, internal tooling) where a brief mixed-version state is acceptable. The rest of this article focuses on canary.

Argo Rollouts vs Flagger: both tools implement automated canary analysis on Kubernetes, and both are production-grade. The architectural difference is where they sit. Argo Rollouts replaces the Deployment object entirely with a Rollout CRD — it owns the pod lifecycle, the traffic weight, and the promotion logic. Flagger wraps an existing Deployment, creates canary and primary Deployments behind the scenes, and delegates traffic to a service mesh (Istio, Linkerd) or ingress controller. If you already have ArgoCD and want a unified GitOps toolchain, Argo Rollouts is the natural choice. If you already have Istio managing traffic and want a lightweight operator that does not change how your Deployments are structured, Flagger fits better.

Argo Rollouts Canary Strategy

Argo Rollouts replaces the Deployment resource with a Rollout CRD. The Rollout defines the canary strategy: the step weights, the pause durations, the analysis template references, and the traffic routing integration. When ArgoCD syncs a new image tag to the Rollout, the controller executes the steps in sequence.

payments-api-rollout.yaml — Argo Rollouts canary strategyyaml
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: payments-api
  namespace: payments
spec:
  replicas: 6
  selector:
    matchLabels:
      app: payments-api
  template:
    metadata:
      labels:
        app: payments-api
    spec:
      containers:
        - name: payments-api
          image: registry.saib.com/payments/payments-api:$(IMAGE_TAG)
  strategy:
    canary:
      # Traffic routing via Istio VirtualService + DestinationRule
      trafficRouting:
        istio:
          virtualService:
            name: payments-api-vs
          destinationRule:
            name: payments-api-dr
            canarySubsetName: canary
            stableSubsetName: stable
      # Analysis runs throughout canary steps
      analysis:
        startingStep: 2            # start analysis after first weight step
        templates:
          - templateName: payments-api-canary-analysis
        args:
          - name: canary-hash
            valueFrom:
              podTemplateHashValue: Latest
      steps:
        - setWeight: 5              # 5% to canary; wait 3 min
        - pause: {duration: 3m}
        - setWeight: 20             # 20%; analysis starts here
        - pause: {duration: 5m}
        - setWeight: 50             # 50%; manual gate for major releases
        - pause: {}                 # indefinite pause; requires `argo rollouts promote`
        - setWeight: 80
        - pause: {duration: 5m}
The indefinite pause at 50% is the change advisory gate

The pause: {} step without a duration holds the rollout until a human promotes it with argo rollouts promote payments-api. This maps directly to the “go/no-go” decision point in SAMA TRM change management. The first two automated steps (5% and 20%) give operations a 3–5 minute window to observe behaviour before the major exposure occurs; the manual gate at 50% is where the change manager signs off. Wire the promote command to your ITSM approval workflow so that the ITSM ticket closure triggers the promote rather than a manual kubectl command.

The corresponding Istio resources for traffic routing. Argo Rollouts will update the weights field in the VirtualService as the canary progresses:

payments-api-vs.yaml — Istio VirtualService for canary routingyaml
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
  name: payments-api-vs
  namespace: payments
spec:
  hosts:
    - payments-api
  http:
    - name: primary
      route:
        - destination:
            host: payments-api
            subset: stable
          weight: 100       # Argo Rollouts controller overwrites this during canary
        - destination:
            host: payments-api
            subset: canary
          weight: 0
---
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: payments-api-dr
  namespace: payments
spec:
  host: payments-api
  subsets:
    - name: stable
      labels:
        rollouts-pod-template-hash: "stable-hash"   # managed by Argo Rollouts
    - name: canary
      labels:
        rollouts-pod-template-hash: "canary-hash"   # managed by Argo Rollouts

Metric-Gated Analysis

The AnalysisTemplate defines the success criteria for canary promotion. Each metric query runs at a configurable interval; if the metric returns a value outside the threshold for a consecutive failure count, the analysis fails and the rollout is aborted. The template below queries Prometheus for three signals: HTTP error rate, p99 latency, and payment success rate. The third metric is the critical one for a payment service — a canary that technically has a low HTTP error rate but is failing more ISO 20022 payment validations than the stable version is a regression that a purely infrastructure-level analysis would miss.

payments-api-canary-analysis.yaml — AnalysisTemplateyaml
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: payments-api-canary-analysis
  namespace: payments
spec:
  args:
    - name: canary-hash
  metrics:
    # Metric 1: HTTP error rate must stay below 1%
    - name: http-error-rate
      interval: 60s
      failureLimit: 3           # fail after 3 consecutive failed intervals
      successCondition: result[0] < 0.01
      failureCondition: result[0] >= 0.05  # immediate abort at 5%
      provider:
        prometheus:
          address: http://prometheus.monitoring.svc.cluster.local:9090
          query: |
            sum(rate(http_requests_total{
              job="payments-api",
              rollouts_pod_template_hash="{{args.canary-hash}}",
              status=~"5.."
            }[5m]))
            /
            sum(rate(http_requests_total{
              job="payments-api",
              rollouts_pod_template_hash="{{args.canary-hash}}"
            }[5m]))
    # Metric 2: p99 latency must stay below 800ms
    - name: p99-latency
      interval: 60s
      failureLimit: 2
      successCondition: result[0] < 0.8  # seconds
      provider:
        prometheus:
          address: http://prometheus.monitoring.svc.cluster.local:9090
          query: |
            histogram_quantile(0.99, sum by (le) (
              rate(http_request_duration_seconds_bucket{
                job="payments-api",
                rollouts_pod_template_hash="{{args.canary-hash}}"
              }[5m])
            ))
    # Metric 3: business KPI — payment success rate must be >= 99.5%
    - name: payment-success-rate
      interval: 120s           # longer interval: business metric needs more data points
      count: 5                 # must pass 5 consecutive intervals
      failureLimit: 1
      successCondition: result[0] >= 0.995
      provider:
        prometheus:
          address: http://prometheus.monitoring.svc.cluster.local:9090
          query: |
            sum(increase(payment_transactions_total{
              job="payments-api",
              rollouts_pod_template_hash="{{args.canary-hash}}",
              outcome="success"
            }[5m]))
            /
            sum(increase(payment_transactions_total{
              job="payments-api",
              rollouts_pod_template_hash="{{args.canary-hash}}"
            }[5m]))

The payment_transactions_total metric requires that the payments-api service exposes a counter with outcome label (success or failure) and a rollouts_pod_template_hash label that allows Prometheus to distinguish canary pods from stable pods. In a Spring Boot 3.x application, use Micrometer to expose this counter with a tag that reads the pod template hash from an environment variable injected by the Rollout controller:

Low-traffic canaries need a minimum request count guard

A canary analysis interval that runs against a pod receiving 5% of traffic will have far fewer data points than the same interval against the stable version. A 5% canary serving 10 payment transactions per minute means 2–3 transactions per analysis interval — statistically meaningless. Add a successCondition that guards on the denominator: result[0] >= 0.995 and result[1] > 20 where result[1] is the transaction count. Without this guard, a canary that serves zero traffic in an interval will show 100% success rate and promote without evidence.

Flagger with Istio

Flagger takes a different approach: you keep your existing Deployment, and Flagger’s operator watches it. When the Deployment’s pod template changes (a new image, a config change), Flagger creates a -canary Deployment for the new version, progressively shifts traffic from the primary Service to the canary Service using an Istio VirtualService, and promotes (or rolls back) based on the metric thresholds in its Canary CRD.

payments-api-flagger.yaml — Flagger Canary resourceyaml
apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: payments-api
  namespace: payments
spec:
  provider: istio
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: payments-api
  service:
    port: 8080
    targetPort: 8080
    gateways:
      - payments/payments-gateway
    hosts:
      - payments-api.payments.svc.cluster.local
  analysis:
    interval: 2m
    threshold: 5               # abort after 5 failed metric checks
    maxWeight: 50              # cap canary at 50% before manual gate
    stepWeight: 10             # increment by 10% per interval
    metrics:
      - name: request-success-rate
        thresholdRange:
          min: 99             # percent; Flagger uses Istio telemetry by default
        interval: 1m
      - name: request-duration
        thresholdRange:
          max: 800            # milliseconds; Flagger reads from Istio histograms
        interval: 1m
      - name: payment-success-rate  # custom metric from Prometheus
        templateRef:
          name: payment-success-rate
          namespace: flagger-system
        thresholdRange:
          min: 99.5
        interval: 2m
    webhooks:
      - name: acceptance-test    # smoke tests run before first traffic step
        type: pre-rollout
        url: http://smoke-runner.payments/test/payments-api
      - name: itsm-gate          # ITSM approval required at 50%
        type: confirm-rollout    # Flagger pauses and calls this URL for go/no-go
        url: http://itsm-bridge.ops/flagger/gate/payments-api

The confirm-rollout webhook is the bridge between Flagger’s automated progression and your ITSM change management workflow. Flagger calls the webhook URL with a JSON body describing the canary state; the ITSM bridge service returns 200 only when the ITSM ticket attached to this deployment has been approved by the change manager. The ITSM ticket is opened automatically at the start of the rollout (via the pre-rollout webhook) with the canary metrics snapshot attached. The change manager reviews actual production metrics, not a hypothetical risk assessment, before approving.

Feature Flags: Unleash and OpenFeature

Feature flags decouple deployment from release. A new version of a payment service can be deployed to 100% of traffic while the new payment flow it contains is toggled off for all users. The toggle is flipped when the business decides the feature is ready — not when the next deployment slot becomes available. For a regulated bank this is valuable: it separates the SAMA TRM change management window (required for every deployment) from the product release decision (which happens on business timing).

The OpenFeature specification (CNCF incubating project) provides a vendor-neutral SDK that wraps any feature flag backend. Unleash is the open-source backend: self-hosted, with role-based access control for flag management and a full audit log of flag changes. The combination gives you flagging capability that is both auditable and not locked to a SaaS vendor — an important consideration for a bank with data residency requirements.

PaymentFeatureConfig.java — OpenFeature SDK with Unleash providerjava
package info.saib.payments.config;

import dev.openfeature.sdk.OpenFeatureAPI;
import dev.openfeature.contrib.providers.unleash.UnleashProvider;
import io.getunleash.UnleashContext;
import io.getunleash.strategy.Strategy;

@Configuration
public class PaymentFeatureConfig {

    @Value("${unleash.api-url}")
    private String unleashApiUrl;

    @Value("${unleash.api-token}")
    private String unleashApiToken;

    @Bean
    public OpenFeatureAPI openFeatureAPI() {
        var config = new UnleashConfig.Builder()
            .unleashAPI(unleashApiUrl)
            .apiKey(unleashApiToken)
            .appName("payments-api")
            .instanceId(System.getenv("POD_NAME"))    // k8s pod name for per-instance context
            .environment(System.getenv("APP_ENV"))    // prod | staging | uat
            .build();
        var provider = new UnleashProvider(config);
        var api = OpenFeatureAPI.getInstance();
        api.setProvider(provider);
        return api;
    }
}

// Usage in payment service — flag check at the business logic boundary
@Service
public class PaymentRoutingService {

    private final Client featureClient;

    public PaymentRoutingService(OpenFeatureAPI api) {
        this.featureClient = api.getClient("payments-api");
    }

    public PaymentRoute routePayment(PaymentRequest request) {
        // Context carries bank identifier; Unleash strategy can target by bankId
        var ctx = EvaluationContext.builder()
            .targetingKey(request.getBankId())
            .build();
        boolean useIso20022v2 = featureClient.getBooleanValue(
            "iso20022-v2-payment-routing", false, ctx
        );
        return useIso20022v2
            ? routeViaIso20022V2(request)
            : routeViaLegacy(request);
    }
}

The Unleash feature flag iso20022-v2-payment-routing can be targeted at specific bank IDs (or user segments, or percentage rollouts) without touching the deployment. When the product team is ready to roll out ISO 20022 v2 routing to all banks, they flip the flag from the Unleash console — no deployment required, no change advisory board, and every flag change is timestamped and attributed in the Unleash audit log.

Stale flags are technical debt with a compliance dimension

A feature flag that has been fully enabled for six months and is never queried in the false branch is dead code wrapped in an SDK call. More importantly, in a regulated environment it is an unmarked code path that auditors may ask about. Set a flag expiry policy in Unleash: flags older than 90 days with no recent variant switches trigger a review notification. Dead flags should be removed in the next sprint with a corresponding code cleanup, not left as a permanent toggle that obscures the actual code path.

SAMA TRM Alignment

SAMA’s Technology Risk Management framework (2017, updated 2022) requires that technology changes follow a structured change management process covering risk assessment, testing evidence, rollback capability, and post-implementation review. Progressive delivery aligns with each of these requirements — but the alignment must be made explicit in your change management documentation, not assumed.

SAMA TRM RequirementProgressive Delivery Implementation
Risk assessment before changeRollout strategy document describes step weights, analysis metrics and thresholds, manual gate position, and automatic rollback trigger conditions. Attached to the ITSM change ticket as a PDF.
Testing evidencePre-rollout smoke test webhook result, plus canary analysis run record (Argo Rollouts AnalysisRun or Flagger Canary event history) attached to the post-implementation review ticket.
Rollback capability verifiedAutomated rollback demonstrated in the change plan: the Prometheus metric threshold that would trigger rollback, the rollback time (seconds for traffic reset), and the rollback test result from the most recent DR exercise.
Post-implementation reviewArgo Rollouts rollout status history or Flagger Canary events provide a complete record: when each step completed, which metrics were evaluated, what values were observed, and when promotion occurred. Export to ITSM ticket at promotion completion.
Segregation of dutiesThe ITSM-gated pause at 50% ensures the deployment engineer and the change manager are two different people. The promote command requires a separate ITSM approval action, not just a code push.
Progressive delivery does not replace change management — it makes it verifiable

A canary rollout that runs fully automated from 0% to 100% without a human gate satisfies the technical risk management objective but fails the SAMA segregation of duties requirement for significant changes to production payment systems. For any change classified as significant (new payment flow, changes to IPS or BUNA connectivity, modifications to SAMA Open Banking API flows), retain the manual gate at 50% and wire it to ITSM approval. Reserve fully automated rollouts for low-risk changes (dependency updates, non-payment-path features, internal tooling). Document the classification criteria in your change management policy.

Rollback Automation

Rollback in Argo Rollouts is a traffic weight reset: the controller sets the canary VirtualService weight to 0%, scales down the canary ReplicaSet, and marks the Rollout as Degraded. The stable version continues serving 100% of traffic without any pod restarts — it was never removed from the pool. This is the fundamental advantage over a standard rolling update rollback, which requires a new pod bring-up. The rollback time is bounded by the time for the VirtualService update to propagate through Istio — typically under 5 seconds on a well-configured mesh.

  1. Argo Rollouts AnalysisRun detects that the payment-success-rate metric has returned below threshold for failureLimit consecutive intervals. The AnalysisRun is marked Failed.
  2. The Rollout controller reads the failed AnalysisRun result and aborts the rollout. It patches the Istio VirtualService: stable.weight: 100, canary.weight: 0. Istio envoy sidecars pick up the new VirtualService within 2–5 seconds.
  3. The canary ReplicaSet is scaled to zero over the next 30 seconds. The stable ReplicaSet continues unchanged. No pods in the stable path are restarted.
  4. The Rollout is marked Degraded. An alert fires to the on-call team via Alertmanager with the AnalysisRun detail: which metric failed, the threshold value, the observed value, and a link to the Grafana dashboard showing the canary metric time series.
  5. The ITSM change ticket associated with the rollout is automatically transitioned to “Rollback Complete” by the ITSM bridge webhook. The change record preserves the full canary analysis run history for the post-incident review.
  6. The on-call engineer diagnoses the root cause using the Grafana canary dashboard (available even post-rollback because Prometheus retains the metric time series). No production logs are lost; canary pods were not forcibly terminated.
  7. A new rollout (with the bug fixed) is opened after the root cause is resolved. The canary analysis provides evidence that the fix resolves the metric regression before promoting again.

Pitfalls

Canary metric latency means analysis lags reality by several minutes

Prometheus scrapes at 15-second intervals by default, and the analysis queries use a 5-minute rate window. This means the canary analysis is always observing behaviour from 2–5 minutes ago, not the current second. A canary that starts failing at T+0 may not trip the analysis until T+5 to T+8 minutes, depending on the failureLimit and interval configuration. For a payment service processing 1,000 transactions per minute, that is up to 8,000 transactions at the canary weight before rollback begins. Size the stepWeight with this lag in mind: a 5% canary weight with a 5-minute analysis window means at most 400 transactions on the canary before rollback — acceptable. A 20% canary weight with a 5-minute lag means 1,600 transactions before rollback.

  • Sticky sessions break canary analysis. If your API gateway or load balancer uses session affinity (cookie-based sticky sessions), the “5% canary” weight applies to new sessions, not to all requests. Existing sessions will stay pinned to their assigned pod version, which may be stable or canary depending on when the session was established. This distorts the canary analysis: the canary metric data is actually from a mix of new sessions (5%) and pre-existing sessions that happened to land on canary pods. Disable session affinity for services that run canary deployments, or use header-based canary routing (a fixed x-canary: true header for a specific test user pool) instead of weighted traffic split.
  • Schema-incompatible changes need blue/green, not canary. A canary that runs database schema migrations (ALTER TABLE, new index) alongside application changes will leave the stable pods running against a schema they were not tested against. Use blue/green with a maintenance window for any change that alters the database schema visible to the running application. Progressive delivery applies to the application tier; schema changes are a separate concern that requires a different deployment pattern.
  • Unleash flag evaluation latency in the critical payment path. The Unleash client SDK caches the flag state and refreshes it in the background at a configurable interval (default 15 seconds). A flag toggle in the Unleash console will not propagate to all running pods within the same second. For features that must toggle instantaneously across all pods simultaneously — an emergency circuit breaker that must disable a payment flow across the entire fleet within 2 seconds — the Unleash background polling model is too slow. Use an Istio EnvoyFilter for sub-second traffic manipulation in emergency scenarios, and reserve Unleash for business feature toggles that can tolerate a 15–30 second propagation window.
  • Argo Rollouts and ArgoCD conflict on Deployment ownership. If you have an existing ArgoCD Application managing a Deployment and you convert it to a Rollout, ArgoCD will still see the Deployment resource (which Argo Rollouts creates alongside the Rollout) as the source of truth and try to sync it back to the spec in Git. Configure the ArgoCD Application to ignore the Deployment that Argo Rollouts creates: add ignoreDifferences for the managed Deployment’s spec.replicas and pod template hash fields, and add the Rollout CRD to the Application’s tracked resources.

Production Checklist

  1. Verify that the application exposes the required Prometheus metrics before the first canary rollout: HTTP error rate with status code label, p99 request duration histogram, and at least one business-level counter (payment transactions with outcome label). Confirm these metrics appear in Prometheus with correct labels before configuring the AnalysisTemplate.
  2. Configure the Rollout or Flagger Canary with a conservative first step weight: 5% for payment-critical services, 10% for supporting services. The first step exists to detect gross failures before significant traffic is exposed; reserve larger weights for later steps after analysis has run.
  3. Wire the manual gate (indefinite pause or Flagger confirm-rollout webhook) to your ITSM change management workflow for changes classified as significant by your change management policy. Do not bypass the gate for convenience; it is the segregation-of-duties control that satisfies SAMA TRM.
  4. Set failureCondition thresholds for the business KPI metric to trigger immediate abort (not just a failed interval) for values that represent an unacceptable user impact: a payment success rate below 95% should abort immediately, not after 5 consecutive intervals at 94.9%.
  5. Deploy Unleash in HA mode on OpenShift with a PostgreSQL backend. Run at least two Unleash proxy instances; the proxy is what the SDK connects to, and a single-instance proxy taking 5 minutes of downtime means all feature flag evaluations fall back to their default values for 5 minutes — which may be fine for most flags but is a production incident for an emergency circuit breaker flag.
  6. Set flag expiry dates in Unleash at creation time for all release flags. A release flag that enabled a new payment flow should expire within 90 days; if the flag is still being evaluated in the default-on path at day 91, it should trigger a review notification and a Jira ticket to remove it from the code.
  7. Document the rollback procedure in the ITSM change template: “Automated rollback triggered by Argo Rollouts if Prometheus analysis fails; rollback time <5 seconds. Manual rollback: argo rollouts abort payments-api or flag-off via Unleash console.” This is the rollback evidence SAMA TRM requires attached to every significant change ticket.
  8. Test the rollback path in the pre-production environment before the first production canary. Deliberately trigger a threshold breach (inject synthetic errors via a Chaos Monkey or a test flag that increments the error counter), confirm the AnalysisRun fails, confirm the traffic resets to 100% stable within the expected window, and confirm the ITSM ticket is transitioned by the bridge webhook.