Overview
The multi-cluster problem arrives quietly. One day you have a dev cluster and a production cluster. Six months later you have dev, test, staging, prod-zone-1, prod-zone-2 (required by SAMA’s Business Continuity Policy for Tier 1 systems), a dedicated cluster for the payments team, and a platform services cluster. That is eight clusters. Each cluster needs every shared platform application deployed to it — cert-manager, the observability stack, ExternalDNS, External Secrets Operator, the Vault agent injector — and each application team’s services deployed to the applicable subset of clusters. With a naive ArgoCD setup you write one Application manifest per app per cluster: 8 clusters × 30 apps = 240 Application objects to maintain manually.
Adding a new cluster means writing 30 new Application manifests. Deleting an app means hunting down 8 Application manifests. Changing a Helm values override that should apply to all prod clusters means editing 2 manifests. At 240 objects this is tedious; at 400 it is the primary reason platform engineers leave.
The ArgoCD ApplicationSet controller eliminates this by generating Application objects from a template and a generator. Add a cluster to ArgoCD with the right label: the ApplicationSet automatically generates an Application for every app that should run on it. Add an app directory to the repo: the ApplicationSet generates an Application for every cluster that should run it. The git repo is still the source of truth; the ApplicationSet is the factory that turns the repo structure into Application objects without manual authoring.
When you add a new cluster to the ArgoCD cluster inventory with the appropriate labels, it automatically receives all apps managed by matching ApplicationSets — no new manifests to write. When you add a new app directory to the git repo, all clusters in the matching generator scope receive it at the next sync. The platform team writes one ApplicationSet per logical app group; fleet growth is a label on a cluster, not a YAML authoring event.
ArgoCD ApplicationSets
An ApplicationSet is an ArgoCD CRD processed by the ApplicationSet controller (bundled with ArgoCD 2.3+ and now GA in 2.12). It has two required fields: a generators list (which produces parameter sets) and a template (which defines the ArgoCD Application to generate for each parameter set). For every combination of parameters the generator produces, the controller creates one ArgoCD Application object. When the generator produces a parameter set that no longer matches (a cluster removed from inventory, a directory deleted from the repo), the controller deletes the Application and ArgoCD garbage-collects its resources from the spoke cluster according to the configured pruning policy.
The five generator types available in ArgoCD 2.12 serve distinct use cases:
| Generator | Use case | Dynamic discovery | Scale ceiling | When to use |
|---|---|---|---|---|
| Cluster | Deploy an app to many clusters | Yes — reads ArgoCD cluster secrets | Hundreds of clusters | Platform apps that run on every cluster (cert-manager, ESO) |
| Git (files) | One app per config file in a directory | Yes — walks the repo glob | Thousands of files | Tenant-per-file patterns; bank subsidiary per env config |
| Git (dirs) | One app per subdirectory | Yes — walks the repo tree | Hundreds of directories | One app per team folder; microservice per directory |
| Matrix | Cross-product of two generators | Yes — inherited from child generators | Product of both scales; watch for combinatorial explosion | Apps × Clusters; directories × Clusters |
| Pull Request | Ephemeral preview per open PR | Yes — polls SCM API | Open PR count | Preview environments; integration team feature branches |
| SCM Provider | Discover repositories in an org | Yes — polls SCM org | Repository count in org | Auto-onboarding repos that match a naming convention |
Cluster Generator Pattern
The Cluster generator reads ArgoCD’s cluster registry (cluster secrets in the argocd namespace) and produces one parameter set per matching cluster. Labels on the cluster secret drive selection — a cluster labelled tier: prod and region: riyadh matches a generator selector for tier: prod but not one for tier: dev. This label-based selection is the only mechanism that prevents a dev-targeted ApplicationSet from deploying to production clusters, so treat cluster labels as a security boundary: only the platform team can apply or change labels on cluster secrets.
Per-cluster Helm values overrides are passed via the cluster secret’s metadata annotations and referenced in the ApplicationSet template as {{ metadata.annotations.<key> }}. This is how you express “prod-zone-1 uses 3 replicas and the ssd-retain StorageClass, while dev uses 1 replica and the default StorageClass” without duplicating Helm charts or templating logic.
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: platform-apps
namespace: argocd
spec:
generators:
- clusters:
selector:
matchLabels:
platform.saib.internal/managed: "true" # only platform-managed clusters
values: # static values available in template as {{ values.* }}
revision: main
template:
metadata:
name: "platform-cert-manager-{{ name }}" # unique per cluster
spec:
project: platform
source:
repoURL: https://gitops.saib.internal/platform/apps.git
targetRevision: "{{ values.revision }}"
path: platform/cert-manager
helm:
valueFiles:
- values.yaml
# per-cluster overlay resolved from cluster secret annotation
- "values-{{ metadata.annotations.tier }}.yaml"
parameters:
# per-cluster values injected directly from cluster secret metadata
- name: replicaCount
value: "{{ metadata.annotations.replicaCount }}"
- name: global.storageClass
value: "{{ metadata.annotations.storageClass }}"
- name: global.clusterName
value: "{{ name }}"
destination:
server: "{{ server }}" # cluster API URL from the generator
namespace: cert-manager
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
- ServerSideApply=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
Git Generator for Configuration Files
The Git generator walks a git repository and produces one parameter set per file or per directory that matches a glob pattern. The files generator is ideal for tenant-per-file patterns: each bank subsidiary or integration team has a config/<env>/<tenant>.json file in the GitOps repo. When a new tenant is onboarded, the team submits a PR adding their config file; ArgoCD detects the new file on the next refresh and generates a new Application for that tenant automatically, without any change to the ApplicationSet itself.
The configuration file’s JSON fields become template parameters. A file at config/prod/payments-north.json containing { "tenant": "payments-north", "namespace": "payments-north", "tier": "prod", "replicaCount": "2" } produces parameters {{ tenant }}, {{ namespace }}, {{ tier }}, and {{ replicaCount }} available in the Application template.
Matrix Generator
The Matrix generator computes the cross-product of two child generators. The canonical use case: deploy every app directory in the repo to every cluster in the inventory — one Application per app per cluster. A Matrix generator whose first child is a Git (directories) generator returning 10 apps and whose second child is a Cluster generator returning 8 clusters produces 80 Applications.
The Matrix generator is powerful and therefore dangerous. An unguarded Matrix over many apps and many clusters produces a large number of Applications that all attempt to sync simultaneously when the ApplicationSet is first applied or when a cluster is restored after maintenance. Use sync wave annotations to stagger syncs: wave 0 for namespace and RBAC objects, wave 1 for platform services (cert-manager, ESO), wave 2 for application services. Without waves, a restored cluster receiving 80 simultaneous syncs will saturate the API server and likely fail several of them.
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: integration-services
namespace: argocd
spec:
generators:
- matrix:
generators:
# First: discover app directories under apps/ in the gitops repo
- git:
repoURL: https://gitops.saib.internal/platform/integration-apps.git
revision: HEAD
directories:
- path: "apps/*" # one dir per integration app
- path: "apps/archived"
exclude: true # skip the archived directory
# Second: all clusters labelled for integration workloads
- clusters:
selector:
matchLabels:
platform.saib.internal/integration: "true"
template:
metadata:
name: "{{ path.basenameNormalized }}-{{ name }}"
annotations:
# Sync waves: stagger across clusters to prevent thundering herd
argocd.argoproj.io/sync-wave: "2"
spec:
project: integration
source:
repoURL: https://gitops.saib.internal/platform/integration-apps.git
targetRevision: HEAD
path: "{{ path }}"
helm:
valueFiles:
- values.yaml
- "values-{{ metadata.annotations.tier }}.yaml"
parameters:
- name: global.env
value: "{{ metadata.annotations.tier }}"
- name: global.clusterName
value: "{{ name }}"
destination:
server: "{{ server }}"
namespace: "{{ path.basenameNormalized }}"
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions: [CreateNamespace=true, ServerSideApply=true]
A Matrix ApplicationSet with 10 clusters and 10 apps creates 100 Applications. When a cluster is restored after maintenance, ArgoCD queues all 100 Applications for sync simultaneously. Without sync wave annotations, the spoke cluster’s API server receives 100 concurrent apply operations that can overload the admission controllers and cause cascading sync failures. Add argocd.argoproj.io/sync-wave annotations to stagger syncs: wave 0 for namespaces and RBAC, wave 1 for platform services, wave 2 for application services, wave 3 for integration services. Each wave must reach Healthy before the next begins.
Hub-Spoke vs Mesh Topology
Two ArgoCD topologies emerge at multi-cluster scale: hub-spoke (one ArgoCD instance on a management cluster manages all spoke clusters) and mesh (each cluster runs its own ArgoCD, all syncing from a shared git repo). The hub-spoke model is operationally simpler — one control plane to upgrade, one place to review sync status, one set of RBAC policies to audit. The mesh model provides cluster isolation: a compromised or unavailable management cluster does not block deployments to spoke clusters.
For most regulated bank deployments, hub-spoke is the correct starting point. The management cluster is a platform team concern; its availability SLA should be equivalent to the highest-tier spoke cluster it manages. The spoke clusters have ArgoCD Application objects applied to them by the hub — the spoke does not run an ArgoCD server itself.
The primary objection to hub-spoke is blast radius: if the management cluster is unavailable, no new deployments can reach any spoke. This is acceptable if your deployment cadence is measured in hours (batch CI/CD pipeline) but not if you need continuous autonomous reconciliation per cluster (security patching pipelines that run independently of the management cluster). In that case, run a lightweight Argo CD Application Set instance on each spoke that manages only security-patch-critical applications; retain the hub for integration service deployments where coordination across clusters is a feature, not a problem.
Environment Promotion Workflow
Promotion in a GitOps model is a git operation, not a kubectl command. A CI pipeline that successfully builds and tests an image tags it and opens a pull request against the env-specific Helm values file in the GitOps repo. The PR changes exactly one thing: the image.tag value in config/test/values.yaml (or whichever environment is the promotion target). A human or an automated gate reviews and merges the PR. ArgoCD detects the change via its 3-minute polling interval (or a webhook push) and syncs the Application. The deployment to that environment happens as a side effect of the PR merge — no deploy script, no CI runner with kubectl credentials, no manual rollout.
The audit trail is the git history: who approved the PR, when it was merged, what the previous and new values were. For a SAMA change management audit, the git history of the GitOps repo is the change log. Production deployments require a PR from a protected branch reviewed by at least two engineers — enforced by the git platform’s branch protection rules, not by an ArgoCD configuration.
ArgoCD’s sync windows provide time-based deployment gates at the platform layer: no automated syncs to production clusters on Friday between 14:00 and 24:00 (SAMA change freeze window), or outside business hours for non-emergency changes. Sync windows are defined on the AppProject and apply to all Applications within the project, regardless of how they were generated.
SAMA Multi-Zone Requirements
SAMA’s Business Continuity Policy (BCP) requires that Tier 1 systems (which include payment processing, core banking interfaces, and SARIE/IPS connectivity) be active at two geographically separate sites at least 30 kilometres apart, with each site capable of sustaining full operational load independently. For OpenShift this means two separate prod clusters — not a stretched cluster, not a cluster with nodes in two datacentres — in two physically distinct facilities.
The ApplicationSet with a Cluster generator selecting tier: prod deploys identically to both prod-zone-1 (Riyadh) and prod-zone-2 (Jeddah) automatically. The per-cluster annotation carries site-specific values: storage class names differ by site, the Vault server URL points to the local Vault cluster, and the SIEM endpoint is the regional collector. The application image, the Helm chart, and the business logic are identical — only infrastructure coordinates differ, expressed as Helm parameter overrides.
Cross-zone health comparison is a platform SLA obligation: an ArgoCD health dashboard should show sync state for both zones side by side. A zone where ArgoCD reports OutOfSync for more than 30 minutes is a BCP violation in progress. Alert on this condition from the Prometheus metrics ArgoCD exposes (argocd_app_info with sync_status="OutOfSync" filtered to the prod project) and route to the on-call platform engineer, not just the application team.
Multi-Cluster RBAC
ArgoCD’s RBAC model has three scopes: the AppProject (which repos, which clusters, which namespaces an Application can target), the RBAC policy (which ArgoCD actions a user or group can perform on Applications in a project), and the sync window (time-based deployment gates on the project). Together they allow you to grant an integration team the ability to deploy to dev and test clusters while requiring platform team review for production.
SSO integration via SAML or OIDC maps the bank’s Active Directory groups (via Entra ID or LDAP) to ArgoCD roles. A member of the payments-integration-team AD group gets the integration-dev ArgoCD role, which can sync Applications in the payments-dev AppProject targeting the dev and test clusters. They cannot sync Applications in the payments-prod AppProject — that requires membership in the platform-engineering group, which maps to the platform-admin role.
apiVersion: argoproj.io/v1alpha1
kind: AppProject
metadata:
name: payments-integration
namespace: argocd
spec:
description: Payments integration team — dev and test only; prod requires platform team PR
# Restrict which source repos are allowed
sourceRepos:
- https://gitops.saib.internal/payments/*
- https://gitops.saib.internal/platform/integration-apps.git
# Integration team can only deploy to dev and test cluster namespaces
destinations:
- server: https://dev-cluster.saib.internal:6443
namespace: "payments-*"
- server: https://test-cluster.saib.internal:6443
namespace: "payments-*"
# prod clusters NOT listed — integration team has no destination on prod
# Block deploy of cluster-scoped resources (ClusterRole, CRDs)
clusterResourceWhitelist: []
namespaceResourceBlacklist:
- group: ""
kind: ResourceQuota # platform team owns quota; integration teams cannot change it
- group: networking.k8s.io
kind: NetworkPolicy # NetworkPolicy is injected by Crossplane, not app teams
# Sync window: no automated syncs to test on Friday afternoon
syncWindows:
- kind: deny
schedule: "0 14 * * 5" # Friday 14:00
duration: 10h # until Friday midnight
applications: ["*"]
clusters: ["test-cluster"]
namespaces: ["*"]
manualSync: true # allow manual sync override for hotfixes
# RBAC: integration team can sync dev/test; platform team can sync anywhere
roles:
- name: integration-developer
description: Payments integration team — dev and test only
policies:
- p, proj:payments-integration:integration-developer, applications, get, payments-integration/*, allow
- p, proj:payments-integration:integration-developer, applications, sync, payments-integration/*, allow
- p, proj:payments-integration:integration-developer, applications, action/*, payments-integration/*, deny
groups:
- payments-integration-team # maps from SAML group assertion
The ArgoCD hub uses a service account (bound to a cluster secret) to apply resources to each spoke cluster. Granting cluster-admin to this service account means a misconfigured or malicious Application can delete any resource on the spoke, including its own RBAC objects. Instead, create a namespace-scoped role for each AppProject’s namespace prefix, and a minimal cluster-scoped role for CRDs and ClusterRoles that only the platform AppProject can request. Audit the ArgoCD service account RBAC on every spoke quarterly; cluster-admin should appear nowhere in the audit output for integration team projects.
Production Checklist
- ArgoCD running in HA mode on the management cluster: 3 application controller replicas, 2 repo server replicas, Redis Sentinel; PodDisruptionBudget on each component.
- Management cluster has the same SLA as the highest-tier spoke it manages; it is not a dev or shared cluster.
- All spoke clusters registered in ArgoCD with labels encoding tier (
dev/test/prod), region (riyadh/jeddah), and function (integration/platform). - Platform apps (cert-manager, ESO, Vault agent injector, monitoring stack) managed by a Cluster-generator ApplicationSet; no per-cluster Application manifests for platform components.
- Integration services managed by a Matrix-generator ApplicationSet (Git dirs × Clusters); sync wave annotations prevent thundering herd on cluster restore.
- AppProject per team scoping destinations to dev and test only; platform team AppProject for prod destinations; no overlap.
- ArgoCD service account on each spoke has namespace-scoped RBAC; cluster-admin absent from all integration team project service accounts.
- SSO configured via SAML/OIDC from bank IdP (Entra ID); AD group → ArgoCD role mapping reviewed quarterly.
- Sync windows configured: deny automated syncs to prod and test on Friday 14:00–24:00; manual sync override allowed for platform team.
- Promotion workflow documented and enforced: image tag change via PR to GitOps repo; no
kubectl rolloutfrom CI pipelines; CI runner has no kubeconfig. - SAMA BCP: prod-zone-1 and prod-zone-2 are geographically separate (≥30 km); ApplicationSet deploys identically to both; sync state compared in Prometheus dashboard.
- Alert on
argocd_app_info{sync_status="OutOfSync", project=~"payments.*|platform.*"}for more than 30 minutes; routes to platform on-call, not application teams.