Overview
Static secrets fail in regulated environments in predictable, auditable ways. The same Oracle DB password appears in twelve different ConfigMaps across six namespaces, each copied by a different engineer at a different time. When a DBA leaves the organisation, the credential survives in every ConfigMap that was never updated — because updating it requires a coordinated deployment across a dozen services, which nobody schedules. SAMA’s Technology Risk Management framework mandates annual rotation for operational keys; manual rotation across that many ConfigMaps is not an annual ceremony, it is a perennial risk that accumulates until an examiner finds it.
The audit trail problem compounds the rotation problem. With static secrets there is no record of which pod read which credential at what time. A SAMA examination that asks “who has had access to the payments database credentials in the past 12 months?” cannot be answered by inspecting Kubernetes Secrets: everyone who can read the namespace, everyone who has ever deployed to it, everyone whose CI runner touched the image. With Vault every credential lease is logged: which Kubernetes service account, from which pod, on which node, at what time, for what TTL.
Dynamic secrets solve both problems simultaneously: every credential is generated fresh at pod start time with a short TTL, so rotation is implicit, the blast radius of a leaked credential is bounded by its TTL, and every issuance is logged. This article covers the full production stack: Vault HA on OpenShift, Kubernetes auth method, dynamic database and PKI engines, External Secrets Operator for GitOps-safe delivery, and SAMA’s key management requirements.
Every new dynamic credential is already “rotated” because it was never used before. The database engine generates a new Oracle user with a random password at mount time, grants it the required roles, and sets an expiry. When the TTL expires Vault revokes the credential and Oracle drops the user. The application never held a long-lived password; rotation is not an event you schedule, it is a property of every credential issuance.
Vault Architecture on OpenShift
Vault runs in HA mode using Integrated Storage (Raft) — three StatefulSet replicas, each holding a copy of the Raft log on a dedicated PersistentVolumeClaim. Raft eliminates the external Consul dependency that earlier Vault deployments required; one active node handles writes, performance standby nodes serve reads, and any replica can be promoted to active if the current leader fails. Do not run Raft on a shared StorageClass that is also used by application workloads — Raft is I/O sensitive and a noisy neighbour on shared storage will cause election timeouts.
Seal mechanism: A freshly started Vault is sealed and refuses all requests until enough Shamir key shares are provided to reconstruct the master key. In a bank environment, Shamir shares are held by named key custodians (typically the CISO, platform engineering lead, and a third executive) who must each present their share during a planned unseal or an unplanned restart. Auto-unseal via an HSM (AWS CloudHSM or Thales on-premise) removes the need for human intervention on restart — the HSM holds the unseal key, Vault calls it at boot. SAMA’s requirement that root CA and payment signing keys reside on a FIPS 140-2 Level 3 HSM makes HSM-based auto-unseal the correct choice for a bank deployment, because the HSM infrastructure will be there regardless.
Three delivery mechanisms exist for getting secrets from Vault into pods; a fourth (Direct SDK) is valid but rarely the right choice for most services:
| Mechanism | Secret refresh | ArgoCD visibility | Dev experience | Failure mode |
|---|---|---|---|---|
| Vault Agent sidecar | Continuous (in-memory template) | No secret value in git | File or env var; Agent handles renewal | Sidecar crash = no new leases; pod must restart |
| CSI driver | On pod restart only | No secret value in git | Mounted as file; simple mental model | CSI mount failure blocks pod start; no cached value |
| External Secrets Operator | Configurable refreshInterval (1h–1m) | ExternalSecret CR only; no value | Standard Kubernetes Secret after sync; familiar to devs | ESO down → last known Secret cached in etcd; pod starts |
| Direct Vault SDK | Application-managed renewal | No secret value in git | Full control; significant implementation burden | Application must handle Vault downtime and token renewal |
External Secrets Operator is the recommended delivery path for integration services because it produces a standard Kubernetes Secret that ArgoCD never sees the value of, it survives a Vault outage via the cached Secret in etcd, and it integrates naturally with existing Kubernetes-native tooling. Vault Agent is preferred for high-frequency secret access (PKI certificate renewal under a 1-hour TTL) where per-pod in-memory caching reduces Vault request volume significantly.
Kubernetes Auth Method
The Kubernetes auth method is the correct auth mechanism for pods on OpenShift. It works as follows: the pod presents its projected service account JWT (mounted at /var/run/secrets/kubernetes.io/serviceaccount/token) to Vault’s login endpoint. Vault calls the Kubernetes TokenReview API to validate the JWT’s authenticity against the API server, verifies the pod is running as the expected service account in the expected namespace, and if the JWT is valid, issues a Vault token scoped to the policies defined for that Vault role. The pod never has a static Vault token; the JWT it presents is rotated by the Kubernetes API server automatically (every 24 hours by default on OpenShift 4.15).
The mapping chain is: ServiceAccount → Vault role → Vault policy → secret paths. A payments service running as payments-ips-adapter in the payments namespace is mapped to the payments-ips-adapter Vault role, which is bound to the payments-ips-adapter-policy, which grants read on database/creds/payments-ips-adapter and pki/issue/payments-mesh-identity. Nothing else. The policy is minimal by construction.
Vault Agent handles token renewal: it performs the initial Kubernetes auth login, obtains a Vault token, and renews it before expiry for as long as the pod is alive. When the pod terminates, the token expires and all dynamic leases associated with it are revoked.
# Vault Kubernetes auth role — payments-ips-adapter service
# vault write auth/kubernetes/role/payments-ips-adapter \
bound_service_account_names = ["payments-ips-adapter"]
bound_service_account_namespaces = ["payments"]
token_ttl = "1h"
token_max_ttl = "4h"
token_policies = ["payments-ips-adapter-policy"]
token_bound_cidrs = ["10.128.0.0/14"] # OpenShift pod CIDR
# Vault policy — payments-ips-adapter-policy.hcl
# Grants minimum required access only
# Dynamic DB credentials for the IPS Oracle schema
path "database/creds/payments-ips-adapter" {
capabilities = ["read"]
}
# PKI — issue short-lived mTLS cert for service mesh identity
path "pki/issue/payments-mesh-identity" {
capabilities = ["create", "update"]
}
# Static secrets — API gateway client credentials (rotated quarterly)
path "secret/data/payments/ips-adapter/apigw-credentials" {
capabilities = ["read"]
}
# Deny everything else — explicit deny is belt-and-suspenders
path "secret/*" {
capabilities = ["deny"]
}
path "database/*" {
capabilities = ["deny"]
}
path "pki/*" {
capabilities = ["deny"]
}
Dynamic Secret Engines
Dynamic secret engines generate credentials on demand, set a TTL, and revoke them automatically when the lease expires. The killer property of dynamic credentials is that revocation does not require touching every application: you revoke the lease in Vault and the credential is gone at the database layer, regardless of how many places once had a copy. Compare this to rotating a static password: you must update every ConfigMap, Secret, and environment variable that holds it and restart every pod simultaneously — a coordination problem that almost always results in either downtime or a window where old and new credentials coexist.
Database secrets engine (Oracle, DB2, PostgreSQL): Vault connects to the database using a privileged “management” account and dynamically creates a new database user per Vault role per lease request. For Oracle the creation statement grants the user the appropriate roles from a predefined set; for PostgreSQL a GRANT statement assigns schema-level permissions. The credential is unique per pod at startup time — a deployment with 3 replicas produces 3 distinct DB users, each with a TTL of 1 hour. When the pod terminates or the TTL expires, Vault drops the user from the database.
AWS secrets engine generates IAM credentials (access key + secret) for pods that need to call AWS services. In a bank that is not on AWS this engine is unused, but it serves as the model for understanding how dynamic engines work with external systems.
PKI secrets engine acts as an intermediate certificate authority and issues short-lived X.509 certificates. It is covered in depth in the PKI section below.
# Enable and configure the database secrets engine for Oracle
vault secrets enable database
# Configure the Oracle plugin connection
# Requires the vault-plugin-database-oracle binary in the plugin directory
vault write database/config/payments-oracle \
plugin_name=vault-plugin-database-oracle \
allowed_roles="payments-ips-adapter,payments-ach-processor" \
connection_url="{{username}}/{{password}}@payments-db.saib.internal:1521/PAYPRD" \
username="VAULT_MGR" \
password="{{ env \"ORACLE_VAULT_MGR_PASSWORD\" }}"
# Rotate the management credential immediately — Vault owns it from here
vault write -force database/rotate-root/payments-oracle
# Define a role: dynamic user per lease, 1-hour default TTL, 4-hour max
vault write database/roles/payments-ips-adapter \
db_name=payments-oracle \
creation_statements="
CREATE USER {{name}} IDENTIFIED BY {{password}};
GRANT IPS_APP_ROLE TO {{name}};
GRANT CREATE SESSION TO {{name}};
ALTER USER {{name}} PROFILE IPS_PROFILE;
" \
revocation_statements="
DROP USER {{name}} CASCADE;
" \
default_ttl="1h" \
max_ttl="4h"
# DB2 role configuration (db2 plugin)
vault write database/config/lending-db2 \
plugin_name=vault-plugin-database-db2 \
allowed_roles="lending-loan-origination" \
connection_url="SERVER=lending-db.saib.internal:50000;DATABASE=LENDPRD" \
username="VAULTMGR" \
password="{{ env \"DB2_VAULT_MGR_PASSWORD\" }}"
vault write database/rotate-root/lending-db2
External Secrets Operator
External Secrets Operator (ESO) runs as a controller in the cluster and reconciles ExternalSecret CRDs into standard Kubernetes Secrets. The ArgoCD operator sees only the ExternalSecret manifest — the source of truth in git — and never the credential value. ESO fetches the secret from Vault on behalf of the namespace and writes it to a Kubernetes Secret that the pod mounts normally. This is the GitOps-safe delivery pattern: no plaintext credentials in git, no Sealed Secrets public-key ceremony, no Helm chart Secret values to redact.
SecretStore vs ClusterSecretStore: a SecretStore is namespace-scoped and authenticates to Vault as the namespace’s service account; a ClusterSecretStore is cluster-scoped and authenticates as a dedicated ESO service account. Use SecretStore per namespace — it enforces the same service-account-to-Vault-role mapping you would use for Vault Agent, and the blast radius of a misconfigured ClusterSecretStore credential is the entire cluster.
refreshInterval determines how often ESO polls Vault and updates the Kubernetes Secret. For static credentials set to 1h; for dynamic credentials approaching their TTL, set the refreshInterval to 80% of the TTL so the pod always has a valid credential. ESO updates the Kubernetes Secret in place; pods that mount the Secret as a volume see the new value without restarting (file watch). Pods that load the Secret as environment variables at startup do not see the update until they restart — design accordingly.
External Secrets Operator with a refreshInterval of 1m across 200 ExternalSecrets generates 200 Vault API reads per minute per cluster — 12 000 per hour — before accounting for any application load. Enable Vault Agent caching (a Vault Agent instance per node that caches responses for its TTL) or increase the interval for low-volatility secrets. For secrets that change only on explicit rotation, 1h or even 24h is appropriate; reserve short intervals for PKI certificates with TTLs under 2 hours.
# SecretStore — namespace-scoped Vault auth via Kubernetes SA
apiVersion: external-secrets.io/v1beta1
kind: SecretStore
metadata:
name: vault
namespace: payments
spec:
provider:
vault:
server: "https://vault.platform.svc:8200"
path: "database"
version: "v1"
caBundle: "{{ internal_ca_bundle_base64 }}"
auth:
kubernetes:
mountPath: "kubernetes"
role: "payments-ips-adapter"
serviceAccountRef:
name: "payments-ips-adapter"
---
# ExternalSecret — fetches dynamic DB creds and writes a K8s Secret
apiVersion: external-secrets.io/v1beta1
kind: ExternalSecret
metadata:
name: payments-ips-adapter-db-creds
namespace: payments
spec:
refreshInterval: "45m" # 75% of 1h TTL — ensures renewal before expiry
secretStoreRef:
name: vault
kind: SecretStore
target:
name: payments-ips-adapter-db # name of the K8s Secret ESO creates
creationPolicy: Owner # ESO owns the Secret; deleted with ExternalSecret
template:
type: Opaque
data:
# rewrite Vault field names to app-expected env var names
ORACLE_USER: "{{ .username }}"
ORACLE_PASSWORD: "{{ .password }}"
ORACLE_URL: "jdbc:oracle:thin:@payments-db.saib.internal:1521/PAYPRD"
data:
- secretKey: username
remoteRef:
key: "database/creds/payments-ips-adapter"
property: username
- secretKey: password
remoteRef:
key: "database/creds/payments-ips-adapter"
property: password
PKI Secret Engine for mTLS
Vault’s PKI secret engine acts as a certificate authority within a larger PKI hierarchy. The recommended topology for a bank: the root CA is an offline HSM (never on a network-connected system), an intermediate CA is signed by the root CA and loaded into Vault’s PKI engine, and Vault issues end-entity certificates for workloads on demand. The intermediate CA private key never leaves the HSM during the signing ceremony; after signing, the intermediate CA certificate is loaded into Vault and the HSM private key for the intermediate is optionally also loaded if Vault will sign with it (or Vault generates its own intermediate key and submits a CSR for offline signing).
cert-manager’s Vault issuer integrates with the PKI engine via a Certificate resource. cert-manager watches the certificate’s expiry and automatically submits a new CertificateRequest to Vault when the certificate is within the renewal window (typically at 2/3 of its lifetime). For 24-hour certificates, renewal triggers at 16 hours. This makes certificate rotation fully automatic — no cron job, no human intervention, and because the certificate is short-lived, a compromised certificate has a bounded useful lifetime.
SAMA’s PKI integration requirements mandate that: (1) the root CA key is held on a FIPS 140-2 Level 3 HSM and its private key is never exported; (2) Vault’s intermediate CA certificate must be revocable via the root CA’s CRL; (3) certificate issuance for payment systems must be logged to the central audit trail with the requesting workload identity. Vault’s PKI audit log satisfies requirement (3) automatically when Vault audit logging is enabled and forwarded to the SIEM.
SAMA’s Technology Risk Management §5.3 requires annual rotation for operational keys. For TLS certificates, issuing 24-hour certificates via Vault PKI means the key is effectively rotated every 24 hours — exceeding the annual requirement by orders of magnitude. An examiner who sees a cert-manager + Vault PKI configuration producing 24-hour certificates will view this as a significantly stronger control than a manual annual rotation process, which relies on humans remembering to rotate on schedule.
SAMA Key Management Requirements
SAMA’s Technology Risk Management framework §5.3 and the Cyber Security Framework both establish a key lifecycle policy for regulated financial institutions. The requirements that directly constrain the Vault architecture are:
- Annual rotation for operational keys. Database credentials, API signing keys, and inter-service TLS certificates are operational keys. Vault’s dynamic credential and PKI engines satisfy this with far shorter rotation periods — 1-hour DB credentials and 24-hour certificates never accumulate the staleness that annual manual rotation requires remediating.
- Quarterly rotation for high-sensitivity keys. Payment signing keys (used by the IPS/SARIE gateway adapter) are classified as high-sensitivity and require quarterly rotation. These are not dynamic credentials — they are HSM-held asymmetric keys. Vault’s Transit engine can act as an abstraction layer: the signing key is never exported from the HSM; the application calls Vault Transit to sign a payload, and Vault delegates to the HSM backend.
- HSM requirement for root CA and payment signing keys. SAMA requires FIPS 140-2 Level 3 hardware for root CA key storage and payment signing key storage. Vault integrates with Thales Luna HSM and AWS CloudHSM as auto-unseal backends; the same HSM can hold the Vault auto-unseal key, the root CA key, and the payment signing key — consolidating the HSM footprint.
- Audit log for every key usage. Every Vault operation — secret read, credential lease, PKI issuance, Transit sign — is written to Vault’s audit log in JSON format. Forward the audit log to the SIEM via a syslog exporter. The SIEM must retain audit events for the SAMA-mandated retention period (typically 5 years for financial transaction-related events).
Automated Rotation Patterns
Vault’s architecture makes most rotation patterns automatic, but some require explicit design decisions.
Database credential rotation: the database engine’s rotate-root command generates a new password for Vault’s management account at the database and immediately updates Vault’s stored connection credentials. After this command, even Vault’s own engineers do not know the management password — Vault owns it. Dynamic credential leases issued to pods are not affected by a root rotation; they continue to work until their TTL expires, at which point Vault revokes them using the new management credentials.
Static secret rotation: for credentials that cannot be made dynamic (legacy application that requires a static password), store them in Vault KV v2 and define a rotation policy. Vault Enterprise supports automatic rotation of KV secrets; Vault CE requires a rotation pipeline (a CronJob that generates a new credential, writes it to Vault, and triggers a pod restart). ESO picks up the new secret on the next poll cycle and updates the Kubernetes Secret.
Break-glass procedure: every Vault deployment must have a documented break-glass procedure for when Vault itself is unavailable. ESO caches the last known Kubernetes Secret in etcd, so pods can restart and start successfully using the cached credential even if Vault is down. The break-glass scenario is when the cached credential has also expired — in this case an authorised engineer injects a temporary static secret via a Kubernetes Secret override, generates a SIEM alert, and opens an incident. The temporary secret is revoked as the first act after Vault returns to service.
-
Inventory all static secrets. Find every Kubernetes ConfigMap and Secret containing a credential. Use
kubectl get secret -A -o json | jqto extract secret names; grep the values for known credential patterns. Build a registry with secret name, namespace, consuming service, and rotation frequency. -
Enable and configure the Vault database engine for the target database. Use a dedicated management account with the minimum grants required to create/drop users. Run
vault write -force database/rotate-root/<config-name>immediately after configuration to take ownership of the management credential. -
Create Vault roles and policies for each service. Map each service’s Kubernetes ServiceAccount to a Vault role. Set the policy to allow
readondatabase/creds/<role-name>and nothing else for the migration. Test the role manually withvault read database/creds/<role-name>. - Deploy ExternalSecret and SecretStore to the service namespace. Verify ESO syncs a Kubernetes Secret containing valid credentials. Test the credentials against the database directly before changing the application configuration.
- Update the application Deployment to mount the ESO-managed Secret. Remove the static Secret or ConfigMap reference; add a reference to the ESO-managed Secret. Deploy to the dev environment first and validate connectivity. Watch Vault audit logs to confirm credential issuance.
- Promote to test and production. Apply the ExternalSecret and updated Deployment to each environment in sequence via ArgoCD. Vault roles and policies are environment-specific (separate Vault namespaces per environment, or separate path prefixes). Do not share dynamic credential roles between environments.
- Delete the static Secret and ConfigMap. Only after the service has been running successfully on dynamic credentials for one full business day in each environment. Archive the secret name and hash in the rotation registry as “migrated to Vault dynamic” with the date.
The Vault root token has unrestricted access to all secrets, all policies, and all auth methods. It should be generated once during initial setup, used to configure the Kubernetes auth method and initial policies, and then revoked immediately. Any subsequent administrative work must use a human operator token with specific policies. If the root token must be regenerated for a break-glass operation, that action must trigger an alert to the SIEM and require approval from at least two named key custodians. Storing the root token in a password manager, a ConfigMap, or any persistent medium is a material security deficiency that a SAMA examiner will classify as a critical finding.
Pitfalls
Vault’s Raft storage is sensitive to disk I/O latency. A Raft leader that cannot commit a write within the heartbeat window triggers a leader election; if I/O is consistently slow, the cluster oscillates between leaders and eventually becomes read-only while elections cycle. Provision Vault’s PersistentVolumeClaims on a dedicated StorageClass backed by SSDs, separate from application PVCs. Set Raft’s performance_multiplier to 2 or 3 in environments with higher latency storage. Monitor Vault’s own vault.raft.leader.lastContact metric; alert when it exceeds 200 ms.
A Vault cluster that issues short-lived credentials to many pods without lease TTL limits accumulates millions of lease entries in its storage. Vault’s lease management background job processes these on a configurable interval, but at very high volumes the job falls behind and Vault performance degrades. Set explicit max_ttl on every role and enable lease count quotas (vault write sys/quotas/lease-count/global max_leases=300000) to prevent runaway accumulation. Review lease counts monthly via the Vault sys/leases API.
If Vault is sealed (planned maintenance, an unplanned restart without auto-unseal, or an HSM connectivity failure), new pods that depend on ESO for their credentials will start successfully because ESO’s Kubernetes Secret is already in etcd. Pods that use the Vault Agent sidecar directly for credential injection will fail to start — the Agent cannot obtain a token from a sealed Vault, so it cannot inject credentials, and the init container exits non-zero. Design all critical payment services to use ESO rather than the Vault Agent sidecar, and ensure the ESO-managed Secret has a TTL long enough to survive a planned Vault maintenance window.
Production Checklist
- Vault running in HA mode: 3 Raft replicas, dedicated PersistentVolumeClaims on isolated StorageClass, PodDisruptionBudget allowing at most 1 unavailable replica.
- Auto-unseal configured via FIPS 140-2 Level 3 HSM; Shamir share custodians documented for break-glass manual unseal.
- Vault root token revoked immediately after initial setup; root token regeneration requires dual approval and triggers SIEM alert.
- Kubernetes auth method enabled; every service mapped to a minimal-scope policy (no wildcard path grants).
- Database secrets engine enabled for Oracle and DB2; management credential rotated with
rotate-root;max_ttlset on every role. - External Secrets Operator deployed; refreshInterval set to 75% of credential TTL for all ExternalSecrets.
- Vault PKI intermediate CA signed by offline HSM root CA; cert-manager Vault issuer configured per namespace; 24-hour certificate TTL.
- Vault audit log forwarded to SIEM; 5-year retention configured; alert on audit log pipeline failure (missing events from expected sources).
- Lease count quota configured; Raft lastContact metric alerted at 200 ms.
- Break-glass procedure documented: ESO-cached Secret fallback, temporary static Secret injection process, SIEM alert, incident ticket template.
- SAMA key rotation registry maintained: every key or certificate mapped to its TTL, rotation mechanism, and last verified rotation date.
- HSM holding payment signing keys: Transit engine mount delegating sign operations to HSM; quarterly rotation tested in non-production.