Overview

Every bank that has deployed mTLS for partner API access has also, somewhere, a spreadsheet. It tracks certificate Common Names, expiry dates, the partner contact email, and whoever last rotated the cert — in free-text fields, maintained manually, updated when someone remembers. When a cert expires at 2am on a Thursday before the Eid holiday and the spreadsheet says it was renewed six months ago by someone who left the organisation, the incident is already in progress.

This article is about eliminating that spreadsheet. The tools are HashiCorp Vault’s PKI secrets engine (issuing short-lived certs from an intermediate CA it controls) and cert-manager (reconciling Kubernetes Certificate objects, automating renewal before expiry, and injecting rotated secrets without a pod restart). Together, they move certificate lifecycle from a manual calendar event to an automated control loop — one that runs continuously, produces audit evidence, and satisfies the SAMA Cybersecurity Framework’s cryptographic key management requirements.

Scope

This article covers TLS certificates used for API mTLS, ingress/egress TLS, and OAuth client authentication (RFC 8705). It does not cover code-signing certificates, document-signing (PDF), or HSM-backed root CA operations — those deserve separate treatment. For the mTLS protocol itself and the FAPI 2.0 requirements that mandate sender-constrained tokens, see OAuth 2.1 & mTLS and FAPI 2.0 & Open Banking Security.

Architecture

The production architecture has three tiers: Vault owns the private CA hierarchy, cert-manager translates Certificate objects into API calls against Vault, and applications consume TLS secrets mounted as projected volumes. The boundary between Vault and cert-manager is clean: Vault issues, cert-manager requests and reconciles, applications are passive consumers.

A few key decisions baked into this architecture:

  • Short-lived leaf certs (90 days max). Vault issues 90-day certificates by default with renewal triggered at 30 days remaining. At SAIB, we run 30-day certs for internal east-west mTLS and 90-day certs for external partner APIs, matching the SAMA-recommended rotation cadence for transient credentials.
  • Intermediate CA in Vault, root CA offline. The root CA private key never lives in Vault’s storage. The root signs the intermediate CA cert once; thereafter the intermediate issues all leaf certs. If Vault is compromised, the root CA can revoke the intermediate and reissue.
  • No pod restarts on cert rotation. cert-manager updates the Kubernetes TLS Secret in-place. The application either uses projected volumes (which update automatically) or an inotify-triggered reload hook. No deployment rollout, no downtime, no change window required.

Vault PKI engine

Vault’s PKI secrets engine is a certificate authority built into Vault’s secret management framework. The setup sequence: enable the mount, generate or import the intermediate CA, configure issuing and CRL URLs, and define roles that constrain what cert-manager can request.

vault-pki-setup.shbash
# Enable PKI mount for intermediate CA
vault secrets enable -path=pki_int -max-lease-ttl=87600h pki

# Generate intermediate CA CSR (Vault holds the private key)
vault write -format=json pki_int/intermediate/generate/internal \
  common_name="SAIB Intermediate CA 2026" \
  organization="Saudi Investment Bank" \
  country="SA" \
  key_type="rsa" \
  key_bits=4096 \
  | jq -r '.data.csr' > intermediate.csr

# Sign the CSR with the offline root CA (separate secure process)
# openssl ca -config root_ca.conf -in intermediate.csr -out intermediate.crt

# Import the signed intermediate cert back into Vault
vault write pki_int/intermediate/set-signed certificate=@intermediate.crt

# Configure issuing and CRL URLs (internal hostnames, not public)
vault write pki_int/config/urls \
  issuing_certificates="https://vault.internal.saib.sa/v1/pki_int/ca" \
  crl_distribution_points="https://vault.internal.saib.sa/v1/pki_int/crl" \
  ocsp_servers="https://vault.internal.saib.sa/v1/pki_int/ocsp"

# Define an issuing role for internal service certs (max 30 days)
vault write pki_int/roles/internal-services \
  allowed_domains="svc.cluster.local,internal.saib.sa" \
  allow_subdomains=true \
  allow_glob_domains=false \
  max_ttl="720h" \
  key_usage="DigitalSignature,KeyEncipherment" \
  ext_key_usage="ServerAuth,ClientAuth" \
  require_cn=true

# Define an issuing role for external partner certs (max 90 days)
vault write pki_int/roles/partner-apis \
  allowed_domains="api.saib.sa" \
  allow_subdomains=true \
  max_ttl="2160h" \
  key_usage="DigitalSignature,KeyEncipherment" \
  ext_key_usage="ServerAuth,ClientAuth"
CRL and OCSP must be reachable from workloads

If a service validates its peer’s client certificate and performs CRL/OCSP checks, the distribution point URL must be reachable from the workload’s network segment. In a KSA bank with strict network segmentation between payment zones and management networks, verify that the Vault OCSP URL is routable from every zone that does mTLS before deploying cert validation that fails hard on revocation check failure.

cert-manager issuer

cert-manager’s ClusterIssuer resource wraps the Vault API call. The issuer authenticates to Vault using an AppRole (role ID + secret ID stored in a Kubernetes Secret), calls the PKI sign endpoint, and writes the result into a TLS Secret in any namespace. A ClusterIssuer — rather than a namespace-scoped Issuer — is appropriate here because the same Vault PKI serves multiple teams across the cluster.

vault-clusterissuer.yamlyaml
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: vault-pki-int
spec:
  vault:
    server: https://vault.internal.saib.sa
    path: pki_int/sign/internal-services   # role name matches vault write pki_int/roles/...
    caBundle: LS0tLS1CRUd...                 # base64 of Vault’s TLS CA cert
    auth:
      appRole:
        path: approle
        roleId: cert-manager-role-id
        secretRef:
          name: vault-approle-secret-id      # Secret in cert-manager namespace
          key: secretId
---
# The AppRole secret ID lives in cert-manager’s own namespace
apiVersion: v1
kind: Secret
metadata:
  name: vault-approle-secret-id
  namespace: cert-manager
type: Opaque
stringData:
  secretId: "<secret-id from vault write auth/approle/role/cert-manager/secret-id>"

The Vault AppRole policy must be scoped to only what cert-manager needs. Issue certs, read CA chain, read CRL — nothing else.

vault-certmanager-policy.hclhcl
# cert-manager Vault policy — least-privilege
path "pki_int/sign/internal-services" {
  capabilities = ["create", "update"]
}
path "pki_int/sign/partner-apis" {
  capabilities = ["create", "update"]
}
path "pki_int/cert/ca" {
  capabilities = ["read"]
}
path "pki_int/ca" {
  capabilities = ["read"]
}
path "pki_int/crl" {
  capabilities = ["read"]
}

Certificate CRD

cert-manager’s Certificate resource is the control-plane object that binds a workload’s certificate requirements to an issuer. When the controller detects that a Certificate is approaching its renewBefore window, it calls the issuer’s API and updates the backing TLS Secret. No human action required.

payments-api-cert.yamlyaml
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: payments-api-tls
  namespace: payments
spec:
  secretName: payments-api-tls-secret         # Kubernetes TLS Secret to create/update
  duration: 720h                               # 30 days — internal service cert
  renewBefore: 168h                            # Renew at 7 days remaining (23 days in)
  subject:
    organizations: ["Saudi Investment Bank"]
    organizationalUnits: ["Payments Platform"]
    countries: ["SA"]
  commonName: payments-api.payments.svc.cluster.local
  dnsNames:
    - payments-api.payments.svc.cluster.local
    - payments-api.payments.svc
    - payments-api.internal.saib.sa
  usages:
    - digital signature
    - key encipherment
    - server auth
    - client auth                               # needed for mTLS as both server and client
  issuerRef:
    name: vault-pki-int
    kind: ClusterIssuer
    group: cert-manager.io
  secretTemplate:
    annotations:
      reloader.stakater.com/match: "true"        # triggers Stakater Reloader on change
    labels:
      cert-source: vault-pki

The comparison below captures the trade-off between different certificate duration strategies — a decision the team makes once and encodes in the Certificate CRD, not something to revisit every 90 days:

DurationrenewBeforeUse caseTrade-off
7 days24hInternal east-west (Istio sidecar)Minimal blast radius on compromise; higher Vault API call rate
30 days168h (7d)Internal service mTLS (non-Istio)Balanced rotation cadence; low Vault load; renews while service is live
90 days720h (30d)External partner APIs; ingress TLSPartner systems get 30-day notice; SAMA-acceptable rotation cadence
1 year2160h (90d)Legacy system client certs (mainframe)Long-lived cert for systems that cannot hot-reload; revocation must be tested

Zero-restart rotation

cert-manager writes the new certificate and private key into the Kubernetes TLS Secret before the old cert expires. The question is how the application picks up the new material without a pod restart — because in production, especially for stateful payment services, a pod restart means a pod disruption budget event, and that requires a change window under SAMA TRM.

There are three patterns, in order of preference:

  1. Projected volume mount (best). Mount the TLS Secret as a projected volume, not a plain volume. Projected volumes propagate Secret updates to the filesystem in-place without remounting. The file at /etc/tls/tls.crt is updated atomically — the application reads the new cert on its next TLS handshake if it re-reads the cert file per connection (Envoy, Nginx with a reload hook).
  2. Stakater Reloader (for stateless services). The reloader.stakater.com/match: "true" annotation on the TLS Secret triggers Stakater Reloader to roll the Deployment when the Secret changes. This is a pod restart, so it only applies to stateless services (API gateways, microservices) where a rolling restart is acceptable.
  3. Vault Agent Sidecar (for legacy apps). For applications that cannot consume dynamic file changes, Vault Agent (running as a sidecar) writes the cert to a shared emptyDir volume and can signal the main container via a lifecycle hook. This works for legacy Java apps that load the cert from a JKS store at startup: the Vault Agent can re-generate the JKS and signal the app to reload via kill -HUP.
  4. Istio / Envoy SDS (for service mesh). If the service is running inside Istio, the Envoy sidecar receives cert updates via the Secret Discovery Service (SDS). cert-manager integrated with istio-csr delivers Istio workload certificates with 24-hour TTL. The service never touches the cert directly; Envoy handles all TLS and cert rotation transparently.
  5. Application-level reload (last resort). Some applications (Kong Gateway, Nginx Plus, HAProxy) expose a reload API or a signal-based reload that re-reads the cert on disk without dropping connections. Wire an inotify watch on the cert file to trigger the reload API call. This pattern is documented in the Kong cert hot-reload playbook.
Projected volume update latency

Projected volume updates are not instantaneous. Kubernetes’ kubelet syncs Secret changes to mounted volumes on its syncFrequency interval (default 1 minute) plus the time for the inotify event to propagate. In practice, expect 60–90 seconds between cert-manager writing the new Secret and the file appearing in the pod. Plan your renewBefore window accordingly — the 7-day head start on a 30-day cert easily absorbs this.

Partner cert lifecycle

External partner certs are different from internal service certs in one critical respect: the partner holds the private key, not Vault. The bank’s role is to trust the partner’s CA, register the partner’s client certificate, and maintain a revocation capability.

The operational workflow for a new partner (TPP onboarding under SAMA Open Banking, for example) is:

  1. Partner submits a CSR or self-signed certificate through the developer portal.
  2. Operations verifies the partner’s identity (out-of-band, per SAMA TPP registration policy).
  3. The certificate is added to the Kong Gateway’s mTLS trusted CA bundle or the API Management platform’s partner trust store.
  4. cert-manager is not involved for the partner’s client cert — only for the bank’s server cert that the partner validates.
  5. An expiry record is created in the cert lifecycle tracker (not a spreadsheet — see below).

For automated expiry tracking of external partner certs, the x509-certificate-exporter (below) can scrape the TLS handshake expiry from the partner’s presented certificate, not from a local file. This works by dialling the partner’s mTLS endpoint and reading the cert from the TLS handshake, exporting the expiry as a Prometheus metric. An alert fires 30 days before expiry, giving the partner a notification lead time that matches most enterprise PKI rotation windows.

Revocation must be tested, not assumed

When a partner cert needs to be revoked — on a suspected compromise, on partner offboarding, or on a SAMA directive to suspend a TPP — the operational path must be pre-tested. Add the partner cert to the revocation list and verify that the next request from that cert is rejected at the gateway within the CRL refresh interval (typically 5 minutes with Kong). A revocation that only works in the post-incident review is not a revocation capability.

OAuth mTLS token binding

RFC 8705 (OAuth 2.0 Mutual-TLS Client Authentication and Certificate-Bound Access Tokens) binds an OAuth access token to a specific client TLS certificate. A token stolen from the wire is useless without the matching private key. This is the sender-constrained token requirement in FAPI 2.0.

The mechanism: the authorisation server embeds the SHA-256 thumbprint of the client’s mTLS cert into the access token as a cnf.x5t#S256 claim. The resource server validates that the TLS cert presented in the mTLS handshake matches the thumbprint in the token.

cert-bound-token-claim.jsonjson
{
  "iss": "https://idp.saib.sa/realms/openbanking",
  "sub": "svc:payments-aggregator",
  "aud": "accounts-api",
  "scope": "accounts.read payments.initiate",
  "exp": 1758000000,
  "cnf": {
    "x5t#S256": "bwcK0esc3ACC3DB2Y5_lESsXE8o9ltc05O89jdN-dg2"  // SHA-256 of DER-encoded client cert
  }
}

In Keycloak 24, cert-bound tokens are enabled per-client in the Advanced Settings tab by setting Client Authentication to X.509 and enabling Token endpoint authentication method as tls_client_auth. The Keycloak X.509 authenticator validates the client cert presented at the token endpoint and attaches the thumbprint claim automatically.

On the resource server side (Kong, Spring Security, custom middleware), the validation is:

  1. Extract the client cert from the TLS handshake (Kong’s $ssl_client_fingerprint or Nginx’s $ssl_client_s_dn header after mTLS offload).
  2. Compute SHA-256 of the DER-encoded cert.
  3. Compare to cnf.x5t#S256 in the validated JWT.
  4. Reject if they differ.

When cert-manager rotates the client cert, the thumbprint changes. The service must re-request an access token with the new cert before the next API call — which is why keeping access token TTL at 5 minutes for cert-bound tokens matters: the window between rotation and re-auth is bounded.

Expiry monitoring

The x509-certificate-exporter (enix-io/x509-certificate-exporter) is a Kubernetes DaemonSet that reads TLS Secrets and certificate files from node paths and exports expiry metrics to Prometheus. At SAIB we run it in two modes: scanning Kubernetes TLS Secrets across all namespaces, and dialling external partner endpoints to read their presented certs from the TLS handshake.

x509-cert-exporter-values.yamlyaml
# Helm values for x509-certificate-exporter
secretsExporter:
  enabled: true
  includeLabels:
    cert-source: vault-pki                     # only scan Vault-issued certs
  namespaces: []                              # empty = all namespaces

hostPathsExporter:
  enabled: false                              # only needed for node-level certs

prometheusRules:
  enabled: true
  alertOnReadErrors: true
  warningDaysBeforeExpiry: 30
  criticalDaysBeforeExpiry: 7

# External partner endpoint scraping (separate deployment)
extraEnv:
  - name: WATCH_REMOTE_ENDPOINTS
    value: "tpp-a.partner.sa:443,tpp-b.partner.sa:443"

The Prometheus alert rules generated automatically by the Helm chart cover 30-day warning and 7-day critical thresholds. Add a third alert for 1-day remaining — this fires when rotation failed for some reason and a manual intervention is needed before expiry:

Grafana dashboard panel

The key metric is x509_cert_expires_in_seconds labelled by secret_name and secret_namespace. A stat panel showing “certs expiring in < 30 days” as a count, red when non-zero, is the VP-visible indicator. The ops team monitors the time-series version; leadership sees the count. A count permanently at zero means automation is working.

SAMA CSF alignment

SAMA’s Cybersecurity Framework (CSF) Domain 3 (Cybersecurity Operations) and Domain 4 (Third-Party Cybersecurity) both touch certificate management. The specific controls that automated cert lifecycle directly satisfies:

CSF ControlRequirement (summary)How automation satisfies it
3-9-1 Cryptographic key management procedure covering generation, storage, rotation, and revocation Vault PKI policy is the documented procedure; AppRole auth log is the audit trail; Certificate CRDs are the rotation schedule
3-9-2 Certificates rotated at defined intervals; expired certificates prohibited cert-manager’s renewBefore guarantees rotation before expiry; x509-exporter Prometheus alert fires before any cert can expire undetected
3-9-4 Revocation capability tested and documented Vault CRL is immediately updated on vault write pki_int/revoke; Kong CRL refresh interval is 5 minutes; revocation runbook is the test evidence
4-2-3 Third-party (TPP/partner) certificate expiry tracked and communicated with advance notice x509-exporter scrapes partner endpoints; 30-day Prometheus alert triggers partner notification workflow in ServiceNow
SAMA audit artefact: Certificate inventory report

Auditors ask for a certificate inventory showing every cert in scope, its expiry date, and evidence of the last rotation. The kubectl get certificates -A -o wide output, combined with kubectl get certificaterequests -A for the rotation history, is the audit artefact. Export it to a PDF report as part of the quarterly PKI review. Vault’s audit log (vault audit list) provides the sign request history with timestamps, satisfying the “how was each cert issued?” question.

Production checklist

The migration from manual cert management to Vault PKI + cert-manager runs in phases. Each phase can be delivered in a two-week sprint:

  1. Vault PKI engine and intermediate CA (Sprint 1). Enable the pki_int mount. Generate the intermediate CA CSR. Work with the InfoSec team to sign it against the offline root CA. Import and configure issuing/CRL URLs. Define the internal-services and partner-apis roles. Deploy the Vault AppRole for cert-manager. Exit criteria: vault write pki_int/sign/internal-services csr=... | openssl x509 -text produces a valid cert signed by the intermediate CA.
  2. cert-manager ClusterIssuer and pilot Certificate (Sprint 2). Install cert-manager 1.15 (if not already present; it comes bundled with OpenShift 4.15). Create the ClusterIssuer. Write one Certificate CRD for a non-critical service (the internal metrics endpoint is a safe candidate). Verify the TLS Secret is created, the cert is valid, and the expiry is as configured. Exit criteria: kubectl describe certificate <name> shows Certificate is up to date and has not expired.
  3. Rotation without restart (Sprint 3). For each service category (Nginx-based, Spring Boot, Kong plugin), implement and test the reload mechanism: projected volume for Envoy/Istio, Stakater Reloader for stateless services, inotify hook for legacy. Run a forced renewal (kubectl cert-manager renew <cert-name>) and verify the new cert appears in the pod without a restart. Exit criteria: Prometheus metric certmanager_certificate_expiration_timestamp_seconds updates without a pod restart event in the namespace.
  4. Partner cert monitoring (Sprint 4). Deploy x509-certificate-exporter with the Helm values above. Import the partner endpoint list from the existing spreadsheet. Validate that expiry metrics appear in Grafana. Wire the 30-day warning alert to the ServiceNow incident workflow that sends partner notification emails. Exit criteria: One test partner cert with a short TTL triggers the alert and creates a ServiceNow ticket.
  5. RFC 8705 cert-bound tokens (Sprint 5, FAPI-required services only). Enable tls_client_auth on the Keycloak 24 client config for each FAPI 2.0-scoped API consumer. Update the Kong OIDC plugin configuration to pass the client cert thumbprint to the JWT validator. Write a negative test: token acquired with cert A, presented with cert B, verify 401. Exit criteria: FAPI Conformance Suite test fapi1-advanced-final-client-clientAuthentication-mtls passes.
  6. Decommission the spreadsheet and SAMA audit evidence package (Sprint 6). Export the final cert inventory from kubectl get certificates -A -o wide. Add it to the PKI governance document as the new authoritative source. Schedule the quarterly cert review as a Grafana dashboard review rather than a spreadsheet audit. Archive the spreadsheet. File the Vault audit log excerpt in the SAMA TRM evidence folder. Exit criteria: Zero manually tracked certs in scope; SAMA evidence folder updated; spreadsheet access removed from the shared drive.