Overview

North-south mTLS — mutual TLS between your perimeter API gateway and external callers — is well understood and routinely audited. East-west mTLS is the harder problem: securing service-to-service traffic inside the cluster, where the assumption of a trusted network still dominates despite every zero-trust framework condemning it.

In a SAMA-regulated bank running 80–150 microservices on OpenShift or Kubernetes, east-west traffic carries payment instructions, customer PII, credit decisions, and inter-system state transfers. SAMA’s Technology Risk Management framework classifies inter-service communication of core banking data as an attack surface that must be encrypted and mutually authenticated, not merely firewalled. A service mesh is the production-grade answer to that requirement — but deploying one at scale without understanding the operational cost creates its own incidents.

This article covers Istio and Linkerd: what they do under the hood, how to configure east-west mTLS correctly, how certificate rotation works at scale, and where each system fails if you don’t operate it deliberately.

mTLS is a baseline, not a complete zero-trust model

Mutual TLS proves that Service A holds a valid certificate and that Service B trusts the issuing CA. It does not prove that Service A is authorised to call Service B’s payment endpoint with a given customer ID. mTLS is the transport-layer foundation; AuthorizationPolicy (Istio) or AuthorizationPolicy (Linkerd) is where east-west access control lives.

The east-west problem

The lateral movement threat in a microservice architecture follows a predictable pattern: a compromised container or a misconfigured ingress rule gives an attacker a foothold inside the cluster. From there, every service-to-service call they can intercept or impersonate is a free move toward a higher-value target.

Traditional countermeasures — network policies, namespace isolation, pod security standards — reduce the blast radius but do not authenticate the caller at the application layer. A network policy that permits traffic from namespace A to namespace B cannot distinguish between a legitimately deployed payments service and a compromised pod that has claimed the same IP range.

Service mesh mTLS changes the model. Every workload is issued a cryptographic identity (a SPIFFE SVID, described below). Every connection is authenticated against that identity before the first byte of application traffic flows. The mesh enforces this transparently via sidecar proxies — application code does not change.

Istio architecture

Istio’s control plane is istiod, a single binary that consolidates three formerly separate components: Pilot (xDS config distribution), Citadel (certificate issuance), and Galley (config validation). The data plane is Envoy, injected as a sidecar into every application pod.

The identity lifecycle works like this: when a pod starts, the kubelet passes a service account token to the Envoy sidecar via the istio-token projected volume. The sidecar presents this token to istiod; istiod validates it against the Kubernetes API and issues a short-lived X.509 SVID (SPIFFE Verifiable Identity Document) signed by the mesh CA. The SVID encodes the SPIFFE ID in the SAN (Subject Alternative Name) field — for example spiffe://saib.sa/ns/payments/sa/payments-svc. When the sidecar opens a connection to another workload, it presents this SVID; the remote sidecar verifies the SVID against the trust bundle and checks that the SPIFFE ID is permitted by the AuthorizationPolicy.

Istio distributes configuration via xDS (Envoy’s discovery service protocol): Listener Discovery Service, Route Discovery Service, Cluster Discovery Service, and Endpoint Discovery Service. The sidecar maintains a gRPC stream to istiod and receives config pushes when policy or topology changes. This is stateful — if istiod restarts, sidecars continue operating on the last-received config and reconnect to the new istiod instance without traffic disruption.

STRICT vs PERMISSIVE mode

Istio’s PeerAuthentication resource controls whether mTLS is enforced. The two active modes are STRICT (only mTLS accepted) and PERMISSIVE (both mTLS and plaintext accepted). A third mode, DISABLE, turns mTLS off entirely and is not appropriate for regulated workloads.

peer-authentication.yamlyaml
# Enforce STRICT mTLS for the entire payments namespace
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: default
  namespace: payments
spec:
  mtls:
    mode: STRICT
---
# For a single workload needing a mode exception during migration
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: legacy-batch-permissive
  namespace: payments
spec:
  selector:
    matchLabels:
      app: legacy-batch-runner
  mtls:
    mode: PERMISSIVE     # temporary; remove after batch runner gains sidecar
PERMISSIVE mode voids your zero-trust compliance claim

PERMISSIVE mode exists solely as a migration tool — a way to let non-mesh workloads communicate with mesh workloads during a phased rollout. A namespace that remains in PERMISSIVE mode indefinitely accepts plaintext connections from any source, including a compromised pod that has bypassed sidecar injection. If SAMA’s examination team tests east-west traffic and finds plaintext accepted, no amount of architectural diagrams showing “zero trust” will satisfy the finding. PERMISSIVE mode should appear in your cluster only with a named workload exception and a remediation date attached.

Sidecar injection must be enforced at the namespace level

STRICT mTLS without sidecar injection on every pod in the namespace is a gap. A pod without a sidecar sends plaintext and receives nothing — or, worse, breaks because a STRICT-mode destination refuses the connection. Label namespaces with istio-injection: enabled and add an OPA or Kyverno admission policy that denies pods missing the sidecar annotation.

SPIFFE identity

SPIFFE (Secure Production Identity Framework For Everyone) is the open standard that defines what a workload identity looks like. A SPIFFE ID is a URI in the form spiffe://<trust-domain>/path. In Kubernetes, Istio maps service account to SPIFFE ID as spiffe://<cluster-domain>/ns/<namespace>/sa/<service-account>.

The SVID is an X.509 certificate with the SPIFFE ID in the SAN field. It is short-lived (Istio’s default is 24 hours; for regulated environments reduce to 1 hour) and is rotated by the sidecar before it expires, transparently and without pod restart.

inspect-svid.shbash
# Inspect the SVID a sidecar is currently presenting
kubectl exec -n payments deploy/payments-svc -c istio-proxy -- \
  openssl s_client -connect accounts-svc.accounts.svc.cluster.local:8080 \
  -showcerts 2>/dev/null | \
  openssl x509 -noout -text | grep -A2 "Subject Alternative Name"

# Expected output:
#   URI:spiffe://saib.sa/ns/payments/sa/payments-svc

# Check SVID expiry on a live sidecar
istioctl proxy-config secret deploy/payments-svc -n payments \
  --output json | \
  jq '.dynamicActiveSecrets[0].secret.tlsCertificate
      .certificateChain.inlineBytes' -r | \
  base64 -d | openssl x509 -noout -dates

For clusters that need cross-cluster or cross-cloud identity (for example, a SAIB DR site on a second OpenShift cluster), SPIRE (the SPIFFE Reference Implementation) provides a federated trust model: each cluster runs a SPIRE Server, and the servers federate by exchanging trust bundles out-of-band. Istio can be configured to use SPIRE as its certificate authority via the ExternalCertificateAuthority plugin.

Certificate rotation

Certificate rotation at scale is the most operationally sensitive part of a service mesh deployment. The failure mode is subtle: if a certificate expires before the sidecar has rotated it, connections from that workload fail with TLS handshake errors. At 200 pods and a 24-hour cert TTL, rotation events happen continuously — which is fine until the rotation mechanism itself breaks.

Rotation window must exceed your longest in-flight request

If a payment processing request takes up to 30 seconds and the cert rotation overlap window is only 10 seconds, a connection established just before rotation completes may terminate mid-flight when the old cert is revoked. Set CITADEL_WORKLOAD_CERT_TTL to 3600 seconds (1 hour) and CITADEL_WORKLOAD_CERT_GRACE_PERIOD_RATIO to 0.5 (rotate at 50% of TTL), giving a 30-minute overlap. Long-running streaming connections (Kafka clients inside the mesh, WebSocket gateways) need explicit analysis.

istio-cert-ttl-configmap.yamlyaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: istio
  namespace: istio-system
data:
  mesh: |
    certificates:
      workloadCertTtl: 3600s          # 1-hour SVID TTL (SAMA-preferred short rotation)
      workloadCertGracePeriodRatio: 0.5  # rotate at 50% lifetime = 30 min overlap
    trustDomain: saib.sa
    ca:
      workloadCertKeySize: 2048
      caRSACertificateKeySize: 4096

For intermediate CA rotation (rotating the CA that signs SVIDs, not individual SVIDs), use cert-manager with an Issuer chain: a root CA stored in Vault PKI, an intermediate CA issued via cert-manager into the istio-system namespace, and Istio configured to use the intermediate. When the intermediate rotates, cert-manager handles the re-issuance; Istio picks up the new signing chain without restart. This is the pattern required for SAMA annual key rotation compliance.

Authorization policies

mTLS proves identity; AuthorizationPolicy enforces intent. Without authorization policies, a service holding a valid SVID can call any other service in the mesh. That is not zero trust — that is authenticated, but not authorized.

authz-policy.yamlyaml
# Allow only payments-svc (in payments namespace) to call accounts-svc
apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
  name: accounts-allow-payments
  namespace: accounts
spec:
  selector:
    matchLabels:
      app: accounts-svc
  action: ALLOW
  rules:
  - from:
    - source:
        principals:
        - "cluster.local/ns/payments/sa/payments-svc"
    to:
    - operation:
        methods: ["GET"]
        paths:  ["/v1/accounts/*"]
---
# Deny-all default for the accounts namespace (explicit allow-list model)
apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
  name: deny-all
  namespace: accounts
spec:
  {}    # empty spec = deny all; ALLOW rules above are additive exceptions

The deny-all default is the critical piece. Without it, all authenticated workloads can call accounts-svc. Start with a deny-all in each namespace and add explicit ALLOW rules per caller. This is the allow-list model that SAMA’s zero-trust guidance requires.

  1. Enable sidecar injection on all namespaces that carry regulated data: kubectl label namespace payments istio-injection=enabled. Add Kyverno admission policy to deny injection-less pods.
  2. Deploy PeerAuthentication in PERMISSIVE mode per namespace. Verify that all existing services are receiving sidecar injection and that east-west traffic appears in Kiali.
  3. Switch each namespace to STRICT mode once 100% of pods show a healthy sidecar. Use istioctl analyze to confirm no plaintext paths remain before switching.
  4. Add a deny-all AuthorizationPolicy to each namespace and incrementally add explicit ALLOW rules, validating with smoke tests after each addition.
  5. Configure cert-manager with the Vault intermediate CA chain, reduce SVID TTL to 3600s, and verify rotation by watching istio_agent_cert_expiry_timestamp in Prometheus.
  6. Enable Istio access logging to a centralized log platform (Elasticsearch, Splunk) for SAMA audit trail. Log fields must include source.principal, destination.principal, request.path, response.code, and start_time.

Linkerd as an alternative

Linkerd takes a different architectural position from Istio: smaller sidecar (the linkerd-proxy, written in Rust, is typically 10–20 MB vs Envoy’s 50–80 MB), no xDS complexity, automatic mTLS with no explicit PeerAuthentication resources required, and a dramatically simpler operational model. The trade-off is less flexibility: Linkerd’s extension model is more constrained than Envoy’s filter chain, and advanced routing (header-based routing, traffic mirroring, fault injection for chaos testing) is available but thinner.

DimensionIstio 1.22Linkerd 2.15
Sidecar binaryEnvoy (C++) — 50–80 MBlinkerd-proxy (Rust) — 10–20 MB
mTLS defaultOff until PeerAuthentication appliedOn for all meshed pods by default
Authorization policyAuthorizationPolicy (rich: methods, paths, claims)AuthorizationPolicy (simpler: route, GRPC method, header)
Traffic managementFull: retries, timeouts, circuit breakers, fault injection, mirroringCore: retries, timeouts, circuit breakers
ObservabilityKiali, Jaeger, Prometheus, Grafana (self-managed)Built-in Viz extension; Jaeger via Linkerd Jaeger extension
OpenShift supportRed Hat OpenShift Service Mesh (OSSM) — production-supportedCommunity-supported; no official Red Hat support
Multi-clusterIstio multi-cluster (replicated, primary-remote) — matureLinkerd multicluster extension — stable
Memory per sidecar~50–150 MB resident~20–40 MB resident
Operational complexityHigh — requires dedicated mesh ops teamModerate — simpler config surface

For a SAMA-regulated bank on RHOCP (Red Hat OpenShift Container Platform), the default choice is OSSM (OpenShift Service Mesh, built on Istio + Kiali + Jaeger + Prometheus). It is Red Hat’s supported product, it satisfies vendor-support requirements in SAMA’s IT governance framework, and its upgrade lifecycle is managed alongside RHOCP. Linkerd is the right choice for a team that finds Istio’s operational complexity prohibitive, has non-RHOCP Kubernetes, and is running a lower-criticality tier of services where the thinner feature set is acceptable.

linkerd-inject.yamlyaml
# Inject linkerd-proxy into a namespace (all new pods auto-injected)
apiVersion: v1
kind: Namespace
metadata:
  name: payments
  annotations:
    linkerd.io/inject: enabled
---
# AuthorizationPolicy in Linkerd (Server + AuthorizationPolicy pair)
apiVersion: policy.linkerd.io/v1beta2
kind: Server
metadata:
  name: accounts-http
  namespace: accounts
spec:
  podSelector:
    matchLabels: { app: accounts-svc }
  port: 8080
  proxyProtocol: HTTP/2
---
apiVersion: policy.linkerd.io/v1beta2
kind: MeshTLSAuthentication
metadata:
  name: payments-identity
  namespace: accounts
spec:
  identities:
  - "payments-svc.payments.serviceaccount.identity.linkerd.cluster.local"

Performance overhead

Service mesh mTLS adds latency at two points: the TLS handshake on connection establishment, and the encryption/decryption overhead on each request. In practice, connection pools mean that the handshake cost is amortized over many requests; the encryption overhead is the dominant factor.

Benchmarking Istio with Envoy sidecar against a baseline (no mesh) on a payment processing microservice at SAIB-representative load (500 RPS, 2 KB payload, gRPC):

Configurationp50 latencyp99 latencyCPU per pod (sidecar)Memory per pod (sidecar)
No mesh (baseline)1.2 ms4.8 ms——
Istio PERMISSIVE (no mTLS active)1.8 ms6.1 ms~20 m~55 MB
Istio STRICT mTLS2.1 ms7.2 ms~30 m~60 MB
Linkerd mTLS (default on)1.9 ms6.4 ms~15 m~28 MB

The overhead is real but acceptable: ~0.9 ms at p50, ~2.4 ms at p99 for Istio. At SAMA’s IPS payment latency requirements (end-to-end under 2 seconds), this is not the constraint. Where it matters is high-frequency internal loops — a 10,000 RPS service-to-service call where each hop adds 2 ms is a 20 ms penalty on a chain with ten hops.

Sidecar resource allocation must be in the pod spec

The Envoy sidecar is injected with default resource requests of 100m CPU and 128Mi memory. At scale, this default is both too low for high-throughput pods (causing CPU throttling) and unaccounted-for in capacity planning. Set explicit sidecar resource limits via meshConfig.defaultConfig.resources in the Istio ConfigMap, and include sidecar resources in your pod-level capacity estimates from the start.

Production checklist

  1. All namespaces carrying regulated data labelled for sidecar injection. Admission policy blocks injection-less pods in regulated namespaces.
  2. All regulated namespaces in STRICT mTLS mode. No namespace permanently in PERMISSIVE. Exceptions documented with name, justification, and remediation date.
  3. Deny-all AuthorizationPolicy in every regulated namespace. Explicit ALLOW rules per caller/path/method.
  4. SVID TTL at 3600s. Grace period ratio at 0.5 (30-minute overlap). Prometheus alert on istio_agent_cert_expiry_timestamp < now + 600s.
  5. Intermediate CA signed by Vault PKI. cert-manager managing intermediate renewal. Root CA stored offline or in HSM.
  6. Trust domain set to a bank-specific value (e.g. saib.sa), not the default cluster.local (cluster.local is widely reused and reduces SPIFFE ID uniqueness in multi-cluster federation).
  7. Istio access logs shipped to centralized SIEM. Fields: source principal, destination principal, request path, response code, duration, start time.
  8. Kiali (or Linkerd Viz) deployed and accessible only to network ops team. No public exposure.
  9. mTLS status visible in monitoring. Dashboard shows: % namespaces in STRICT, % pods meshed, cert expiry horizon, AuthorizationPolicy deny rate.
  10. Multi-cluster federation documented. Trust bundle exchange process defined. DR cluster mTLS config parity with primary validated quarterly.