Most integration platforms at Saudi banks are monitored. They are not observed. The distinction matters more than it sounds. Monitoring tells you when something has already failed: a queue depth alert fires, a health-check returns 503, a dashboard goes red. Observability tells you why something is failing, across system boundaries, in the specific request that triggered the failure, before the second failure in the cascade makes the first one unrecoverable. The gap between those two postures — between reactive monitoring and proactive observability — is the gap between an operations team that closes incidents in hours and one that chases symptoms for days.
The technical solution is well-defined: OpenTelemetry for instrumentation, Prometheus for metrics, Loki for logs, Tempo for distributed traces, Grafana as the unified query layer. None of these are new technologies, and none of them are expensive relative to the outage costs they prevent. The leadership decision is not whether to adopt the stack. It is whether your organisation is willing to treat observability as infrastructure rather than a tool, to fund it accordingly, and to hold teams accountable for the signal quality their services emit.
What the regulator will ask
SAMA’s Technology Risk Management framework and the Cyber Security Framework both treat operational visibility as a regulated control, not a convenience. The examination question is not “do you have dashboards?” — it is “can you produce, on demand, a complete audit trail of which system accessed which data, at what time, in response to which user action?”
A distributed tracing platform that correlates every payment instruction from API gateway to core banking system, with timestamps, service identities, and response codes on every hop, answers that question completely. It also answers the question that follows: “When the IBAN transfer at 14:32 on the 3rd was delayed by 4 seconds, which system caused the delay?” Without correlated traces, that answer requires manually correlating log files across five systems, mapping timestamps that may differ by milliseconds, and hoping that the right log level was enabled at the time. With traces, it is a single query.
There is also a PDPL dimension. The same instrumentation that makes your integration estate observable can inadvertently capture personal data — account numbers in HTTP query strings, IBANs in database statements, customer IDs in span attributes. An observability platform that is not governed is a personal data exposure waiting to happen. The correct response is not to avoid instrumentation; it is to treat attribute sanitisation as a mandatory configuration step and to scope access to the telemetry platform using the same team-level controls applied to any other sensitive system.
An observability platform that captures everything is a liability without governance. An observability platform governed like a data asset is a regulatory advantage.
The trade-off the business needs to own
Observability infrastructure has a cost that is easy to underestimate in the initial business case. The storage backends — Prometheus TSDB, Loki object store, Tempo object store — grow with the number of services instrumented and the retention period required. A 30-service integration estate running at modest load, with 90-day log retention for regulated logs, can accumulate 5–10 TiB of telemetry data over a year. Object storage is cheap; the index, the compaction jobs, and the query capacity for Loki are not trivially sized.
The growth is manageable if cost governance is built in from day one: log retention tiered by classification, high-cardinality metrics suppressed at the collector before they reach Prometheus, and trace sampling designed to keep 100% of error traces while sampling normal traffic at 10%. The integration platforms that end up with runaway observability costs are the ones that instrumented everything at full fidelity, set uniform 90-day retention, and never audited volume growth. The quarterly storage invoice is how they discover the problem.
The second cost is operational. The observability platform is itself a platform that needs to be operated. Someone owns Prometheus capacity, Loki compaction failures, Tempo query timeouts, Grafana upgrades, and OTel Collector version management. If that ownership is assumed rather than assigned — if the implicit assumption is that “the platform team will handle it” when the platform team is already at capacity running six other systems — the observability platform will be the first thing to degrade when it is most needed, which is during an incident.
The decision that leadership has to make
Three commitments determine whether an observability programme succeeds or produces expensive infrastructure that teams learn to ignore.
First: instrument as a mandate, not a recommendation. Teams that are asked to add OTel instrumentation will do it eventually, unevenly, and with inconsistent attribute naming. Teams that are told every service must emit standard OpenTelemetry signals — with a validated baseline configuration, enforced by a CI pipeline check, with a named person accountable for signal quality in each domain — will do it completely. The difference in observability utility between 80% coverage and 100% coverage is not 20%; it is the difference between traces that stop at the boundary of the unmonitored service and traces that span the full call graph.
Second: resolve the data residency and access control design before instrumentation begins. Retrofitting multi-tenancy onto Loki after 60 days of log ingestion means re-labelling six weeks of data or starting again. Retrofitting attribute sanitisation after discovering IBAN leakage in traces means an incident report and a root-cause remediation under time pressure. Neither of these is a platform engineering problem; both are a consequence of governance decisions deferred until it was too late to design them in cleanly.
Third: fund the on-call rotation. An observability platform that pages on its own failures at 2 a.m. without an engineer who knows how to respond is worse than no observability platform, because it erodes trust in alerts and creates alert fatigue. The platform engineering team that builds the observability stack needs at least one engineer with deep expertise in each component — someone who can diagnose a Loki ingester OOM or a Tempo compactor stall under incident pressure. That expertise is not acquired in passing; it requires dedicated learning time and, ideally, a staging environment that is allowed to fail.
For the engineering depth behind this decision — OTel Collector pipeline configuration, Prometheus recording rules for ACE and Kafka, Loki log parsing for IBM MQ, tail-based trace sampling, cardinality control, and the production checklist — see the companion Lab article: DevOps & Platform: Observability Pipeline Architecture.