Overview

Observability in integration platforms is harder than in microservice applications because the signals cross more system boundaries. A payment message enters an IBM ACE integration server, publishes to IBM MQ, is consumed by a Kafka Streams processor, triggers an ISO 20022 transformation, calls a core banking REST endpoint, and writes the result to an Oracle database — all within a single business transaction. If end-to-end latency is 2.3 seconds instead of 400 milliseconds, you need to know which hop is slow and why. Logs from each system alone don’t answer that. Metrics alone don’t either. Only correlated traces, enriched with metrics and logs, do.

The stack that holds in production: OTel everywhere for instrumentation and collection, Prometheus for metrics, Loki for logs, Tempo for traces, Grafana as the unified query and dashboard layer. This is not a Grafana Labs advocacy article — it is the stack that works on OpenShift without a cloud-managed backend, without license cost per node, and without sending regulated data to a SaaS endpoint outside the SAMA regulatory perimeter.

OpenTelemetry is the instrumentation standard, not a product

OpenTelemetry (OTel) defines APIs, SDKs, and a wire protocol (OTLP). It does not store data — the Collector routes signals to storage backends. Adopting OTel means your instrumentation is backend-agnostic: swap Prometheus for Thanos, Loki for Elastic, or Tempo for Jaeger without touching application code. This portability matters when SAMA or internal security requirements force a backend change mid-programme.

Three Signals

Observability is built on three complementary signal types. They answer different questions and have different cost profiles. Picking the wrong signal for a question is expensive — querying Loki for a latency answer that Prometheus could return in milliseconds will teach you this immediately.

SignalToolResolutionHot retentionQueryCost driver
MetricsPrometheus 2.5315 s scrape15–30 daysPromQLSeries cardinality
LogsLoki 3.1On ingestion60–90 daysLogQLGiB ingested / day
TracesTempo 2.5Per request7–14 daysTraceQLSpan volume × retention

Practical guidance: instrument everything with OTel (all three signals), store metrics long, store logs medium, store traces short with head-based sampling for volume and tail-based sampling for errors. Never delete traces for transactions that resulted in errors or regulatory events — those go to cold storage.

OpenTelemetry

OTel provides two entry points for Java-based integration services: the Java Agent (zero-code instrumentation via -javaagent JAR) and the SDK (manual instrumentation for custom spans). IBM ACE integration servers use the Java Agent approach via the server’s jvmArgs in server.conf.yaml; Kafka Streams applications use the SDK directly.

The OTel Java Agent auto-instruments 60+ libraries at attach time: JDBC, HTTP clients, Kafka producer/consumer, JAX-RS, Spring, gRPC. Spans for each ACE message flow node are created automatically once the agent is attached. For MQ, the IBM MQ JMS client is instrumented by the agent; context propagation across MQ message boundaries requires ACE 12.0.10+ with the OTel trace propagation policy enabled — more on this in pitfalls.

Metrics: Prometheus

Prometheus is pull-based: it scrapes HTTP endpoints that expose metrics in text format, stores them in its TSDB, and evaluates alerting and recording rules on a configurable interval. In a regulated environment the pull model is an advantage — the Prometheus server controls what it collects, reducing the risk of a compromised application pushing fabricated metrics.

Critical metric categories for an integration platform:

  • Message flow throughput. Messages per second entering and leaving each ACE integration server, by flow name. Sudden drops indicate upstream stalls or application errors.
  • End-to-end flow latency. Time from message receipt to downstream acknowledgement. Track as a histogram — not a gauge — so you can compute p50/p95/p99 without re-scraping raw data.
  • MQ queue depth. Exported via the IBM MQ Prometheus exporter (community). A growing depth is an early warning of a consumer falling behind; alert at 10 000 msgs, page at 50 000.
  • Kafka consumer group lag. Per topic, per partition, per consumer group. The payments-bridge consumer group lag is the most operationally critical metric in most integration estates.
  • Pipeline CI/CD metrics. Tekton pipeline duration histogram and success rate by pipeline name. A p95 build time exceeding 15 minutes erodes developer trust faster than failures do.
recording-rules.yamlyaml
groups:
  - name: integration_platform_slos
    interval: 30s
    rules:

      # 5-minute rate of ACE message flow ingestion
      - record: job:ace_flow_msgs:rate5m
        expr: rate(ace_message_flow_input_count_total[5m])

      # p99 end-to-end flow latency
      - record: job:ace_flow_latency_p99:rate5m
        expr: |
          histogram_quantile(0.99, sum by (flow_name, le) (
            rate(ace_message_flow_duration_seconds_bucket[5m])
          ))

      # MQ queue depth breach
      - alert: MQQueueDepthHigh
        expr: ibmmq_queue_depth > 10000
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: 'MQ {{ $labels.queue }} on {{ $labels.qmgr }}: {{ $value }} msgs'
          runbook_url: "https://wiki.internal/runbooks/mq-queue-depth"

      # Kafka consumer lag for payments bridge
      - alert: KafkaConsumerLagCritical
        expr: kafka_consumer_group_lag{group="payments-bridge"} > 5000
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: 'payments-bridge lag {{ $value }} msgs on {{ $labels.topic }}'
          runbook_url: "https://wiki.internal/runbooks/kafka-consumer-lag"
High cardinality kills Prometheus

Every unique label combination creates a new time series. Adding customer_id or transaction_id as a Prometheus label is the most common cause of OOM kills in this stack. Use these values as Loki log fields or Tempo span attributes — never as metric labels. Budget 500–2 000 active series per service, enforce it via the OTel Collector’s filter/metric processor, and review the cardinality dashboard weekly for the first 90 days after rollout.

Logs: Loki

Loki is intentionally minimalist: it indexes only stream labels, not log line content. This makes it dramatically cheaper than Elasticsearch at scale, but changes how you query — you filter by label first, then full-text-search the matching chunks. For integration platforms this is the right trade-off: you almost always know the application, the environment, and the flow name before you need to search content.

ACE produces structured JSON logs when userTrace is enabled and the IBM ACE OTel log exporter is configured. MQ channel initiator logs are less structured; parse them in the Collector using a regex_parser operator in the filelog receiver pipeline.

otel-collector-logs-pipeline.yamlyaml
receivers:
  otlp:
    protocols:
      grpc: { endpoint: "0.0.0.0:4317" }
  filelog:
    include: ["/var/log/ace/*.log"]
    operators:
      - type: json_parser
        timestamp:
          parse_from: attributes.timestamp
          layout: "%Y-%m-%dT%H:%M:%S.%fZ"
      - type: regex_parser   # extract ACE flow name
        regex: 'Flow=(?P<flow_name>[^,;]+)'
        parse_from: body
        on_error: send   # don't drop lines that don't match

processors:
  resource:
    attributes:
      - { action: insert, key: loki.resource.labels,
          value: "service.name,k8s.namespace.name,flow_name" }
  filter/drop_health:    # suppress heartbeat lines before they leave the node
    error_mode: ignore
    logs:
      log_record:
        - 'body contains "HealthCheck"'
        - 'body contains "BIPE3788I"'  # ACE broker ready heartbeat
  batch:
    timeout: 5s
    send_batch_size: 8192

exporters:
  loki:
    endpoint: "http://loki-gateway.monitoring.svc:3100/loki/api/v1/push"
    tls:
      insecure: false
      ca_file: "/etc/otelcol/certs/internal-ca.crt"

service:
  pipelines:
    logs:
      receivers:  [otlp, filelog]
      processors: [resource, filter/drop_health, batch]
      exporters:  [loki]

Loki’s chunk-based storage backs onto object storage — OpenShift Data Foundation (ODF/Ceph) or an on-premise S3-compatible endpoint. Configure retention per-tenant in Loki’s limits_config: 90 days for regulated logs (transaction, audit), 30 days for application INFO, 7 days for debug-level output.

Distributed Traces: Tempo

Distributed tracing answers the question metrics and logs cannot: which specific code path, in which service, in which call graph, accounted for the latency spike at 14:32 today? A trace is a tree of spans, each representing a unit of work — a JDBC query, an MQ send, an ACE node execution, an HTTP call. Every span carries a traceId linking it to the root transaction and a parentSpanId linking it to its caller.

Tempo stores traces in object storage with no indexing of span content — only TraceID and service name — making it significantly cheaper than Jaeger+Elasticsearch at comparable trace volumes. TraceQL (Tempo 2.0+) allows structured queries over span attributes: find all traces where db.statement took over 500 ms, or where rpc.method="ISO20022Transform" returned an error.

PII in span attributes is a data classification incident

The OTel Java Agent captures HTTP request headers, query parameters, and JDBC statement parameters by default — all of which may carry customer IDs, account numbers, or IBANs. Before enabling any integration service for tracing in production, audit and configure attribute sanitisation. Suppress db.statement, http.url query strings, and any custom attribute that carries personal data. Under SAMA’s data classification policy and Saudi Arabia’s PDPL, a trace buffer containing IBANs that is accessible to operations engineers is a breach, not an acceptable operational trade-off.

Use head-based sampling at 10% for high-volume services and tail-based sampling to keep 100% of error traces, regardless of head sample decision. The OTel Collector’s tail_sampling processor implements this by buffering spans for a configurable window (10–30 s) before evaluating policies.

Collector Pipeline Architecture

Run the Collector in two deployment modes simultaneously: a DaemonSet on every node (for pod-local file log and host metrics collection) and a Gateway Deployment with horizontal scaling (for OTLP ingest from all pods). This avoids the fan-out problem where every pod opens a direct connection to each backend.

otel-collector-gateway.yamlyaml
processors:
  memory_limiter:          # NEVER skip this — prevents OOM on traffic spike
    check_interval: 1s
    limit_mib: 1500         # 75% of container memory request
    spike_limit_mib: 400
  batch:
    timeout: 5s
    send_batch_size: 4096
  tail_sampling:
    decision_wait: 15s
    num_traces: 50000
    policies:
      - { name: always-errors, type: status_code,
          status_code: { status_codes: [ERROR] } }
      - { name: probabilistic-10pct, type: probabilistic,
          probabilistic: { sampling_percentage: 10 } }
  filter/drop_healthz:
    error_mode: ignore
    traces:
      span:
        - 'attributes["http.route"] == "/healthz"'
        - 'attributes["http.route"] == "/readyz"'
        - 'attributes["http.route"] == "/metrics"'

exporters:
  prometheusremotewrite:
    endpoint: "http://prometheus.monitoring.svc:9090/api/v1/write"
    resource_to_telemetry_conversion: { enabled: true }
  loki:
    endpoint: "http://loki-gateway.monitoring.svc:3100/loki/api/v1/push"
  otlp/tempo:
    endpoint: "tempo-distributor.monitoring.svc:4317"
    tls: { insecure: true }   # mTLS handled by service mesh sidecar

service:
  pipelines:
    traces:
      receivers:  [otlp]
      processors: [memory_limiter, filter/drop_healthz, tail_sampling, batch]
      exporters:  [otlp/tempo]
    metrics:
      receivers:  [otlp]
      processors: [memory_limiter, resource, batch]
      exporters:  [prometheusremotewrite]
    logs:
      receivers:  [otlp]
      processors: [memory_limiter, resource, batch]
      exporters:  [loki]

Set the Collector container memory request to match expected peak ingest. A rough budget: 500 MiB per 50 000 spans/second at the Gateway; 256 MiB per node at the DaemonSet. Size the DaemonSet filequeue buffer to hold at least 60 seconds of peak log volume so the DaemonSet survives a Gateway restart without data loss.

Grafana: Unified Pane

Grafana 11’s value is correlation. From a Prometheus alert panel you drill into Loki logs for the same time window and service, then pivot to a Tempo trace for the specific failing request. This reduces mean-time-to-diagnosis from hours of tab-switching to a single page interaction.

Three dashboard types cover 80% of operational use on an integration platform:

  • Platform SLO dashboard. One row per integration domain: traffic rate, error rate, latency p99, and SLO burn rate. The burn rate panel alerts before the SLO window expires — not after the SLO has already been breached for the week.
  • Service map. Tempo’s service-graph processor emits inter-service call metrics that Grafana renders as a live dependency graph. Essential for understanding blast radius when a downstream core banking system degrades.
  • Runbook-linked alerts. Every Prometheus alert annotation carries a runbook_url. Grafana renders the link in the alert detail panel; on-call engineers don’t need to remember the procedure under pressure.
Grafana Alerting replaces separate Alertmanager config

Grafana 9+ ships a unified alerting layer that evaluates rules against any datasource — Prometheus, Loki, or Tempo — and routes through a single notification policy. You no longer need a separate Alertmanager instance per backend. Route all alerts through Grafana Alerting to a single PagerDuty or OpsGenie integration; Alertmanager instances on each backend become optional forwarding receivers only.

Cardinality and Cost Control

Observability platform cost is determined almost entirely by two variables: Prometheus active series count and Loki ingest volume (GiB/day). Trace volume is manageable once tail-based sampling is in place. The cost conversation with finance is almost always about uncontrolled log volume growth compounding quarterly.

  1. Audit existing volume. Run sum(rate(loki_ingester_samples_ingested_total[24h])) by (namespace) to rank namespaces by ingestion rate. Three namespaces typically produce 70% of total volume — focus cost reduction there first.
  2. Drop at the Collector, not at Loki. Filter health-check, heartbeat, and DEBUG-level log records in the Collector’s filter processor before they leave the node. Bytes dropped at source cost nothing; bytes ingested into Loki always consume storage and index.
  3. Cap series per application. Set a Prometheus limits.max_samples_per_send on remote write, or enforce it via the Collector’s transform processor to drop labels with unbounded cardinality before they reach Prometheus.
  4. Tier log retention by classification. 7 days for DEBUG, 30 days for INFO, 90+ days for audit, security-event, and transaction logs. Configure per-tenant retention via Loki compactor with retention_enabled: true.
  5. Index only what you query by. Loki’s stream index stores label values. Every new unique combination is an index entry. Do not use log content fields as Loki labels — use the json pipeline stage and query fields with | json | .field == "value" in LogQL.
  6. Review monthly. Loki’s built-in usage dashboard breaks down ingestion by namespace, pod, and label. Treat this as a standing monthly review in platform engineering; uncontrolled growth is invisible until the quarterly storage invoice arrives.

Compliance and Data Governance

In a Saudi bank regulated by SAMA and subject to the Personal Data Protection Law (PDPL), the observability pipeline is a data processor. It receives, stores, and potentially exposes customer-adjacent data. Design compliance in from the start; retrofitting access control onto a 90-day log corpus after a SAMA examination request is neither fast nor cheap.

  • Data residency. All backends must run on infrastructure within Saudi Arabia. Do not route OTLP to a SaaS collector, even temporarily for testing. ODF/Ceph object storage backing Loki and Tempo must be on-premise or in a SAMA-approved cloud region.
  • Datasource scoping in Grafana. Use Grafana’s organisation/team model and Loki multi-tenancy (X-Scope-OrgID header) to enforce namespace isolation: the payments team can query the payments namespace; they cannot query the treasury namespace. Prometheus has no built-in authorization — enforce it at the Grafana datasource proxy layer.
  • Audit log for queries. Grafana’s audit log (or Nginx access log in front of Grafana) captures every dashboard view and data query. Alert on any Loki query returning more than 10 000 log lines — it is a likely data harvesting pattern rather than legitimate troubleshooting.
  • Compactor verification. Add a Prometheus rule that alerts if Loki/Tempo object store usage grows faster than net ingestion rate — it means the compactor’s retention job is not deleting expired chunks. Verify compaction on a quarterly schedule, not just at initial deployment.

Common Pitfalls

OTel Collector as a single point of failure

If the Gateway Collector crashes, all telemetry stops. Run the Gateway as a Deployment with at least 3 replicas behind a Kubernetes Service, and set a PodDisruptionBudget allowing at most 1 unavailable pod. The DaemonSet Collectors must buffer to disk (filequeue exporter) if the Gateway is unreachable; without a filequeue, a Gateway restart drops minutes of telemetry silently.

TraceID not propagated across MQ message boundaries

W3C TraceContext propagates via HTTP headers or Kafka record headers — not through IBM MQ message properties by default. ACE 12.0.10+ supports OTel context propagation through a custom trace propagation policy; without it, traces break at every MQ boundary. You get a forest of disconnected single-service traces rather than end-to-end trees. Verify propagation explicitly before declaring tracing production-ready on any integration flow that crosses an MQ boundary.

Clock skew breaks trace waterfall views

Spans are stitched together by timestamp as well as TraceID. If node clocks diverge by more than 500 milliseconds, trace waterfall views show child spans starting before their parent — confusing enough to make traces useless for diagnosis. Enforce NTP synchronisation across all OpenShift nodes as a platform prerequisite. Alert on clock drift above 200 ms using the node_timex_offset_seconds metric from node exporter.

Production Checklist

  1. OTel Java Agent attached to all ACE and Spring Boot JVMs; OTel SDK in Kafka Streams apps.
  2. PII sanitisation config reviewed and tested per service before traces are enabled in prod.
  3. Gateway Collector: 3 replicas, PodDisruptionBudget, memory_limiter at 75% of memory request.
  4. DaemonSet Collectors buffering to disk (filequeue); buffer sized for 60 s of peak log ingest.
  5. Tail-based sampling: 100% error traces kept, 10% of normal traffic.
  6. Prometheus recording rules for ACE throughput, MQ depth, Kafka lag, pipeline duration.
  7. Loki retention tiered: 90 days audit/transaction, 30 days INFO, 7 days DEBUG.
  8. Tempo retention: 14 days hot on ODF; cold archive for error/regulatory traces.
  9. Grafana team datasource scopes enforced: payment team cannot query treasury namespace.
  10. Alert on Loki queries returning >10 000 lines (data harvesting indicator).
  11. Compactor retention verified quarterly; object-store growth rate alert configured.
  12. NTP enforced across all nodes; clock drift alert at 200 ms on node_timex_offset_seconds.
  13. W3C TraceContext propagation tested across every MQ boundary in the integration estate.
  14. Runbook URL set in every Prometheus alert annotation; on-call playbook tested quarterly.