Overview

GraphQL has won the developer-experience argument for internal API platforms — strong typing, self-documenting schemas, and client-driven query shaping are genuine productivity improvements over REST for frontend and mobile teams. The hard question is not whether to adopt GraphQL. It is whether to adopt it as a monolith or as a federated supergraph, and that decision is fundamentally about how domain ownership is structured across your engineering teams.

A monolithic GraphQL schema is owned by whoever last touched it. Every team that needs to expose a new entity or field must either wait for a central team to merge their change or accept co-ownership of a schema that spans domains they do not understand. This works at low team count; it breaks at scale. When your accounts domain team, your payments domain team, and your notifications team each have independent release cycles, independent deployment pipelines, and independent on-call rotations, they cannot all be blocked by a shared schema gateway that only one team owns.

Apollo Federation 2 resolves this by splitting the schema. Each domain publishes a subgraph — a GraphQL schema covering only the types that domain owns — and the Apollo Router composes those subgraphs into a unified supergraph at runtime. Consumers query the supergraph and receive a single response; the Router plans and executes the query across however many subgraphs are involved, transparently.

REST still serves edge traffic. GraphQL is not a replacement for REST for SAMA Open Banking APIs, ISO 20022 payment initiation, or machine-to-machine integrations where the consumer is a third-party fintech with a fixed contract. Federation sits inside the perimeter — between Kong at the edge and the domain microservices behind it. Kong handles authentication, rate limiting, and WAF for all traffic; the Apollo Router handles GraphQL routing for the subset of that traffic that arrives as GraphQL operations.

Federation is an API contract, not just a routing problem

Apollo Federation requires domain teams to publish schema SDL as a deployment artifact, register it with a schema registry, and coordinate on cross-domain types using directives. This is a governance and contract model, not a technical routing optimisation. Teams that adopt federation without the governance framework — no breaking-change policy, no schema registry, no ownership mapping — will find it harder to operate than a monolith. The technical architecture only holds if the organisational architecture supports it.

Apollo Federation 2 Architecture

Federation 2 is built on four concepts: the subgraph schema, the supergraph schema, the schema registry, and the Router.

Each domain team publishes a subgraph schema — a standard SDL file annotated with federation directives — to a schema registry (Apollo GraphOS Managed Federation, or a self-hosted Rover CLI pipeline feeding a custom registry). The registry runs composition: it validates that all subgraph schemas are compatible, resolves cross-subgraph type references, and emits a compiled supergraph schema. The Apollo Router reads this supergraph schema and uses it to build a query plan for every incoming operation at request time.

The key federation 2 directives that make this work:

  • @key(fields: "id") — marks the primary key of an entity type. Any subgraph that defines a type with @key can resolve that entity by its key. The Router uses this to stitch cross-subgraph references.
  • @external — marks a field as owned by another subgraph. A subgraph that references an external field declares it with @external so the Router knows to fetch it from the owning subgraph.
  • @provides(fields: "...") — tells the Router that a subgraph can locally satisfy a specific set of external fields for a type, avoiding an extra fetch to the owning subgraph when those fields are in scope.
  • @requires(fields: "...") — the inverse: declares that a resolver for a given field requires specific fields from the entity (potentially owned by another subgraph) to be fetched first.
  • @shareable — marks a non-entity type as resolvable by multiple subgraphs. Useful for value types shared between domain teams (e.g. a Money scalar or a Currency enum used by both payments and accounts).

Query planning is deterministic given the same supergraph schema and operation. The Router computes a plan, executes fetch operations against subgraphs in parallel where there are no dependencies, and sequences them where @requires chains create a dependency order. At SAIB scale — <500 rps on the internal GraphQL surface — plan computation adds under 2 ms of overhead per request; this becomes relevant at 5–10K rps if plan caching is not enabled.

Domain Subgraph Design

The cardinal rule of subgraph design: one domain team owns one subgraph. The payments team owns the payments subgraph. They define its schema, they deploy it on their pipeline, and they are on-call when it fails. The Apollo Router does not care which namespace the subgraph runs in or which team deployed it — it only needs a reachable HTTPS endpoint and a registered SDL.

For a banking integration platform, three subgraphs cover the primary internal API surface:

  • Accounts subgraph — owns Account, Balance, and Statement. The entity key is iban (the IBAN string, validated against Saudi IBAN format). Downstream: core banking balance enquiry service via synchronous REST. Latency SLO: p99 <250 ms.
  • Payments subgraph — owns Transaction, PaymentLimit, and the ISO 20022 pain.001/camt.053 wrapper types. Entity key: transactionId (UUID). Cross-reference to Account via @external + @requires: a transaction resolver needs the source account’s IBAN, which the accounts subgraph owns, before calling the payment execution engine. The @provides directive on Transaction.sourceAccount tells the Router it can skip a round-trip to Accounts if the account’s IBAN is already in the fetch result.
  • Notifications subgraph — owns Alert, NotificationPreference, and Channel. Entity key: userId. Mutations here are fire-and-forget; the subgraph publishes to an internal Kafka topic rather than waiting for delivery confirmation. The Router treats the mutation as resolved once the subgraph returns the Kafka publish acknowledgement.

The @shareable directive applies to value types that multiple subgraphs need to embed inline without a cross-subgraph fetch. A Money type (amount + currency + ISO 4217 code) is used by both Accounts and Payments; declaring it @shareable means both subgraphs can define and resolve it independently without violating the “one owner” principle.

payments-subgraph.graphqlgraphql
extend schema
  @link(url: "https://specs.apollo.dev/federation/v2.3",
        import: ["@key", "@external", "@requires", "@provides", "@shareable"])

# Entity owned by this subgraph — resolvable by transactionId
type Transaction @key(fields: "transactionId") {
  transactionId:  ID!
  amount:         Money!
  status:         TransactionStatus!
  initiatedAt:    DateTime!
  # Cross-domain reference: Account is owned by the Accounts subgraph
  sourceAccount:  Account! @provides(fields: "iban")
}

# Reference type — key defined here, fields resolved by Accounts subgraph
type Account @key(fields: "iban", resolvable: false) {
  iban: String! @external
}

# Value type shared across subgraphs — no cross-subgraph fetch needed
type Money @shareable {
  amount:   Decimal!
  currency: String!   # ISO 4217, e.g. "SAR"
}

type PaymentLimit @key(fields: "limitId") {
  limitId:   ID!
  channel:   String!   # SARIE, SWIFT, IPS
  daily:     Money!
  perTx:     Money!
}

enum TransactionStatus {
  PENDING
  CLEARED
  REJECTED
  REVERSED
}

type Query {
  transaction(transactionId: ID!):       Transaction
  transactionHistory(iban: String!, first: Int): [Transaction!]!
  paymentLimits(channel: String!):        PaymentLimit
}

type Mutation {
  initiatePayment(input: PaymentInput!): Transaction!
}

Apollo Router Configuration on OpenShift

The Apollo Router runs as a Kubernetes Deployment on OpenShift. Unlike the Apollo Gateway (the Node.js predecessor), the Router is a compiled Rust binary with significantly lower latency and memory overhead. A single Router pod handles approximately 2 000 rps at p99 <10 ms plan + execution overhead; run at least 3 replicas with a PodDisruptionBudget of max-unavailable 1.

The Router reads its configuration from a router.yaml mounted as a ConfigMap. The supergraph schema comes from one of two sources: Managed Federation (Router polls GraphOS for schema updates — requires outbound connectivity to Apollo’s cloud, which may conflict with SAMA data-residency requirements) or schema file mount (the supergraph SDL is compiled by Rover CLI in the CI pipeline and mounted as a ConfigMap, updated on every subgraph push). The self-hosted path is preferred in a regulated environment where external SaaS dependencies at the data plane are controlled.

router.yamlyaml
supergraph:
  listen: "0.0.0.0:4000"
  path: "/graphql"
  introspection: false   # never expose introspection without auth in production

health_check:
  listen: "0.0.0.0:8088"

sandbox:
  enabled: false         # disable GraphQL sandbox in non-dev environments

authentication:
  router:
    jwt:
      jwks:
        - url: "https://idp.internal.saib.com.sa/.well-known/jwks.json"
          issuer: "https://idp.internal.saib.com.sa"
      header_name: "Authorization"
      header_value_prefix: "Bearer "

authorization:
  require_authentication: true  # reject unauthenticated requests at the router
  preview_directives:
    enabled: true             # enable @authenticated / @requiresScopes directives

headers:
  all:
    request:
      # Propagate correlation headers from client into subgraph requests
      - propagate: { named: "x-correlation-id" }
      - propagate: { named: "x-request-id" }
      - propagate: { named: "x-channel" }   # mobile / web / branch
      # Inject resolved JWT claims into subgraph requests
      - insert:
          name: "x-user-id"
          value: "{{ $claims.sub }}"
      - insert:
          name: "x-user-roles"
          value: "{{ $claims.roles | join(\",\") }}"

limits:
  max_depth: 10             # reject queries deeper than 10 levels
  max_aliases: 15
  max_root_fields: 5        # max root-level selections per operation
  max_height: 50             # total field count across all selections

persisted_queries:
  enabled: true
  safelist:
    enabled: true
    require_id: true          # in production: only registered operations accepted

telemetry:
  tracing:
    propagation:
      trace_context: true    # W3C TraceContext into subgraph requests
    exporters:
      otlp:
        - endpoint: "http://otel-collector.monitoring.svc:4317"
          protocol: grpc
  metrics:
    exporters:
      prometheus:
        - listen: "0.0.0.0:9090"
          path: "/metrics"

traffic_shaping:
  all:
    timeout: 30s
    compression: true
  subgraphs:
    payments:
      timeout: 15s   # tighter SLO for payment subgraph
    accounts:
      timeout: 10s

GraphQL-Aware Rate Limiting

Standard rate limiting counts requests. GraphQL rate limiting must count cost. A single GraphQL operation can be trivial (fetch one scalar field) or catastrophic (recursive fragments fetching thousands of nested objects). Treating both the same at the IP or token level misses the actual load profile.

The Kong graphql-proxy-cache-advanced plugin and the upstream graphql-rate-limiting-advanced plugin (Kong Enterprise) implement field-level complexity scoring: each field in the schema is assigned a cost weight. The plugin computes the total cost of an incoming operation before execution and rejects it if the cost exceeds the per-client budget.

Cost configuration principles for a banking API:

  • Scalar fields: cost 1. A query for a single account balance costs roughly as much as the number of scalar fields selected.
  • List fields: cost = multiplier × child cost. transactionHistory(first: 100) with 5 scalar fields costs 500 points. Set a default multiplier of 10 for unspecified list fields; require clients to pass explicit first/limit arguments and reject operations that omit them.
  • Mutation fields: minimum cost 10 regardless of field count. Mutations have side effects and should not be treated as cheap read operations.
  • Per-client budget: internal BFF clients 5 000 points/minute; third-party portal clients 1 000 points/minute; machine-to-machine service accounts 10 000 points/minute with a dedicated tier.

Depth and alias limits in the Router (limits.max_depth, limits.max_aliases) complement complexity scoring. Depth limits prevent deeply nested query attacks; alias limits prevent clients from requesting the same expensive field 50 times under different names to multiply their effective cost.

kong-graphql-rate-limit-plugin.yamlyaml
name: graphql-rate-limiting-advanced
service: apollo-router-svc
config:
  limit_by: consumer         # rate-limit per authenticated consumer (JWT sub)
  window_size: 60            # sliding window in seconds
  window_type: sliding
  cost_strategy: default     # use schema-introspected field complexity
  max_cost: 5000             # internal BFF tier: 5 000 points/min
  score_factor: 1.0
  sync_rate: -1              # use Redis for distributed counting (cluster mode)
  strategy: redis
  redis:
    cluster_addresses:
      - "redis-0.redis-headless.infra.svc:6379"
      - "redis-1.redis-headless.infra.svc:6379"
      - "redis-2.redis-headless.infra.svc:6379"
    ssl: true
    ssl_verify: true
    ssl_server_name: "redis.infra.svc"
  hide_client_headers: false  # expose X-RateLimit-* headers to clients
  error_message: "Query cost budget exceeded. Reduce field count or request frequency."

---

# Separate tier for third-party portal consumers
name: graphql-rate-limiting-advanced
service: apollo-router-svc
consumer_group: third-party-portal
config:
  max_cost: 1000             # 1 000 points/min for external portal clients
  window_type: sliding
  strategy: redis

Security at the Federation Layer

The security surface of a federated GraphQL API is wider than REST because a single operation can reach multiple subgraphs. A JWT validated at the Router does not automatically constrain what each subgraph will return for that token’s claims. Security must be enforced at both the Router and the subgraph level, with the Router handling authentication and coarse-grained authorisation, and subgraphs handling data-level access control for their own types.

Key controls at each layer:

  • Router JWT validation. The Router validates the JWT signature against the IdP’s JWKS endpoint on every request. Claims (sub, roles) are extracted and injected as HTTP headers into every subgraph request (see router.yaml above). Subgraphs trust these headers — they do not re-validate the JWT; they rely on the Router’s assertion. This only works if subgraph endpoints are not publicly reachable; enforce this with Kubernetes NetworkPolicy restricting subgraph pod ingress to the Router’s service account only.
  • Schema-level authorisation with @requiresScopes and @authenticated. Federation 2.4+ supports these directives natively in the Router. Annotate sensitive fields with @requiresScopes(scopes: [["payments:write"]]) and the Router rejects any operation that selects that field if the JWT does not carry the required scope — before the query reaches a subgraph.
  • Subgraph-to-subgraph mTLS. The Router connects to subgraphs over HTTPS. On OpenShift with Istio, inject the Router pod into the mesh and configure PeerAuthentication to enforce STRICT mTLS on all subgraph namespaces. This prevents any pod other than the Router from reaching subgraph endpoints, regardless of NetworkPolicy gaps.
  • Persisted queries. In production, enable persisted_queries.safelist.require_id: true in router.yaml. Only operations pre-registered in the persisted query list (by ID) are executed. Arbitrary query strings from the wire are rejected. This eliminates query injection as an attack vector and makes schema introspection useless to an attacker who cannot submit arbitrary operations.
  • Operation allow-listing. In regulated contexts, persisted queries serve as an allow-list. Combine with Kong’s request-validator plugin to reject any request that does not carry a valid extensions.persistedQuery.sha256Hash header.
Never expose supergraph introspection in production without authentication

GraphQL introspection returns the complete supergraph schema — every type, every field, every argument, every directive, across all subgraphs. In a banking API this is a detailed map of your data model, your domain structure, and your regulatory data classifications. Set supergraph.introspection: false in router.yaml for all non-development environments. If developer tooling (Apollo Studio, Insomnia) requires introspection, expose a separate authenticated introspection endpoint behind a VPN or bastion that is not reachable from the internet or from production workloads.

Federation Observability

Federated GraphQL observability differs from microservice observability in two ways. First, a single client request generates multiple subgraph fetches — you need traces that span the Router and all subgraphs. Second, errors can be partial: a GraphQL response can include both data and errors, meaning a 200 HTTP status can still represent a degraded query. Your error-rate metrics must count response.errors, not HTTP 5xx, as the primary error signal.

Trace propagation: the Router injects W3C TraceContext headers (traceparent, tracestate) into every subgraph fetch. Subgraphs instrumented with OTel automatically pick up these headers and create child spans under the Router’s root span. The result is an end-to-end trace that shows plan time, fetch time per subgraph, and merge time in the Router as a unified waterfall view in Grafana Tempo.

Key metrics to track for federation health:

  • Operation cost distribution. The Kong rate-limiting plugin exposes per-operation cost as a log field. Route these to Loki and build a Grafana histogram of cost per consumer tier — a sudden shift in the cost distribution (operations becoming more expensive) predicts a rate-limit incident before it happens.
  • Subgraph fetch duration by subgraph name. The Router emits apollo_router_http_requests_duration_seconds labelled by subgraph. Track p99 per subgraph; a payments subgraph p99 rising above 2 s while accounts is normal identifies the degraded service without looking at application logs.
  • Field-level error rates. Parse the Router’s structured access log (JSON) for graphql.errors[].extensions.code grouped by path. This gives per-field error rates — the Grafana panel equivalent of per-endpoint HTTP error rates for REST APIs.
  • Slow operation detection. Set a Router span attribute filter: operations with total duration >5 s emit a slow_operation: true attribute. Alert when this attribute appears more than 5 times in a 5-minute window on a production consumer.

Response Caching

GraphQL caching is more nuanced than REST caching because a response may contain data from multiple subgraphs with different freshness requirements. The Apollo Federation approach is per-type cache hints via the @cacheControl directive, which annotates each type and field with a maximum age and scope.

The Router collects @cacheControl hints from all subgraph responses, computes the minimum maxAge across the merged result, and emits a Cache-Control response header. Upstream CDN (Kong proxy cache or Nginx) then caches the response for that duration.

  • Reference data (currency list, bank codes): maxAge: 3600, scope: PUBLIC. Safe to cache at the CDN layer.
  • Account balances: maxAge: 30, scope: PRIVATE. The 30-second TTL is the minimum that prevents hammering the core banking system; the PRIVATE scope prevents CDN caching and restricts to client-side or Kong per-consumer cache.
  • Transaction history: maxAge: 60, scope: PRIVATE. Slightly stale history is acceptable for the account statement view; real-time transaction status uses a separate query without caching.
  • Mutations: no cache headers. The Router does not cache mutation responses.

For stale-while-revalidate, configure Kong’s proxy-cache-advanced plugin with cache_control: true and storage_ttl set to maxAge + revalidate_window. The Kong cache serves the stale response immediately while it revalidates in the background, eliminating the latency spike at TTL expiry for high-traffic reference data queries.

Circular subgraph dependencies break query planning

If Subgraph A @requires a field from Subgraph B, and Subgraph B @requires a field from Subgraph A, the Router’s query planner cannot resolve the dependency order and fails to compose. This is not a runtime error — it is a composition error that prevents the supergraph from being built at all. Circular dependencies almost always indicate a domain boundary design problem: two subgraphs are co-owning a type that should belong to one of them exclusively. Resolve ownership before wiring the dependency, not after.

Common Pitfalls

Federation is more operationally complex than a monolithic GraphQL server. Most failures in production federated APIs trace back to one of three root causes: N+1 query patterns across subgraph boundaries, schema changes that break the supergraph without warning, or deployment sequencing errors between the supergraph and a subgraph.

  • N+1 across subgraph boundaries. A query that returns a list of 50 Transaction objects, each resolving the associated Account from the Accounts subgraph, will generate 50 separate fetch requests to the Accounts subgraph unless the DataLoader pattern is implemented. The Router batches entity fetches by default (the _entities query), but only if the subgraph resolver correctly handles a batch of keys and returns results in the same order. Failing to handle _entities batches efficiently is the single most common performance problem in federated deployments.
  • Schema breaking changes without coordination. Removing a field from a subgraph SDL is a breaking change if any registered persisted query references that field — even if no consumer appears to use it. The schema registry must run breaking-change checks on every subgraph push and block the push if any registered operations become invalid. Do not rely on manual coordination between subgraph teams.
  • Supergraph deployed before subgraph. If a new supergraph schema references a field that a subgraph has not yet deployed, the Router will attempt to fetch that field and receive a resolver-not-found error. The correct deployment sequence is always subgraph first, then supergraph registration. Automate this in the CI pipeline: publish the subgraph, wait for readiness probe, then trigger supergraph recomposition.
  • Mutation without idempotency key. Network retries in the Router (traffic_shaping.retries) apply to all requests including mutations unless explicitly excluded. Enabling retries on mutation subgraphs without idempotency key handling doubles the risk of duplicate payment initiations. Disable retries for the payments subgraph; implement idempotency at the application layer using a client-supplied x-idempotency-key header validated by the payments subgraph before execution.

Production Checklist

Steps to introduce federation incrementally

  1. Start with the monolith. If you have an existing monolithic GraphQL server, run it as a single subgraph behind the Apollo Router without splitting anything. This validates the Router deployment, the JWT plugin configuration, and the persisted query pipeline before you have multi-subgraph complexity.
  2. Extract the highest-ownership domain first. Identify the domain team with the clearest schema ownership (typically payments or accounts) and extract their types into a separate subgraph. Keep all other types in the monolith subgraph. Validate cross-domain references using @external and @requires directives.
  3. Wire the schema registry. Before extracting a third subgraph, configure automated schema registration in CI and enable breaking-change checks. This step must happen before the subgraph count makes manual coordination impossible.
  4. Enforce NetworkPolicy and mTLS. Once the Router is handling real traffic, lock down subgraph pod ingress to Router-only. Enable Istio STRICT mode on all subgraph namespaces. Verify with a penetration test from a pod outside the Router’s namespace before the next subgraph is added.
  5. Enable persisted queries. Register all existing operations in the persisted query list. Set require_id: false initially (allow both registered and ad-hoc queries) during a transition window. After all clients have migrated to sending operation IDs, flip require_id: true to enforce the allow-list.
  6. Implement DataLoader in every subgraph entity resolver. Before any list-returning query goes to production, verify that the subgraph’s _entities resolver batches key lookups. Load-test with 50 and 200 entity references; check subgraph fetch count in Tempo traces — it must be 1, not N.
  7. Establish schema governance. Document the breaking-change policy (deprecation window, consumer notification process), assign a schema owner per subgraph, and schedule a quarterly schema review to remove deprecated fields that have zero usage in the persisted query registry.

Approach comparison

Approach Ownership Schema coordination Breaking change impact Runtime complexity
Apollo Federation 2 Per-domain subgraph team Automated via schema registry & composition checks Isolated to subgraph boundary; registry blocks invalid pushes High — Router + per-subgraph services + registry
Schema stitching Gateway team owns all merge config Manual gateway config changes per schema update Gateway config must be updated in sync; no automated checks Medium — single gateway but complex merge rules
Monolithic GraphQL Single shared codebase Merge conflicts; all teams touch the same schema file Any field removal breaks all consumers immediately Low — one process, one schema, one deployment

Production checklist

  1. Apollo Router deployed as a 3-replica Deployment with PodDisruptionBudget max-unavailable 1.
  2. Supergraph schema compiled by Rover CLI in CI, mounted as ConfigMap — no runtime dependency on Apollo GraphOS cloud.
  3. supergraph.introspection: false and sandbox.enabled: false in router.yaml for all non-development environments.
  4. JWT plugin configured with JWKS URL pointing to internal IdP; authorization.require_authentication: true set.
  5. Correlation and channel headers propagated to all subgraphs; JWT sub and roles injected as subgraph request headers.
  6. Depth limit (10), alias limit (15), and root field limit (5) configured in limits block.
  7. Persisted queries enabled; require_id: true enforced in production.
  8. Subgraph pod ingress locked to Router service account via Kubernetes NetworkPolicy; Istio STRICT mTLS on all subgraph namespaces.
  9. Kong graphql-rate-limiting-advanced plugin configured with per-consumer-tier cost budgets backed by Redis Cluster.
  10. DataLoader pattern implemented and verified in every subgraph entity resolver; batch fetch count confirmed as 1 in Tempo traces.
  11. Breaking-change checks run on every subgraph SDL push; CI blocks registration if any persisted query becomes invalid.
  12. Deployment pipeline enforces subgraph-first sequencing: subgraph readiness confirmed before supergraph recomposition triggers.