Overview

The aggregation problem appears the moment a bank has more than one consumer channel and more than one domain service. The mobile app needs an account summary screen — three fields from the accounts API, one field from the payments API, and a notification badge from the notifications API. The desktop web app needs all fields from all three services plus transaction history from a fourth. The corporate portal needs a bulk view that calls the accounts API in a loop for 200 accounts. The open banking API needs a consent-filtered view with only the fields the customer has authorised the third party to see.

The naive solution is to let each channel call each service directly. The result is over-fetching on mobile (the accounts API returns 40 fields; mobile uses 3), under-fetching on desktop (desktop needs account and payment data in a single screen transition, which requires two sequential API calls with a dependency between them), and performance coupling between channels — the slowest downstream service determines the latency of every channel’s most critical screen.

The Backend-for-Frontend (BFF) pattern addresses this by placing a thin aggregation layer between each channel and the downstream domain services. The BFF is not a domain service — it has no business logic, no database, and no state. It is an aggregator and transformer that owns the shape of the API contract a specific channel needs. Each channel gets its own BFF, owned by the channel team, not the domain service teams.

BFF is a channel concern, not a domain concern

The BFF for the mobile channel should be owned and deployed by the mobile channel team. It knows what the mobile app needs. Domain logic — rules about when a payment can be initiated, what constitutes a valid IBAN, how account limits work — belongs in the domain services. The moment a BFF starts enforcing domain rules, it has become a second domain layer, which multiplies the places where domain logic must be changed when business rules evolve.

BFF Architecture

The BFF architecture at a KSA bank with three primary channels — mobile, web, and open banking — produces three BFF services. Each is independently deployable, independently scalable, and independently versioned. The domain services are unaware of the BFF layer; they expose their APIs to any authenticated caller and the BFF is just another caller.

The BFF owns: aggregation (calling multiple downstream services and assembling a single response), transformation (renaming fields, converting types, filtering optional fields not needed by the channel), partial result handling (deciding what to return when a downstream is degraded), and caching (storing reference data that the channel reads frequently).

The BFF does not own: authentication of end users (that stays at the API gateway, Kong 3.7 in this stack), rate limiting (Kong plugin), TLS termination (OpenShift route or Kong), or domain validation (the domain service). Kong terminates the inbound TLS and enforces the OAuth 2.0 / FAPI 2.0 token validation before the request reaches the BFF. The BFF receives an already-authenticated request carrying the token claims and acts on them.

ACE Aggregator Flow

IBM ACE 12.0.10’s Fan Out and Fan In message flow nodes are the native mechanism for parallel aggregation in an ACE integration. The Fan Out node clones the inbound message, routes each copy to a separate parallel branch, and the Fan In node reassembles the results when all branches have completed or the timeout has expired.

WebBffAggregator_OverrideBuiltin.esql — ACE Fan-Out/Fan-In configurationesql
-- ACE flow: WebBffAggregator.msgflow
-- Nodes: MQInput → FanOut → [AccountsBranch, PaymentsBranch, NotificationsBranch] → FanIn → Aggregator → HTTPReply

-- FanOut node configuration (set in message flow properties)
-- Distribution mode: Fan out to all terminals
-- Timeout: 800ms (total SLA budget: 1000ms minus 200ms for aggregation and reply serialisation)

CREATE COMPUTE MODULE AggregateResults
  CREATE FUNCTION Main() RETURNS BOOLEAN
  BEGIN
    DECLARE accountsResult   REFERENCE TO InputRoot.XMLNSC.AggregateReply.Reply[1];
    DECLARE paymentsResult   REFERENCE TO InputRoot.XMLNSC.AggregateReply.Reply[2];
    DECLARE notifyResult     REFERENCE TO InputRoot.XMLNSC.AggregateReply.Reply[3];

    -- Check for partial results (Fan In with timeout: some branches may have timed out)
    DECLARE accountsOk  BOOLEAN (FIELDNAME(accountsResult) IS NOT NULL);
    DECLARE paymentsOk  BOOLEAN (FIELDNAME(paymentsResult) IS NOT NULL);
    DECLARE notifyOk    BOOLEAN (FIELDNAME(notifyResult) IS NOT NULL);

    -- Accounts is mandatory: fail fast if absent
    IF NOT accountsOk THEN
      SET OutputRoot.XMLNSC.BffError.code    = 'UPSTREAM_TIMEOUT';
      SET OutputRoot.XMLNSC.BffError.service = 'accounts-api';
      PROPAGATE TO TERMINAL 'out_error';
      RETURN FALSE;
    END IF;

    -- Assemble response; use degraded values for optional services
    SET OutputRoot.XMLNSC.WebDashboard.accountSummary  = accountsResult.AccountSummary;
    SET OutputRoot.XMLNSC.WebDashboard.paymentStatus   =
        CASE paymentsOk WHEN TRUE
          THEN paymentsResult.LatestPaymentStatus
          ELSE 'UNAVAILABLE'   -- documented in API contract; consumer renders grey badge
        END;
    SET OutputRoot.XMLNSC.WebDashboard.notificationCount =
        CASE notifyOk WHEN TRUE
          THEN notifyResult.UnreadCount
          ELSE NULL   -- null badge: consumer hides notification icon when null
        END;
    SET OutputRoot.XMLNSC.WebDashboard.partial =
        NOT (paymentsOk AND notifyOk);

    RETURN TRUE;
  END;
END MODULE;

Spring Boot Composition Service

For the Mobile BFF, Spring Boot 3.3 with CompletableFuture.allOf provides parallel composition without the operational overhead of an ACE integration server. Resilience4j circuit breakers per downstream service prevent a single downstream failure from cascading to all channels.

MobileBffCompositionService.java — CompletableFuture + Resilience4jjava
package info.saib.bff.mobile.service;

import io.github.resilience4j.circuitbreaker.annotation.CircuitBreaker;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;

@Service
public class MobileBffCompositionService {

    private final AccountsApiClient      accountsClient;
    private final PaymentsApiClient      paymentsClient;
    private final NotificationsApiClient notifyClient;
    private final MeterRegistry         meterRegistry;

    // Total SLA budget: 800ms (Kong adds ~50ms; serialisation ~50ms; leaves 700ms for calls)
    private static final long DEADLINE_MS = 700L;

    public MobileDashboardDto compose(String customerId, String correlationId) {
        Sample timer = Timer.start(meterRegistry);

        // Parallel calls with per-downstream circuit breakers
        CompletableFuture<AccountSummaryDto> accountsFuture =
            CompletableFuture.supplyAsync(() -> fetchAccounts(customerId, correlationId))
                .orTimeout(DEADLINE_MS, TimeUnit.MILLISECONDS);

        CompletableFuture<PaymentStatusDto> paymentsFuture =
            CompletableFuture.supplyAsync(() -> fetchPaymentStatus(customerId, correlationId))
                .orTimeout(DEADLINE_MS, TimeUnit.MILLISECONDS)
                .exceptionally(ex -> PaymentStatusDto.degraded());  // optional field: degrade gracefully

        CompletableFuture<Integer> notifyFuture =
            CompletableFuture.supplyAsync(() -> fetchUnreadCount(customerId, correlationId))
                .orTimeout(DEADLINE_MS, TimeUnit.MILLISECONDS)
                .exceptionally(ex -> null);  // badge absent when notifications unavailable

        // Wait for all; accountsFuture failure propagates (mandatory field)
        CompletableFuture.allOf(accountsFuture, paymentsFuture, notifyFuture)
            .orTimeout(DEADLINE_MS + 50L, TimeUnit.MILLISECONDS)
            .join();

        timer.stop(Timer.builder("bff.mobile.compose.duration")
            .tag("partial", String.valueOf(paymentsFuture.isCompletedExceptionally() || notifyFuture.join() == null))
            .register(meterRegistry));

        return MobileDashboardDto.builder()
            .account(accountsFuture.join())              // throws if accounts failed (mandatory)
            .paymentStatus(paymentsFuture.join())        // returns degraded DTO if payments failed
            .notificationBadge(notifyFuture.join())      // returns null if notifications failed
            .partial(!paymentsFuture.join().isHealthy() || notifyFuture.join() == null)
            .build();
    }

    @CircuitBreaker(name = "accounts-api", fallbackMethod = "accountsFallback")
    private AccountSummaryDto fetchAccounts(String customerId, String correlationId) {
        return accountsClient.getSummary(customerId, correlationId);
    }

    @CircuitBreaker(name = "payments-api", fallbackMethod = "paymentsFallback")
    private PaymentStatusDto fetchPaymentStatus(String customerId, String correlationId) {
        return paymentsClient.getLatestStatus(customerId, correlationId);
    }

    private AccountSummaryDto accountsFallback(String cid, String cid2, Exception ex) {
        throw new BffMandatoryServiceException("accounts-api", ex);  // bubble up: mandatory
    }
    private PaymentStatusDto paymentsFallback(String cid, String cid2, Exception ex) {
        return PaymentStatusDto.degraded();  // silent degrade: optional
    }
}

Open Banking BFF

The Open Banking BFF has compliance requirements that the mobile and web BFFs do not. SAMA’s Open Banking framework mandates FAPI 2.0 security profiles for third-party provider (TPP) access. The BFF layer is the correct place to enforce consent scope filtering, not the downstream domain services.

When a TPP presents an AIS (Account Information Service) token, the Open Banking BFF must return only the fields the customer has consented to share, regardless of what the Accounts API would return for a full authenticated session. The Accounts API has no concept of per-TPP consent — it returns all data for the authenticated customer. The BFF reads the consent record from the consent store and applies field-level filtering before the response leaves the BFF.

openapi-mobile-bff.yaml — BFF endpoint with composed response schemayaml
openapi: "3.1.0"
info:
  title: Mobile BFF API
  version: "1.0.0"
  description: Mobile-channel aggregation layer. All fields from multiple domain services composed into one response.

paths:
  /v1/dashboard:
    get:
      summary: Mobile dashboard composition
      description: Returns composed dashboard data from Accounts, Payments, and Notifications APIs. When partial=true, one or more optional fields are absent due to downstream degradation.
      security:
        - bearerAuth: []
      parameters:
        - name: X-Correlation-ID
          in: header
          required: true
          schema: { type: string, format: uuid }
        - name: X-Request-Deadline
          in: header
          description: Absolute deadline as Unix epoch milliseconds; BFF respects this as an external constraint on internal timeout budget
          schema: { type: integer, format: int64 }
      responses:
        "200":
          description: Full or partial dashboard response
          content:
            application/json:
              schema:
                $ref: "#/components/schemas/MobileDashboard"
        "206":
          description: Partial Content — mandatory fields present, one or more optional fields absent due to downstream degradation. Consumer must check partial=true and render gracefully.
          content:
            application/json:
              schema:
                $ref: "#/components/schemas/MobileDashboard"
        "503":
          description: Mandatory downstream (Accounts API) unavailable; BFF cannot return a usable response

components:
  schemas:
    MobileDashboard:
      type: object
      required: [account, partial]
      properties:
        account:
          $ref: "#/components/schemas/AccountSummary"   # mandatory
        paymentStatus:
          $ref: "#/components/schemas/PaymentStatus"     # optional; null when unavailable
        notificationBadge:
          type: integer
          nullable: true                                # null: hide badge icon
        partial:
          type: boolean
          description: true when one or more optional downstream services were degraded
Partial results must be documented in the API contract, not discovered in production

If the BFF returns a response with some fields absent due to downstream degradation and the API contract does not explicitly describe this behaviour — which fields are mandatory, which are optional, what value is returned when optional fields are absent — mobile or web clients will interpret the absent field as a bug and file an incident. Document partial result behaviour in the OpenAPI spec (nullable: true on optional fields, HTTP 206 for degraded responses, an explicit partial: boolean flag in the response body) and validate it in consumer-driven contract tests before the BFF reaches production.

Timeout Budget Management

The timeout budget is the total time available from the moment a request enters the BFF to the moment the BFF must begin sending its response. Every hop inside the BFF consumes from this budget. The budget is not negotiable — it is determined by the channel SLA, typically 1,000ms for mobile and 2,000ms for web. Allocating it carefully is the difference between a fast BFF and a fast-failing BFF.

For a 1,000ms mobile SLA: subtract Kong overhead (~50ms), BFF framework overhead (~30ms), response serialisation (~30ms). That leaves ~890ms for upstream calls. If downstream services are called in parallel, the wall-clock time consumed is the maximum of all call durations, not the sum. Budget 700ms for upstream calls, which gives 190ms of margin for p99 latency variance.

Propagate the deadline to downstream services via a request header: X-Request-Deadline: <unix epoch ms>. A downstream service that receives an X-Request-Deadline header can short-circuit expensive operations if the deadline has already passed, avoiding wasted work. This is the gRPC context deadline model applied to HTTP. Not all downstream services honour this header initially — add it now, enforce it as downstream services are upgraded.

For Kong 3.7, set the proxy connect timeout and send timeout to match the BFF’s external SLA, not the internal timeout budget. Kong’s read timeout for the BFF upstream should be 2,000ms (web SLA + buffer); the BFF’s own internal calls use 700ms. This ensures Kong doesn’t time out the BFF before the BFF has had a chance to return a partial result.

Error Handling Strategies

Two error handling philosophies apply at the BFF level, and the choice between them must be made per field in the composed response, not globally for the BFF.

Fail-fast. When a mandatory downstream service (Accounts API for a dashboard screen) returns 503 or times out, the BFF immediately returns 503 to the channel. There is no degraded view that is useful to the mobile user without account data. The circuit breaker for the Accounts API trips after 50% failure rate in a 10-second window; while open, the BFF returns 503 without attempting the call, reducing latency of the failure response from 700ms (timeout) to under 10ms.

Partial results. When an optional downstream service (Notifications API for a badge count) returns 503, the BFF returns 200 or 206 with partial: true and a null notificationBadge. The mobile client is responsible for rendering the badge-absent state gracefully — typically by hiding the notification icon. This behaviour must be tested in the mobile client’s contract test suite, not assumed.

BFF without circuit breakers turns a single downstream failure into a full-channel outage

Without circuit breakers, a Notifications API that is responding slowly (800ms instead of 50ms) causes every BFF request to block for 800ms before returning, even though the Notifications data is optional. Under load, all BFF threads are consumed waiting for a slow optional service; the mobile channel becomes unresponsive. Configure Resilience4j circuit breakers for every downstream service in the BFF, regardless of whether the service is mandatory or optional. A circuit breaker in open state returns immediately with a fallback value; a slow service without a circuit breaker holds every thread until the request timeout.

BFF-Layer Caching

Reference data that the BFF reads on every request and that changes infrequently is a strong caching candidate at the BFF layer. Account metadata — IBAN, account name, product type — is stable for hours between changes; caching it at the BFF eliminates a downstream call on every dashboard load.

The caching strategy for the Mobile BFF: Caffeine in-process cache for per-customer reference data with a 5-minute TTL and a stale-while-revalidate window. For the Web BFF running in IBM ACE, use ACE’s built-in GlobalCache node backed by an embedded WebSphere eXtreme Scale partition; this cache is shared across all ACE instances in the integration server group.

Cache invalidation via event: when the Accounts API publishes an account.updated Kafka event, the BFF consumes it and evicts the cached entry for the affected customerId. This keeps the cache coherent without a short TTL that would increase downstream load. The event consumer in the BFF is a lightweight Spring Kafka consumer; it does not participate in the request path.

Per-channel cache TTL: the mobile BFF caches reference data for 5 minutes; the open banking BFF caches reference data for 1 minute (TPP-visible data has a tighter freshness requirement under SAMA’s Open Banking framework). Never cache consent records at the BFF — consent changes must be reflected immediately, and a 5-minute-stale consent cache would return data a customer has revoked.

Pitfalls

BFF becoming a second domain layer through business logic creep

The most common way a BFF degrades in production: a business rule is added to it because it is convenient (“the BFF already has the account data, so let’s add the eligibility check here”). Over 18 months, the BFF accumulates eligibility rules, limit calculations, and product-type logic that should live in domain services. The BFF is now a domain service that also aggregates. When the domain rule changes, it must be changed in the BFF and in the domain service. Enforce a rule: the BFF transforms and selects data; it never derives data. Any computation that could be expressed as a domain service method belongs in a domain service.

  • N+1 calls from naive composition. A BFF that calls the Accounts API once per account in a list — instead of calling a bulk-fetch endpoint — produces N+1 round trips per request. For a customer with 10 accounts, this is 10 serial calls taking 10× the latency of a single call. Audit every BFF call pattern during design; if the domain service does not have a bulk-fetch endpoint, build one rather than doing N+1 at the BFF layer.
  • BFF not propagating correlation IDs. Every request that enters the BFF carries a X-Correlation-ID or traceparent header from Kong (or from the client, validated by Kong). The BFF must propagate this header to every downstream call. Without correlation ID propagation, distributed traces break at the BFF boundary and the downstream service logs cannot be correlated to the originating channel request — which makes diagnosing BFF-originating issues in downstream services impossible.
  • Single BFF for multiple channels. A team that creates one BFF for both mobile and web because “they need similar data” ends up with a service that neither channel fully owns. When the mobile team needs to add a field, they need coordination with the web team. When the web BFF changes its error contract, the mobile client breaks. The operational overhead of two BFFs is modest; the coupling overhead of one shared BFF is not.
  1. Identify the channel and its data requirements

    For each channel that will have a BFF, document every field the channel needs on its most critical screens, which domain service owns each field, and whether each field is mandatory or optional for the channel to be usable.

  2. Define the BFF API contract first

    Write the OpenAPI specification for the BFF endpoint before writing any code. Mark mandatory fields, document optional fields and their degraded values, and specify the partial flag behaviour. Share the spec with the channel team for validation before implementation.

  3. Classify each downstream service as mandatory or optional per screen

    Apply fail-fast to mandatory services, partial-result handling to optional ones. Wire Resilience4j circuit breakers for all downstream services with different thresholds for mandatory (lower failure rate tolerance) and optional (higher failure rate tolerance).

  4. Deploy the BFF and its circuit breakers to a lower environment before adding caching

    Caching makes circuit breaker behaviour harder to test. Validate that the fail-fast and partial-result paths work correctly without caching before adding the cache layer.

  5. Add BFF-layer caching with cache invalidation events

    Cache reference data with per-channel TTLs; wire Kafka event consumers to invalidate cache entries when upstream data changes. Never cache consent data or anything the customer has a right to revoke immediately.

  6. Write consumer-driven contract tests between the channel and the BFF

    Pact contract tests ensure that when the BFF’s response schema changes, the channel’s tests fail before the BFF is deployed. Without contract tests, schema changes at the BFF are discovered in production.

  7. Establish a governance process for domain logic in the BFF

    Define and enforce a rule distinguishing BFF logic (field selection, field renaming, partial result assembly) from domain logic (business rules, calculations, eligibility). Review BFF pull requests for domain logic creep during the code review phase.

Production Checklist

  1. One BFF per channel (mobile, web, open banking); each owned and deployed by the channel team.
  2. OpenAPI spec published before implementation; mandatory vs optional fields documented; partial: boolean flag in response body.
  3. Resilience4j circuit breakers configured for every downstream service: mandatory services fail-fast, optional services degrade gracefully.
  4. Timeout budget calculated and allocated: channel SLA minus Kong overhead minus framework overhead leaves headroom for upstream calls.
  5. X-Request-Deadline header propagated to all downstream calls with the absolute epoch timestamp.
  6. Correlation ID (X-Correlation-ID and traceparent) propagated to all downstream calls from every BFF request.
  7. BFF-layer reference data cache configured with per-channel TTLs; cache invalidation wired to upstream change events via Kafka consumer.
  8. Consent records never cached; Open Banking BFF applies consent scope filtering on every request using a freshly fetched consent record.
  9. FAPI 2.0 token validation at Kong; BFF reads token claims and applies scope filtering (AIS vs PIS); SAMA rate plans enforced at Kong plugin, not in BFF code.
  10. N+1 call patterns audited and eliminated; bulk-fetch endpoints requested from domain service teams before BFF goes to production.
  11. Consumer-driven contract tests (Pact) between each channel and its BFF; CI pipeline breaks on contract violation.
  12. No domain logic (eligibility rules, limit calculations, business validations) in any BFF; code review checklist includes explicit domain-logic check.