RFC: Order Gateway Latency on the Admin Console

Date: 2026-09-17

Status: Draft

Scope: surfacing the order-gateway latency series as min / p50 / p99 timeseries on the admin console homepage. Builds on the as-built record Order gateway latency statistics and is the admin-console counterpart to the Grafana deliverable in Performance Tracking & Regression Protection. EP3 edition only; btnl-order-gateway follows once it emits equivalent series.

1. Summary

The order gateway now records six latency histograms and exports them to the observability ClickHouse. The admin homepage shows none of them: every number on it comes from the exchange ClickHouse, Postgres, or Redis, and api-gateway has no connection to the observability cluster at all.

Wiring a chart to the data as it stands would produce a misleading chart. Two problems at the source have to be fixed first:

  1. The buckets are too coarse to separate p50 from p99. On ax-demo today, 100% of og_ws_order_to_ack_ms samples land in two buckets (5–10 ms and 10–25 ms). Every minute of the last hour has p50 and p99 in the same bucket, so both lines would be interpolations inside 10–25 ms and move only with the bucket's fill ratio.
  2. There is no min. A Prometheus histogram carries no minimum; the Min/Max columns of otel_metrics_histogram are always 0.

The plan has three parts:

part change
Source (order-gateway) Replace the three hand-picked bucket sets with one log-spaced set (~9% worst-case quantile error). Record each series a second time as a rolling summary, used only for its exact min and max.
Read path (api-gateway) Optional read-only HTTPS connection to the observability ClickHouse. One new admin endpoint, GET /admin/latency-stats, serving a fixed catalog of series. No metric names or SQL cross the API.
GUI (admin) A "Order gateway latency" section on the homepage: four tiles (p50 headline, p99 and min beside it) and a min/p50/p99 line chart, driven by the existing timeframe tabs.

Nothing is written to the exchange database, and the gateway's hot path gains one lock-free record call per sample. The as-built's three rules still hold.

2. Background

2.1 What the homepage does today

gui/apps/admin/src/components/Dashboard.tsx makes one call, useExchangeStats(timeframe, { groupBy: 'day' }) → GET /api/admin/exchange-stats, and renders MetricCard tiles plus two recharts bar charts (TradingVolumeChart, DailyFillsChart). Timeframe tabs are 1h / 24h / 30d / YTD. There is no polling; the query runs once per mount or tab change.

The handler (rs/api-gateway/src/admin_routes.rs, get_exchange_stats) reads the exchange ClickHouse (AppState.klickhouse), Postgres, and Redis tickers. Request and response types live in rs/sdk-internal/src/protocol.rs. Auth is by router: anything registered in admin_routes::read_router gets authorize_admin.

api-gateway has exactly one ClickHouse pool, configured by CLICKHOUSE_*. No Rust service reads the observability cluster. The only consumers of otel_metrics_* are Grafana (configs/grafana/) and the alerter, and neither queries otel_metrics_histogram. The admin /monitor page is three external links.

2.2 How the latency data arrives

order-gateway  --/metrics on REST port 4000-->  OTel collector (15 s scrape)
               --HTTPS 8443, PrivateLink-->     ClickHouse Cloud, default.otel_metrics_histogram

Facts that shape the design, checked against the live ax-demo cluster:

2.3 The resolution problem, measured

Per-minute bucket deltas for og_ws_order_to_ack_ms{req="place",result="acked"} on ax-demo, 2026-09-17 17:13–17:18 UTC:

minute samples ≤ 10 ms ≤ 25 ms p50 bucket p99 bucket
17:13 321 60 321 10–25 10–25
17:14 239 60 239 10–25 10–25
17:15 264 58 264 10–25 10–25
17:16 312 70 309 10–25 10–25
17:17 262 74 262 10–25 10–25
17:18 251 36 251 10–25 10–25

The as-built states the limit plainly: a quantile's error bound is the width of its bucket. A 15 ms-wide bucket is fine for "are we in the right decade" and useless for a chart whose purpose is to show p99 pulling away from p50.

3. Goals

4. Non-goals

5. Design

5.1 Source: bucket resolution

Replace GATEWAY_INTERNAL_BUCKETS_MS, ORDER_PATH_BUCKETS_MS, and DELIVERY_BUCKETS_MS in rs/order-gateway/src/in_situ_latency_measurement.rs with one generated set: four buckets per octave (ratio 2^¼ ≈ 1.19) from 0.05 ms to 30 s, about 78 bounds.

A bound change is safe for readers. It ships in a release, the restart changes StartTimeUnix, and the query never differences rows across a StartTimeUnix boundary. The query groups by ExplicitBounds, so a range that spans the change comes back as two histograms. The endpoint does not rebucket one onto the other. For the one step that contains both, and for the range headline, it uses the newer set only.

Alternative considered: OTLP exponential histograms. Better resolution per byte, and the collector already has the table. The metrics Prometheus exporter cannot emit them, so this means replacing the exporter for one service. Not worth it for this feature.

5.2 Source: min and max

Record each of the six series a second time under <name>_window with no configured buckets. The exporter renders it as a rolling summary (60 s window), which lands in otel_metrics_summary with quantile 0 (min) and quantile 1 (max).

Only those two values are used. Min-of-mins and max-of-maxes are exact under any re-aggregation (over time, labels, or replicas), which is the property that the summary's p50/p99 lack and the reason those still come from the histogram.

Alternative considered: read min from the lowest non-empty bucket. No second record, but the value is a bucket bound, so the min line is a staircase with ~19% steps and is always an underestimate. Rejected because min is the one statistic where operators read the absolute value ("what is our floor").

Alternative considered: an atomic min per series, swapped to +∞ when /metrics renders, exported as a gauge. Exact per scrape window, but any second reader of /metrics (a curl during an incident) silently steals a window. Rejected.

5.3 Read path: api-gateway → observability ClickHouse

Optional config read from the O11Y_CLICKHOUSE_* variables, which are already provisioned on ax-prod. The config struct's fields follow the existing variable names rather than introducing new ones; it needs the HTTPS endpoint (port 8443, the PrivateLink host the collector also uses; api_gateway runs on the same host with host networking), a username, and a password. The credentials must be read-only. If the provisioned user can write, swap it for one granted SELECT on default.otel_metrics_histogram and default.otel_metrics_summary before the endpoint ships.

The client is reqwest against the ClickHouse HTTP interface with FORMAT JSONEachRow, parameters bound with param_* query arguments, readonly=1, and max_execution_time=5. klickhouse is not reused: the exchange pool speaks the native protocol without TLS, and a handful of read-only aggregate queries do not justify a second native pool.

When the config is absent, AppState.o11y_clickhouse is None. This keeps local dev, CI, and the AIEX edition unchanged.

Failure isolation. The observability cluster is a separate vendor-hosted system with weaker availability expectations than the exchange. The new endpoint is the only code that touches it, it is a separate request from exchange-stats, and it has a 5 s timeout. A failure degrades one homepage section.

Alternative considered: have the gateway write per-minute rollups to the exchange ClickHouse, so the homepage reads them through the existing pool. Rejected. It contradicts the as-built's rule that these numbers use the observability pipeline, it needs a second in-process recorder because the Prometheus handle only renders text, and it adds an insert path to the most latency-sensitive service to feed a dashboard.

Alternative considered: embed Grafana panels. Grafana is Tailscale-only and anonymous; the admin console is Clerk-authenticated and IP-allowlisted. Bridging the two is more work than one endpoint, and the result would not look or behave like the rest of the homepage.

5.4 Endpoint

GET /admin/latency-stats, registered in admin_routes::read_router. Types in rs/sdk-internal/src/protocol.rs:

pub struct GetLatencyStatsRequest {
    pub start_timestamp_ns: u64,
    pub end_timestamp_ns: Option<u64>,
}

pub struct GetLatencyStatsResponse {
    pub step_secs: u32,
    pub retention_clamped: bool,
    pub series: Vec<LatencySeries>,
}

pub struct LatencySeries {
    pub key: LatencySeriesKey,
    pub window: LatencyPoint,
    pub points: Vec<LatencyPoint>,
}

pub struct LatencyPoint {
    pub timestamp: DateTime<Utc>,
    pub count: u64,
    pub min_ms: Option<f64>,
    pub p50_ms: Option<f64>,
    pub p99_ms: Option<f64>,
    pub max_ms: Option<f64>,
}

The server owns the catalog. LatencySeriesKey is an enum; each variant maps to a metric name and a label filter:

key series labels reads as
ws_place_to_ack og_ws_order_to_ack_ms req=place, result=acked WebSocket place, frame read to ack written
ws_place_to_ep3 og_ws_order_to_ep3_ms req=place gateway work before the EP3 RPC
ws_cancel_to_out og_ws_cancel_to_out_ms none WebSocket cancel, frame read to OrderCanceled written
dropcopy_to_ws og_dropcopy_to_ws_ms all event values merged drop-copy arrival to socket write

ws_place_to_ack minus og_ws_order_to_ep3_ack_ms (gateway delivery) and the REST series are left to Grafana for now. Adding a key is a one-line catalog change plus a GUI title.

Behaviour:

5.5 Query

One query per call for the histogram side, one for the summary side. The histogram side, for one catalog entry:

WITH per_step AS (
    SELECT toStartOfInterval(TimeUnix, toIntervalSecond({step:UInt32})) AS t,
           Attributes, StartTimeUnix, ExplicitBounds,
           argMax(BucketCounts, TimeUnix) AS bc
    FROM otel_metrics_histogram
    WHERE ServiceName = 'order_gateway'
      AND MetricName = {metric:String}
      AND TimeUnix >= {start:DateTime64(9)} - toIntervalSecond({step:UInt32})
      AND TimeUnix <  {end:DateTime64(9)}
      /* label filter from the catalog */
    GROUP BY t, Attributes, StartTimeUnix, ExplicitBounds
),
deltas AS (
    SELECT t, ExplicitBounds,
           arrayMap((a, b) -> toInt64(a) - toInt64(b), bc,
                    lagInFrame(bc, 1, arrayWithConstant(length(bc), toUInt64(0)))
                        OVER w) AS d,
           row_number() OVER w AS rn,
           StartTimeUnix >= {start:DateTime64(9)} AS born_in_range
    FROM per_step
    WINDOW w AS (PARTITION BY Attributes, StartTimeUnix, ExplicitBounds ORDER BY t)
)
SELECT t, ExplicitBounds, sumForEach(d) AS buckets
FROM deltas
WHERE rn > 1 OR born_in_range
GROUP BY t, ExplicitBounds
ORDER BY t

Quantiles are computed in Rust from buckets and ExplicitBounds: linear interpolation inside the bucket, lower edge 0 for the first bucket, the last finite bound for the overflow bucket. This is ~30 lines, is where the "newer bounds win" rule lives, and is unit-tested against known distributions. quantilePrometheusHistogram is available on 26.4 but expects running-total buckets and would still need the delta step; doing the last step in Rust keeps one tested implementation for both points and window.

The summary side takes, per step, min of quantile 0 and max of quantile 1 from otel_metrics_summary for <name>_window, ignoring rows whose window was empty.

The histogram query was run by hand against ax-demo for 15-minute ranges and returns the expected per-minute deltas; it has not been timed. A 30-day range reads about 170k rows per label set. If that proves slow, see §7.

5.6 GUI

6. Rollout

phase contents ships in
0 Deploy 16.1.x to ax-prod so the six series exist there at all. No code. ops
1 Log-spaced buckets; _window summary twins; probe-cost bench. Update the as-built record's bucket table. Update the scrape integration test to assert the twins render as summaries. order-gateway
2 O11Y_CLICKHOUSE_* config, HTTP client, catalog, GET /admin/latency-stats, quantile function with unit tests, integration test against a ClickHouse testcontainer seeded with cumulative rows including a restart and a bounds change. api-gateway, sdk-internal
3 Confirm the O11Y_CLICKHOUSE_* credentials on ax-prod are read-only; provision the same variables on ax-demo. deploy repos, Doppler
4 GUI section. admin
5 Add the same min/p50/p99 panels to configs/grafana/grafana/dashboards/order-gateway.json using the §5.5 SQL, which closes the perf RFC's M0 dashboard item against the real series names. configs/grafana

Phases 1 and 2 are independent. Phase 4 can be built against demo as soon as 1–3 are there.

Disconnect and restart paths to test before prod: gateway restart mid-range (new StartTimeUnix), gateway down for a whole step (gap, not zero), collector down (gap), observability cluster unreachable (502, homepage otherwise intact), config absent (503, section hidden).

7. Later: a rollup table

Deferred until either the 30-day query is too slow or someone needs more than 30 days.

A refreshable materialized view in the observability cluster computes the §5.5 deltas once a minute into og_latency_1m (t, metric, attributes, bounds, buckets, min_ms, max_ms) with a 13-month TTL. The endpoint switches its FROM; nothing else changes, and YTD becomes real. The observability schema is created by the collector and is not under Atlas, so this needs a home for its DDL first (see open questions).

8. Decisions

# decision choice
D1 Where p50/p99 come from The histogram, with log-spaced buckets. Mergeable across time, labels, and replicas.
D2 Where min/max come from A _window summary twin, quantiles 0 and 1 only. Exact and mergeable.
D3 How the admin backend reaches the data api-gateway reads the observability ClickHouse over HTTPS, optional and failure-isolated. The gateway writes nothing new.
D4 API surface Fixed server-side catalog. No metric names, labels, or SQL from the client.
D5 Where quantiles are computed Rust, from delta buckets, one implementation for points and the window headline.

9. Open questions

  1. Prod dependency. Is a read-only dependency from api-gateway to the vendor-hosted logs cluster acceptable on ax-prod, given the isolation in §5.3? The alternative is the rejected gateway-writes-rollups design.
  2. Bucket density. Four per octave (±9%, 78 bounds) or eight (±4.4%, 155 bounds)? The exposition and storage cost doubles; nothing else changes.
  3. Headline statistic. p50 as the tile's big number with p99 beneath, or p99 as the headline because that is what the market-maker conversation is about?
  4. DDL home for the observability cluster. Needed for the §7 rollup, and for any read-only grant phase 3 turns out to need. Proposal: db/observability/*.sql, applied by hand, outside Atlas.
  5. REST on the homepage. og_rest_request_to_response_ms is left out of the catalog because route × status makes "the" REST number ambiguous. Add rest_place_order (route /place-order, status 200) if REST order flow is material.
  6. AIEX. btnl-order-gateway has no equivalent series. Same names with a btnl_og_ prefix (as the perf RFC suggests) would make the catalog edition-keyed rather than duplicated.