Decision records

ADR-0034: PressureScore v1 and bottleneck classification

Status: accepted (design, D0) — 2026-10-07

Context

Adding scalwsd instances helps only when the web tier itself is the bottleneck. Measured on the development host: behind scalws, PHP-FPM, Node.js and uvicorn saturate their own CPUs long before scalws does (Node 100 % of one CPU at 22–30k req/s while scalws used ~1.4 of 2 CPUs). Cloning scalws there costs CPU and RAM and changes nothing.

Decision

Signals per instance (sampled every second by the controller from the instance’s metrics endpoint and /proc): CPU seconds (process), event-loop lag (time from a timer’s deadline to its run, per loop, p99 over the window), in-flight requests and open connections, accept rate, request rate, latency p50/p95/p99 (from the existing histograms), 5xx and timeouts, RSS, FastCGI queue wait (php.queue_wait), upstream head time, CPU of each managed application process group.

Components, each normalised to 0–1:

ComponentDefinition
cpuinstance CPU / CPUs assigned to it
lagclamp(loop-lag p99 / lag_saturated (default 5 ms))
inflightclamp(in-flight / (loops × inflight_per_loop, default 256))
latencyclamp(p99 / p99_target) when a target is configured, else 0

PressureScore = max(cpu_lag, inflight, latency) with cpu_lag = min(cpu, lag + 0.5): high CPU alone is not pressure (a busy-but-keeping-up instance has little loop lag); high CPU with rising lag is. All constants are configuration with these defaults, and every component is exported, not only the result.

Classification (evaluated over the scale-out window, first match wins):

ClassEvidence
MemoryPressurehost free RAM < min_free_ram or instance RSS > budget
PhpSaturated / NodeSaturated / PythonSaturatedthat application’s CPU ≥ 90 % of its allotment, or its queue wait rising, while web-tier cpu_lag < 0.75
WebTierSaturatedweb-tier PressureScore > scale_out_above (0.75) with cpu_lag the dominant component and no application class matching
ConnectionPressurein-flight/connections near limits while CPU is low (slow clients, long-lived streams)
Unknown/InsufficientEvidenceanything else, or fewer samples than the window needs

Actions. Only WebTierSaturated may act (D4), and only within min/max/resource budgets and cooldowns (ADR-0035). Every other class yields a diagnosis and a recommendation. Every evaluation logs its evidence, e.g. pressure=0.84 cpu=0.97 lag=0.62 class=WebTierSaturated action=ScaleOut from=2 to=3, including no-action reasons (class=NodeSaturated action=None reason="Node CPU 1.00, web tier 0.41", action=None reason="host CPU budget exhausted").

Scale-in when PressureScore < scale_in_below (0.30) for scale_in_window (60 s) and the remaining instances would stay below scale_out_above at the current load (projected by CPU share), never below min_instances.

Consequences

  • Event-loop lag is a new data-plane metric (a periodic timer per loop; cost measured in D2 against the regression gates).
  • No learning component: decisions are explainable from logged numbers; thresholds are tuned from D3 observe-only runs on synthetic and real workloads.