Decision records
ADR-0034: PressureScore v1 and bottleneck classification
Status: accepted (design, D0) — 2026-10-07
Context
Adding scalwsd instances helps only when the web tier itself is the bottleneck.
Measured on the development host: behind scalws, PHP-FPM, Node.js and uvicorn saturate
their own CPUs long before scalws does (Node 100 % of one CPU at 22–30k req/s while scalws used
~1.4 of 2 CPUs). Cloning scalws there costs CPU and RAM and changes nothing.
Decision
Signals per instance (sampled every second by the controller from the instance’s
metrics endpoint and /proc): CPU seconds (process), event-loop lag (time from a timer’s
deadline to its run, per loop, p99 over the window), in-flight requests and open
connections, accept rate, request rate, latency p50/p95/p99 (from the existing
histograms), 5xx and timeouts, RSS, FastCGI queue wait (php.queue_wait), upstream
head time, CPU of each managed application process group.
Components, each normalised to 0–1:
| Component | Definition |
|---|---|
cpu | instance CPU / CPUs assigned to it |
lag | clamp(loop-lag p99 / lag_saturated (default 5 ms)) |
inflight | clamp(in-flight / (loops × inflight_per_loop, default 256)) |
latency | clamp(p99 / p99_target) when a target is configured, else 0 |
PressureScore = max(cpu_lag, inflight, latency) with cpu_lag = min(cpu, lag + 0.5):
high CPU alone is not pressure (a busy-but-keeping-up instance has little loop lag);
high CPU with rising lag is. All constants are configuration with these defaults, and
every component is exported, not only the result.
Classification (evaluated over the scale-out window, first match wins):
| Class | Evidence |
|---|---|
MemoryPressure | host free RAM < min_free_ram or instance RSS > budget |
PhpSaturated / NodeSaturated / PythonSaturated | that application’s CPU ≥ 90 % of its allotment, or its queue wait rising, while web-tier cpu_lag < 0.75 |
WebTierSaturated | web-tier PressureScore > scale_out_above (0.75) with cpu_lag the dominant component and no application class matching |
ConnectionPressure | in-flight/connections near limits while CPU is low (slow clients, long-lived streams) |
Unknown/InsufficientEvidence | anything else, or fewer samples than the window needs |
Actions. Only WebTierSaturated may act (D4), and only within min/max/resource
budgets and cooldowns (ADR-0035). Every other class yields a diagnosis and a
recommendation. Every evaluation logs its evidence, e.g.
pressure=0.84 cpu=0.97 lag=0.62 class=WebTierSaturated action=ScaleOut from=2 to=3,
including no-action reasons (class=NodeSaturated action=None reason="Node CPU 1.00, web tier 0.41", action=None reason="host CPU budget exhausted").
Scale-in when PressureScore < scale_in_below (0.30) for scale_in_window (60 s) and
the remaining instances would stay below scale_out_above at the current load (projected
by CPU share), never below min_instances.
Consequences
- Event-loop lag is a new data-plane metric (a periodic timer per loop; cost measured in D2 against the regression gates).
- No learning component: decisions are explainable from logged numbers; thresholds are tuned from D3 observe-only runs on synthetic and real workloads.