Architecture
One Rust process, scalwsd, per instance. An optional controller adds and removes instances on one host. Every request is served by an instance; the controller only observes and manages them.
Control plane and data plane
With listener_mode: reuseport, every instance binds its own socket per listener address and the kernel spreads new connections across the group by hash. A controller acting as a reverse proxy was rejected: an extra hop for every byte and a single choke point.
- The controller never accepts, proxies or terminates client traffic and holds no TLS keys.
- If the controller stops, instances keep serving. On restart it re-adopts them after checking PID, start time and executable.
- Scaling adds capacity for new connections. In-flight requests are never migrated.
- With the default
listener_mode: shared, scalwsd’s request path is unchanged; regression gates compare it with a frozen baseline.
control plane
data plane
Request pipeline
No LLM, network control-plane call or other optional subsystem is on the request path. It depends only on in-process, immutable state: one atomic snapshot load, one hash lookup for the host, a short scan of the app’s routes, then the handler.
- Accept + connection capSemaphore; accept pauses when full, the kernel backlog absorbs the burst
- TLS handshakerustls, SNI resolver, ALPN h2 and http/1.1, handshake timeout
- HTTP/1.1 and HTTP/2Header size and count limits, read and idle timeouts, h2 stream and reset limits
- Host and tenant lookupExact and single-label wildcard hosts; duplicate or missing Host rejected
- Request limitsURI length, Content-Length precheck, streaming body cap, per-frame read timeout
- SecurityPer-source rate and connection limits, strict SNI (421), deny paths (404)
- CacheOpt-in: explicit freshness only, Vary variants, stale-while-revalidate, purge
- RoutingLongest segment-aligned prefix on the normalised path
- HandlerStatic files, reverse proxy, PHP over FastCGI, Node and Python pools, fixed responses
- Metrics and logsPrometheus exposition, JSON access logs without query strings
Configuration reloads build a complete snapshot first (validate, then prepare: files, certificates, upstreams) and swap it atomically. In-flight requests keep the snapshot they started with. A failed reload changes nothing. Architecture document →
Bottleneck classifier
Behind ScalWS, PHP-FPM, Node.js and uvicorn usually saturate their own CPUs long before the web tier does. Cloning scalwsd there costs CPU and memory and changes nothing, so the controller classifies before it acts.
PressureScore v1 (ADR-0034)
cpu = instance CPU / CPUs assigned
lag = clamp(loop lag p99 / lag_saturated)
inflight = clamp(in-flight / (loops × 256))
latency = clamp(p99 / p99_target) # optional
cpu_lag = min(cpu, lag + 0.5)
PressureScore = max(cpu_lag, inflight, latency)Defaults: lag_saturated 5 ms, 256 in-flight requests per loop, no latency target. High CPU alone is not pressure: a busy instance that keeps up has little loop lag. High CPU with rising lag is. Every component is exported and every evaluation is logged with its evidence, including the reason for doing nothing. No learning component.
- Throughput
- 13.3k req/s
- p99
- rising
- Class
- PhpSaturated
- Action
- None
class=PhpSaturated action=None reason="php CPU 1.00/1, web tier cpu_lag 0.41: adding scalws instances would not help"
PHP-FPM is at its CPU allotment. The web tier is busy but keeping up: its event loops show no lag.
Measured: D1, php, 64 connections. Pressure bars and values inside log lines are illustrative.
Scaling buys elasticity, not free CPU efficiency
Web-tier paths scale close to linearly with instances, but every instance brings its own CPU. Against one instance given the same CPUs, N instances deliver 0.92× to 1.14×. The value is adding and removing capacity as load changes. Each instance costs about 26 to 30 MiB.
1 instance · 1 CPU
2 instances · 2 CPUs
3 instances · 3 CPUs
Bars share one axis per metric, starting at zero. Lower p99 is better.
| Scenario (connections) | 1 instance | 2 instances | efficiency | vs 1 instance, 2 loops | 3 instances | efficiency | vs 1 instance, 3 loops |
|---|---|---|---|---|---|---|---|
| respond (64) | 51.4k | 92.5k | 0.90 | 1.04× | 166.8k | 1.08 | 1.14× |
| respond (256) | 52.4k | 102.6k | 0.98 | 1.08× | 159.0k | 1.01 | 0.97× |
| static-small (64) | 43.5k | 75.9k | 0.87 | 0.92× | 127.3k | 0.97 | 1.02× |
| static-small (256) | 43.2k | 82.2k | 0.95 | 1.00× | 144.2k | 1.11 | 1.08× |
| static-1m (64) | 4.22k | 8.96k | 1.06 | 1.07× | 13.1k | 1.03 | 1.06× |
| cache-hit (64) | 46.6k | 87.9k | 0.94 | 0.95× | 152.3k | 1.09 | 1.05× |
| TLS + HTTP/2 (256) | 34.0k | 65.4k | 0.96 | 0.92× | 115.8k | 1.13 | 1.08× |
| proxy (256) | 25.4k | 47.3k | 0.93 | 1.08× | 65.6k | 0.86 | 1.06× |
| php (64) | 13.3k | 15.4k | 0.58 | 0.95× | 15.8k | 0.40 | 1.04× |
| Path | connections | req/s | p99 ms | errors | instances | RSS MiB |
|---|---|---|---|---|---|---|
| respond | 8 | 43,608 | 0.28 | 0 | 1 | 26 |
| respond | 32 | 49,266 | 1.04 | 0 | 2 | 54 |
| respond | 64 | 99,109 | 1.10 | 0 | 3 | 81 |
| respond | 128 | 143,317 | 1.48 | 0 | 3 | 82 |
| respond | 256 | 154,618 | 3.00 | 0 | 3 | 84 |
| respond | 512 | 158,828 | 5.38 | 0 | 3 | 89 |
| respond | 1024 | 167,002 | 11.24 | 0 | 3 | 102 |
| static-small | 8 | 60,242 | 0.25 | 0 | 3 | 79 |
| static-small | 64 | 127,346 | 0.95 | 0 | 3 | 80 |
| static-small | 1024 | 121,023 | 15.18 | 0 | 3 | 98 |
Host: development VM (8 vCPU EPYC 7642, Ubuntu 24.04, kernel 6.8), tcp_migrate_req=1. Topology separate from the frozen 2-CPU baseline: instance i on CPU i (one event loop each), upstream and PHP-FPM (4 workers) on CPU 3, oha on CPUs 4-7.
Lifecycle, generations and recovery
Readiness gates traffic: an instance is bound but not listening until the controller activates it. Removal drains, it does not kill.
Stopped
No process. The controller keeps a record per slot (web-0, web-1, ...) in its state directory: PID, CPU set, generation, state, timestamps. Files are 0600 and written atomically.
Starting
Process spawned with its instance id, an explicit configuration generation and its CPU set. It binds its reuseport sockets only when told to activate, so it receives no traffic.
Warming
Configuration prepared, runtimes and caches initialised. The readiness probe answers with the generation number and state “warming”.
Ready
Readiness passed. Still not in the SO_REUSEPORT group. A start that does not reach Ready within start_timeout is a failure.
Active
The controller activates it (SIGUSR2): listeners join the group and the kernel starts handing it new connections. Connections stay on the instance that accepted them.
Draining
Listening sockets close first; queued connections migrate to the remaining instances (net.ipv4.tcp_migrate_req=1). Keep-alive connections are told to close (Connection: close, HTTP/2 GOAWAY). In-flight requests finish until drain_timeout; a forced kill only after drain_timeout + kill_grace.
Failed is reachable from every state: process exit, start timeout, or drain timeout followed by a forced stop. Failed slots restart with exponential backoff (1 s to 60 s); more than max_restarts in restart_window leaves the slot failed instead of starting a fork storm. Liveness, readiness and saturation are separate: a saturated instance is healthy and is never restarted because of load.
Configuration generations
A generation is a validated, immutable copy of the configuration, written to a temporary name, fsynced and renamed. A reload rolls instances one at a time: start on the new generation, Ready, Active, then drain one old instance. Capacity never drops below the current count. Rollback is the same procedure towards the previous generation.
Anti-thrashing
Scale out after 3 to 5 s above the threshold, scale in after 60 s below it, a cooldown after every action (default 30 s), one action in flight, and hysteresis between scale_out_above (0.75) and scale_in_below (0.30). autoscaling: off freezes the instance count immediately.
Tested under load
Scale 1 → 2 → 3 → 1, rolling reload, reload while scaling, controller restart and a killed instance: zero failed requests. HTTP/2 (8 × 32 streams) across scale-in 3 → 1: no processed stream lost; the client’s queued, never-sent requests were canceled, which RFC 9113 §6.8 makes safe to retry.
Known limitations
- Multi-instance is not more CPU-efficient than one instance with the same CPUs.
- Kernel hash distribution is uneven with few connections: up to 2:1 at 64 connections, even at 256.
- HTTP/3 with several instances needs the controller’s QUIC demux (
listen: inherited); it is refused in reuseport mode, where QUIC packets of one connection could reach different processes. - Managed PHP-FPM, Node.js, Python and container runtimes are refused under the controller. Application tiers must be external endpoints; their scaling is recommendation-only.
- Scale-in requires
net.ipv4.tcp_migrate_req=1(Linux 5.14 or later), persisted in/etc/sysctl.d/. - Caches are per instance and start cold on new instances; purges via the admin API go to every instance.
- WebSocket connections across scale-in have not been tested. Slow-client drains were not conclusively exercised.
- Node.js: on the 2026-10-06 VM baseline ScalWS was behind nginx; on bare metal (2026-10-11) it is 1.00× nginx 1.31.6and 0.93× OpenLiteSpeed, which is ahead of both. Benchmarks