Architecture

One Rust process, scalwsd, per instance. An optional controller adds and removes instances on one host. Every request is served by an instance; the controller only observes and manages them.

Control plane and data plane

With listener_mode: reuseport, every instance binds its own socket per listener address and the kernel spreads new connections across the group by hash. A controller acting as a reverse proxy was rejected: an extra hop for every byte and a single choke point.

  • The controller never accepts, proxies or terminates client traffic and holds no TLS keys.
  • If the controller stops, instances keep serving. On restart it re-adopts them after checking PID, start time and executable.
  • Scaling adds capacity for new connections. In-flight requests are never migrated.
  • With the default listener_mode: shared, scalwsd’s request path is unchanged; regression gates compare it with a frozen baseline.
ADR-0033 →

control plane

scalws-controller observe → classify → decide → lifecycle state dir: generations, instance records (0600) · local Unix socket API, no remote mode

data plane

ClientsTCP :80 / :443
Linux kernelSO_REUSEPORT group
scalwsd × N each terminates TLS, serves HTTP, cache, static, proxy, FastCGI
ApplicationsPHP-FPM, upstreams (external endpoints)
No controller proxy hop in the normal HTTP data path. The kernel hands each new connection to one instance; the connection, its keep-alive requests and its HTTP/2 streams stay there.

Request pipeline

No LLM, network control-plane call or other optional subsystem is on the request path. It depends only on in-process, immutable state: one atomic snapshot load, one hash lookup for the host, a short scan of the app’s routes, then the handler.

  1. Accept + connection capSemaphore; accept pauses when full, the kernel backlog absorbs the burst
  2. TLS handshakerustls, SNI resolver, ALPN h2 and http/1.1, handshake timeout
  3. HTTP/1.1 and HTTP/2Header size and count limits, read and idle timeouts, h2 stream and reset limits
  4. Host and tenant lookupExact and single-label wildcard hosts; duplicate or missing Host rejected
  5. Request limitsURI length, Content-Length precheck, streaming body cap, per-frame read timeout
  6. SecurityPer-source rate and connection limits, strict SNI (421), deny paths (404)
  7. CacheOpt-in: explicit freshness only, Vary variants, stale-while-revalidate, purge
  8. RoutingLongest segment-aligned prefix on the normalised path
  9. HandlerStatic files, reverse proxy, PHP over FastCGI, Node and Python pools, fixed responses
  10. Metrics and logsPrometheus exposition, JSON access logs without query strings

Configuration reloads build a complete snapshot first (validate, then prepare: files, certificates, upstreams) and swap it atomically. In-flight requests keep the snapshot they started with. A failed reload changes nothing. Architecture document →

Bottleneck classifier

Behind ScalWS, PHP-FPM, Node.js and uvicorn usually saturate their own CPUs long before the web tier does. Cloning scalwsd there costs CPU and memory and changes nothing, so the controller classifies before it acts.

PressureScore v1 (ADR-0034)

cpu      = instance CPU / CPUs assigned
lag      = clamp(loop lag p99 / lag_saturated)
inflight = clamp(in-flight / (loops × 256))
latency  = clamp(p99 / p99_target)  # optional

cpu_lag       = min(cpu, lag + 0.5)
PressureScore = max(cpu_lag, inflight, latency)

Defaults: lag_saturated 5 ms, 256 in-flight requests per loop, no latency target. High CPU alone is not pressure: a busy instance that keeps up has little loop lag. High CPU with rising lag is. Every component is exported and every evaluation is logged with its evidence, including the reason for doing nothing. No learning component.

Adaptive Runtime: how ScalWS decides when to add instances

Throughput
13.3k req/s
p99
rising
Class
PhpSaturated
Action
None

class=PhpSaturated action=None reason="php CPU 1.00/1, web tier cpu_lag 0.41: adding scalws instances would not help"

PHP-FPM is at its CPU allotment. The web tier is busy but keeping up: its event loops show no lag.

Measured: D1, php, 64 connections. Pressure bars and values inside log lines are illustrative.

Scaling buys elasticity, not free CPU efficiency

Web-tier paths scale close to linearly with instances, but every instance brings its own CPU. Against one instance given the same CPUs, N instances deliver 0.92× to 1.14×. The value is adding and removing capacity as load changes. Each instance costs about 26 to 30 MiB.

Burst, idle to 1024 connections Autoscaling in web-tier mode, min 1, max 3, one CPU per instance. Three bursts in a row, zero errors.

1 instance · 1 CPU

req/s 37.2k
p99 46 ms

2 instances · 2 CPUs

req/s 81.5k
p99 19 ms

3 instances · 3 CPUs

req/s 143.2k
p99 14 ms

Bars share one axis per metric, starting at zero. Lower p99 is better.

D1: N instances on SO_REUSEPORT vs one instance. 3 runs × 10 s, medians, one CPU per instance. Efficiency = rps(N) / (rps(1) × N). “vs 1 instance” compares with one instance given the same number of event loops and CPUs.
Scenario (connections)1 instance2 instancesefficiencyvs 1 instance, 2 loops3 instancesefficiencyvs 1 instance, 3 loops
respond (64)51.4k92.5k0.901.04×166.8k1.081.14×
respond (256)52.4k102.6k0.981.08×159.0k1.010.97×
static-small (64)43.5k75.9k0.870.92×127.3k0.971.02×
static-small (256)43.2k82.2k0.951.00×144.2k1.111.08×
static-1m (64)4.22k8.96k1.061.07×13.1k1.031.06×
cache-hit (64)46.6k87.9k0.940.95×152.3k1.091.05×
TLS + HTTP/2 (256)34.0k65.4k0.960.92×115.8k1.131.08×
proxy (256)25.4k47.3k0.931.08×65.6k0.861.06×
php (64)13.3k15.4k0.580.95×15.8k0.401.04×
D5: concurrency sweep with autoscaling on (web-tier mode, min 1, max 3), 20 s per point. Instances = active instances at the end of the point.
Pathconnectionsreq/sp99 mserrorsinstancesRSS MiB
respond843,6080.280126
respond3249,2661.040254
respond6499,1091.100381
respond128143,3171.480382
respond256154,6183.000384
respond512158,8285.380389
respond1024167,00211.2403102
static-small860,2420.250379
static-small64127,3460.950380
static-small1024121,02315.180398

Host: development VM (8 vCPU EPYC 7642, Ubuntu 24.04, kernel 6.8), tcp_migrate_req=1. Topology separate from the frozen 2-CPU baseline: instance i on CPU i (one event loop each), upstream and PHP-FPM (4 workers) on CPU 3, oha on CPUs 4-7.

Lifecycle, generations and recovery

Readiness gates traffic: an instance is bound but not listening until the controller activates it. Removal drains, it does not kill.

Stopped

No process. The controller keeps a record per slot (web-0, web-1, ...) in its state directory: PID, CPU set, generation, state, timestamps. Files are 0600 and written atomically.

Starting

Process spawned with its instance id, an explicit configuration generation and its CPU set. It binds its reuseport sockets only when told to activate, so it receives no traffic.

Warming

Configuration prepared, runtimes and caches initialised. The readiness probe answers with the generation number and state “warming”.

Ready

Readiness passed. Still not in the SO_REUSEPORT group. A start that does not reach Ready within start_timeout is a failure.

Active

The controller activates it (SIGUSR2): listeners join the group and the kernel starts handing it new connections. Connections stay on the instance that accepted them.

Draining

Listening sockets close first; queued connections migrate to the remaining instances (net.ipv4.tcp_migrate_req=1). Keep-alive connections are told to close (Connection: close, HTTP/2 GOAWAY). In-flight requests finish until drain_timeout; a forced kill only after drain_timeout + kill_grace.

Failed is reachable from every state: process exit, start timeout, or drain timeout followed by a forced stop. Failed slots restart with exponential backoff (1 s to 60 s); more than max_restarts in restart_window leaves the slot failed instead of starting a fork storm. Liveness, readiness and saturation are separate: a saturated instance is healthy and is never restarted because of load.

Configuration generations

A generation is a validated, immutable copy of the configuration, written to a temporary name, fsynced and renamed. A reload rolls instances one at a time: start on the new generation, Ready, Active, then drain one old instance. Capacity never drops below the current count. Rollback is the same procedure towards the previous generation.

Anti-thrashing

Scale out after 3 to 5 s above the threshold, scale in after 60 s below it, a cooldown after every action (default 30 s), one action in flight, and hysteresis between scale_out_above (0.75) and scale_in_below (0.30). autoscaling: off freezes the instance count immediately.

Tested under load

Scale 1 → 2 → 3 → 1, rolling reload, reload while scaling, controller restart and a killed instance: zero failed requests. HTTP/2 (8 × 32 streams) across scale-in 3 → 1: no processed stream lost; the client’s queued, never-sent requests were canceled, which RFC 9113 §6.8 makes safe to retry.

ADR-0035 →

Known limitations

Adaptive Runtime final report →