Adaptive Runtime
ScalWS Adaptive Runtime — design summary, configuration and benchmark plan
Design ADRs: 0033 (control/data plane, reuseport), 0034 (PressureScore, classification),
0035 (lifecycle, drain, generations, recovery), 0036 (multi-instance state).
Rollback point: pre-adaptive-runtime-2026-10-06.
Architecture
clients
│ TCP :80/:443 (kernel SO_REUSEPORT group)
┌───────────┼────────────┬─────────────┐
▼ ▼ ▼ │
scalwsd scalwsd scalwsd … │ data plane: TLS, HTTP, cache,
web-0 web-1 web-2 │ static, proxy, FastCGI
CPUs 0-1 CPUs 2-3 CPUs 4-5 │
│ metrics/readiness (loopback/unix) │
└───────────┬──────────────────────────┘
▼
scalws-controller — observe → classify → decide → lifecycle
(no client traffic, no TLS termination)
│ start/activate/drain/stop, generations, state dirConfiguration reference (controller.yaml)
controller:
state_dir: /var/lib/scalws-controller # generations/, instances/, logs/, limit-share (0700)
socket: /run/scalws-controller/control.sock # local control API (0600); no remote mode
webd_binary: /usr/local/bin/scalwsd
config: /etc/scalws/scalws.yaml # validated into a new generation on start/reload
listen: inherited # inherited (default): one accept queue opened by the controller and
# shared by all instances; reuseport: one SO_REUSEPORT socket per instance
instances:
min: 1 # required, >= 1
max: 3 # required, >= min; max × cpus_per_instance CPUs must exist
cpus_per_instance: 2 # required; also the instance's event-loop count
cpus: "0-5" # required; instance i gets the i-th slice of this list
metrics_base_port: 9910 # instance i: /metrics, /healthz, /readyz on 127.0.0.1:(base+i)
min_free_ram_mib: 512 # no new instance below this MemAvailable
start_timeout: 20 # seconds (starting -> ready), else failed
drain_timeout: 30 # seconds for open connections after SIGTERM
kill_grace: 5 # seconds after drain_timeout before SIGKILL
max_restarts: 5 # per slot and restart_window; then the slot stays failed
restart_window: 300 # seconds
allow_drain_without_migration: false # listen: reuseport only — scale-in needs net.ipv4.tcp_migrate_req=1 otherwise
autoscaling:
mode: off # off (emergency switch) | observe | web-tier
scale_out_above: 0.75 # hysteresis: 0 < scale_in_below < scale_out_above <= 1
scale_out_window: 4 # seconds above the threshold before scaling out
scale_in_below: 0.30
scale_in_window: 60 # seconds below before scaling in
cooldown: 30 # seconds between two actions
lag_saturated: 5 # milliseconds of event-loop lag that count as saturated
inflight_per_loop: 256
p99_target_ms: null # optional latency component
applications: # application tiers to watch (classification only)
- {name: php, kind: php, comm: php-fpm8.3, cpus: 1}The source configuration must reach application tiers as external endpoints (php: fastcgi, others through proxy): a generation with managed PHP-FPM/Node/Python/container
runtimes is refused (every instance would start its own pool, ADR-0036; proposal:
docs/d8-managed-runtimes-proposal.md). HTTP/3 works with several instances through the
controller’s QUIC demux (listen: inherited); it is refused with listen: reuseport.
scalwsd flags used by the controller (also usable by hand):
--wait-for-activation (bind but do not listen until SIGUSR2), --listen-socket <path>
(receive the controller’s listening sockets; listen: inherited), --reuseport,
--instance-id <id> (per-instance runtime dir, admin socket, disk cache dir; turns on
the loop-lag probe; no remote admin), --metrics-address <addr>,
--limit-share-file <path> (per-IP, route and tenant limits divided by the number in the
file, re-read on SIGHUP), --worker-threads <n>.
Controller commands: scalws-controller run --config controller.yaml; status,
scale <n>, reload, autoscaling off|observe|web-tier, purge QUERY against
--socket; quic-lb-status --demux-socket <state_dir>/quic-lb.sock.
status fields: instances and phases, generation, last_decision,
distribution (per-instance work and imbalance, docs/connection-distribution.md),
listeners (sockets per generation, docs/listener-generations.md), quic (demux,
docs/http3-multi-instance.md), last_rollback.
Host requirements: Linux, taskset (util-linux), root or the capabilities scalwsd
needs for its ports. With listen: reuseport also Linux ≥ 5.14 and
net.ipv4.tcp_migrate_req=1 (the package installs it in /usr/lib/sysctl.d/).
Draining WebSockets: after the HTTP connections, each tunnel sends the client a close
frame 1001 Going Away at the next frame boundary (a frame that does not end within 2 s
is cut) and forwards the client’s close reply upstream; clients reconnect to an instance
that stays.
Benchmark plan
Separate topology from the frozen 2-CPU baseline (that one stays the single-instance regression reference). On the 8-vCPU host: oha on 2 CPUs, applications on 1–2 CPUs, instances on 2 CPUs each.
- Gate (every stage): single instance, Adaptive Runtime disabled, against the frozen baseline (interleaved A/B): hot-path RPS within 5 %, cache/H2/PHP within 5 %, p99 within 10–15 %, RSS within 15 %.
- D1 scaling: 1, 2, 3 instances × respond, static-small, static-1m, cache-hit,
tls-h2, proxy, php; connections 8/32/64/128/256/512/1024; report RPS, p99,
CPU/request, RSS, per-instance connection share (imbalance), and
efficiency(N) = throughput(N) / (throughput(1) × N). - Lifecycle: start-to-ready and drain times, add/remove under load with zero failed requests, long-lived HTTP/2 and WebSocket connections across scale-in, slow clients.
- D3 observe-only: web-tier-bound (respond/static/cache) vs application-bound (PHP, Node, Python) loads — the classifier must name the right tier.
- D5 overload: concurrency sweeps beyond host capacity, bursts, instance crash, controller restart, reload while scaling, memory pressure.