Adaptive Runtime

ScalWS Adaptive Runtime — design summary, configuration and benchmark plan

Design ADRs: 0033 (control/data plane, reuseport), 0034 (PressureScore, classification), 0035 (lifecycle, drain, generations, recovery), 0036 (multi-instance state). Rollback point: pre-adaptive-runtime-2026-10-06.

Architecture

                 clients
                    │  TCP :80/:443 (kernel SO_REUSEPORT group)
        ┌───────────┼────────────┬─────────────┐
        ▼           ▼            ▼             │
   scalwsd     scalwsd     scalwsd   …      │  data plane: TLS, HTTP, cache,
   web-0        web-1        web-2             │  static, proxy, FastCGI
   CPUs 0-1     CPUs 2-3     CPUs 4-5          │
        │ metrics/readiness (loopback/unix)    │
        └───────────┬──────────────────────────┘
                    ▼
             scalws-controller    — observe → classify → decide → lifecycle
             (no client traffic, no TLS termination)
                    │ start/activate/drain/stop, generations, state dir

Configuration reference (controller.yaml)

controller:
  state_dir: /var/lib/scalws-controller       # generations/, instances/, logs/, limit-share (0700)
  socket: /run/scalws-controller/control.sock # local control API (0600); no remote mode
  webd_binary: /usr/local/bin/scalwsd
  config: /etc/scalws/scalws.yaml                # validated into a new generation on start/reload
  listen: inherited            # inherited (default): one accept queue opened by the controller and
                               # shared by all instances; reuseport: one SO_REUSEPORT socket per instance
instances:
  min: 1                       # required, >= 1
  max: 3                       # required, >= min; max × cpus_per_instance CPUs must exist
  cpus_per_instance: 2         # required; also the instance's event-loop count
  cpus: "0-5"                  # required; instance i gets the i-th slice of this list
  metrics_base_port: 9910      # instance i: /metrics, /healthz, /readyz on 127.0.0.1:(base+i)
  min_free_ram_mib: 512        # no new instance below this MemAvailable
  start_timeout: 20            # seconds (starting -> ready), else failed
  drain_timeout: 30            # seconds for open connections after SIGTERM
  kill_grace: 5                # seconds after drain_timeout before SIGKILL
  max_restarts: 5              # per slot and restart_window; then the slot stays failed
  restart_window: 300          # seconds
  allow_drain_without_migration: false   # listen: reuseport only — scale-in needs net.ipv4.tcp_migrate_req=1 otherwise
autoscaling:
  mode: off                    # off (emergency switch) | observe | web-tier
  scale_out_above: 0.75        # hysteresis: 0 < scale_in_below < scale_out_above <= 1
  scale_out_window: 4          # seconds above the threshold before scaling out
  scale_in_below: 0.30
  scale_in_window: 60          # seconds below before scaling in
  cooldown: 30                 # seconds between two actions
  lag_saturated: 5             # milliseconds of event-loop lag that count as saturated
  inflight_per_loop: 256
  p99_target_ms: null          # optional latency component
  applications:                # application tiers to watch (classification only)
    - {name: php, kind: php, comm: php-fpm8.3, cpus: 1}

The source configuration must reach application tiers as external endpoints (php: fastcgi, others through proxy): a generation with managed PHP-FPM/Node/Python/container runtimes is refused (every instance would start its own pool, ADR-0036; proposal: docs/d8-managed-runtimes-proposal.md). HTTP/3 works with several instances through the controller’s QUIC demux (listen: inherited); it is refused with listen: reuseport.

scalwsd flags used by the controller (also usable by hand): --wait-for-activation (bind but do not listen until SIGUSR2), --listen-socket <path> (receive the controller’s listening sockets; listen: inherited), --reuseport, --instance-id <id> (per-instance runtime dir, admin socket, disk cache dir; turns on the loop-lag probe; no remote admin), --metrics-address <addr>, --limit-share-file <path> (per-IP, route and tenant limits divided by the number in the file, re-read on SIGHUP), --worker-threads <n>.

Controller commands: scalws-controller run --config controller.yaml; status, scale <n>, reload, autoscaling off|observe|web-tier, purge QUERY against --socket; quic-lb-status --demux-socket <state_dir>/quic-lb.sock.

status fields: instances and phases, generation, last_decision, distribution (per-instance work and imbalance, docs/connection-distribution.md), listeners (sockets per generation, docs/listener-generations.md), quic (demux, docs/http3-multi-instance.md), last_rollback.

Host requirements: Linux, taskset (util-linux), root or the capabilities scalwsd needs for its ports. With listen: reuseport also Linux ≥ 5.14 and net.ipv4.tcp_migrate_req=1 (the package installs it in /usr/lib/sysctl.d/).

Draining WebSockets: after the HTTP connections, each tunnel sends the client a close frame 1001 Going Away at the next frame boundary (a frame that does not end within 2 s is cut) and forwards the client’s close reply upstream; clients reconnect to an instance that stays.

Benchmark plan

Separate topology from the frozen 2-CPU baseline (that one stays the single-instance regression reference). On the 8-vCPU host: oha on 2 CPUs, applications on 1–2 CPUs, instances on 2 CPUs each.

  1. Gate (every stage): single instance, Adaptive Runtime disabled, against the frozen baseline (interleaved A/B): hot-path RPS within 5 %, cache/H2/PHP within 5 %, p99 within 10–15 %, RSS within 15 %.
  2. D1 scaling: 1, 2, 3 instances × respond, static-small, static-1m, cache-hit, tls-h2, proxy, php; connections 8/32/64/128/256/512/1024; report RPS, p99, CPU/request, RSS, per-instance connection share (imbalance), and efficiency(N) = throughput(N) / (throughput(1) × N).
  3. Lifecycle: start-to-ready and drain times, add/remove under load with zero failed requests, long-lived HTTP/2 and WebSocket connections across scale-in, slow clients.
  4. D3 observe-only: web-tier-bound (respond/static/cache) vs application-bound (PHP, Node, Python) loads — the classifier must name the right tier.
  5. D5 overload: concurrency sweeps beyond host capacity, bursts, instance crash, controller restart, reload while scaling, memory pressure.