Architecture
Architecture
Status: living document. Reflects the code as of milestone M9 (operations).
The engineering contract is SCALWS_NextGen_Web_Server_Claude_Code_Handoff.docx; this file
explains how the code realizes it and where it does not yet.
1. Goals and non-goals
Goals (from the handoff, §1): a Rust data plane that is fast, deterministic, observable, secure and multi-tenant; PHP / Node.js / Python / arbitrary upstreams as first-class runtimes via adapters; declarative, validated, atomically reloadable configuration.
Non-goals for the current milestones: dashboard, AI, eBPF, clustering, HTTP/3, ACME, containers. These are explicitly sequenced after the serving core (handoff §15, §18).
Hard rule: no LLM, network control-plane call or other optional subsystem is on the request path. The request path depends only on in-process, immutable state.
2. Process model
One process, scalwsd, running a multi-threaded Tokio runtime.
┌────────────────────────── scalwsd ───────────────────────────┐
SIGHUP ───► │ reload: load → validate → prepare ─(ok)─► ArcSwap<ActiveState> │
SIGTERM ──► │ graceful shutdown: stop accept → drain → exit │
│ │
:80/:443 ─► │ listener tasks ─► connection tasks ─► request pipeline │
127.0.0.1 ► │ metrics listener (/metrics, /healthz) │
└────────────────────────────────────────────────────────────────┘Managed application processes (PHP-FPM pools, Node, Python application servers) are
supervised children of scalwsd (ADR-0008), reached over private Unix sockets, never
linked in-process, so an application crash cannot crash the server. Placement into
per-application cgroups and per-tenant Unix users is M3.
scalwsd ─┬─ supervisor ── php-fpm master ── workers (nobody when scalws runs as root)
├─ supervisor ── node server.js (one per worker, own socket)
└─ supervisor ── uvicorn / gunicorn (one per instance, own workers)3. Request pipeline
Handoff §2: socket → TLS → HTTP parser → tenant/domain lookup → limits → security/WAF → cache → router → runtime/proxy/static handler → response filters → metrics/traces/logs.
Implemented today (M1):
| Stage | Where | Notes |
|---|---|---|
| socket accept + connection cap | scalws-http::listener | Semaphore; accept pauses when full (kernel backlog absorbs the burst). |
| TLS handshake | scalws-http + scalws-tls | rustls, SNI resolver, ALPN h2,http/1.1, handshake timeout. |
| HTTP/1.1 + HTTP/2 | hyper via hyper-util auto builder | Header size/count limits, h1 header-read (and keep-alive idle) timeout, h2 stream/reset limits, per-connection idle timeout. |
| host normalisation + tenant/app lookup | scalws-server::pipeline, scalws-router | Exact and single-label wildcard hosts. Duplicate/missing Host rejected. |
| request limits | scalws-server::pipeline, scalws-core::body | URI length, Content-Length precheck, streaming body cap and per-frame read timeout. |
| security | scalws-server::pipeline, scalws-security | Per-source request rate (429) and connection limits (at accept), strict SNI (421), profile + configured deny paths (404). Content-inspecting WAF: not implemented. |
| cache | scalws-server::cache_stage, scalws-cache | Opt-in bounded memory cache (ADR-0012): explicit freshness only, Vary variants, profile bypass rules, stale-while-revalidate, purge by URL/prefix/tag. |
| path routing | scalws-router | Longest segment-aligned prefix. |
| handlers | scalws-static, scalws-proxy, scalws-runtime-*, scalws-server::respond | Static files, HTTP/1.1 reverse proxy (TCP, TLS, Unix socket, WebSocket tunnels), PHP via FastCGI, Node/Python worker pools, fixed responses. |
| response filters | scalws-server::pipeline | Server, X-Request-Id; byte counting. |
| metrics / logs | scalws-observe | Prometheus text exposition; JSON access logs. Traces: not yet (OTLP export planned). |
The hot path performs: one ArcSwap load, one hash lookup for the host, a short linear
scan of the app’s routes, and the handler. No locks are taken on the hot path except
inside prometheus-client metric families (read-locked map lookups).
4. State and reload
ActiveState is an immutable snapshot built entirely before activation:
- parsed + validated
Config HostRouter<AppEntry>; eachAppEntryholds tenant/app identity and its routes (Arc<dyn Handler>)- TLS certificate resolvers (one per TLS listener)
- request limits
prepare(config) performs every fallible step (file existence, certificate parsing and
key/cert matching, SAN coverage, upstream address parsing). Only if it succeeds is the
snapshot published with ArcSwap::store. In-flight requests keep the Arc of the
snapshot they started with. A failed reload changes nothing and is logged and counted
(scalws_config_reloads_total{result="failure"}).
Listener topology (addresses, TLS on/off) is fixed at start; a reload that changes it is rejected with an explicit error (“requires restart”). Certificates, routes, handlers and limits are reloadable.
scalwsctl check runs exactly the same prepare code path without binding sockets.
Persisted last-known-good snapshots and scalwsctl rollback require the admin API (M9);
today the last-known-good is the in-memory active snapshot.
5. Crates
| Crate | Responsibility |
|---|---|
scalws-core | Shared identities (TenantId, AppId), body types, Handler trait, RequestContext, body limit wrapper, small HTTP helpers. No config, no I/O. |
scalws-config | YAML schema (serde, deny_unknown_fields), units (sizes, durations, CPU), semantic validation. Pure: no filesystem access. |
scalws-runtime | RuntimeAdapter / RuntimeInstance contract and registry; process supervisor, endpoint allocation, log capture and the HTTP worker pool shared by managed runtimes. |
scalws-runtime-node | Node.js adapter: supervised processes (command) or external upstream. |
scalws-runtime-python | Python ASGI/WSGI adapter: uvicorn / hypercorn / gunicorn launch templates, virtualenvs, explicit commands, external upstreams. |
scalws-runtime-container | Container adapter (ADR-0020): podman run/docker run workers in a scalable pool, tenant cgroup placement, cleanup. |
scalws-runtime-php | PHP adapter: managed PHP-FPM pools or external FastCGI, PHP/static/front-controller routing, CGI response parsing (ADR-0009). |
scalws-router | Host and path routing tables, host normalisation. Generic over the routed value. |
scalws-tls | Loading PEM material, validating cert/key/SAN, SNI resolver, rustls::ServerConfig. |
scalws-static | Confined static file handler: traversal/symlink/dotfile protection, conditional and range requests. |
scalws-proxy | Reverse-proxy handler: HTTP/1.1 upstream over TCP, TLS or Unix socket, hop-by-hop stripping, forwarding headers, timeouts, WebSocket tunnelling. |
scalws-observe | Metrics registry/definitions and logging initialisation. |
scalws-http | Listener accept loop, TLS handshake, hyper connection setup and limits, connection idle tracking, graceful drain. |
scalws-server | Composition root: builds ActiveState, the request pipeline, the configuration controller (reload, rollback, restart, history), the admin API (Unix socket) and read-only metrics surface, per-application telemetry and the analysis task. |
scalws-cache | Response cache: Cache-Control parsing, storability, variants, tee body, purge (moka storage). |
scalws-security | Deny-path matching and per-source limits with bounded, sharded state. |
scalws-ai | Local LLM advisor (ADR-0016): redaction, input preparation, advice schema and validation, OpenAI-compatible provider. |
scalws-ebpf | Optional (feature ebpf): cgroup_skb network accounting per application cgroup (ADR-0025). |
scalws-diag | Deterministic diagnostics (ADR-0015): pure mapping of evidence to correlated findings. |
scalws-policy | Deterministic optimizer rules (ADR-0014): pure evaluation of telemetry windows into bounded, audited decisions. |
scalws-profiles | Profile catalog, detection engine, fact extraction, proposals, profile: resolution during configuration loading. |
scalws-cgroups | cgroup v2 hierarchy setup, delegation, tenant limits, placement trampoline, usage reading, pruning. |
cmd/scalwsd | Daemon binary: CLI flags, signals, logging. |
cmd/scalwsctl | Operator CLI: offline check, detect, profiles, backup, restore, completion; admin socket status, apps, tenants, top, reload, rollback, history, diff, restart, scale, purge, recommendations, diagnose, explain. |
Crates from the handoff layout that do not exist yet (scalws-tenant; the admin API lives in scalws-server::admin instead of a
separate scalws-api because it needs the server’s internals) are created when their milestone starts, not as empty shells. The built-in
static, proxy and respond runtimes are registered as adapters in scalws-server.
6. Runtime adapter contract
See ADR-0004. Two levels:
Handler(inscalws-core) — the request-path interface:handle(Request, &RequestContext) -> Response.RuntimeAdapter/RuntimeInstance(inscalws-runtime) — the lifecycle interface: prepare an instance from a spec, start/drain it, report lifecycle state and hand out aHandler. Every runtime type goes through the registry; managed runtimes return a proxy (Node, Python) or FastCGI (PHP) handler pointed at Unix sockets they own.
Reload semantics: an instance whose tenant/app/route, spec and application root are
unchanged is reused (its processes keep running); new or changed instances are
started and given their startup_timeout to become ready before the swap; instances
the new generation no longer uses are drained in the background after it. While no worker
is ready a managed handler answers 503 + Retry-After immediately.
The core router only ever sees Arc<dyn Handler>; it never knows the language.
7. Tenancy (M3, ADR-0010)
- Process limits live in cgroup v2:
<base>/tenants/<tenant>/carriescpu.max,memory.max,memory.swap.max,pids.max(andio.weightwhere supported) for all of the tenant’s managed applications; each application has a leaf<base>/tenants/<tenant>/<app>/. Workers enter their leaf through a fixed shell trampoline before their first instruction, so forks can never escape the limits. Limits are rewritten on every activation (live updates on reload); groups of removed applications are pruned after drain. - Server work done on a tenant’s behalf (in the shared scalws process) is not in any
tenant cgroup; it is bounded by per-tenant request quotas in the pipeline
(
max_concurrent_requests→ 503,requests_per_second/burst→ 429). - Accounting: a 5 s sampler exports per-tenant and per-application CPU seconds, memory, pids and OOM kills from the cgroup files.
- Not yet: per-tenant Unix users and namespaces (design in ADR-0010), per-tenant
connection limits, I/O bandwidth limits (
io.max), cache/disk quotas (M5).
8. Application profiles (M4, ADR-0011)
Profiles are YAML data in profiles/ (embedded at build time): detection rules with
weights, a runtime template, health check, static paths, cache-bypass and security rules.
scalwsctl detect <dir> evaluates every profile against a directory (bounded reads, no
symlinks followed, nothing executed), prints evidence and confidence, and proposes an
application block filled with facts extracted from the tree (Python application object,
Django WSGI module, virtualenv, npm start command). Nothing is changed by detection.
In configuration, profile: is explicit and static: it supplies the runtime when omitted
(profiles whose template needs no facts: PHP, WordPress, WooCommerce, Laravel) and the
health check path of Node/Python runtimes. Loading configuration never reads application
files. Cache and security rules are carried as data for M5.
9. Deterministic optimizer (M6, ADR-0014)
The request path only increments per-application atomics (requests, 5xx, latency bucket,
cache hit/miss) in finish(). A background task samples in-flight gauges every second,
drains the counters once per server.optimizer.interval, and calls the pure
scalws_policy::AppPolicy::evaluate per application. Decisions with disposition applied
or rolled_back call RuntimeInstance::set_worker_target; every decision goes to the
scalws::audit log target, the scalws_optimizer_decisions_total counter and the bounded
journal at GET /optimizer on the admin listener (scalwsctl recommendations). Nothing in
the optimizer is on the request path, and a disabled or failing optimizer changes nothing.
10. Diagnostics (M7, ADR-0015)
The same analysis task builds an scalws_diag::Evidence per application and window from
the request counters (5xx, upstream failures by reason, quota rejections, latency, cache),
the runtime (lifecycle state, ready workers, restarts) and cgroup v2 files (memory and
limit, OOM kills, PSI some avg10 for memory/CPU/I/O). scalws_diag::diagnose returns
findings sorted by severity, with resource causes attached to symptoms of the same window
as related. New findings increment scalws_diagnostic_findings_total and, unless info,
are logged under scalws::diag. At the end of every response body the pipeline observes the
transfer time and keeps failed or slow requests in a bounded per-application ring buffer.
GET /diagnose[?app=tenant/app] and scalwsctl diagnose expose both. eBPF is not used.
11. Local AI advisor (M8, ADR-0016)
Disabled by default. When enabled, scalws-server::advisor owns a bounded queue and one
worker. New warning/critical findings (rate limited per application) and operator
requests (POST /explain, scalwsctl explain --refresh) enqueue an application; the worker
builds the input from the latest diagnosis, recent failed/slow requests, optimizer
decisions and a runtime summary, passes it through scalws_ai::prepare (redaction, size
cap), calls the operator’s OpenAI-compatible endpoint with a JSON schema and a timeout,
and validates the answer (scalws_ai::advice::parse). Results — ok, invalid,
timeout, error — are stored per application and served by GET /explain. The advisor
has no handle to configuration, rules or runtimes; the request path never waits for it.
12. Operations (M9, ADR-0017)
control::Control performs every configuration change — SIGHUP, POST /v1/reload,
rollback, application restart — through one serialised path (validate → prepare → swap)
and keeps the last 10 applied configurations for rollback and diff. A restart prepares
fresh runtime instances for one application only, starts them, swaps and drains the old
ones. admin::route serves the Unix admin socket (all endpoints, mutations audited with
peer credentials) and the loopback metrics listener (read-only views, dashboard; POST
answers 403). scalwsctl talks HTTP/1.1 over the socket. Packaging (packaging/nfpm.yaml,
scripts/package.sh) produces DEB/RPM with the hardened systemd unit.
13. ACME (M10, ADR-0018)
acme::run checks applications with acme: true at start, every check_interval and
after every applied configuration (Control::changed). Orders use HTTP-01: tokens are
placed in acme::Challenges, which the pipeline answers on plain-HTTP listeners before
routing. Issued certificates are written atomically to state_dir; Control::reapply
then runs the normal prepare step, where acme::load_all loads them as
scalws_tls::ManagedCerts into the SNI resolvers of listeners with tls.acme: true
(configured certificates win on conflicts; unusable files are skipped, never failing a
reload).
14. HTTP/3 (M10, ADR-0019, feature http3)
scalws_http::h3 binds a quinn endpoint on the UDP port of a TLS listener with
http3: true, using scalws_tls::server_config_h3 (TLS 1.3, ALPN h3, the listener’s
reloadable resolver). Each request is converted to Request<RequestBody> (a channel body
fed by the stream reader) and handed to RequestService::call_boxed, i.e. the same
generic pipeline as hyper’s connections; the response body is streamed back frame by
frame. TCP responses of such listeners carry Alt-Svc.
15. Platform
Linux is the production target (Ubuntu 24.04, AlmaLinux/Rocky 9). Development may happen
on other OSes (ADR-0007): code builds everywhere, Linux-only features (Unix-socket
upstreams, signals-based reload, openat2 confinement, cgroups) are cfg-gated, and CI
plus the integration suite run on Linux.