Architecture

Architecture

Status: living document. Reflects the code as of milestone M9 (operations). The engineering contract is SCALWS_NextGen_Web_Server_Claude_Code_Handoff.docx; this file explains how the code realizes it and where it does not yet.

1. Goals and non-goals

Goals (from the handoff, §1): a Rust data plane that is fast, deterministic, observable, secure and multi-tenant; PHP / Node.js / Python / arbitrary upstreams as first-class runtimes via adapters; declarative, validated, atomically reloadable configuration.

Non-goals for the current milestones: dashboard, AI, eBPF, clustering, HTTP/3, ACME, containers. These are explicitly sequenced after the serving core (handoff §15, §18).

Hard rule: no LLM, network control-plane call or other optional subsystem is on the request path. The request path depends only on in-process, immutable state.

2. Process model

One process, scalwsd, running a multi-threaded Tokio runtime.

             ┌────────────────────────── scalwsd ───────────────────────────┐
 SIGHUP ───► │ reload: load → validate → prepare ─(ok)─► ArcSwap<ActiveState> │
 SIGTERM ──► │ graceful shutdown: stop accept → drain → exit                  │
             │                                                                │
 :80/:443 ─► │ listener tasks ─► connection tasks ─► request pipeline         │
 127.0.0.1 ► │ metrics listener (/metrics, /healthz)                          │
             └────────────────────────────────────────────────────────────────┘

Managed application processes (PHP-FPM pools, Node, Python application servers) are supervised children of scalwsd (ADR-0008), reached over private Unix sockets, never linked in-process, so an application crash cannot crash the server. Placement into per-application cgroups and per-tenant Unix users is M3.

 scalwsd ─┬─ supervisor ── php-fpm master ── workers (nobody when scalws runs as root)
           ├─ supervisor ── node server.js      (one per worker, own socket)
           └─ supervisor ── uvicorn / gunicorn  (one per instance, own workers)

3. Request pipeline

Handoff §2: socket → TLS → HTTP parser → tenant/domain lookup → limits → security/WAF → cache → router → runtime/proxy/static handler → response filters → metrics/traces/logs.

Implemented today (M1):

StageWhereNotes
socket accept + connection capscalws-http::listenerSemaphore; accept pauses when full (kernel backlog absorbs the burst).
TLS handshakescalws-http + scalws-tlsrustls, SNI resolver, ALPN h2,http/1.1, handshake timeout.
HTTP/1.1 + HTTP/2hyper via hyper-util auto builderHeader size/count limits, h1 header-read (and keep-alive idle) timeout, h2 stream/reset limits, per-connection idle timeout.
host normalisation + tenant/app lookupscalws-server::pipeline, scalws-routerExact and single-label wildcard hosts. Duplicate/missing Host rejected.
request limitsscalws-server::pipeline, scalws-core::bodyURI length, Content-Length precheck, streaming body cap and per-frame read timeout.
securityscalws-server::pipeline, scalws-securityPer-source request rate (429) and connection limits (at accept), strict SNI (421), profile + configured deny paths (404). Content-inspecting WAF: not implemented.
cachescalws-server::cache_stage, scalws-cacheOpt-in bounded memory cache (ADR-0012): explicit freshness only, Vary variants, profile bypass rules, stale-while-revalidate, purge by URL/prefix/tag.
path routingscalws-routerLongest segment-aligned prefix.
handlersscalws-static, scalws-proxy, scalws-runtime-*, scalws-server::respondStatic files, HTTP/1.1 reverse proxy (TCP, TLS, Unix socket, WebSocket tunnels), PHP via FastCGI, Node/Python worker pools, fixed responses.
response filtersscalws-server::pipelineServer, X-Request-Id; byte counting.
metrics / logsscalws-observePrometheus text exposition; JSON access logs. Traces: not yet (OTLP export planned).

The hot path performs: one ArcSwap load, one hash lookup for the host, a short linear scan of the app’s routes, and the handler. No locks are taken on the hot path except inside prometheus-client metric families (read-locked map lookups).

4. State and reload

ActiveState is an immutable snapshot built entirely before activation:

  • parsed + validated Config
  • HostRouter<AppEntry>; each AppEntry holds tenant/app identity and its routes (Arc<dyn Handler>)
  • TLS certificate resolvers (one per TLS listener)
  • request limits

prepare(config) performs every fallible step (file existence, certificate parsing and key/cert matching, SAN coverage, upstream address parsing). Only if it succeeds is the snapshot published with ArcSwap::store. In-flight requests keep the Arc of the snapshot they started with. A failed reload changes nothing and is logged and counted (scalws_config_reloads_total{result="failure"}).

Listener topology (addresses, TLS on/off) is fixed at start; a reload that changes it is rejected with an explicit error (“requires restart”). Certificates, routes, handlers and limits are reloadable.

scalwsctl check runs exactly the same prepare code path without binding sockets.

Persisted last-known-good snapshots and scalwsctl rollback require the admin API (M9); today the last-known-good is the in-memory active snapshot.

5. Crates

CrateResponsibility
scalws-coreShared identities (TenantId, AppId), body types, Handler trait, RequestContext, body limit wrapper, small HTTP helpers. No config, no I/O.
scalws-configYAML schema (serde, deny_unknown_fields), units (sizes, durations, CPU), semantic validation. Pure: no filesystem access.
scalws-runtimeRuntimeAdapter / RuntimeInstance contract and registry; process supervisor, endpoint allocation, log capture and the HTTP worker pool shared by managed runtimes.
scalws-runtime-nodeNode.js adapter: supervised processes (command) or external upstream.
scalws-runtime-pythonPython ASGI/WSGI adapter: uvicorn / hypercorn / gunicorn launch templates, virtualenvs, explicit commands, external upstreams.
scalws-runtime-containerContainer adapter (ADR-0020): podman run/docker run workers in a scalable pool, tenant cgroup placement, cleanup.
scalws-runtime-phpPHP adapter: managed PHP-FPM pools or external FastCGI, PHP/static/front-controller routing, CGI response parsing (ADR-0009).
scalws-routerHost and path routing tables, host normalisation. Generic over the routed value.
scalws-tlsLoading PEM material, validating cert/key/SAN, SNI resolver, rustls::ServerConfig.
scalws-staticConfined static file handler: traversal/symlink/dotfile protection, conditional and range requests.
scalws-proxyReverse-proxy handler: HTTP/1.1 upstream over TCP, TLS or Unix socket, hop-by-hop stripping, forwarding headers, timeouts, WebSocket tunnelling.
scalws-observeMetrics registry/definitions and logging initialisation.
scalws-httpListener accept loop, TLS handshake, hyper connection setup and limits, connection idle tracking, graceful drain.
scalws-serverComposition root: builds ActiveState, the request pipeline, the configuration controller (reload, rollback, restart, history), the admin API (Unix socket) and read-only metrics surface, per-application telemetry and the analysis task.
scalws-cacheResponse cache: Cache-Control parsing, storability, variants, tee body, purge (moka storage).
scalws-securityDeny-path matching and per-source limits with bounded, sharded state.
scalws-aiLocal LLM advisor (ADR-0016): redaction, input preparation, advice schema and validation, OpenAI-compatible provider.
scalws-ebpfOptional (feature ebpf): cgroup_skb network accounting per application cgroup (ADR-0025).
scalws-diagDeterministic diagnostics (ADR-0015): pure mapping of evidence to correlated findings.
scalws-policyDeterministic optimizer rules (ADR-0014): pure evaluation of telemetry windows into bounded, audited decisions.
scalws-profilesProfile catalog, detection engine, fact extraction, proposals, profile: resolution during configuration loading.
scalws-cgroupscgroup v2 hierarchy setup, delegation, tenant limits, placement trampoline, usage reading, pruning.
cmd/scalwsdDaemon binary: CLI flags, signals, logging.
cmd/scalwsctlOperator CLI: offline check, detect, profiles, backup, restore, completion; admin socket status, apps, tenants, top, reload, rollback, history, diff, restart, scale, purge, recommendations, diagnose, explain.

Crates from the handoff layout that do not exist yet (scalws-tenant; the admin API lives in scalws-server::admin instead of a separate scalws-api because it needs the server’s internals) are created when their milestone starts, not as empty shells. The built-in static, proxy and respond runtimes are registered as adapters in scalws-server.

6. Runtime adapter contract

See ADR-0004. Two levels:

  • Handler (in scalws-core) — the request-path interface: handle(Request, &RequestContext) -> Response.
  • RuntimeAdapter / RuntimeInstance (in scalws-runtime) — the lifecycle interface: prepare an instance from a spec, start/drain it, report lifecycle state and hand out a Handler. Every runtime type goes through the registry; managed runtimes return a proxy (Node, Python) or FastCGI (PHP) handler pointed at Unix sockets they own.

Reload semantics: an instance whose tenant/app/route, spec and application root are unchanged is reused (its processes keep running); new or changed instances are started and given their startup_timeout to become ready before the swap; instances the new generation no longer uses are drained in the background after it. While no worker is ready a managed handler answers 503 + Retry-After immediately.

The core router only ever sees Arc<dyn Handler>; it never knows the language.

7. Tenancy (M3, ADR-0010)

  • Process limits live in cgroup v2: <base>/tenants/<tenant>/ carries cpu.max, memory.max, memory.swap.max, pids.max (and io.weight where supported) for all of the tenant’s managed applications; each application has a leaf <base>/tenants/<tenant>/<app>/. Workers enter their leaf through a fixed shell trampoline before their first instruction, so forks can never escape the limits. Limits are rewritten on every activation (live updates on reload); groups of removed applications are pruned after drain.
  • Server work done on a tenant’s behalf (in the shared scalws process) is not in any tenant cgroup; it is bounded by per-tenant request quotas in the pipeline (max_concurrent_requests → 503, requests_per_second/burst → 429).
  • Accounting: a 5 s sampler exports per-tenant and per-application CPU seconds, memory, pids and OOM kills from the cgroup files.
  • Not yet: per-tenant Unix users and namespaces (design in ADR-0010), per-tenant connection limits, I/O bandwidth limits (io.max), cache/disk quotas (M5).

8. Application profiles (M4, ADR-0011)

Profiles are YAML data in profiles/ (embedded at build time): detection rules with weights, a runtime template, health check, static paths, cache-bypass and security rules. scalwsctl detect <dir> evaluates every profile against a directory (bounded reads, no symlinks followed, nothing executed), prints evidence and confidence, and proposes an application block filled with facts extracted from the tree (Python application object, Django WSGI module, virtualenv, npm start command). Nothing is changed by detection.

In configuration, profile: is explicit and static: it supplies the runtime when omitted (profiles whose template needs no facts: PHP, WordPress, WooCommerce, Laravel) and the health check path of Node/Python runtimes. Loading configuration never reads application files. Cache and security rules are carried as data for M5.

9. Deterministic optimizer (M6, ADR-0014)

The request path only increments per-application atomics (requests, 5xx, latency bucket, cache hit/miss) in finish(). A background task samples in-flight gauges every second, drains the counters once per server.optimizer.interval, and calls the pure scalws_policy::AppPolicy::evaluate per application. Decisions with disposition applied or rolled_back call RuntimeInstance::set_worker_target; every decision goes to the scalws::audit log target, the scalws_optimizer_decisions_total counter and the bounded journal at GET /optimizer on the admin listener (scalwsctl recommendations). Nothing in the optimizer is on the request path, and a disabled or failing optimizer changes nothing.

10. Diagnostics (M7, ADR-0015)

The same analysis task builds an scalws_diag::Evidence per application and window from the request counters (5xx, upstream failures by reason, quota rejections, latency, cache), the runtime (lifecycle state, ready workers, restarts) and cgroup v2 files (memory and limit, OOM kills, PSI some avg10 for memory/CPU/I/O). scalws_diag::diagnose returns findings sorted by severity, with resource causes attached to symptoms of the same window as related. New findings increment scalws_diagnostic_findings_total and, unless info, are logged under scalws::diag. At the end of every response body the pipeline observes the transfer time and keeps failed or slow requests in a bounded per-application ring buffer. GET /diagnose[?app=tenant/app] and scalwsctl diagnose expose both. eBPF is not used.

11. Local AI advisor (M8, ADR-0016)

Disabled by default. When enabled, scalws-server::advisor owns a bounded queue and one worker. New warning/critical findings (rate limited per application) and operator requests (POST /explain, scalwsctl explain --refresh) enqueue an application; the worker builds the input from the latest diagnosis, recent failed/slow requests, optimizer decisions and a runtime summary, passes it through scalws_ai::prepare (redaction, size cap), calls the operator’s OpenAI-compatible endpoint with a JSON schema and a timeout, and validates the answer (scalws_ai::advice::parse). Results — ok, invalid, timeout, error — are stored per application and served by GET /explain. The advisor has no handle to configuration, rules or runtimes; the request path never waits for it.

12. Operations (M9, ADR-0017)

control::Control performs every configuration change — SIGHUP, POST /v1/reload, rollback, application restart — through one serialised path (validate → prepare → swap) and keeps the last 10 applied configurations for rollback and diff. A restart prepares fresh runtime instances for one application only, starts them, swaps and drains the old ones. admin::route serves the Unix admin socket (all endpoints, mutations audited with peer credentials) and the loopback metrics listener (read-only views, dashboard; POST answers 403). scalwsctl talks HTTP/1.1 over the socket. Packaging (packaging/nfpm.yaml, scripts/package.sh) produces DEB/RPM with the hardened systemd unit.

13. ACME (M10, ADR-0018)

acme::run checks applications with acme: true at start, every check_interval and after every applied configuration (Control::changed). Orders use HTTP-01: tokens are placed in acme::Challenges, which the pipeline answers on plain-HTTP listeners before routing. Issued certificates are written atomically to state_dir; Control::reapply then runs the normal prepare step, where acme::load_all loads them as scalws_tls::ManagedCerts into the SNI resolvers of listeners with tls.acme: true (configured certificates win on conflicts; unusable files are skipped, never failing a reload).

14. HTTP/3 (M10, ADR-0019, feature http3)

scalws_http::h3 binds a quinn endpoint on the UDP port of a TLS listener with http3: true, using scalws_tls::server_config_h3 (TLS 1.3, ALPN h3, the listener’s reloadable resolver). Each request is converted to Request<RequestBody> (a channel body fed by the stream reader) and handed to RequestService::call_boxed, i.e. the same generic pipeline as hyper’s connections; the response body is streamed back frame by frame. TCP responses of such listeners carry Alt-Svc.

15. Platform

Linux is the production target (Ubuntu 24.04, AlmaLinux/Rocky 9). Development may happen on other OSes (ADR-0007): code builds everywhere, Linux-only features (Unix-socket upstreams, signals-based reload, openat2 confinement, cgroups) are cfg-gated, and CI plus the integration suite run on Linux.