Decision records

ADR-0008: Managed application processes

  • Status: Accepted (2026-10-05)

Context

Handoff §5/§13: Node, Python and PHP-FPM processes are supervised by scalwsd with lifecycle state, readiness, crash backoff, graceful drain/restart, rate-limited stdout/stderr capture and restart-loop protection. A runtime crash must never crash scalwsd. Unix sockets are preferred.

Decision

  • One supervisor task per managed process (scalws_runtime::supervisor, Unix only). Processes are spawned without a shell from an explicit argv, in their own process group, with a cleared environment (only PATH, LANG, HOME, the configured env and the SCALWS_* variables), stdin from /dev/null, cwd = application root.
  • Each worker listens on its own Unix socket under server.runtime_dir (<runtime_dir>/s/<16-hex>-<n>.sock, short to respect the 108-byte sun_path limit). The path is passed as SCALWS_SOCKET and, for HTTP runtimes, as PORT (Node’s listen(process.env.PORT) accepts a path). listen: tcp gives a loopback port instead for frameworks that only accept numbers.
  • Readiness = the socket accepts a connection (optionally an HTTP health_path answers with a non-5xx status) within startup_timeout.
  • Crash handling: exponential backoff 0.5 s → 30 s, reset after 60 s of healthy running; ≥ 5 crashes in 60 s marks the worker Failed (it keeps retrying at the maximum backoff, and the state is visible in metrics and logs).
  • Stop/drain: SIGTERM to the process group, stop_timeout, then SIGKILL.
  • stdout/stderr are read line by line (max 8 KiB per line) and logged under target scalws::app with tenant/app/worker; a token bucket (100 lines/s, burst 200) drops excess lines and reports the drop count.
  • Reload reuses a running instance when its tenant/app/route and runtime spec are unchanged; changed instances are started (and given up to startup_timeout to become ready) before the configuration swap, and the old ones are drained afterwards.
  • While no worker is ready, the handler answers 503 with Retry-After; it never blocks.

Consequences

  • Per-tenant Unix users and cgroup placement plug into the spawn step (M3).
  • Windows/macOS cannot run managed runtimes (prepare fails with an explicit error).