Decision records

ADR-0010: cgroup v2 placement, tenant resource limits and request quotas

  • Status: Accepted (2026-10-05)

Context

Handoff §7: Tenant → Application → Domain; CPU, memory, PIDs, I/O, concurrent requests, connections and request rate policies; cgroups v2 for application workloads; do not claim cgroups isolate CPU used inside the shared web-server process. Per-tenant Unix users and namespaces are a later hardening feature (design only).

Decision

Hierarchy

<base>/                         scalws's delegated cgroup (systemd Delegate=yes, or server.cgroups.path)
<base>/server                   scalwsd itself (only when <base> is scalws's own cgroup)
<base>/tenants/<tenant>/        tenant limits: cpu.max, memory.max, memory.swap.max, pids.max, io.weight
<base>/tenants/<tenant>/<app>/  leaf: the application's worker processes

cgroup v2 forbids processes in a group that delegates controllers to children, so when <base> is scalws’s own cgroup scalws first moves itself to <base>/server. Controllers cpu memory pids io are enabled where available (cgroup.controllers); missing ones are reported once and skipped (io.weight needs a weight-capable I/O scheduler).

Placement without races

A worker must be inside its leaf before its first instruction: PHP-FPM and npm start fork immediately, and children forked before a later move would stay outside the limits. Without unsafe (pre_exec) or clone3(CLONE_INTO_CGROUP) support in std/tokio, the supervisor starts workers through a fixed POSIX shell trampoline:

/bin/sh -c 'echo $$ > "$0/cgroup.procs" && exec "$@"' <leaf> <argv...>

The script is a constant; the leaf path and the user argv are positional parameters and are never interpreted by the shell, so the “no shell for user commands” rule (ADR-0008) holds. If the write fails the worker does not start (it is never run unconfined).

Limits

  • cpu: "N%" → cpu.max = N*1000 100000; memory → memory.max and memory.swap.max = swap (default 0, otherwise memory limits are absorbed by swap); pids → pids.max; io_weight → io.weight (best effort; 100 when unset); io_bandwidth → io.max rbps/wbps on every disk holding the tenant’s application roots (partitions are mapped to their disk).
  • Limits are written on every activation, so changing a tenant’s resources takes effect on reload without restarting its applications. Groups of removed applications/tenants are removed after their instances have drained.
  • Modes: server.cgroups.mode: auto (default; enforce when delegation works, otherwise warn once and serve without enforcement), required (refuse to start), off.

Run-time limits

POST /v1/tenants/{tenant}/limits (admin socket; scalwsctl limit, scalws-limitcpu, scalws-limitram, scalws-limitio) overrides cpu, memory, io_weight and io_bandwidth of one tenant at once. Overrides live in scalwsd’s memory: they are written again after every reload and end when scalwsd restarts (under scalws-controller: when its instances are replaced). --persist writes the value into the tenant’s resources in scalws.yaml (block-style YAML edited in place, comments kept, previous file kept as *.bak.<time>) and reloads, through scalws-controller when it runs.

Accounting

A sampler reads cpu.stat, memory.current, pids.current, memory.events of every tenant and application group every 5 s into gauges (scalws_tenant_*, scalws_app_*). GET /v1/tenants/usage?interval_ms= (scalwsctl high, scalws-highcpu, scalws-highram, scalws-highio) samples the tenant groups twice and reports CPU (percent of one CPU), memory, disk read/write rates (io.stat) and I/O pressure with the limits that apply.

Request quotas (shared server work)

Per tenant: max_concurrent_requests (→ 503 + Retry-After) and requests_per_second + burst token bucket (→ 429 + Retry-After). State is one semaphore and one bucket per configured tenant, so memory is bounded by configuration. Per-source and per-route limits are M5 (security); per-tenant connection limits need SNI-time tenant attribution and are deferred.

Not done (design for later)

Per-tenant Unix users: spawn workers through setpriv --reuid --regid --clear-groups (util-linux; std’s CommandExt::groups is unstable and uid() alone would keep root’s supplementary groups), chown the tenant’s socket directory, check document-root ownership. User and mount namespaces after that.

Consequences

  • Delegation must be configured in production (Delegate=yes in the systemd unit).
  • CPU used by scalwsd on behalf of a tenant is not attributed to the tenant’s cgroup; request quotas are the control for that, as the handoff requires.