Decision records
ADR-0010: cgroup v2 placement, tenant resource limits and request quotas
- Status: Accepted (2026-10-05)
Context
Handoff §7: Tenant → Application → Domain; CPU, memory, PIDs, I/O, concurrent requests, connections and request rate policies; cgroups v2 for application workloads; do not claim cgroups isolate CPU used inside the shared web-server process. Per-tenant Unix users and namespaces are a later hardening feature (design only).
Decision
Hierarchy
<base>/ scalws's delegated cgroup (systemd Delegate=yes, or server.cgroups.path)
<base>/server scalwsd itself (only when <base> is scalws's own cgroup)
<base>/tenants/<tenant>/ tenant limits: cpu.max, memory.max, memory.swap.max, pids.max, io.weight
<base>/tenants/<tenant>/<app>/ leaf: the application's worker processescgroup v2 forbids processes in a group that delegates controllers to children, so when
<base> is scalws’s own cgroup scalws first moves itself to <base>/server. Controllers
cpu memory pids io are enabled where available (cgroup.controllers); missing ones are
reported once and skipped (io.weight needs a weight-capable I/O scheduler).
Placement without races
A worker must be inside its leaf before its first instruction: PHP-FPM and npm start
fork immediately, and children forked before a later move would stay outside the limits.
Without unsafe (pre_exec) or clone3(CLONE_INTO_CGROUP) support in std/tokio, the
supervisor starts workers through a fixed POSIX shell trampoline:
/bin/sh -c 'echo $$ > "$0/cgroup.procs" && exec "$@"' <leaf> <argv...>The script is a constant; the leaf path and the user argv are positional parameters and are never interpreted by the shell, so the “no shell for user commands” rule (ADR-0008) holds. If the write fails the worker does not start (it is never run unconfined).
Limits
cpu: "N%"→cpu.max = N*1000 100000;memory→memory.maxandmemory.swap.max = swap(default 0, otherwise memory limits are absorbed by swap);pids→pids.max;io_weight→io.weight(best effort; 100 when unset);io_bandwidth→io.max rbps/wbpson every disk holding the tenant’s application roots (partitions are mapped to their disk).- Limits are written on every activation, so changing a tenant’s resources takes effect on reload without restarting its applications. Groups of removed applications/tenants are removed after their instances have drained.
- Modes:
server.cgroups.mode: auto(default; enforce when delegation works, otherwise warn once and serve without enforcement),required(refuse to start),off.
Run-time limits
POST /v1/tenants/{tenant}/limits (admin socket; scalwsctl limit, scalws-limitcpu,
scalws-limitram, scalws-limitio) overrides cpu, memory, io_weight and
io_bandwidth of one tenant at once. Overrides live in scalwsd’s memory: they are written
again after every reload and end when scalwsd restarts (under scalws-controller: when its
instances are replaced). --persist writes the value into the tenant’s resources in
scalws.yaml (block-style YAML edited in place, comments kept, previous file kept as
*.bak.<time>) and reloads, through scalws-controller when it runs.
Accounting
A sampler reads cpu.stat, memory.current, pids.current, memory.events of every
tenant and application group every 5 s into gauges (scalws_tenant_*, scalws_app_*).
GET /v1/tenants/usage?interval_ms= (scalwsctl high, scalws-highcpu,
scalws-highram, scalws-highio) samples the tenant groups twice and reports CPU (percent
of one CPU), memory, disk read/write rates (io.stat) and I/O pressure with the limits
that apply.
Request quotas (shared server work)
Per tenant: max_concurrent_requests (→ 503 + Retry-After) and
requests_per_second + burst token bucket (→ 429 + Retry-After). State is one
semaphore and one bucket per configured tenant, so memory is bounded by configuration.
Per-source and per-route limits are M5 (security); per-tenant connection limits need
SNI-time tenant attribution and are deferred.
Not done (design for later)
Per-tenant Unix users: spawn workers through setpriv --reuid --regid --clear-groups
(util-linux; std’s CommandExt::groups is unstable and uid() alone would keep root’s
supplementary groups), chown the tenant’s socket directory, check document-root ownership.
User and mount namespaces after that.
Consequences
- Delegation must be configured in production (
Delegate=yesin the systemd unit). - CPU used by scalwsd on behalf of a tenant is not attributed to the tenant’s cgroup; request quotas are the control for that, as the handoff requires.