Decision records

ADR-0025: Optional eBPF enrichment — per-application network accounting

  • Status: Accepted (2026-10-06)

Context

Handoff: eBPF (Aya) only where it materially improves diagnosis, with a non-eBPF fallback for everything important; never make basic diagnosis depend on it. CPU, memory and I/O pressure per tenant/application already come from cgroup v2 (PSI, ADR-0015). What cgroup v2 does not offer is network traffic per cgroup.

Decision

  • Crate scalws-ebpf (Linux; cargo feature ebpf, off by default): two cgroup_skb programs (ingress, egress) in C (bpf/net.bpf.c, compiled with clang at build time), loaded with Aya (stable Rust), attached once to <cgroups base>/tenants as BPF links. They count bytes and packets per socket-owner cgroup (bpf_skb_cgroup_id, i.e. the application’s leaf group) in a per-CPU hash map and always allow the packet.
  • No unsafe in our code: map keys/values are plain u64 / [u64; 2].
  • server.ebpf.network: true (requires cgroup enforcement) publishes scalws_app_network_bytes / scalws_app_network_packets{tenant,app,direction} every 5 s.
  • Failure to load or attach (old kernel, missing CAP_BPF/CAP_NET_ADMIN) is a warning; the server runs unchanged. A build without the feature refuses the setting at start.
  • Links require attach flags 0 (found in testing: AllowMultiple → EINVAL); links coexist with other cgroup programs anyway.

Limits

  • Unix-socket traffic (scalws ↔ applications by default) is not IP traffic and is not counted; the metric covers the applications’ own TCP/UDP (upstream APIs, databases, listen: tcp).
  • Process and block-I/O correlation via eBPF are not implemented; PSI covers pressure.
  • The packaged systemd unit does not grant CAP_BPF; enabling it is an operator decision.

Verification

Kernel test (crates/scalws-ebpf/tests/net.rs: a process in a test cgroup downloads 100 KiB and is counted) and end-to-end test (tests/integration/tests/ebpf.rs) in the privileged Linux test container (WSL2 kernel 6.18 with BTF).