Decision records

ADR-0047: Profile-guided release builds

Status: accepted — 2026-10-10

Context

On bare metal (2× Xeon E5-2650 v2, docs/benchmarks/2026-10-10-xeon-bare-metal-pgo.md) scalws stayed behind nginx 1.31 where little besides the server itself runs: respond 0.67×, proxy 0.71×, static-small 0.85×. Hardware counters put the gap in code layout, not in work: per respond request scalws executed 1.7× nginx’s user-space instructions but spent 2.3× its user cycles (IPC 0.53 vs 0.69), with a third more L1 instruction-cache misses. The kernel executed the same instructions for both servers but needed 24 % more cycles for scalws, whose larger code footprint evicts the kernel’s.

A profile-guided build (-Cprofile-generate, a training load, -Cprofile-use) lays out the hot paths together, moves cold branches out of line and inlines along the measured call paths. Without a code change it moved respond to 0.91×, proxy to 0.99×, static-small to 1.11× and TLS + HTTP/2 to 1.36× of nginx; user cycles per respond request fell by 44 % and kernel cycles by 14 %.

Decision

  • Release binaries of scalwsd are built with tools/pgo-build.sh: an instrumented build, a training run, llvm-profdata merge, and the optimized build. A plain cargo build --release stays valid (development, platforms without the tooling); it is only slower.
  • The training load is defined by the script, not by the benchmarks: static files of several sizes, HEAD, conditional and range requests, misses, the response cache, the reverse proxy with GET and POST bodies, fixed responses, HTTP/1.1 and HTTP/2 over TLS, PHP when PHP-FPM is installed — with both the default and the experimental HTTP/1 server and upstream client, so neither path is laid out as cold. Benchmark numbers of a PGO build are only reported for builds trained this way.
  • The profile is generated from the source being built, at build time. No profile data is committed: a stale profile silently loses its effect as code changes.
  • BOLT is an optional second pass (PGO_BOLT=1): the PGO binary is linked with --emit-relocs, run on the same training load under perf record -j any,u, and laid out again by llvm-bolt (block and function reordering, hot/cold splitting, identical code folding). On the Xeon bench it added +4 % to static-small and +3 % to proxy over PGO alone (respond unchanged). It needs last-branch records, so a bare-metal build host; release builds without one use PGO only.

Consequences

  • Release builds take about three times as long (two builds and a ~2 min training run) and need rustup component add llvm-tools, oha, openssl and curl on the build host.
  • PGO changes inlining and code layout, not semantics; the test suite runs on plain builds. Release candidates are checked with the integration tests and a smoke run of the PGO binary itself before packaging.
  • Builds are not bit-for-bit reproducible across training runs (counter values differ); the source revision and the training script identify a build.
  • BOLT rewrites the linked binary, including where landing pads live (-split-eh); tokio catches task panics by unwinding. A BOLT binary is accepted only after the benchmark scenarios (static, proxy, cache, TLS + HTTP/2, PHP) ran on it without errors.
  • Paths the training does not exercise (HTTP/3, Node/Python runtimes, admin API) are optimized as cold code: correct, but without the layout benefit. Extending the training load is the way to cover them.