Benchmarks

On bare metal against nginx 1.31.6, the latest release-mode run has ScalWS ahead in 4 of 9 scenarios, level in 2 and behind in 3. For example: 1.37× nginx 1.31.6 in the tested TLS + HTTP/2 small-file workload, 0.92× on fixed responses and 0.98× as a reverse proxy. Peak server memory in the 4 scenarios where it was recorded: ScalWS 36 to 41 MiB, nginx 1.31.6 79 to 98 MiB, OpenLiteSpeed 38 to 45 MiB (medians).

2× Xeon E5-2650 v2, server pinned to four physical cores, load generator on the other socket, oha with 64 connections; MODE=release: 10 s warm-up, 30 s measurement, 5 runs, medians. ScalWS here is a profile-guided (PGO + BOLT) release build with the experimental own HTTP/1 server and upstream client switched on; both are opt-in, not the default. The earlier VM results (2026-10-06) are superseded and kept below as history. Full methodology →

Benchmark explorer

Run
Metric

Bare metal, release mode, 2026-10-11, ScalWS vs nginx 1.31.6

  • Small static file1.13×
  • 1 MiB static file1.00×
  • Fixed response0.92×
  • Reverse proxy0.98×
  • Cache hit1.93×
  • TLS + HTTP/2 small file1.37×
  • PHP (PHP-FPM)1.86×
  • Node.js1.00×
  • Python (uvicorn)0.93×

ScalWS requests/s divided by nginx 1.31.6 requests/s, as published. Right of the 1.00× line: ScalWS faster. Symmetric log scale, so 0.80× and 1.25× are the same length.

Exact results

Bare metal, release mode, 2026-10-11latest verified

Requests/s are medians. Ratios are ScalWS divided by the peer, as published. Source: docs/benchmarks/2026-10-11-xeon-release-nginx-openlitespeed.md.
ScenarioScalWS req/snginx 1.31.6 req/sOpenLiteSpeed req/svs nginx 1.31.6vs OpenLiteSpeedScalWS defaults req/snginx 1.28.3 req/sCPU µs/req ScalWS / nginx / OLSRSS MiB ScalWS / nginx / OLSCV % ScalWSScalWS min to max
Small static file109.3k96.8k93.0k1.13×1.18×103.9k94.2k36.6 / 41.3 / 40.836 / 79 / 38n/an/a
1 MiB static file9.88k9.88k8.71k1.00×1.13×9.87k9.87kn/a / n/a / n/an/a / n/a / n/an/an/a
Fixed response196.9k214.4kn/a0.92×n/a186.2k200.3k20.3 / 18.7 / n/an/a / n/a / n/an/an/a
Reverse proxy89.7k91.7k64.9k0.98×1.38×78.6k77.1k44.5 / 43.5 / 57.636 / 83 / 39n/an/a
Cache hit163.9k84.8k45.7k1.93×3.58×159.7k81.3kn/a / n/a / n/a37 / 80 / 39n/an/a
TLS + HTTP/2 small file93.2k68.2k64.5k1.37×1.44×91.7k66.1kn/a / n/a / n/a41 / 98 / 45n/an/a
PHP (PHP-FPM)31.0k16.7k20.5k1.86×1.51×34.7k16.4kn/a / n/a / n/an/a / n/a / n/a25.017.2k to 36.8k
Node.js11.2k11.2k12.1k1.00×0.93×11.5k11.4kn/a / n/a / n/an/a / n/a / n/an/an/a
Python (uvicorn)6.10k6.54k5.94k0.93×1.03×6.01k6.53kn/a / n/a / n/an/a / n/a / n/an/an/a
Host
Bare metal: 2× Intel Xeon E5-2650 v2 (Ivy Bridge EP, 8 cores / 16 threads per socket, no AVX2), 121 GB, Ubuntu 26.04, kernel 7.0
CPU allocation
Server on CPUs 0-3 (four physical cores of socket 0, HT siblings idle) · upstreams and applications on CPUs 4-7 · load generator oha on socket 1
Tuning
performance governor, turbo off
Load
oha with 64 connections; MODE=release: 10 s warm-up, 30 s measurement, 5 runs, medians
ScalWS build
PGO + BOLT release build from tools/pgo-build.sh (ADR-0047), commit 9eaebe7. Main column: with the experimental own HTTP/1 server and upstream client enabled (ADR-0046, ADR-0045). “ScalWS defaults”: the same binary with the default hyper-based HTTP/1 paths.
Peers
nginx 1.31.6 (built from the signed source tarball; every ratio is against it) and nginx 1.28.3 (Ubuntu 26.04 package); OpenLiteSpeed 1.9.3 (official release tarball, SHA-256 checked)
Answers
Every measurement of every server answered 2xx only

Caveats

  • Fixed responses: 0.92× nginx (8 % behind). Reverse proxy: 0.98× nginx in release mode. Both remain open.
  • PHP under ScalWS varies between runs (17.2k to 36.8k req/s, CV 25 %). Root cause found the same day: the machine's firmware lowered the clock of the server cores under partial load (1197 MHz instead of 2593 MHz in slow runs) despite the performance governor; instructions and cycles per request were unchanged. Application-bound scenarios (PHP, Node.js, Python) carry this noise until the BIOS power profile is changed and they are re-measured.
  • Python is application-bound (two uvicorn workers): 0.93× nginx with the same application CPU. The same nginx + uvicorn setup measured 5.95k req/s in the follow-up versus 6.53k here, so the gap is likely within the frequency noise; to be re-measured. Node.js: OpenLiteSpeed is ahead of both ScalWS and nginx (ScalWS at 0.93× OpenLiteSpeed).
  • The 1 MiB static file is bandwidth-bound: ScalWS and nginx are equal.
  • OpenLiteSpeed has no fixed-response directive, so respond is not measured for it.
  • The own HTTP/1 server and upstream client are opt-in and experimental, not the default. With the defaults ScalWS is slower on proxy and respond (“ScalWS defaults” column).
  • Server CPU per request was reported for 3 scenarios and peak RSS for 4.

Frozen baseline (VM), 2026-10-06superseded

Superseded by the bare-metal results of 2026-10-10/11. Kept as history: the VM drifted between runs and its numbers are not comparable with them.

Requests/s are medians. Ratios are ScalWS divided by the peer, as published. Source: docs/pre-adaptive-runtime-baseline.md; results3.txt (Faz B ratios, cross-checked). Rows labelled scb in the source are ScalWS.
ScenarioScalWS req/snginx 1.24 req/svs nginx 1.24p99 ms ScalWS / nginxCPU µs/req ScalWS / nginxRSS MiB ScalWS / nginxCV % ScalWS
Small static file77.7k71.9k1.08×1.52 / 1.5425.7 / 27.833 / 469.6
1 MiB static file8.28k8.61k0.96×12.04 / 11.81241.4 / 232.133 / 425.2
Fixed response97.1k114.7k0.85×1.15 / 0.9220.6 / 17.433 / 426.1
Reverse proxy46.7k50.4k0.93×2.66 / 2.0042.7 / 39.636 / 439.2
Cache hit80.8k64.6k1.25×1.35 / 1.6124.7 / 31.036 / 437.1
TLS + HTTP/2 small file64.7k44.2k1.47×1.76 / 2.4430.8 / 45.241 / 507.9
HTTPS/1.1 1 MiB (kTLS)1.56k1.53k1.02×79.10 / 65.301280.7 / 1305.142 / 504.7
PHP (PHP-FPM)28.2k26.2k1.07×3.50 / 4.0664.3 / 68.742 / 505.2
Node.js22.2k28.1k0.79×5.62 / 4.7761.5 / 45.742 / 504.2
Python (uvicorn)12.2k11.9k1.02×11.56 / 12.2892.4 / 80.442 / 503.5
Host
VMware VM (not bare metal), AMD EPYC 7642, 8 vCPU, 31 GiB, Ubuntu 24.04, kernel 6.8.0-146
CPU allocation
Server CPUs 0-1 (2 CPUs) · applications 2-3 · load generator (oha 1.16.0) 4-7
Load
64 connections, 10 s warm-up, 30 s measurement
Runs
5 runs, median reported, server order rotated per run
Versions
ScalWS commit 71fc873 (release, mimalloc), nginx 1.24.0, PHP-FPM 8.3 (4 static workers), Node 18, uvicorn (2 workers)
File limits
NOFILE 65535 for every process
Config parity
Open file cache with a check per hit and kTLS on both servers; nginx sends the same forwarding headers
Test suite
scripts/linux-test.sh on the same host: 301 passed, 0 failed, 1 ignored (the explicit measurement tool)

Caveats

  • The same build drifts by up to ±30 % over minutes on this VM. Only interleaved A/B runs decide; absolute numbers from different times are not comparable.
  • nginx python has 4 valid runs: run 4 answered only 502 because the external uvicorn was not ready.
  • respond 0.85× and proxy 0.93×: the gap is user-space CPU; closing the proxy remainder would mean replacing hyper.
  • node 0.79×: Node.js spends more CPU per request behind ScalWS than behind nginx. Open investigation.
  • Recorded under the previous name (SCB, scb-webd). Rows labelled scb are ScalWS.

Release-mode run (VM), 2026-10-06superseded

Superseded by the bare-metal results of 2026-10-10/11. Kept as history: the VM drifted between runs and its numbers are not comparable with them.

Requests/s are medians. Ratios are ScalWS divided by the peer, as published. Source: results2.txt. Rows labelled scb in the source are ScalWS.
ScenarioScalWS req/snginx 1.24 req/svs nginx 1.24CPU µs/req ScalWS / nginxCV % ScalWS
Small static file73.7k74.5k0.99×27.1 / 26.81.6
1 MiB static file7.90k9.00k0.88×n/a / n/a3.4
Fixed response97.4k111.8k0.87×20.5 / 17.93.8
Reverse proxy47.4k52.8k0.90×42.1 / 37.92.2
Cache hit82.4k69.0k1.19×24.3 / 28.97.9
TLS + HTTP/2 small file64.2k44.6k1.44×31.1 / 44.83.4
HTTPS/1.1 1 MiB (kTLS)1.57k1.56k1.00×n/a / n/a1.4
PHP (PHP-FPM)29.7k26.8k1.11×59.8 / 68.35.6
Node.js21.8k26.0k0.84×n/a / n/a3.4
Python (uvicorn)11.6k12.2k0.96×n/a / n/a3.8

Peak RSS range: ScalWS 33 to 42 MiB, nginx 42 to 50 MiB.

Host
VMware VM (not bare metal), AMD EPYC 7642, 8 vCPU, 31 GiB, Ubuntu 24.04, kernel 6.8.0-146
CPU allocation
Server CPUs 0-1 (2 CPUs) · applications 2-3 · load generator (oha 1.16.0) 4-7
Load
64 connections, 10 s warm-up, 30 s measurement
Runs
5 runs, median reported, server order rotated per run
File limits
NOFILE 65535 for every process
Config parity
Open file cache and kTLS on both servers; nginx sends the same X-Forwarded-* and X-Request-Id headers

Caveats

  • Earlier build than the frozen baseline. Superseded by it.
  • CPU/request was reported for six scenarios only.
  • The cache-hit drop against the morning run was checked with an A/B (86.4k vs 85.3k): host variance, not a regression.

First server run with Caddy (VM), 2026-10-06superseded

Superseded by the bare-metal results of 2026-10-10/11. Kept as history: the VM drifted between runs and its numbers are not comparable with them.

Requests/s are medians. Ratios are ScalWS divided by the peer, as published. Source: results.txt. Rows labelled scb in the source are ScalWS.
ScenarioScalWS req/snginx 1.24 req/sCaddy req/svs nginx 1.24vs CaddyScalWS min to max
Small static file62.7k73.7k22.7k0.85×2.8×62.1k to 66.2k
1 MiB static file8.10k8.80k6.20k0.92×1.3×7.70k to 8.20k
Fixed response94.5k102.9k36.0k0.92×2.6×91.5k to 95.8k
Reverse proxy49.1k55.8k16.4k0.88×3.0×46.3k to 49.1k
Cache hit91.4k68.5k18.3k1.33×5.0×80.6k to 93.6k
TLS + HTTP/2 small file61.3k45.4k13.4k1.35×4.6×57.4k to 64.9k
HTTPS/1.1 1 MiB (kTLS)1.47k1.57k1.19k0.94×1.2×1.45k to 1.61k
PHP (PHP-FPM)22.2k26.2k8.00k0.85×2.8×21.6k to 23.0k
Node.js19.7k29.1k16.7k0.68×1.2×17.4k to 20.8k
Python (uvicorn)12.6k13.6k9.60k0.93×1.3×11.3k to 12.8k

Peak RSS range: ScalWS 33 to 40 MiB, nginx 41 to 49 MiB, Caddy 44 to 58 MiB.

Host
VMware VM (not bare metal), AMD EPYC 7642, 8 vCPU, 31 GiB, Ubuntu 24.04, kernel 6.8.0-146
CPU allocation
Server CPUs 0-1 (2 CPUs) · applications 2-3 · load generator (oha 1.16.0) 4-7 (ScalWS, nginx and Caddy on the same 2 CPUs)
Load
64 connections, 3 s warm-up, 10 s measurement
Runs
3 runs, median reported
Versions
nginx 1.24.0, Caddy 2.6.2, PHP-FPM 8.3, Node 18, uvicorn

Caveats

  • Short runs (3 × 10 s). nginx ran with a file-descriptor limit of 1024 and sent fewer forwarding headers to Node/Python/proxy upstreams; both were corrected in later runs.
  • The Node.js 0.68× here was later shown to be run-to-run noise (8 × 30 s A/B: 29.9k vs 29.8k).
  • The only run that includes Caddy.

Follow-up on the open points (same day)

Not part of the release-mode table above, which stays the reference. These are a quick A/B and diagnoses from the same machine on 2026-10-11.

Source: docs/benchmarks/2026-10-11-xeon-release-nginx-openlitespeed.md, “Follow-up on the open points”.

Where the gap to nginx was, and what closed it

Before profile-guided builds, ScalWS was behind nginx on bare metal wherever little besides the server runs. Hardware counters put the gap in code layout, not in work: instruction-cache misses on ScalWS’s larger code footprint, which also slowed the kernel’s share of each request. A profile-guided build lays the hot paths out together; with PGO, user cycles per fixed-response request fell by 44 % and kernel cycles by 14 %. BOLT is an optional second pass over the linked binary.

ScalWS (experimental HTTP/1 server and upstream client) divided by nginx 1.31 requests/s on the same bare-metal host. The first three columns are MODE=quick (3 runs, spread about 1 %); the last is the release-mode run above (5 runs × 30 s). Sources: docs/benchmarks/2026-10-10-xeon-bare-metal-pgo.md and the 2026-10-11 report; the before/after ratios are cross-checked against ADR-0047.
ScenarioNo PGO (commit 4415c31)PGO, benchmark as trainingPGO + BOLT, own training loadPGO + BOLT, release mode
Small static file0.85×1.11×1.12×1.13×
1 MiB static file1.00×1.00×1.00×1.00×
Fixed response0.67×0.91×0.96×0.92×
Reverse proxy0.71×0.99×1.09×0.98×
Cache hit1.28×1.85×1.89×1.93×
TLS + HTTP/2 small file0.88×1.36×1.39×1.37×
PHP (PHP-FPM)1.72×1.83×1.89×1.86×

Release binaries are built with tools/pgo-build.sh; its training load is defined by the script, not by the benchmarks, and exercises both the default and the experimental HTTP/1 paths. A plain cargo build --release stays valid, only slower. Paths the training does not exercise (HTTP/3, Node.js/Python runtimes, admin API) get no layout benefit. ADR-0047 →

Soak

The same PGO + BOLT build with both experimental settings, all scenarios at once (16 connections each, including HTTP/3), on the bare-metal host.

PASS Overnight, 8 h 2026-10-10/11

Requests
789 M
Errors / non-2xx
0 / 0
Server RSS
63.7 MiB at start and end (slope -0.23 MiB/h)
Descriptors
312 throughout
HTTP/3
No result: the HTTP/3 load generator ran out of memory at the end (it keeps every result). The harness now runs the load in segments.

PASS Segmented, including HTTP/3, 2 h 2026-10-11

Requests
285 M
Errors / non-2xx
0 / 0
Server RSS
59.7 → 62.4 MiB
Descriptors
321 → 337
HTTP/3
76.7 M requests, 10.6k req/s, p99 5.5 ms

Source: docs/benchmarks/2026-10-11-xeon-release-nginx-openlitespeed.md (gate 6, part 1).

Scenarios

static-small Small static file
GET /hello.txt (13 B), HTTP/1.1 keep-alive
static-1m 1 MiB static file
GET /1m.bin (1 MiB)
respond Fixed response
GET /respond (13 B, in memory)
proxy Reverse proxy
GET /api/x to an upstream (13 B)
cache-hit Cache hit
Upstream sends max-age=3600; served from the response cache
tls-h2-static-small TLS + HTTP/2 small file
GET /hello.txt over TLS 1.3 + HTTP/2
tls-static-1m HTTPS/1.1 1 MiB (kTLS)
GET /1m.bin over TLS, kernel TLS on both servers
php PHP (PHP-FPM)
Identical PHP-FPM pool behind each server (PHP-FPM 8.3 on the VM runs)
node Node.js
Identical Node.js server behind each server (Node 18 on the VM runs)
python Python (uvicorn)
Identical uvicorn server behind each server

Load generator oha at fixed concurrency, on CPUs the servers never share. Access logs off everywhere, upstream keep-alive on everywhere. Security and correctness features stay on for ScalWS; nothing is disabled to improve numbers. The HTTPS 1 MiB (kTLS) scenario was measured on the VM only. Benchmark harness →

Adaptive Runtime scaling

Measured on the VM (2026-10-07), in a separate topology from the comparisons above: instance i on CPU i, upstream and PHP-FPM on CPU 3, oha on CPUs 4 to 7. More instances means more CPUs; read the efficiency columns before the throughput. Not yet repeated on bare metal.

Burst, idle to 1024 connections Autoscaling in web-tier mode, min 1, max 3, one CPU per instance. Three bursts in a row, zero errors.

1 instance · 1 CPU

req/s 37.2k
p99 46 ms

2 instances · 2 CPUs

req/s 81.5k
p99 19 ms

3 instances · 3 CPUs

req/s 143.2k
p99 14 ms

Bars share one axis per metric, starting at zero. Lower p99 is better.

D1: N instances on SO_REUSEPORT vs one instance. 3 runs × 10 s, medians, one CPU per instance. Efficiency = rps(N) / (rps(1) × N). “vs 1 instance” compares with one instance given the same number of event loops and CPUs.
Scenario (connections)1 instance2 instancesefficiencyvs 1 instance, 2 loops3 instancesefficiencyvs 1 instance, 3 loops
respond (64)51.4k92.5k0.901.04×166.8k1.081.14×
respond (256)52.4k102.6k0.981.08×159.0k1.010.97×
static-small (64)43.5k75.9k0.870.92×127.3k0.971.02×
static-small (256)43.2k82.2k0.951.00×144.2k1.111.08×
static-1m (64)4.22k8.96k1.061.07×13.1k1.031.06×
cache-hit (64)46.6k87.9k0.940.95×152.3k1.091.05×
TLS + HTTP/2 (256)34.0k65.4k0.960.92×115.8k1.131.08×
proxy (256)25.4k47.3k0.931.08×65.6k0.861.06×
php (64)13.3k15.4k0.580.95×15.8k0.401.04×
D5: concurrency sweep with autoscaling on (web-tier mode, min 1, max 3), 20 s per point. Instances = active instances at the end of the point.
Pathconnectionsreq/sp99 mserrorsinstancesRSS MiB
respond843,6080.280126
respond3249,2661.040254
respond6499,1091.100381
respond128143,3171.480382
respond256154,6183.000384
respond512158,8285.380389
respond1024167,00211.2403102
static-small860,2420.250379
static-small64127,3460.950380
static-small1024121,02315.180398

Host: development VM (8 vCPU EPYC 7642, Ubuntu 24.04, kernel 6.8), tcp_migrate_req=1. Topology separate from the frozen 2-CPU baseline: instance i on CPU i (one event loop each), upstream and PHP-FPM (4 workers) on CPU 3, oha on CPUs 4-7.