On bare metal against nginx 1.31.6, the latest release-mode run has ScalWS ahead in 4 of 9 scenarios, level in 2 and behind in 3. For example: 1.37× nginx 1.31.6 in the tested TLS + HTTP/2 small-file workload, 0.92× on fixed responses and 0.98× as a reverse proxy. Peak server memory in the 4 scenarios where it was recorded: ScalWS 36 to 41 MiB, nginx 1.31.6 79 to 98 MiB, OpenLiteSpeed 38 to 45 MiB (medians).
2× Xeon E5-2650 v2, server pinned to four physical cores, load generator on the other socket, oha with 64 connections; MODE=release: 10 s warm-up, 30 s measurement, 5 runs, medians. ScalWS here is a profile-guided (PGO + BOLT) release build with the experimental own HTTP/1 server and upstream client switched on; both are opt-in, not the default. The earlier VM results (2026-10-06) are superseded and kept below as history. Full methodology →
Benchmark explorer
Bare metal, release mode, 2026-10-11, ScalWS vs nginx 1.31.6
Small static file1.13×
1 MiB static file1.00×
Fixed response0.92×
Reverse proxy0.98×
Cache hit1.93×
TLS + HTTP/2 small file1.37×
PHP (PHP-FPM)1.86×
Node.js1.00×
Python (uvicorn)0.93×
0.67×0.80×1.00×1.25×1.50×
ScalWS requests/s divided by nginx 1.31.6 requests/s, as published. Right of the 1.00× line: ScalWS faster. Symmetric log scale, so 0.80× and 1.25× are the same length.
Exact results
Bare metal, release mode, 2026-10-11latest verified
Requests/s are medians. Ratios are ScalWS divided by the peer, as published. Source: docs/benchmarks/2026-10-11-xeon-release-nginx-openlitespeed.md.
Scenario
ScalWS req/s
nginx 1.31.6 req/s
OpenLiteSpeed req/s
vs nginx 1.31.6
vs OpenLiteSpeed
ScalWS defaults req/s
nginx 1.28.3 req/s
CPU µs/req ScalWS / nginx / OLS
RSS MiB ScalWS / nginx / OLS
CV % ScalWS
ScalWS min to max
Small static file
109.3k
96.8k
93.0k
1.13×
1.18×
103.9k
94.2k
36.6 / 41.3 / 40.8
36 / 79 / 38
n/a
n/a
1 MiB static file
9.88k
9.88k
8.71k
1.00×
1.13×
9.87k
9.87k
n/a / n/a / n/a
n/a / n/a / n/a
n/a
n/a
Fixed response
196.9k
214.4k
n/a
0.92×
n/a
186.2k
200.3k
20.3 / 18.7 / n/a
n/a / n/a / n/a
n/a
n/a
Reverse proxy
89.7k
91.7k
64.9k
0.98×
1.38×
78.6k
77.1k
44.5 / 43.5 / 57.6
36 / 83 / 39
n/a
n/a
Cache hit
163.9k
84.8k
45.7k
1.93×
3.58×
159.7k
81.3k
n/a / n/a / n/a
37 / 80 / 39
n/a
n/a
TLS + HTTP/2 small file
93.2k
68.2k
64.5k
1.37×
1.44×
91.7k
66.1k
n/a / n/a / n/a
41 / 98 / 45
n/a
n/a
PHP (PHP-FPM)
31.0k
16.7k
20.5k
1.86×
1.51×
34.7k
16.4k
n/a / n/a / n/a
n/a / n/a / n/a
25.0
17.2k to 36.8k
Node.js
11.2k
11.2k
12.1k
1.00×
0.93×
11.5k
11.4k
n/a / n/a / n/a
n/a / n/a / n/a
n/a
n/a
Python (uvicorn)
6.10k
6.54k
5.94k
0.93×
1.03×
6.01k
6.53k
n/a / n/a / n/a
n/a / n/a / n/a
n/a
n/a
Host
Bare metal: 2× Intel Xeon E5-2650 v2 (Ivy Bridge EP, 8 cores / 16 threads per socket, no AVX2), 121 GB, Ubuntu 26.04, kernel 7.0
CPU allocation
Server on CPUs 0-3 (four physical cores of socket 0, HT siblings idle) · upstreams and applications on CPUs 4-7 · load generator oha on socket 1
Tuning
performance governor, turbo off
Load
oha with 64 connections; MODE=release: 10 s warm-up, 30 s measurement, 5 runs, medians
ScalWS build
PGO + BOLT release build from tools/pgo-build.sh (ADR-0047), commit 9eaebe7. Main column: with the experimental own HTTP/1 server and upstream client enabled (ADR-0046, ADR-0045). “ScalWS defaults”: the same binary with the default hyper-based HTTP/1 paths.
Peers
nginx 1.31.6 (built from the signed source tarball; every ratio is against it) and nginx 1.28.3 (Ubuntu 26.04 package); OpenLiteSpeed 1.9.3 (official release tarball, SHA-256 checked)
Answers
Every measurement of every server answered 2xx only
Caveats
Fixed responses: 0.92× nginx (8 % behind). Reverse proxy: 0.98× nginx in release mode. Both remain open.
PHP under ScalWS varies between runs (17.2k to 36.8k req/s, CV 25 %). Root cause found the same day: the machine's firmware lowered the clock of the server cores under partial load (1197 MHz instead of 2593 MHz in slow runs) despite the performance governor; instructions and cycles per request were unchanged. Application-bound scenarios (PHP, Node.js, Python) carry this noise until the BIOS power profile is changed and they are re-measured.
Python is application-bound (two uvicorn workers): 0.93× nginx with the same application CPU. The same nginx + uvicorn setup measured 5.95k req/s in the follow-up versus 6.53k here, so the gap is likely within the frequency noise; to be re-measured. Node.js: OpenLiteSpeed is ahead of both ScalWS and nginx (ScalWS at 0.93× OpenLiteSpeed).
The 1 MiB static file is bandwidth-bound: ScalWS and nginx are equal.
OpenLiteSpeed has no fixed-response directive, so respond is not measured for it.
The own HTTP/1 server and upstream client are opt-in and experimental, not the default. With the defaults ScalWS is slower on proxy and respond (“ScalWS defaults” column).
Server CPU per request was reported for 3 scenarios and peak RSS for 4.
Frozen baseline (VM), 2026-10-06superseded
Superseded by the bare-metal results of 2026-10-10/11. Kept as history: the VM drifted between runs and its numbers are not comparable with them.
Requests/s are medians. Ratios are ScalWS divided by the peer, as published. Source: docs/pre-adaptive-runtime-baseline.md; results3.txt (Faz B ratios, cross-checked). Rows labelled scb in the source are ScalWS.
Scenario
ScalWS req/s
nginx 1.24 req/s
vs nginx 1.24
p99 ms ScalWS / nginx
CPU µs/req ScalWS / nginx
RSS MiB ScalWS / nginx
CV % ScalWS
Small static file
77.7k
71.9k
1.08×
1.52 / 1.54
25.7 / 27.8
33 / 46
9.6
1 MiB static file
8.28k
8.61k
0.96×
12.04 / 11.81
241.4 / 232.1
33 / 42
5.2
Fixed response
97.1k
114.7k
0.85×
1.15 / 0.92
20.6 / 17.4
33 / 42
6.1
Reverse proxy
46.7k
50.4k
0.93×
2.66 / 2.00
42.7 / 39.6
36 / 43
9.2
Cache hit
80.8k
64.6k
1.25×
1.35 / 1.61
24.7 / 31.0
36 / 43
7.1
TLS + HTTP/2 small file
64.7k
44.2k
1.47×
1.76 / 2.44
30.8 / 45.2
41 / 50
7.9
HTTPS/1.1 1 MiB (kTLS)
1.56k
1.53k
1.02×
79.10 / 65.30
1280.7 / 1305.1
42 / 50
4.7
PHP (PHP-FPM)
28.2k
26.2k
1.07×
3.50 / 4.06
64.3 / 68.7
42 / 50
5.2
Node.js
22.2k
28.1k
0.79×
5.62 / 4.77
61.5 / 45.7
42 / 50
4.2
Python (uvicorn)
12.2k
11.9k
1.02×
11.56 / 12.28
92.4 / 80.4
42 / 50
3.5
Host
VMware VM (not bare metal), AMD EPYC 7642, 8 vCPU, 31 GiB, Ubuntu 24.04, kernel 6.8.0-146
Open file cache with a check per hit and kTLS on both servers; nginx sends the same forwarding headers
Test suite
scripts/linux-test.sh on the same host: 301 passed, 0 failed, 1 ignored (the explicit measurement tool)
Caveats
The same build drifts by up to ±30 % over minutes on this VM. Only interleaved A/B runs decide; absolute numbers from different times are not comparable.
nginx python has 4 valid runs: run 4 answered only 502 because the external uvicorn was not ready.
respond 0.85× and proxy 0.93×: the gap is user-space CPU; closing the proxy remainder would mean replacing hyper.
node 0.79×: Node.js spends more CPU per request behind ScalWS than behind nginx. Open investigation.
Recorded under the previous name (SCB, scb-webd). Rows labelled scb are ScalWS.
Release-mode run (VM), 2026-10-06superseded
Superseded by the bare-metal results of 2026-10-10/11. Kept as history: the VM drifted between runs and its numbers are not comparable with them.
Requests/s are medians. Ratios are ScalWS divided by the peer, as published. Source: results2.txt. Rows labelled scb in the source are ScalWS.
Scenario
ScalWS req/s
nginx 1.24 req/s
vs nginx 1.24
CPU µs/req ScalWS / nginx
CV % ScalWS
Small static file
73.7k
74.5k
0.99×
27.1 / 26.8
1.6
1 MiB static file
7.90k
9.00k
0.88×
n/a / n/a
3.4
Fixed response
97.4k
111.8k
0.87×
20.5 / 17.9
3.8
Reverse proxy
47.4k
52.8k
0.90×
42.1 / 37.9
2.2
Cache hit
82.4k
69.0k
1.19×
24.3 / 28.9
7.9
TLS + HTTP/2 small file
64.2k
44.6k
1.44×
31.1 / 44.8
3.4
HTTPS/1.1 1 MiB (kTLS)
1.57k
1.56k
1.00×
n/a / n/a
1.4
PHP (PHP-FPM)
29.7k
26.8k
1.11×
59.8 / 68.3
5.6
Node.js
21.8k
26.0k
0.84×
n/a / n/a
3.4
Python (uvicorn)
11.6k
12.2k
0.96×
n/a / n/a
3.8
Peak RSS range: ScalWS 33 to 42 MiB, nginx 42 to 50 MiB.
Host
VMware VM (not bare metal), AMD EPYC 7642, 8 vCPU, 31 GiB, Ubuntu 24.04, kernel 6.8.0-146
Short runs (3 × 10 s). nginx ran with a file-descriptor limit of 1024 and sent fewer forwarding headers to Node/Python/proxy upstreams; both were corrected in later runs.
The Node.js 0.68× here was later shown to be run-to-run noise (8 × 30 s A/B: 29.9k vs 29.8k).
The only run that includes Caddy.
Follow-up on the open points (same day)
Not part of the release-mode table above, which stays the reference. These are a quick A/B and diagnoses from the same machine on 2026-10-11.
Reverse proxy: quick A/B after commit 989f486
The own HTTP/1 server now keeps the per-request pipeline future in the connection’s state instead of copying it into a fresh allocation. In a 3-run quick A/B, proxy went from 88.6k to 95.1k req/s (CPU 45 → 41.9 µs per request), above the 91.7k req/s of nginx 1.31.6 in the release-mode table; static-small +3 %. Fixed responses are unchanged at about 0.91× nginx. Quick mode (3 runs) is not release mode; the release-mode number stays 89.7k req/s until it is re-measured.
PHP run-to-run variance: the machine, not ScalWS
In slow runs, instructions and cycles per request were unchanged for both ScalWS and PHP-FPM, but each took twice the time: the server cores ran at 1,197 MHz instead of 2,593 MHz (1,795 MHz in a medium run). The firmware’s power regulator lowers the clock under partial load despite the performance governor. Scenarios that saturate the server cores are not affected; the application-bound ones (PHP, Node.js, Python) carry this noise until the BIOS power profile is changed and they are re-measured.
PHP: ScalWS’s own cost
The FastCGI client zeroed a 64 KiB buffer twice per request (about 15 % of ScalWS’s cycles on the PHP path). The vendored fix starts the buffer at 4 KiB: 197k → 167k cycles per PHP request. Not yet in a release-mode run.
Python
uvicorn receives the same headers and makes the same syscalls per request behind ScalWS and nginx. The same nginx + uvicorn setup measured about 5.95k req/s in the follow-up versus 6.53k in the release run, so the 7 % gap is likely within the frequency noise. To be re-measured.
Source: docs/benchmarks/2026-10-11-xeon-release-nginx-openlitespeed.md, “Follow-up on the open points”.
Where the gap to nginx was, and what closed it
Before profile-guided builds, ScalWS was behind nginx on bare metal wherever little besides the server runs. Hardware counters put the gap in code layout, not in work: instruction-cache misses on ScalWS’s larger code footprint, which also slowed the kernel’s share of each request. A profile-guided build lays the hot paths out together; with PGO, user cycles per fixed-response request fell by 44 % and kernel cycles by 14 %. BOLT is an optional second pass over the linked binary.
ScalWS (experimental HTTP/1 server and upstream client) divided by nginx 1.31 requests/s on the same bare-metal host. The first three columns are MODE=quick (3 runs, spread about 1 %); the last is the release-mode run above (5 runs × 30 s). Sources: docs/benchmarks/2026-10-10-xeon-bare-metal-pgo.md and the 2026-10-11 report; the before/after ratios are cross-checked against ADR-0047.
Scenario
No PGO (commit 4415c31)
PGO, benchmark as training
PGO + BOLT, own training load
PGO + BOLT, release mode
Small static file
0.85×
1.11×
1.12×
1.13×
1 MiB static file
1.00×
1.00×
1.00×
1.00×
Fixed response
0.67×
0.91×
0.96×
0.92×
Reverse proxy
0.71×
0.99×
1.09×
0.98×
Cache hit
1.28×
1.85×
1.89×
1.93×
TLS + HTTP/2 small file
0.88×
1.36×
1.39×
1.37×
PHP (PHP-FPM)
1.72×
1.83×
1.89×
1.86×
Release binaries are built with tools/pgo-build.sh; its training load is defined by the script, not by the benchmarks, and exercises both the default and the experimental HTTP/1 paths. A plain cargo build --release stays valid, only slower. Paths the training does not exercise (HTTP/3, Node.js/Python runtimes, admin API) get no layout benefit. ADR-0047 →
Soak
The same PGO + BOLT build with both experimental settings, all scenarios at once (16 connections each, including HTTP/3), on the bare-metal host.
PASS Overnight, 8 h 2026-10-10/11
Requests
789 M
Errors / non-2xx
0 / 0
Server RSS
63.7 MiB at start and end (slope -0.23 MiB/h)
Descriptors
312 throughout
HTTP/3
No result: the HTTP/3 load generator ran out of memory at the end (it keeps every result). The harness now runs the load in segments.
PASS Segmented, including HTTP/3, 2 h 2026-10-11
Requests
285 M
Errors / non-2xx
0 / 0
Server RSS
59.7 → 62.4 MiB
Descriptors
321 → 337
HTTP/3
76.7 M requests, 10.6k req/s, p99 5.5 ms
Source: docs/benchmarks/2026-10-11-xeon-release-nginx-openlitespeed.md (gate 6, part 1).
Scenarios
static-small Small static file
GET /hello.txt (13 B), HTTP/1.1 keep-alive
static-1m 1 MiB static file
GET /1m.bin (1 MiB)
respond Fixed response
GET /respond (13 B, in memory)
proxy Reverse proxy
GET /api/x to an upstream (13 B)
cache-hit Cache hit
Upstream sends max-age=3600; served from the response cache
tls-h2-static-small TLS + HTTP/2 small file
GET /hello.txt over TLS 1.3 + HTTP/2
tls-static-1m HTTPS/1.1 1 MiB (kTLS)
GET /1m.bin over TLS, kernel TLS on both servers
php PHP (PHP-FPM)
Identical PHP-FPM pool behind each server (PHP-FPM 8.3 on the VM runs)
node Node.js
Identical Node.js server behind each server (Node 18 on the VM runs)
python Python (uvicorn)
Identical uvicorn server behind each server
Load generator oha at fixed concurrency, on CPUs the servers never share. Access logs off everywhere, upstream keep-alive on everywhere. Security and correctness features stay on for ScalWS; nothing is disabled to improve numbers. The HTTPS 1 MiB (kTLS) scenario was measured on the VM only. Benchmark harness →
Adaptive Runtime scaling
Measured on the VM (2026-10-07), in a separate topology from the comparisons above: instance i on CPU i, upstream and PHP-FPM on CPU 3, oha on CPUs 4 to 7. More instances means more CPUs; read the efficiency columns before the throughput. Not yet repeated on bare metal.
Burst, idle to 1024 connectionsAutoscaling in web-tier mode, min 1, max 3, one CPU per instance. Three bursts in a row, zero errors.
1 instance · 1 CPU
req/s37.2k
p9946 ms
2 instances · 2 CPUs
req/s81.5k
p9919 ms
3 instances · 3 CPUs
req/s143.2k
p9914 ms
Bars share one axis per metric, starting at zero. Lower p99 is better.
D1: N instances on SO_REUSEPORT vs one instance. 3 runs × 10 s, medians, one CPU per instance. Efficiency = rps(N) / (rps(1) × N). “vs 1 instance” compares with one instance given the same number of event loops and CPUs.
Scenario (connections)
1 instance
2 instances
efficiency
vs 1 instance, 2 loops
3 instances
efficiency
vs 1 instance, 3 loops
respond (64)
51.4k
92.5k
0.90
1.04×
166.8k
1.08
1.14×
respond (256)
52.4k
102.6k
0.98
1.08×
159.0k
1.01
0.97×
static-small (64)
43.5k
75.9k
0.87
0.92×
127.3k
0.97
1.02×
static-small (256)
43.2k
82.2k
0.95
1.00×
144.2k
1.11
1.08×
static-1m (64)
4.22k
8.96k
1.06
1.07×
13.1k
1.03
1.06×
cache-hit (64)
46.6k
87.9k
0.94
0.95×
152.3k
1.09
1.05×
TLS + HTTP/2 (256)
34.0k
65.4k
0.96
0.92×
115.8k
1.13
1.08×
proxy (256)
25.4k
47.3k
0.93
1.08×
65.6k
0.86
1.06×
php (64)
13.3k
15.4k
0.58
0.95×
15.8k
0.40
1.04×
D5: concurrency sweep with autoscaling on (web-tier mode, min 1, max 3), 20 s per point. Instances = active instances at the end of the point.
Path
connections
req/s
p99 ms
errors
instances
RSS MiB
respond
8
43,608
0.28
0
1
26
respond
32
49,266
1.04
0
2
54
respond
64
99,109
1.10
0
3
81
respond
128
143,317
1.48
0
3
82
respond
256
154,618
3.00
0
3
84
respond
512
158,828
5.38
0
3
89
respond
1024
167,002
11.24
0
3
102
static-small
8
60,242
0.25
0
3
79
static-small
64
127,346
0.95
0
3
80
static-small
1024
121,023
15.18
0
3
98
Host: development VM (8 vCPU EPYC 7642, Ubuntu 24.04, kernel 6.8), tcp_migrate_req=1. Topology separate from the frozen 2-CPU baseline: instance i on CPU i (one event loop each), upstream and PHP-FPM (4 workers) on CPU 3, oha on CPUs 4-7.