Backend Sweep Results#
GLASS ships overlapping thread-, warp-, block-, and vendor-backed execution
tiers. For operations present in more than one tier, the fastest choice depends
on operation, size, dtype, launch shape, and batch regime. The mega sweep
(bench/tune.py’s ladder leg) compares supported contenders — this is
exactly the measurement behind glass-defaults.cuh’s recommend<>()
(see Execution Plans (glass-defaults.cuh)).
Warning
This page is a dated measurement archive, not a universal performance promise. Several legacy sections predate the current warm-up, trial-spread, source-digest, and signed-receipt protocol and are marked accordingly. Use the source capture named beside a claim, and rerun on the target machine.
The figures and table below come from the 2026-08-30 RTX 5090 / sm_120 five-backend sweep, shown across three batch regimes — NPROB=64 (low batch, latency-leaning), NPROB=1024 (mid), and NPROB=8192 (throughput). The committed dispatch table uses a separate 500-repetition throughput replication; the figures use the full three-regime capture. The winner can shift with batch size, so low-batch plots describe that workload rather than overriding the throughput table. These are committed static assets — regenerate them for your own hardware with:
python bench/tune.py --sm auto # remeasures + regenerates tables AND figures
# or just the figures from an existing sweep .txt:
python bench/export_sweep_figures.py bench/mega_sweep_*.txt
bench/explore_sweep.ipynb is the interactive version of the same analysis.
The ladder — ns/problem vs N, per backend#
Lower is faster. Each subplot is one op; the curves are native warp
(green), block (blue), and thread (orange), plus NVIDIA block
(red) and thread (purple). Both thread curves pack one problem per CUDA
thread. The native thread implementation is instantiated through N=64 in
this harness; NVIDIA thread LAPACK is instantiated through N=32. A stopped
curve therefore marks an explicit instantiation or resource boundary, not a
claim that larger sizes are unsupported by CUDA or cuSOLVERDx. NVIDIA block
f64 curves may stop at their shared-memory feasibility limit.
recommend<>() is keyed on the NPROB=8192 throughput regime.
Its policy first applies a ±2% tie band among dependency-free SIMT contenders,
then requires an NVIDIA contender to clear the best native choice by more than
5%. For destructive POTRF/TRSV/POSV, every displayed and selected candidate is
remeasured by the unified solver ladder on fresh valid inputs. The plans are
randomized within paired rounds and every raw sample is retained; all native
and NVIDIA candidates use the same selection rule. These rules keep small,
noisy, or mutation-dependent gaps out of the generated table.
float32#
float64#
Raw ladder winner per (op, N), per regime#
The policy winner at each (op, N) is listed for all three NPROB regimes.
Non-destructive rows come from the general ladder; POTRF/TRSV/POSV rows come
from its fresh-input solver companion. Thus the plots, winner list, and shipped
solver defaults all consume the same authoritative measurements.
The broad high-batch shape is mixed by design: native thread dominates many
small packed problems, native warp/block remain important as work grows, and
the NVIDIA block and thread implementations take measured factor/solve bands.
The NPROB=8192 block shown here is from the full capture; the generated
sm_120 table uses its higher-repetition replication.
NPROB=8192 f32 winner by op x N:
op 4 6 8 12 16 24 32 48 64 96 128
dot thread thread thread thread warp thread warp warp warp warp warp
gemv thread thread warp warp warp warp warp block warp warp warp
gemm warp warp warp warp warp block nvidia block block block block
potrf thread nvidia_thread nvidia_thread warp warp warp nvidia nvidia nvidia nvidia nvidia
trsv thread thread thread thread thread nvidia_thread nvidia warp warp warp warp
posv thread thread nvidia_thread thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia
NPROB=8192 f64 winner by op x N:
op 4 6 8 12 16 24 32 48 64 96 128
dot thread thread thread thread thread thread thread warp warp warp warp
gemv thread thread warp warp warp warp warp warp warp warp warp
gemm warp warp warp block block block block block block block block
potrf nvidia_thread nvidia_thread nvidia_thread thread thread thread block block block block block
trsv nvidia_thread nvidia_thread nvidia_thread nvidia_thread nvidia_thread nvidia_thread nvidia_thread nvidia warp warp warp
posv nvidia_thread nvidia_thread nvidia_thread thread thread thread nvidia block block block block
vs. host-batched cuBLAS/cuSOLVER (and TF32)#
The ladder above compares device-side backends. A separate question is how
one-block-per-problem GLASS compares to the standard host-side recipe — a
single cublas<t>gemmStridedBatched / cusolverDn<t>potrfBatched
(+ potrsBatched) call over the whole batch.
bench/bench_paper_hostblas.cu measures exactly that: gemm / potrf / posv,
N = 4–64, batch B = 1–8192, both precisions (raw capture committed as
paper_hostblas_20260708_0054.txt (archived externally); RTX 5090 / sm_120, quiet GPU).
Above 1.0 = the best swept GLASS configuration faster (fp32 shown; GLASS is
best of block32, block128, and warp8 where available). The vendor side is one
default host-API configuration, so this is not a symmetric tuner comparison or
a universal “GLASS vs vendor” statement. In this historical capture, at robot
sizes, host batching never catches up: gemm at N ≤ 24 and the full
factor-and-solve (posv) through N = 64 are GLASS wins at every batch
size, reaching 2.9–6.3× at saturation. The vendor’s best regime is mid-batch
(B ≈ 64–1024), where it briefly leads standalone potrf at mid sizes; only
gemm at N ≥ 32 is an outright vendor win at scale — the same mid-band the
ladder already routes to glass::nvidia::.
An audited 2026-08-14 confirmation used 500 reps. Only 8/378 selected
GLASS/vendor pairs had a contender above 5% within-run spread, and all eight
had decisive 1.56×–2.19× gaps. One small f64 potrf comparison changed
winner between two quiet captures despite low within-run spread. Accordingly,
small deltas and cross-run winner changes are treated as ties; publication
claims must survive both independent captures.
Permitting TF32 tensor cores (dashed) does not change the story: cuBLAS
declines to engage them below N = 24 (results bit-identical to FP32),
and where they do engage the speed is a wash against FP32 cuBLAS while max
error jumps three orders of magnitude (~1e-7 → ~2e-4) — unusable for the
Cholesky-chain ops, which have no TF32 cuSOLVER path at all.
Fusion: riccati_gain vs. a 7-call vendor chain#
glass::riccati_gain computes the LQR feedback gain
K = (R + BᵀPB)⁻¹(BᵀPA) in one kernel with all intermediates in shared
memory; the host-batched equivalent is seven vendor calls (four gemms, a
batched Cholesky, two triangular solves) with intermediates in global memory.
bench/bench_paper_fusion.cu compares them (capture
paper_fusion_20260708_0055.txt (archived externally)):
In this historical implementation/capture, fusion wins at every batch size at
quadrotor/manipulator scale — 2.5–2.8× at
(nx,nu) = (12,4) and 1.6–1.9× at (14,7) in fp32, more in fp64 — but the
chain wins at (36,12) fp32 and (48,16), where the staged operands outgrow what
one block overlaps profitably. Fusion is a measured choice, not a default:
GLASS composes both forms from the same primitives.
Single-call latency#
For a batch of one (a high-rate MPC tick), wall-clock per-call latency is
what matters. The non-batched vendor calls pin an essentially flat API floor —
~7–10 µs (cublasSgemm), ~15.5 µs (potrf), ~23–30 µs (posv) — while a GLASS
call starts at 5.2 µs and grows with compute, so GLASS wins single-call
latency through N = 32 (gemm), 12 (potrf), and 24 (posv); at N = 8 the
full factor-and-solve is 2.4× faster (9.6 vs 23.5 µs). And inside your own
kernel, composed GLASS calls never pay the API floor again.
Both harnesses live in bench/ and rerun via
python3 bench/paper_sweeps.py (see bench/PAPER_SWEEPS.md).
Choosing among the dense solve paths (measured guidance)#
GLASS deliberately ships no auto-dispatch between its linear-system
solvers — the right choice depends on structure and conditioning, which a
compile-time table cannot see. Instead, bench/tune.py --legs solvers
measures the trade-offs on your GPU and records them in
bench/RESULTS.md (solvers section). The RTX 5090 numbers (NPROB=8192,
ns/problem):
SPD single solve — use ``posv``; the alternatives price as follows.
gesv (pivoted LU, the robustness fallback) costs 1.0–2.1× posv in
fp32 at N ≥ 16 (1.8–2.1× at 32–64); the invert-then-multiply
anti-pattern (inv + gemv) costs 1.9–4.7× at N ≥ 16. Below
N = 16 all three paths are within a few nanoseconds of each other (and
inv/gesv can even edge out posv, especially in fp64) — there,
choose by numerics, not speed: Cholesky is backward-stable on SPD input and
fails loudly (with CHECK) on indefinite input, while an explicit inverse
amplifies conditioning error silently. Speed only ever argues for posv,
never against it.
Block-tridiagonal chains — ``bdsv`` (direct) vs ``pcg`` (iterative) is
problem-dependent; do not hard-code either. On the benchmark’s synthetic
diagonally-dominant SPD system, PCG converges in about three iterations and
wins 11 of 12 cells (up to 9×) at the harness’s rho = rᵀz relative
tolerance of 1e-6. This is an approximate-solve comparison, not matched final
residual accuracy: the direct f64 solves reach near-machine precision while
PCG stops at its requested tolerance. At
(BlockSize=12, Knots=16) fp32 the direct sweep wins 1.8×. PCG’s cost scales
linearly with its iteration count, so an ill-conditioned Riccati/KKT chain
(10–100× more iterations) moves the crossover proportionally toward
bdsv — read the pcg iters column of your own sweep before
generalizing, or measure with your actual matrices.
``syev`` / ``eig_clamp`` — the historical harness compared only this serial
Jacobi family. It did not include the parallel fixed-sweep eigh /
psd_project family, so it cannot support a recommendation between them.
NVIDIA thread LAPACK — measured integration#
The 2026-08-30 wave added cuSOLVERDx’s thread interface as a fifth contender
for chol, trsv, and posv. It is a selective win, not a replacement
for either GLASS’s native thread code or cuSOLVERDx’s block interface:
GPU |
cells |
float32 |
float64 |
|---|---|---|---|
RTX 5090 / sm_120 |
14 / 132 |
|
|
Jetson AGX Orin / sm_87 |
15 / 132 |
|
|
The table and replication paragraph below document the archived 2026-08-30 release method. The next retune replaces its asymmetric confirmation with the unified fresh-input solver ladder described above; regenerate these counts and ranges from that accepted capture before treating them as current.
Those ranges enumerate the measured sizes 4, 6, 8, 12, 16, 24, 32; they
do not imply testing every intervening integer. Against the fastest native
tier on independent valid inputs, selected NVIDIA-thread cells range from
about 1.09× to 2.34× on the RTX 5090 and 1.08× to 3.53× on Orin. The largest
gains are concentrated in trsv; other operations and sizes still select
another tier.
The 5090 throughput leg was run twice independently at 250 and 500
repetitions. The raw ladder policy agreed in 131 of 132 cells, agreed on all
17 pre-veto NVIDIA-thread selections, and disagreed only on dot f32 N=16 inside the
native-SIMT tie band. Across the NVIDIA-thread measurements, the second/first
median time ratio was 0.999 (10th–90th percentile 0.984–1.026). The Orin
throughput capture used the Tegra profile (50 repetitions, pinned 50 W mode).
The subsequent valid-input captures on both machines had sub-margin spread at
every retained or vetoed NVIDIA-thread decision and reduced the shipped bands
to the 14 and 15 cells above. These checks support the selected bands, not a
portable speedup claim—rerun both ladder components on a new architecture.
The native thread tier — historical characterization#
The 2026-07-19 full-domain sweep (quiet RTX 5090) first added the native thread
contender at every (op, N) point. Throughput regime (NPROB=8192), thread
vs the best other tier (ratio > 1 = thread faster; bold cells shipped):
op / dtype |
N=4 |
N=6 |
N=8 |
N=12 |
N=16 |
N=24 |
shipped band |
|---|---|---|---|---|---|---|---|
posv f64 |
6.17× |
7.46× |
5.03× |
3.05× |
2.26× |
1.19× |
thread ≤ 24 |
chol f64 |
5.61× |
5.79× |
2.88× |
2.18× |
1.71× |
1.04× |
thread ≤ 24 |
trsv f32/f64 |
1.5/4.6× |
2.2/4.6× |
2.0/2.0× |
2.3/1.5× |
1.7/1.2× |
0.2× |
thread ≤ 16 |
dot f64 |
2.32× |
2.33× |
2.31× |
1.96× |
1.53× |
1.30× |
thread ≤ 32 |
posv f32 |
3.21× |
3.44× |
1.73× |
1.16× |
0.66× |
0.35× |
thread ≤ 12 |
gemm (both) |
0.6–0.9× |
0.4–0.5× |
— |
— |
— |
— |
never |
The factor/solve chain is where the tier earns its keep — a warp-per-problem
potrf at N ≤ 7 idles most lanes on the serial pivot; one-problem-per-thread
keeps all 32 busy. f64 amplifies and extends the win past the register ceiling
(even the spilled thread path beats the alternatives through N=24). gemm
is the anti-case: enough work per element that the parallel tiers always win.
Columns past the ceiling price the local-memory spill honestly (down to 0.1× —
measured, never shipped).
Warp-scope vendor A/B (audited characterization)#
bench_nvwarp_l1.cu A/Bs glass::warp:: against glass::nvidia::warp::
(CUB WarpReduce) for the three ops the vendor warp tier ships (reduce/dot/nrm2)
at identical launch shape. The 2026-08-14 sm_120 quiet run used the shared
warm-up, spread, provenance, and isolation protocol. Of 126 cells, 121 tie
within ±2%. Four high-throughput f64 dot cells favor CUB by 2.0–3.0%; the
sole SIMT verdict is a tiny latency cell with 17.45% SIMT spread and is not
actionable. Five individual contender samples exceeded 5% spread, but no
shipped dispatch decision depends on this characterization.
Robotics micro-ops — audited characterization#
The 2026-08-14 sm_120 quiet run measured 8192 problems and 1000 reps per trial through the shared audited protocol:
The thread tier wins every fixed-size row for both dtypes.
softmax_n16is dtype-sensitive: warp wins f32, thread wins f64.At the best thread launch, fused and composed spatial forms are nearly tied in f32; the fused f64 forms are about 15–22% faster.
Only 2 of 136 contender samples exceed 5% spread, neither changing a clear row verdict.
argmax_fastis block-only and does not beat the ordinary thread-tierargmax.
These are characterization results, not generated dispatch inputs.
The full measured-results archive#
The machine-refreshed verdict tables for every sweep live in one file,
bench/RESULTS.md. The ladder, body, blas2, rect, solvers, and reduced marker
blocks are rewritten by bench/tune.py; nvwarp and robotics are hand-run
audited characterizations and do not rewrite dispatch.
Raw captures are archived externally with the paper materials; the paper
harnesses are documented in bench/PAPER_SWEEPS.md.
See Tuning for Your Hardware for how to emit a per-host override table from a sweep, and Execution Plans (glass-defaults.cuh) for the picker API.