Backend Sweep Results#
GLASS ships overlapping thread-, warp-, block-, and vendor-backed execution
tiers. For operations present in more than one tier, the fastest choice depends
on operation, size, dtype, launch shape, and batch regime. The mega sweep
(bench/tune.py’s ladder leg) compares supported contenders — this is
exactly the measurement behind glass-defaults.cuh’s recommend<>()
(see Execution Plans (glass-defaults.cuh)).
Warning
This page is a dated measurement archive, not a universal performance promise. Several legacy sections predate the current warm-up, trial-spread, source-digest, and signed-receipt protocol and are marked accordingly. Use the source capture named beside a claim, and rerun on the target machine.
The figures and table below come from the 2026-09-02 RTX 5090 / sm_120
five-backend selection capture — the same paired fresh-input A/B campaign
that generated the shipped ideal_sm120 / ideal_sm87 / ideal_sm72
tables (an independent same-protocol second capture confirmed the
selections) — shown across three batch regimes: NPROB=64 (low batch,
latency-leaning), NPROB=1024 (mid), and NPROB=8192 (throughput). The
winner can shift with batch
size, so low-batch plots describe that workload rather than overriding the
throughput table. These are committed static assets — regenerate them for your
own hardware with:
python bench/tune.py --sm auto # remeasures + regenerates tables AND figures
# or just the figures from an existing sweep .txt:
python bench/export_sweep_figures.py bench/mega_sweep_*.txt
bench/explore_sweep.ipynb is the interactive version of the same analysis.
The ladder — ns/problem vs N, per backend#
Lower is faster. Each subplot is one op; the curves are native warp
(green), block (blue), and thread (orange), plus NVIDIA block
(red) and thread (purple). Both thread curves pack one problem per CUDA
thread. The native thread implementation is instantiated through N=64 in
this harness; NVIDIA thread LAPACK is instantiated through N=32. A stopped
curve therefore marks an explicit instantiation or resource boundary, not a
claim that larger sizes are unsupported by CUDA or cuSOLVERDx. NVIDIA block
f64 curves may stop at their shared-memory feasibility limit.
recommend<>() is keyed on the NPROB=8192 throughput regime.
Its policy first applies a ±2% tie band among dependency-free SIMT contenders,
then requires an NVIDIA contender to clear the best native choice by more than
5%. For destructive POTRF/TRSV/POSV, every displayed and selected candidate is
remeasured by the unified solver ladder on fresh valid inputs. The plans are
randomized within paired rounds and every raw sample is retained; all native
and NVIDIA candidates use the same selection rule. These rules keep small,
noisy, or mutation-dependent gaps out of the generated table.
float32#
float64#
Raw ladder winner per (op, N), per regime#
The policy winner at each (op, N) is listed for all three NPROB regimes.
Non-destructive rows come from the general ladder; POTRF/TRSV/POSV rows come
from its fresh-input solver companion. Thus the plots, winner list, and shipped
solver defaults all consume the same authoritative measurements.
The broad high-batch shape is mixed by design: native thread dominates many
small packed problems, native warp/block remain important as work grows, and
the NVIDIA block and thread implementations take measured factor/solve bands.
The NPROB=8192 blocks shown here are from the same capture pair that
generated the shipped sm_120 table.
NPROB=64 f32 winner by op x N:
op 4 6 8 12 16 24 32 48 64 96 128
dot block thread thread thread thread thread warp warp warp warp warp
gemv thread thread thread warp warp nvidia nvidia nvidia nvidia nvidia nvidia
gemm thread warp warp warp warp nvidia nvidia nvidia nvidia block block
potrf thread thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia
trsv thread thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia
posv thread thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia
NPROB=64 f64 winner by op x N:
op 4 6 8 12 16 24 32 48 64 96 128
dot thread thread thread thread thread warp warp warp warp warp warp
gemv thread warp warp warp warp warp warp block block block block
gemm warp warp warp nvidia nvidia block block nvidia block block block
potrf nvidia_thread nvidia_thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia block block
trsv thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia block block
posv nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia block block
NPROB=1024 f32 winner by op x N:
op 4 6 8 12 16 24 32 48 64 96 128
dot warp warp thread thread warp warp warp warp warp warp warp
gemv thread thread warp warp warp warp warp block block block block
gemm warp warp warp block warp block nvidia nvidia nvidia block block
potrf thread nvidia_thread nvidia_thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia
trsv thread thread thread nvidia nvidia nvidia nvidia nvidia nvidia block block
posv thread thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia nvidia
NPROB=1024 f64 winner by op x N:
op 4 6 8 12 16 24 32 48 64 96 128
dot thread thread thread warp warp warp warp warp warp warp warp
gemv warp warp warp warp warp warp warp block block block block
gemm warp warp warp warp warp warp warp warp block block block
potrf nvidia_thread nvidia_thread nvidia_thread block block warp warp block block block block
trsv thread thread nvidia nvidia nvidia nvidia nvidia block block block block
posv nvidia_thread nvidia_thread nvidia_thread warp warp warp warp block block block block
NPROB=8192 f32 winner by op x N:
op 4 6 8 12 16 24 32 48 64 96 128
dot thread thread thread thread warp warp warp warp warp warp warp
gemv thread thread warp warp warp warp warp block warp warp warp
gemm warp warp warp warp warp block nvidia block block block block
potrf nvidia_thread nvidia_thread nvidia_thread warp nvidia nvidia nvidia nvidia nvidia nvidia nvidia
trsv thread thread thread thread thread nvidia nvidia nvidia warp warp warp
posv thread thread thread thread nvidia nvidia nvidia nvidia nvidia nvidia nvidia
NPROB=8192 f64 winner by op x N:
op 4 6 8 12 16 24 32 48 64 96 128
dot thread thread thread thread thread thread thread warp warp warp warp
gemv thread thread warp warp warp warp warp warp warp warp block
gemm warp warp warp warp block block block block block block block
potrf nvidia_thread nvidia_thread nvidia_thread thread thread thread block block block block block
trsv thread thread nvidia_thread nvidia_thread nvidia_thread nvidia_thread nvidia warp warp warp warp
posv nvidia_thread nvidia_thread nvidia_thread thread thread thread warp warp block block block
The same ladder at the edge (Jetson AGX Orin)#
The identical sweep on the Jetson AGX Orin (sm_87, 2026-09-01 selection capture, throughput regime) crowns different winners in 145 of 396 cells versus the RTX 5090 above, and the older AGX Xavier (sm_72, native-only) differs from the Orin in 162 of 396 — placement is a property of the silicon, and getting it wrong is expensive: on the Orin’s raw ladder the best and worst placements for a cell differ by a median of 4.9× (up to 81×). See What a retune actually changes (sm_120 vs sm_87 vs sm_72) for the cross-architecture analysis.
vs. host-batched cuBLAS/cuSOLVER (and TF32)#
The ladder above compares device-side backends. A separate question is how
one-block-per-problem GLASS compares to the standard host-side recipe — a
single cublas<t>gemmStridedBatched / cusolverDn<t>potrfBatched
(+ potrsBatched) call over the whole batch.
bench/bench_paper_hostblas.cu measures exactly that: gemm / potrf / posv,
N = 4–64, batch B = 1–8192, both precisions (raw capture committed as
paper_hostblas_20260708_0054.txt (archived externally); RTX 5090 / sm_120, quiet GPU).
Above 1.0 = the best swept GLASS configuration faster (fp32 shown; GLASS is
best of block32, block128, and warp8 where available). The vendor side is one
default host-API configuration, so this is not a symmetric tuner comparison or
a universal “GLASS vs vendor” statement. In this historical capture, at robot
sizes, host batching never catches up: gemm at N ≤ 24 and the full
factor-and-solve (posv) through N = 64 are GLASS wins at every batch
size, reaching 2.9–6.3× at saturation. The vendor’s best regime is mid-batch
(B ≈ 64–1024), where it briefly leads standalone potrf at mid sizes; only
gemm at N ≥ 32 is an outright vendor win at scale — the same mid-band the
ladder already routes to glass::nvidia::.
An audited 2026-08-14 confirmation used 500 reps. Only 8/378 selected
GLASS/vendor pairs had a contender above 5% within-run spread, and all eight
had decisive 1.56×–2.19× gaps. One small f64 potrf comparison changed
winner between two quiet captures despite low within-run spread. Accordingly,
small deltas and cross-run winner changes are treated as ties; publication
claims must survive both independent captures.
The 2026-09 recapture campaign repeated this comparison under the current
median-of-trials protocol on all three measured architectures and confirmed
the shape while sharpening the edge story: the host-dispatch overhead that
GLASS avoids is far more expensive on an embedded GPU, so the same
comparison that peaks near 16× on the RTX 5090 peaks over 100× on the
Jetson AGX Orin (small-N posv, where per-call host overhead dominates).
Raw captures are archived with the paper materials.
Permitting TF32 tensor cores (dashed) does not change the story: cuBLAS
declines to engage them below N = 24 (results bit-identical to FP32),
and where they do engage the speed is a wash against FP32 cuBLAS while max
error jumps three orders of magnitude (~1e-7 → ~2e-4) — unusable for the
Cholesky-chain ops, which have no TF32 cuSOLVER path at all.
Fusion: riccati_gain vs. a 7-call vendor chain#
glass::riccati_gain computes the LQR feedback gain
K = (R + BᵀPB)⁻¹(BᵀPA) in one kernel with all intermediates in shared
memory; the host-batched equivalent is seven vendor calls (four gemms, a
batched Cholesky, two triangular solves) with intermediates in global memory.
bench/bench_paper_fusion.cu compares them (capture
paper_fusion_20260708_0055.txt (archived externally)):
In this historical implementation/capture, fusion wins at every batch size at
quadrotor/manipulator scale — 2.5–2.8× at
(nx,nu) = (12,4) and 1.6–1.9× at (14,7) in fp32, more in fp64 — but the
chain wins at (36,12) fp32 and (48,16), where the staged operands outgrow what
one block overlaps profitably. Fusion is a measured choice, not a default:
GLASS composes both forms from the same primitives.
The harness now also carries a third, device-side unfused arm — the same GLASS primitives launched one kernel per step with intermediates in global memory — which isolates the fusion effect from the device-vs-host effect. On the Jetson AGX Orin (2026-09 captures), fusion wins by up to ~1.75× at small batches in fp32, while the unfused composition wins by up to ~2.9× at large batches and in fp64, where per-primitive kernels regain the occupancy the fused kernel’s shared-memory footprint caps. Which side of that tradeoff a given (size, batch, precision) lands on is exactly the kind of measured placement decision the rest of this page is about.
Single-call latency#
For a batch of one (a high-rate MPC tick), wall-clock per-call latency is
what matters. The non-batched vendor calls pin an essentially flat API floor —
~7–10 µs (cublasSgemm), ~15.5 µs (potrf), ~23–30 µs (posv) — while a GLASS
call starts at 5.2 µs and grows with compute, so GLASS wins single-call
latency through N = 32 (gemm), 12 (potrf), and 24 (posv); at N = 8 the
full factor-and-solve is 2.4× faster (9.6 vs 23.5 µs). And inside your own
kernel, composed GLASS calls never pay the API floor again.
Both harnesses live in bench/ and rerun via
python3 bench/paper_sweeps.py (see bench/PAPER_SWEEPS.md).
Choosing among the dense solve paths (measured guidance)#
GLASS deliberately ships no auto-dispatch between its linear-system
solvers — the right choice depends on structure and conditioning, which a
compile-time table cannot see. Instead, bench/tune.py --legs solvers
measures the trade-offs on your GPU and records them in
bench/RESULTS.md (solvers section). The RTX 5090 numbers (NPROB=8192,
ns/problem):
SPD single solve — use ``posv``; the alternatives price as follows.
gesv (pivoted LU, the robustness fallback) costs 1.0–2.1× posv in
fp32 at N ≥ 16 (1.8–2.1× at 32–64); the invert-then-multiply
anti-pattern (inv + gemv) costs 1.9–4.7× at N ≥ 16. Below
N = 16 all three paths are within a few nanoseconds of each other (and
inv/gesv can even edge out posv, especially in fp64) — there,
choose by numerics, not speed: Cholesky is backward-stable on SPD input and
fails loudly (with CHECK) on indefinite input, while an explicit inverse
amplifies conditioning error silently. Speed only ever argues for posv,
never against it.
Block-tridiagonal chains — ``bdsv`` (direct) vs ``pcg`` (iterative) is
problem-dependent; do not hard-code either. On the benchmark’s synthetic
diagonally-dominant SPD system, PCG converges in about three iterations and
wins 11 of 12 cells (up to 9×) at the harness’s rho = rᵀz relative
tolerance of 1e-6. This is an approximate-solve comparison, not matched final
residual accuracy: the direct f64 solves reach near-machine precision while
PCG stops at its requested tolerance. At
(BlockSize=12, Knots=16) fp32 the direct sweep wins 1.8×. PCG’s cost scales
linearly with its iteration count, so an ill-conditioned Riccati/KKT chain
(10–100× more iterations) moves the crossover proportionally toward
bdsv — read the pcg iters column of your own sweep before
generalizing, or measure with your actual matrices.
``syev`` / ``eig_clamp`` — the historical harness compared only this serial
Jacobi family. It did not include the parallel fixed-sweep eigh /
psd_project family, so it cannot support a recommendation between them.
NVIDIA thread LAPACK — measured integration#
The 2026-08-30 wave added cuSOLVERDx’s thread interface as a fifth contender
for potrf, trsv, and posv. It is a selective win, not a replacement
for either GLASS’s native thread code or cuSOLVERDx’s block interface:
GPU |
cells |
float32 |
float64 |
|---|---|---|---|
RTX 5090 / sm_120 |
12 / 132 |
|
|
Jetson AGX Orin / sm_87 |
18 / 132 |
|
|
These cells were selected by the unified fresh-input solver ladder described above — every native and NVIDIA candidate measured symmetrically on fresh valid inputs, with a selected NVIDIA plan additionally interval-confirmed (the sm_120 regeneration demoted one ambiguous vendor pick to the capture’s native winner under that rule). The AGX Xavier (sm_72) ships no NVIDIA-thread cells: its CUDA 11.4 toolchain has no device-callable MathDx.
The ranges enumerate the measured sizes 4, 6, 8, 12, 16, 24, 32; they
do not imply testing every intervening integer. Against the fastest native
plan in the same generating captures, the selected NVIDIA-thread cells win
by about 1.09× to 2.35× on the RTX 5090 and 1.09× to 3.54× on Orin. The
largest gains are concentrated in trsv; other operations and sizes still
select another tier. Each table is supported by an independent same-protocol
second capture (the A/B pair), and the archived 2026-08-30 release
measurements — a different confirmation method — selected nearly the same
bands with near-identical win ranges, so the bands are stable across both
methods and across independent quiet-window sessions. These checks support
the selected bands, not a portable speedup claim — rerun both ladder
components on a new architecture.
The native thread tier — historical characterization#
The 2026-07-19 full-domain sweep (quiet RTX 5090) first added the native thread
contender at every (op, N) point. Throughput regime (NPROB=8192), thread
vs the best other tier (ratio > 1 = thread faster; bold cells shipped):
op / dtype |
N=4 |
N=6 |
N=8 |
N=12 |
N=16 |
N=24 |
shipped band |
|---|---|---|---|---|---|---|---|
posv f64 |
6.17× |
7.46× |
5.03× |
3.05× |
2.26× |
1.19× |
thread ≤ 24 |
chol f64 |
5.61× |
5.79× |
2.88× |
2.18× |
1.71× |
1.04× |
thread ≤ 24 |
trsv f32/f64 |
1.5/4.6× |
2.2/4.6× |
2.0/2.0× |
2.3/1.5× |
1.7/1.2× |
0.2× |
thread ≤ 16 |
dot f64 |
2.32× |
2.33× |
2.31× |
1.96× |
1.53× |
1.30× |
thread ≤ 32 |
posv f32 |
3.21× |
3.44× |
1.73× |
1.16× |
0.66× |
0.35× |
thread ≤ 12 |
gemm (both) |
0.6–0.9× |
0.4–0.5× |
— |
— |
— |
— |
never |
The factor/solve chain is where the tier earns its keep — a warp-per-problem
potrf at N ≤ 7 idles most lanes on the serial pivot; one-problem-per-thread
keeps all 32 busy. f64 amplifies and extends the win past the register ceiling
(even the spilled thread path beats the alternatives through N=24). gemm
is the anti-case: enough work per element that the parallel tiers always win.
Columns past the ceiling price the local-memory spill honestly (down to 0.1× —
measured, never shipped).
Warp-scope vendor A/B (audited characterization)#
bench_nvwarp_l1.cu A/Bs glass::warp:: against glass::nvidia::warp::
(CUB WarpReduce) for the three ops the vendor warp tier ships (reduce/dot/nrm2)
at identical launch shape. The 2026-08-14 sm_120 quiet run used the shared
warm-up, spread, provenance, and isolation protocol. Of 126 cells, 121 tie
within ±2%. Four high-throughput f64 dot cells favor CUB by 2.0–3.0%; the
sole SIMT verdict is a tiny latency cell with 17.45% SIMT spread and is not
actionable. Five individual contender samples exceeded 5% spread, but no
shipped dispatch decision depends on this characterization.
Robotics micro-ops — audited characterization#
The 2026-08-14 sm_120 quiet run measured 8192 problems and 1000 reps per trial through the shared audited protocol:
The thread tier wins every fixed-size row for both dtypes.
softmax_n16is dtype-sensitive: warp wins f32, thread wins f64.At the best thread launch, fused and composed spatial forms are nearly tied in f32; the fused f64 forms are about 15–22% faster.
Only 2 of 136 contender samples exceed 5% spread, neither changing a clear row verdict.
argmax_fastis block-only and does not beat the ordinary thread-tierargmax.
These are characterization results, not generated dispatch inputs.
The full measured-results archive#
The machine-refreshed verdict tables for every sweep live in one file,
bench/RESULTS.md. The ladder, body, blas2, rect, solvers, and reduced marker
blocks are rewritten by bench/tune.py; nvwarp and robotics are hand-run
audited characterizations and do not rewrite dispatch.
Raw captures are archived externally with the paper materials; the paper
harnesses are documented in bench/PAPER_SWEEPS.md.
See Tuning for Your Hardware for how to emit a per-host override table from a sweep, and Execution Plans (glass-defaults.cuh) for the picker API.