Backend Sweep Results#

GLASS ships overlapping thread-, warp-, block-, and vendor-backed execution tiers. For operations present in more than one tier, the fastest choice depends on operation, size, dtype, launch shape, and batch regime. The mega sweep (bench/tune.py’s ladder leg) compares supported contenders — this is exactly the measurement behind glass-defaults.cuh’s recommend<>() (see Execution Plans (glass-defaults.cuh)).

Warning

This page is a dated measurement archive, not a universal performance promise. Several legacy sections predate the current warm-up, trial-spread, source-digest, and signed-receipt protocol and are marked accordingly. Use the source capture named beside a claim, and rerun on the target machine.

The figures and table below come from the 2026-09-02 RTX 5090 / sm_120 five-backend selection capture — the same paired fresh-input A/B campaign that generated the shipped ideal_sm120 / ideal_sm87 / ideal_sm72 tables (an independent same-protocol second capture confirmed the selections) — shown across three batch regimes: NPROB=64 (low batch, latency-leaning), NPROB=1024 (mid), and NPROB=8192 (throughput). The winner can shift with batch size, so low-batch plots describe that workload rather than overriding the throughput table. These are committed static assets — regenerate them for your own hardware with:

python bench/tune.py --sm auto      # remeasures + regenerates tables AND figures
# or just the figures from an existing sweep .txt:
python bench/export_sweep_figures.py bench/mega_sweep_*.txt

bench/explore_sweep.ipynb is the interactive version of the same analysis.

The ladder — ns/problem vs N, per backend#

Lower is faster. Each subplot is one op; the curves are native warp (green), block (blue), and thread (orange), plus NVIDIA block (red) and thread (purple). Both thread curves pack one problem per CUDA thread. The native thread implementation is instantiated through N=64 in this harness; NVIDIA thread LAPACK is instantiated through N=32. A stopped curve therefore marks an explicit instantiation or resource boundary, not a claim that larger sizes are unsupported by CUDA or cuSOLVERDx. NVIDIA block f64 curves may stop at their shared-memory feasibility limit.

recommend<>() is keyed on the NPROB=8192 throughput regime. Its policy first applies a ±2% tie band among dependency-free SIMT contenders, then requires an NVIDIA contender to clear the best native choice by more than 5%. For destructive POTRF/TRSV/POSV, every displayed and selected candidate is remeasured by the unified solver ladder on fresh valid inputs. The plans are randomized within paired rounds and every raw sample is retained; all native and NVIDIA candidates use the same selection rule. These rules keep small, noisy, or mutation-dependent gaps out of the generated table.

float32#

f32 ladder, NPROB=64 f32 ladder, NPROB=1024 f32 ladder, NPROB=8192 (throughput — feeds recommend)

float64#

f64 ladder, NPROB=64 f64 ladder, NPROB=1024 f64 ladder, NPROB=8192 (throughput — feeds recommend)

Raw ladder winner per (op, N), per regime#

The policy winner at each (op, N) is listed for all three NPROB regimes. Non-destructive rows come from the general ladder; POTRF/TRSV/POSV rows come from its fresh-input solver companion. Thus the plots, winner list, and shipped solver defaults all consume the same authoritative measurements. The broad high-batch shape is mixed by design: native thread dominates many small packed problems, native warp/block remain important as work grows, and the NVIDIA block and thread implementations take measured factor/solve bands. The NPROB=8192 blocks shown here are from the same capture pair that generated the shipped sm_120 table.

NPROB=64  f32 winner by op x N:
op                  4             6             8            12            16            24            32            48            64            96           128
dot             block        thread        thread        thread        thread        thread          warp          warp          warp          warp          warp
gemv           thread        thread        thread          warp          warp        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia
gemm           thread          warp          warp          warp          warp        nvidia        nvidia        nvidia        nvidia         block         block
potrf          thread        thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia
trsv           thread        thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia
posv           thread        thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia

NPROB=64  f64 winner by op x N:
op                  4             6             8            12            16            24            32            48            64            96           128
dot            thread        thread        thread        thread        thread          warp          warp          warp          warp          warp          warp
gemv           thread          warp          warp          warp          warp          warp          warp         block         block         block         block
gemm             warp          warp          warp        nvidia        nvidia         block         block        nvidia         block         block         block
potrf   nvidia_thread nvidia_thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia         block         block
trsv           thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia         block         block
posv           nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia         block         block

NPROB=1024  f32 winner by op x N:
op                  4             6             8            12            16            24            32            48            64            96           128
dot              warp          warp        thread        thread          warp          warp          warp          warp          warp          warp          warp
gemv           thread        thread          warp          warp          warp          warp          warp         block         block         block         block
gemm             warp          warp          warp         block          warp         block        nvidia        nvidia        nvidia         block         block
potrf          thread nvidia_thread nvidia_thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia
trsv           thread        thread        thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia         block         block
posv           thread        thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia

NPROB=1024  f64 winner by op x N:
op                  4             6             8            12            16            24            32            48            64            96           128
dot            thread        thread        thread          warp          warp          warp          warp          warp          warp          warp          warp
gemv             warp          warp          warp          warp          warp          warp          warp         block         block         block         block
gemm             warp          warp          warp          warp          warp          warp          warp          warp         block         block         block
potrf   nvidia_thread nvidia_thread nvidia_thread         block         block          warp          warp         block         block         block         block
trsv           thread        thread        nvidia        nvidia        nvidia        nvidia        nvidia         block         block         block         block
posv    nvidia_thread nvidia_thread nvidia_thread          warp          warp          warp          warp         block         block         block         block

NPROB=8192  f32 winner by op x N:
op                  4             6             8            12            16            24            32            48            64            96           128
dot            thread        thread        thread        thread          warp          warp          warp          warp          warp          warp          warp
gemv           thread        thread          warp          warp          warp          warp          warp         block          warp          warp          warp
gemm             warp          warp          warp          warp          warp         block        nvidia         block         block         block         block
potrf   nvidia_thread nvidia_thread nvidia_thread          warp        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia
trsv           thread        thread        thread        thread        thread        nvidia        nvidia        nvidia          warp          warp          warp
posv           thread        thread        thread        thread        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia        nvidia

NPROB=8192  f64 winner by op x N:
op                  4             6             8            12            16            24            32            48            64            96           128
dot            thread        thread        thread        thread        thread        thread        thread          warp          warp          warp          warp
gemv           thread        thread          warp          warp          warp          warp          warp          warp          warp          warp         block
gemm             warp          warp          warp          warp         block         block         block         block         block         block         block
potrf   nvidia_thread nvidia_thread nvidia_thread        thread        thread        thread         block         block         block         block         block
trsv           thread        thread nvidia_thread nvidia_thread nvidia_thread nvidia_thread        nvidia          warp          warp          warp          warp
posv    nvidia_thread nvidia_thread nvidia_thread        thread        thread        thread          warp          warp         block         block         block

The same ladder at the edge (Jetson AGX Orin)#

The identical sweep on the Jetson AGX Orin (sm_87, 2026-09-01 selection capture, throughput regime) crowns different winners in 145 of 396 cells versus the RTX 5090 above, and the older AGX Xavier (sm_72, native-only) differs from the Orin in 162 of 396 — placement is a property of the silicon, and getting it wrong is expensive: on the Orin’s raw ladder the best and worst placements for a cell differ by a median of 4.9× (up to 81×). See What a retune actually changes (sm_120 vs sm_87 vs sm_72) for the cross-architecture analysis.

f32 ladder on Jetson AGX Orin, NPROB=8192 (throughput — feeds recommend) f64 ladder on Jetson AGX Orin, NPROB=8192 (throughput — feeds recommend)

vs. host-batched cuBLAS/cuSOLVER (and TF32)#

The ladder above compares device-side backends. A separate question is how one-block-per-problem GLASS compares to the standard host-side recipe — a single cublas<t>gemmStridedBatched / cusolverDn<t>potrfBatched (+ potrsBatched) call over the whole batch. bench/bench_paper_hostblas.cu measures exactly that: gemm / potrf / posv, N = 4–64, batch B = 1–8192, both precisions (raw capture committed as paper_hostblas_20260708_0054.txt (archived externally); RTX 5090 / sm_120, quiet GPU).

host-batched vendor time divided by best GLASS time, vs batch size

Above 1.0 = the best swept GLASS configuration faster (fp32 shown; GLASS is best of block32, block128, and warp8 where available). The vendor side is one default host-API configuration, so this is not a symmetric tuner comparison or a universal “GLASS vs vendor” statement. In this historical capture, at robot sizes, host batching never catches up: gemm at N ≤ 24 and the full factor-and-solve (posv) through N = 64 are GLASS wins at every batch size, reaching 2.9–6.3× at saturation. The vendor’s best regime is mid-batch (B ≈ 64–1024), where it briefly leads standalone potrf at mid sizes; only gemm at N ≥ 32 is an outright vendor win at scale — the same mid-band the ladder already routes to glass::nvidia::.

An audited 2026-08-14 confirmation used 500 reps. Only 8/378 selected GLASS/vendor pairs had a contender above 5% within-run spread, and all eight had decisive 1.56×–2.19× gaps. One small f64 potrf comparison changed winner between two quiet captures despite low within-run spread. Accordingly, small deltas and cross-run winner changes are treated as ties; publication claims must survive both independent captures.

The 2026-09 recapture campaign repeated this comparison under the current median-of-trials protocol on all three measured architectures and confirmed the shape while sharpening the edge story: the host-dispatch overhead that GLASS avoids is far more expensive on an embedded GPU, so the same comparison that peaks near 16× on the RTX 5090 peaks over 100× on the Jetson AGX Orin (small-N posv, where per-call host overhead dominates). Raw captures are archived with the paper materials.

Permitting TF32 tensor cores (dashed) does not change the story: cuBLAS declines to engage them below N = 24 (results bit-identical to FP32), and where they do engage the speed is a wash against FP32 cuBLAS while max error jumps three orders of magnitude (~1e-7 → ~2e-4) — unusable for the Cholesky-chain ops, which have no TF32 cuSOLVER path at all.

Fusion: riccati_gain vs. a 7-call vendor chain#

glass::riccati_gain computes the LQR feedback gain K = (R + BᵀPB)⁻¹(BᵀPA) in one kernel with all intermediates in shared memory; the host-batched equivalent is seven vendor calls (four gemms, a batched Cholesky, two triangular solves) with intermediates in global memory. bench/bench_paper_fusion.cu compares them (capture paper_fusion_20260708_0055.txt (archived externally)):

fused riccati_gain vs 7-call host-batched vendor chain

In this historical implementation/capture, fusion wins at every batch size at quadrotor/manipulator scale — 2.5–2.8× at (nx,nu) = (12,4) and 1.6–1.9× at (14,7) in fp32, more in fp64 — but the chain wins at (36,12) fp32 and (48,16), where the staged operands outgrow what one block overlaps profitably. Fusion is a measured choice, not a default: GLASS composes both forms from the same primitives.

The harness now also carries a third, device-side unfused arm — the same GLASS primitives launched one kernel per step with intermediates in global memory — which isolates the fusion effect from the device-vs-host effect. On the Jetson AGX Orin (2026-09 captures), fusion wins by up to ~1.75× at small batches in fp32, while the unfused composition wins by up to ~2.9× at large batches and in fp64, where per-primitive kernels regain the occupancy the fused kernel’s shared-memory footprint caps. Which side of that tradeoff a given (size, batch, precision) lands on is exactly the kind of measured placement decision the rest of this page is about.

Single-call latency#

For a batch of one (a high-rate MPC tick), wall-clock per-call latency is what matters. The non-batched vendor calls pin an essentially flat API floor — ~7–10 µs (cublasSgemm), ~15.5 µs (potrf), ~23–30 µs (posv) — while a GLASS call starts at 5.2 µs and grows with compute, so GLASS wins single-call latency through N = 32 (gemm), 12 (potrf), and 24 (posv); at N = 8 the full factor-and-solve is 2.4× faster (9.6 vs 23.5 µs). And inside your own kernel, composed GLASS calls never pay the API floor again.

Both harnesses live in bench/ and rerun via python3 bench/paper_sweeps.py (see bench/PAPER_SWEEPS.md).

Choosing among the dense solve paths (measured guidance)#

GLASS deliberately ships no auto-dispatch between its linear-system solvers — the right choice depends on structure and conditioning, which a compile-time table cannot see. Instead, bench/tune.py --legs solvers measures the trade-offs on your GPU and records them in bench/RESULTS.md (solvers section). The RTX 5090 numbers (NPROB=8192, ns/problem):

SPD single solve — use ``posv``; the alternatives price as follows. gesv (pivoted LU, the robustness fallback) costs 1.0–2.1× posv in fp32 at N ≥ 16 (1.8–2.1× at 32–64); the invert-then-multiply anti-pattern (inv + gemv) costs 1.9–4.7× at N ≥ 16. Below N = 16 all three paths are within a few nanoseconds of each other (and inv/gesv can even edge out posv, especially in fp64) — there, choose by numerics, not speed: Cholesky is backward-stable on SPD input and fails loudly (with CHECK) on indefinite input, while an explicit inverse amplifies conditioning error silently. Speed only ever argues for posv, never against it.

Block-tridiagonal chains — ``bdsv`` (direct) vs ``pcg`` (iterative) is problem-dependent; do not hard-code either. On the benchmark’s synthetic diagonally-dominant SPD system, PCG converges in about three iterations and wins 11 of 12 cells (up to 9×) at the harness’s rho = rᵀz relative tolerance of 1e-6. This is an approximate-solve comparison, not matched final residual accuracy: the direct f64 solves reach near-machine precision while PCG stops at its requested tolerance. At (BlockSize=12, Knots=16) fp32 the direct sweep wins 1.8×. PCG’s cost scales linearly with its iteration count, so an ill-conditioned Riccati/KKT chain (10–100× more iterations) moves the crossover proportionally toward bdsv — read the pcg iters column of your own sweep before generalizing, or measure with your actual matrices.

``syev`` / ``eig_clamp`` — the historical harness compared only this serial Jacobi family. It did not include the parallel fixed-sweep eigh / psd_project family, so it cannot support a recommendation between them.

NVIDIA thread LAPACK — measured integration#

The 2026-08-30 wave added cuSOLVERDx’s thread interface as a fifth contender for potrf, trsv, and posv. It is a selective win, not a replacement for either GLASS’s native thread code or cuSOLVERDx’s block interface:

Table 4 Throughput cells selected by the shipped 2026-09-02 tables#

GPU

cells

float32

float64

RTX 5090 / sm_120

12 / 132

potrf N=6,8

potrf N=4–8; trsv N=8–24; posv N=4–8

Jetson AGX Orin / sm_87

18 / 132

potrf N=4,8,12; trsv N=8–32

potrf N=4–8; trsv N=4–32

These cells were selected by the unified fresh-input solver ladder described above — every native and NVIDIA candidate measured symmetrically on fresh valid inputs, with a selected NVIDIA plan additionally interval-confirmed (the sm_120 regeneration demoted one ambiguous vendor pick to the capture’s native winner under that rule). The AGX Xavier (sm_72) ships no NVIDIA-thread cells: its CUDA 11.4 toolchain has no device-callable MathDx.

The ranges enumerate the measured sizes 4, 6, 8, 12, 16, 24, 32; they do not imply testing every intervening integer. Against the fastest native plan in the same generating captures, the selected NVIDIA-thread cells win by about 1.09× to 2.35× on the RTX 5090 and 1.09× to 3.54× on Orin. The largest gains are concentrated in trsv; other operations and sizes still select another tier. Each table is supported by an independent same-protocol second capture (the A/B pair), and the archived 2026-08-30 release measurements — a different confirmation method — selected nearly the same bands with near-identical win ranges, so the bands are stable across both methods and across independent quiet-window sessions. These checks support the selected bands, not a portable speedup claim — rerun both ladder components on a new architecture.

The native thread tier — historical characterization#

The 2026-07-19 full-domain sweep (quiet RTX 5090) first added the native thread contender at every (op, N) point. Throughput regime (NPROB=8192), thread vs the best other tier (ratio > 1 = thread faster; bold cells shipped):

op / dtype

N=4

N=6

N=8

N=12

N=16

N=24

shipped band

posv f64

6.17×

7.46×

5.03×

3.05×

2.26×

1.19×

thread ≤ 24

chol f64

5.61×

5.79×

2.88×

2.18×

1.71×

1.04×

thread ≤ 24

trsv f32/f64

1.5/4.6×

2.2/4.6×

2.0/2.0×

2.3/1.5×

1.7/1.2×

0.2×

thread ≤ 16

dot f64

2.32×

2.33×

2.31×

1.96×

1.53×

1.30×

thread ≤ 32

posv f32

3.21×

3.44×

1.73×

1.16×

0.66×

0.35×

thread ≤ 12

gemm (both)

0.6–0.9×

0.4–0.5×

never

The factor/solve chain is where the tier earns its keep — a warp-per-problem potrf at N ≤ 7 idles most lanes on the serial pivot; one-problem-per-thread keeps all 32 busy. f64 amplifies and extends the win past the register ceiling (even the spilled thread path beats the alternatives through N=24). gemm is the anti-case: enough work per element that the parallel tiers always win. Columns past the ceiling price the local-memory spill honestly (down to 0.1× — measured, never shipped).

Warp-scope vendor A/B (audited characterization)#

bench_nvwarp_l1.cu A/Bs glass::warp:: against glass::nvidia::warp:: (CUB WarpReduce) for the three ops the vendor warp tier ships (reduce/dot/nrm2) at identical launch shape. The 2026-08-14 sm_120 quiet run used the shared warm-up, spread, provenance, and isolation protocol. Of 126 cells, 121 tie within ±2%. Four high-throughput f64 dot cells favor CUB by 2.0–3.0%; the sole SIMT verdict is a tiny latency cell with 17.45% SIMT spread and is not actionable. Five individual contender samples exceeded 5% spread, but no shipped dispatch decision depends on this characterization.

Robotics micro-ops — audited characterization#

The 2026-08-14 sm_120 quiet run measured 8192 problems and 1000 reps per trial through the shared audited protocol:

  • The thread tier wins every fixed-size row for both dtypes.

  • softmax_n16 is dtype-sensitive: warp wins f32, thread wins f64.

  • At the best thread launch, fused and composed spatial forms are nearly tied in f32; the fused f64 forms are about 15–22% faster.

  • Only 2 of 136 contender samples exceed 5% spread, neither changing a clear row verdict. argmax_fast is block-only and does not beat the ordinary thread-tier argmax.

These are characterization results, not generated dispatch inputs.

The full measured-results archive#

The machine-refreshed verdict tables for every sweep live in one file, bench/RESULTS.md. The ladder, body, blas2, rect, solvers, and reduced marker blocks are rewritten by bench/tune.py; nvwarp and robotics are hand-run audited characterizations and do not rewrite dispatch. Raw captures are archived externally with the paper materials; the paper harnesses are documented in bench/PAPER_SWEEPS.md.

See Tuning for Your Hardware for how to emit a per-host override table from a sweep, and Execution Plans (glass-defaults.cuh) for the picker API.