Backend Sweep Results (Historical Archive)
============================================================
.. important::
The current release figures and numerical claims are collected in
:doc:`paper_results` and on the `GLASS project page
`_. This page preserves earlier campaigns and
their original hardware, protocols, and conclusions. In particular, the
July host/fusion figures and August micro-operation results below are not
the paper's Orin evaluation. Historical conclusions should not be used as
current cross-platform recommendations.
GLASS ships overlapping thread-, warp-, block-, and vendor-backed execution
tiers. For operations present in more than one tier, the fastest choice depends
on operation, size, dtype, launch shape, and batch regime. The **mega sweep**
(``bench/tune.py``'s ladder leg) compares supported contenders — this is
exactly the measurement behind ``glass-defaults.cuh``'s ``recommend<>()``
(see :doc:`../../api_reference/defaults`).
.. warning::
This page is a dated measurement archive, not a universal performance
promise. Several legacy sections predate the current warm-up, trial-spread,
source-digest, and signed-receipt protocol and are marked accordingly. Use
the source capture named beside a claim, and rerun on the target machine.
The figures and table below come from the 2026-09-02 RTX 5090 / sm_120
five-backend selection capture — the same paired fresh-input A/B campaign
that generated the shipped ``ideal_sm120`` / ``ideal_sm87`` / ``ideal_sm72``
tables (an independent same-protocol second capture confirmed the
selections) — shown across three batch regimes: **NPROB=64** (low batch,
latency-leaning), **NPROB=1024** (mid), and **NPROB=8192** (throughput). The
winner can shift with batch
size, so low-batch plots describe that workload rather than overriding the
throughput table. These are committed static assets — regenerate them for your
own hardware with::
python bench/tune.py --sm auto # remeasures + regenerates tables AND figures
# or just the figures from an existing sweep .txt:
python bench/export_sweep_figures.py bench/mega_sweep_*.txt
``bench/explore_sweep.ipynb`` is the interactive version of the same analysis.
The ladder — ns/problem vs N, per backend
------------------------------------------
Lower is faster. Each subplot is one op; the curves are native ``warp``
(green), ``block`` (blue), and ``thread`` (orange), plus NVIDIA ``block``
(red) and ``thread`` (purple). Both thread curves pack one problem per CUDA
thread. The native thread implementation is instantiated through ``N=64`` in
this harness; NVIDIA thread LAPACK is instantiated through ``N=32``. A stopped
curve therefore marks an explicit instantiation or resource boundary, not a
claim that larger sizes are unsupported by CUDA or cuSOLVERDx. NVIDIA block
f64 curves may stop at their shared-memory feasibility limit.
``recommend<>()`` is keyed on the **NPROB=8192** throughput regime.
Its policy first applies a ±2% tie band among dependency-free SIMT contenders,
then requires an NVIDIA contender to clear the best native choice by more than
5%. For destructive POTRF/TRSV/POSV, every displayed and selected candidate is
remeasured by the unified solver ladder on fresh valid inputs. The plans are
randomized within paired rounds and every raw sample is retained; all native
and NVIDIA candidates use the same selection rule. These rules keep small,
noisy, or mutation-dependent gaps out of the generated table.
float32
~~~~~~~
.. image:: /_static/mega_sweep_ladder_f32_n64.png
:alt: f32 ladder, NPROB=64
:width: 100%
.. image:: /_static/mega_sweep_ladder_f32_n1024.png
:alt: f32 ladder, NPROB=1024
:width: 100%
.. image:: /_static/mega_sweep_ladder_f32_n8192.png
:alt: f32 ladder, NPROB=8192 (throughput — feeds recommend)
:width: 100%
float64
~~~~~~~
.. image:: /_static/mega_sweep_ladder_f64_n64.png
:alt: f64 ladder, NPROB=64
:width: 100%
.. image:: /_static/mega_sweep_ladder_f64_n1024.png
:alt: f64 ladder, NPROB=1024
:width: 100%
.. image:: /_static/mega_sweep_ladder_f64_n8192.png
:alt: f64 ladder, NPROB=8192 (throughput — feeds recommend)
:width: 100%
Raw ladder winner per (op, N), per regime
------------------------------------------
The policy winner at each ``(op, N)`` is listed for all three NPROB regimes.
Non-destructive rows come from the general ladder; POTRF/TRSV/POSV rows come
from its fresh-input solver companion. Thus the plots, winner list, and shipped
solver defaults all consume the same authoritative measurements.
The broad high-batch shape is mixed by design: native thread dominates many
small packed problems, native warp/block remain important as work grows, and
the NVIDIA block and thread implementations take measured factor/solve bands.
The ``NPROB=8192`` blocks shown here are from the same capture pair that
generated the shipped sm_120 table.
.. literalinclude:: /_static/sweep_winners.txt
:language: text
The same ladder at the edge (Jetson AGX Orin)
----------------------------------------------
The identical sweep on the Jetson AGX Orin (sm_87, 2026-09-01 selection
capture, throughput regime) crowns different winners in **145 of 396** cells
versus the RTX 5090 above, and the older AGX Xavier (sm_72, native-only)
differs from the Orin in **162 of 396** — placement is a property of the
silicon, and getting it wrong is expensive: on the Orin's raw ladder the best
and worst placements for a cell differ by a median of 4.9× (up to 81×). See
:ref:`tuning-per-arch-results` for the cross-architecture analysis.
.. image:: /_static/mega_sweep_ladder_f32_sm87.png
:alt: f32 ladder on Jetson AGX Orin, NPROB=8192 (throughput — feeds recommend)
:width: 100%
.. image:: /_static/mega_sweep_ladder_f64_sm87.png
:alt: f64 ladder on Jetson AGX Orin, NPROB=8192 (throughput — feeds recommend)
:width: 100%
vs. host-batched cuBLAS/cuSOLVER (and TF32)
-------------------------------------------
The ladder above compares *device-side* backends. A separate question is how
one-block-per-problem GLASS compares to the standard host-side recipe — a
single ``cublasgemmStridedBatched`` / ``cusolverDnpotrfBatched``
(+ ``potrsBatched``) call over the whole batch.
``bench/bench_paper_hostblas.cu`` measures exactly that: gemm / potrf / posv,
``N`` = 4–64, batch ``B`` = 1–8192, both precisions (raw capture committed as
``paper_hostblas_20260708_0054.txt`` (archived externally); RTX 5090 / sm_120, quiet GPU).
.. image:: /_static/hostblas_speedup.png
:alt: host-batched vendor time divided by best GLASS time, vs batch size
:width: 100%
Above 1.0 = the best swept GLASS configuration faster (fp32 shown; GLASS is
best of block32, block128, and warp8 where available). The vendor side is one
default host-API configuration, so this is not a symmetric tuner comparison or
a universal "GLASS vs vendor" statement. In this historical capture, at robot
sizes, host batching never catches up: gemm at ``N`` ≤ 24 and the full
factor-and-solve (posv) through ``N`` = 64 are GLASS wins at **every** batch
size, reaching 2.9–6.3× at saturation. The vendor's best regime is mid-batch
(``B`` ≈ 64–1024), where it briefly leads standalone potrf at mid sizes; only
gemm at ``N`` ≥ 32 is an outright vendor win at scale — the same mid-band the
ladder already routes to ``glass::nvidia::``.
An audited 2026-08-14 confirmation used 500 reps. Only 8/378 selected
GLASS/vendor pairs had a contender above 5% within-run spread, and all eight
had decisive 1.56×–2.19× gaps. One small f64 ``potrf`` comparison changed
winner between two quiet captures despite low within-run spread. Accordingly,
small deltas and cross-run winner changes are treated as ties; publication
claims must survive both independent captures.
The 2026-09 recapture campaign repeated this comparison under the current
median-of-trials protocol on all three measured architectures and confirmed
the shape while sharpening the edge story: the host-dispatch overhead that
GLASS avoids is far more expensive on an embedded GPU, so the same
comparison that peaks near 16× on the RTX 5090 peaks over 100× on the
Jetson AGX Orin (small-N ``posv``, where per-call host overhead dominates).
Raw captures are archived with the paper materials.
Permitting TF32 tensor cores (dashed) does not change the story: cuBLAS
*declines to engage them* below ``N`` = 24 (results bit-identical to FP32),
and where they do engage the speed is a wash against FP32 cuBLAS while max
error jumps three orders of magnitude (~1e-7 → ~2e-4) — unusable for the
Cholesky-chain ops, which have no TF32 cuSOLVER path at all.
Fusion: ``riccati_gain`` vs. a 7-call vendor chain
--------------------------------------------------
``glass::riccati_gain`` computes the LQR feedback gain
``K = (R + BᵀPB)⁻¹(BᵀPA)`` in one kernel with all intermediates in shared
memory; the host-batched equivalent is seven vendor calls (four gemms, a
batched Cholesky, two triangular solves) with intermediates in global memory.
``bench/bench_paper_fusion.cu`` compares them (capture
``paper_fusion_20260708_0055.txt`` (archived externally)):
.. image:: /_static/fusion_speedup.png
:alt: fused riccati_gain vs 7-call host-batched vendor chain
:width: 75%
:align: center
In this historical implementation/capture, fusion wins at every batch size at
quadrotor/manipulator scale — 2.5–2.8× at
``(nx,nu)`` = (12,4) and 1.6–1.9× at (14,7) in fp32, more in fp64 — but the
chain wins at (36,12) fp32 and (48,16), where the staged operands outgrow what
one block overlaps profitably. Fusion is a measured choice, not a default:
GLASS composes both forms from the same primitives.
The harness now also carries a third, *device-side unfused* arm — the same
GLASS primitives launched one kernel per step with intermediates in global
memory — which isolates the fusion effect from the device-vs-host effect. On
the Jetson AGX Orin (2026-09 captures), fusion wins by up to ~1.75× at small
batches in fp32, while the unfused composition wins by up to ~2.9× at large
batches and in fp64, where per-primitive kernels regain the occupancy the
fused kernel's shared-memory footprint caps. Which side of that tradeoff a
given (size, batch, precision) lands on is exactly the kind of measured
placement decision the rest of this page is about.
Single-call latency
-------------------
For a batch of **one** (a high-rate MPC tick), wall-clock per-call latency is
what matters. The non-batched vendor calls pin an essentially flat API floor —
~7–10 µs (``cublasSgemm``), ~15.5 µs (potrf), ~23–30 µs (posv) — while a GLASS
call starts at 5.2 µs and grows with compute, so GLASS wins single-call
latency through ``N`` = 32 (gemm), 12 (potrf), and 24 (posv); at ``N`` = 8 the
full factor-and-solve is 2.4× faster (9.6 vs 23.5 µs). And inside your own
kernel, composed GLASS calls never pay the API floor again.
Both harnesses live in ``bench/`` and rerun via
``python3 bench/paper_sweeps.py`` (see ``bench/PAPER_SWEEPS.md``).
Choosing among the dense solve paths (measured guidance)
---------------------------------------------------------
GLASS deliberately ships **no** auto-dispatch between its linear-system
solvers — the right choice depends on structure and conditioning, which a
compile-time table cannot see. Instead, ``bench/tune.py --legs solvers``
measures the trade-offs on your GPU and records them in
``bench/RESULTS.md`` (solvers section). The RTX 5090 numbers (NPROB=8192,
ns/problem):
**SPD single solve — use ``posv``; the alternatives price as follows.**
``gesv`` (pivoted LU, the robustness fallback) costs 1.0–2.1× ``posv`` in
fp32 at ``N`` ≥ 16 (1.8–2.1× at 32–64); the invert-then-multiply
anti-pattern (``inv`` + ``gemv``) costs 1.9–4.7× at ``N`` ≥ 16. **Below**
``N`` = 16 all three paths are within a few nanoseconds of each other (and
``inv``/``gesv`` can even edge out ``posv``, especially in fp64) — there,
choose by *numerics*, not speed: Cholesky is backward-stable on SPD input and
fails loudly (with ``CHECK``) on indefinite input, while an explicit inverse
amplifies conditioning error silently. Speed only ever argues *for* ``posv``,
never against it.
**Block-tridiagonal chains — ``bdsv`` (direct) vs ``pcg`` (iterative) is
problem-dependent; do not hard-code either.** On the benchmark's synthetic
diagonally-dominant SPD system, PCG converges in about three iterations and
wins 11 of 12 cells (up to 9×) at the harness's ``rho = rᵀz`` relative
tolerance of 1e-6. This is an approximate-solve comparison, not matched final
residual accuracy: the direct f64 solves reach near-machine precision while
PCG stops at its requested tolerance. At
(BlockSize=12, Knots=16) fp32 the direct sweep wins 1.8×. PCG's cost scales
linearly with its iteration count, so an ill-conditioned Riccati/KKT chain
(10–100× more iterations) moves the crossover proportionally toward
``bdsv`` — read the ``pcg iters`` column of your own sweep before
generalizing, or measure with your actual matrices.
**``syev`` / ``eig_clamp``** — the historical harness compared only this serial
Jacobi family. It did not include the parallel fixed-sweep ``eigh`` /
``psd_project`` family, so it cannot support a recommendation between them.
NVIDIA thread LAPACK — measured integration
---------------------------------------------
The 2026-08-30 wave added cuSOLVERDx's thread interface as a fifth contender
for ``potrf``, ``trsv``, and ``posv``. It is a selective win, not a replacement
for either GLASS's native thread code or cuSOLVERDx's block interface:
.. list-table:: Throughput cells selected by the shipped 2026-09-02 tables
:header-rows: 1
:widths: 16 14 30 40
* - GPU
- cells
- float32
- float64
* - RTX 5090 / sm_120
- 12 / 132
- ``potrf`` N=6,8
- ``potrf`` N=4–8; ``trsv`` N=8–24; ``posv`` N=4–8
* - Jetson AGX Orin / sm_87
- 18 / 132
- ``potrf`` N=4,8,12; ``trsv`` N=8–32
- ``potrf`` N=4–8; ``trsv`` N=4–32
These cells were selected by the unified fresh-input solver ladder described
above — every native and NVIDIA candidate measured symmetrically on fresh
valid inputs, with a selected NVIDIA plan additionally interval-confirmed
(the sm_120 regeneration demoted one ambiguous vendor pick to the capture's
native winner under that rule). The AGX Xavier (sm_72) ships no
NVIDIA-thread cells: its CUDA 11.4 toolchain has no device-callable MathDx.
The ranges enumerate the measured sizes ``4, 6, 8, 12, 16, 24, 32``; they
do not imply testing every intervening integer. Against the fastest native
plan in the same generating captures, the selected NVIDIA-thread cells win
by about 1.09× to 2.35× on the RTX 5090 and 1.09× to 3.54× on Orin. The
largest gains are concentrated in ``trsv``; other operations and sizes still
select another tier. Each table is supported by an independent same-protocol
second capture (the A/B pair), and the archived 2026-08-30 release
measurements — a different confirmation method — selected nearly the same
bands with near-identical win ranges, so the bands are stable across both
methods and across independent quiet-window sessions. These checks support
the selected bands, not a portable speedup claim — rerun both ladder
components on a new architecture.
The native thread tier — historical characterization
------------------------------------------------------
The 2026-07-19 full-domain sweep (quiet RTX 5090) first added the native ``thread``
contender at every ``(op, N)`` point. Throughput regime (NPROB=8192), thread
vs the best other tier (ratio > 1 = thread faster; **bold** cells shipped):
.. list-table::
:header-rows: 1
* - op / dtype
- N=4
- N=6
- N=8
- N=12
- N=16
- N=24
- shipped band
* - posv f64
- **6.17×**
- **7.46×**
- **5.03×**
- **3.05×**
- **2.26×**
- **1.19×**
- thread ≤ 24
* - chol f64
- **5.61×**
- **5.79×**
- **2.88×**
- **2.18×**
- **1.71×**
- **1.04×**
- thread ≤ 24
* - trsv f32/f64
- **1.5/4.6×**
- **2.2/4.6×**
- **2.0/2.0×**
- **2.3/1.5×**
- **1.7/1.2×**
- 0.2×
- thread ≤ 16
* - dot f64
- **2.32×**
- **2.33×**
- **2.31×**
- **1.96×**
- **1.53×**
- **1.30×**
- thread ≤ 32
* - posv f32
- **3.21×**
- **3.44×**
- **1.73×**
- **1.16×**
- 0.66×
- 0.35×
- thread ≤ 12
* - gemm (both)
- 0.6–0.9×
- 0.4–0.5×
- —
- —
- —
- —
- never
The factor/solve chain is where the tier earns its keep — a warp-per-problem
``potrf`` at N ≤ 7 idles most lanes on the serial pivot; one-problem-per-thread
keeps all 32 busy. f64 amplifies and extends the win past the register ceiling
(even the spilled thread path beats the alternatives through N=24). ``gemm``
is the anti-case: enough work per element that the parallel tiers always win.
Columns past the ceiling price the local-memory spill honestly (down to 0.1× —
measured, never shipped).
Warp-scope vendor A/B (audited characterization)
-------------------------------------------------
``bench_nvwarp_l1.cu`` A/Bs ``glass::warp::`` against ``glass::nvidia::warp::``
(CUB WarpReduce) for the three ops the vendor warp tier ships (reduce/dot/nrm2)
at identical launch shape. The 2026-08-14 sm_120 quiet run used the shared
warm-up, spread, provenance, and isolation protocol. Of 126 cells, 121 tie
within ±2%. Four high-throughput f64 ``dot`` cells favor CUB by 2.0–3.0%; the
sole SIMT verdict is a tiny latency cell with 17.45% SIMT spread and is not
actionable. Five individual contender samples exceeded 5% spread, but no
shipped dispatch decision depends on this characterization.
Robotics micro-ops — audited characterization
----------------------------------------------
The 2026-08-14 sm_120 quiet run measured 8192 problems and 1000 reps per trial
through the shared audited protocol:
- The thread tier wins every fixed-size row for both dtypes.
- ``softmax_n16`` is dtype-sensitive: warp wins f32, thread wins f64.
- At the best thread launch, fused and composed spatial forms are nearly tied
in f32; the fused f64 forms are about 15–22% faster.
- Only 2 of 136 contender samples exceed 5% spread, neither changing a clear
row verdict. ``argmax_fast`` is block-only and does not beat the ordinary
thread-tier ``argmax``.
These are characterization results, not generated dispatch inputs.
The full measured-results archive
----------------------------------
The machine-refreshed verdict tables for every sweep live in one file,
``bench/RESULTS.md``. The ladder, body, blas2, rect, solvers, and reduced marker
blocks are rewritten by ``bench/tune.py``; nvwarp and robotics are hand-run
audited characterizations and do not rewrite dispatch.
Raw captures are archived externally with the paper materials; the paper
harnesses are documented in ``bench/PAPER_SWEEPS.md``.
See :doc:`../concepts/tuning` for how to emit a per-host override table from a
sweep, and :doc:`../../api_reference/defaults` for the picker API.