Benchmarks#
Warning
This page includes dated development measurements, not the release benchmark. Some historical reports compare different timing boundaries or model variants; do not reuse their speedups as headlines. The audited release collection is on the release measurements page.
GRiD ships a benchmark harness that compares GRiD against CPU and GPU
baselines (Pinocchio, MJX, Frax, BARD) and against historical GRiD
reference points. BARD (PyTorch) is timed on both torch CPU and CUDA, so
it contributes bard_cpu / bard_gpu columns to the report. The
harness lives under
test/benchmarks/.
Quick start#
The full multi-baseline sweep:
.venv/bin/python test/benchmarks/run_benchmarks.py
A single GRiD cell for fast iteration (per-exe path: one TU / exe / process per algorithm, RAM-safe and crash-isolated):
.venv/bin/python test/benchmarks/per_algo_bench.py --robot iiwa14 --base fixed
A multi-version comparison sweep (multiple GRiD configurations + baselines) producing the canonical markdown report:
.venv/bin/python test/benchmarks/run_multi_version.py \
--columns glass --robots iiwa14 go2 g1 --bases fixed floating \
--single-call-iters 50000 --batch-iters 500 \
--output-dir test/benchmarks/results/comparison_<tag> \
--report test/benchmarks/benchmark_multi_version_<tag>.md
The full sweep takes ~1.5 hours on an RTX 5090. cuBLASDx (glass_nvidia)
was removed in v2.0 of the codegen; see
cuBLASDx Removal & Any-Thread-Count Library Functions for the rationale and the
archive/last-cublasdx git tag for the historical comparison data.
Algorithms measured#
The bench reports 13 rows per cell:
First-order set (10 algorithms):
inverse_dynamics— inverse dynamics (RNEA).minv— direct mass-matrix inverse.forward_dynamics— forward dynamics viaMinv·(τ − c).aba— articulated body algorithm (independent forward dynamics).crba— composite rigid body algorithm (mass matrix).inverse_dynamics_gradient— ∂(inverse_dynamics)/∂(q, qd).forward_dynamics_gradient— ∂(forward_dynamics)/∂(q, qd).end_effector_pose— end-effector pose.end_effector_pose_gradient— end-effector Jacobian.end_effector_pose_hessian— end-effector Hessian (∂²ee/∂q²).
Second-order set (4 algorithms, 3 rows per cell):
idsva_so— second-order inverse dynamics, codegen-time-dispatched toidsva_so_body_frame(fixed-base) oridsva_so_world_frame(floating-base). The dispatcher row reports whichever variant the codegen picked for that cell.idsva_so_body_frame— body-frame propagation, multi-pass.idsva_so_world_frame— world-frame propagation, single-pass.fdsva_so— second-order forward dynamics.
The two IDSVA-SO variants are mathematically equivalent and ship side-by-side so the report shows the body-vs-world crossover; per cell the standalone variant that is bit-identical to the dispatcher’s pick is deduplicated (its row would time the exact same kernel twice), so the four second-order algorithms contribute 3 rows — 13 rows in all. See IDSVA / IDSVA-SO (Inverse Dynamics, Second-Order) for the dispatch rule details.
Beyond the 13 core rows, the per-algo harness’s PER_ALGO_SPECS table
(test/benchmarks/baselines/grid/run.py) drives ~42 benchable keys in
total — the f_ext gradients, inverse_dynamics_regressor,
forward_dynamics_parameter_gradient, the energy regressors, the
integrator family, multi_target_position/_gradient, the runtime-EE
and frame-Jacobian families, the centroidal set, and the mjx twins — so
any of those can be timed the same way.
GRID_BENCH_EXTRA_NVCC_FLAGS is the compile-flag env seam: extra
nvcc flags for the bench build (used for compile-flag A/Bs). The value
is folded into the bench’s content stamp, so changing the flags correctly
invalidates the cached binaries instead of reusing a stale build.
Output schema#
Each cell produces a JSON file at
<output-dir>/<robot>_<base>_grid_<column>.json. Each row in the JSON
captures three measurements per algorithm:
single-call — wall time for one host-level invocation (one timestep). Anti-LICM machinery prevents the compiler from hoisting the call out of the rep-loop, so this is a real per-call cost, not amortized.
N=16 (compute) — kernel-only time for a 16-timestep batch (excludes host→device / device→host transfer).
N=256 (compute) — kernel-only time for a 256-timestep batch.
Pinocchio numbers are batch with-memory (CPU codegen makes compute/transfer inseparable).
Stability and compile flags#
A flag exists for working around hardware/toolchain quirks; the default is correct on sm_120 (RTX 5090) and likely fine on sm_8x:
--ptxas-opt-level 2— drops ptxas from -O3 to -O2, used when nvcc ptxas itself crashes (rare).
Per-host autotuning (recommended for production)#
For best performance on a given GPU, run the GLASS autotuner once per
host. It writes a per-host override file under
external/GLASS/bench/tuning/<hostname>.cuh (the shipped table is left
untouched):
cd external/GLASS
python3 bench/autotune.py --sm AUTO
Wall time: ~10–15 minutes. Consume the override on subsequent builds
with -DGLASS_TUNING_TABLE_LOCAL='"external/GLASS/bench/tuning/<hostname>.cuh"'.
See Installation for details.
Autotune launch config for your robot / GPU#
GRiD kernels are single-block and thread-count-invariant (same result
at any block size), so the optimal (resource_tier, threads_per_block)
for each algorithm is a pure performance choice that depends on the
robot (DoF / topology) and the GPU. GRiD ships measured-optimal
launch configs under config/launch_configs/<robot>/<gpu>.json; codegen bakes
the matching file into grid_launch_config.cuh and the host launchers
(and therefore the python / jax / torch bindings) default their launch
config from it. With no entry for your (robot, GPU), GRiD falls back to a
conservative — still correct, just slower — default.
To autotune your robot on your GPU and write the override:
bash config/autotune_robot.sh <robot> [fixed floating]
# examples
bash config/autotune_robot.sh iiwa14 # fixed + floating
bash config/autotune_robot.sh go2 floating # floating only
GPU_KEY=a40_sm86 bash config/autotune_robot.sh g1 # force the GPU key
This:
Detects your GPU (
nvidia-sminame + compute capability) and derives the GPU key<model>_sm<arch>(e.g.rtx5090_sm120). If detection fails, setGPU_KEY=<model>_sm<arch>and re-run.Runs the GRiD autotune sweep (
per_algo_bench.py --mode autotune --stage sweep) for that robot + bases. The build is RAM-safe serial (--compile-jobs 1, one per-algo TU at a time) so the big-robot second-order TUs never OOM the box.Converts the swept winners into
config/launch_configs/<robot>/<gpu>.jsonin the documented schema (gpu,cuda_arch,gpu_name,autotune_N,source,bases).
Then rebuild to pick it up: re-run codegen for the robot (codegen
auto-discovers the new launch_configs file) and rebuild GRiD / the
bindings as usual. The host launchers will default to your tuned
(tier, threads).
Note
Single-call timing is OFF by default (it is hard to time and needs
the -rdc shim). The autotuner tunes on batch timing (N=256 by
default; override with AUTOTUNE_N), which is what the host launch
config should optimize for. Opt single-call timing back in with
--single-timing / GRID_BENCH_SINGLE_TIMING=1 only if you
specifically need it.
Run the sweep on a quiet GPU — timing must be isolated, so close other GPU workloads first.
Crowdsource it (please do!): once you have a good config, open a PR
adding config/launch_configs/<robot>/<gpu>.json (one file per (robot, GPU)) so
everyone defaults to fast launches on your hardware. No code changes are
needed — codegen auto-discovers the file. See
config/launch_configs/README.md
for the full contribution checklist (GPU model, driver / CUDA version,
robot DoF / base to include in the PR description).
Cheap re-bake: refresh_launch_configs.sh#
config/refresh_launch_configs.sh re-bakes the live
config/launch_configs/<robot>/<gpu>.json picks from an existing
tier-sweep without the day-long recompile: it harvests the latest
results/tier_sweep_phased_* sweep (or a --sweep-dir you pass),
merges the picks, and bakes the per-robot bases (optionally
ffi_bases with --with-ffi), printing a git diff of
config/launch_configs/ at the end (it never commits).
Which one when: autotune_robot.sh builds a full monolithic
per-tier binary, so big robots (g1 / h2_plus) hit the ~hour
second-order recompile — use it for a robot/GPU with no prior sweep.
refresh_launch_configs.sh reuses the cached split (noSO/SO)
binaries from a phased tier sweep
(bash test/benchmarks/run_tier_sweep_phased.sh), so re-baking after
a sweep costs no recompile at all.
Pre-GLASS regression check#
The multi-version harness supports an automatic regression check
against the pre-GLASS reference commit (d2c0d18). Add pre_glass
to --columns and the harness will:
Create a worktree at
../GRiD-A2R-pre-glass(or$GRID_PRE_GLASS_WORKTREE).Initialize submodules pinned to their pre-GLASS revisions.
Run the same cell list using that frozen harness.
The pre_glass column is fixed-base only — the pre-GLASS bench
harness predates floating-base support.
GPU-resident pipelines (no-transfer timing)#
GRiD’s jax/torch handles can keep an entire control/rollout loop on the GPU —
inputs, dynamics calls, and downstream math never round-trip through host
memory. Two example pairs demonstrate and time this
(bindings/examples/jax_gpu_resident.py / torch_cuda_graphs.py for the
iiwa14, and ..._go2.py twins for the floating-base go2), driven by
test/benchmarks/gpu_resident_timing.py.
Measured on the RTX 5090 (sm_120, 2026-08-09,
results/overnight_20260809/legC_gpu_resident.json):
JAX resident rollout vs host round-trip (a
lax.scanrollout calling GRiD’s dynamics each step): iiwa14 24.3x / 14.5x / 7.3x faster at batch 64 / 256 / 1024 (1.5 ms vs 36.6 ms at B=64); go2-floating 6.8x / 6.9x / 3.6x (the floating rollout drives GRiD’s ownintegratorkernel for the on-manifold base retract inside the scan).Torch CUDA graphs: replay wall-time is roughly break-even with eager (0.91-0.99x) — the win is in CPU submission cost, ~13 us eager vs ~1.9 us replay (~6.6x), which is what matters when the CPU is the bottleneck feeding a real-time control loop. (An earlier “replay slower than eager” reading was a harness artifact — it timed a per-iteration device copy-in; replay-only and copy-in variants are now recorded separately.)
Reproduce with .venv/bin/python test/benchmarks/gpu_resident_timing.py
(quiet GPU; results land under test/benchmarks/results/).
Second-order derivatives: the transfer question#
Second-order derivative outputs scale as nv³ per timestep (g1’s
idsva_so at N=256 is a 175.6 MB result batch), so for these algorithms
the device→host copy — not the kernel — is the dominant with-memory cost.
On the g1 idsva_so N=256 cell the measured transfer delta was 13.2 ms
of a 15.2 ms with-mem total (87 %) with pageable host buffers, and is
still ~80 % of the (much smaller) total after pinning.
Two fixes ship for this:
Pinned host staging —
grid_host_alloc(cudaMallocHost) replaced pageablemallocfor everygridDatahost buffer, raising effective D2H bandwidth on large payloads from 13.3 to ~21 GB/s (the box’s PCIe-lane ceiling). That alone flipped 8 with-mem cells vs Pinocchio (CPU, 24 threads) (11 losses + 2 ties → 5 losses + 8 wins); the per-cell before/after table is the “W2a re-capture” section oftest/benchmarks/results/competitive_render_20260826/competitive_analysis.md.Not copying at all — keep the tensor on the GPU (next section).
Staying resident (jax)#
The grid_rbd jax FFI path is device-in / device-out: under jit /
lax.scan composition the nv³ tensor is produced on-device and consumed
by the next pipeline stage (a Hessian-vector product, a DDP backward pass)
with zero host transfer. Layer-3 timings (jax FFI, medians of 30 iters
after warmup, RTX 5090,
test/benchmarks/results/night4_20260827/so_resident_results.json), in
total-batch µs — resident = device-in/device-out, out_host = outputs
pulled to host, roundtrip = numpy-in → numpy-out:
robot |
op |
B |
resident |
out_host |
roundtrip |
|---|---|---|---|---|---|
iiwa14 |
idsva_so |
16 |
515 |
643 |
790 |
iiwa14 |
fdsva_so |
16 |
526 |
657 |
799 |
iiwa14 |
idsva_so |
256 |
519 |
807 |
990 |
iiwa14 |
fdsva_so |
256 |
529 |
831 |
997 |
go2 |
idsva_so |
16 |
784 |
1081 |
1140 |
go2 |
fdsva_so |
16 |
1018 |
1333 |
1474 |
go2 |
idsva_so |
256 |
990 |
4783 |
4999 |
go2 |
fdsva_so |
256 |
1412 |
4861 |
5137 |
g1 |
idsva_so |
16 |
2003 |
3529 |
3801 |
g1 |
fdsva_so |
16 |
5893 |
8310 |
8417 |
g1 |
idsva_so |
256 |
2866 |
35024 |
34551 |
g1 |
fdsva_so |
256 |
13949 |
45340 |
46514 |
Reading the table:
The resident column carries a ~0.4–0.5 ms fixed per-call FFI dispatch overhead (visible as the iiwa14 floor: kernel compute is ~43–72 µs, resident is ~515–530 µs). It is a per-call cost, so it amortizes when the call sits inside a jitted
scan/rollout rather than being invoked from python once per step.The resident-vs-out_host gap is pure output transfer, and it grows with nv³·B exactly as predicted: negligible for iiwa14, ~3.8 ms for go2
idsva_soat N=256, ~32 ms for g1.The honest headline for the largest tensor: g1
idsva_soN=256 resident 2.87 ms vs Pinocchio (CPU, 24 threads) 2.38 ms total — near parity on wall time, while the result is already on the GPU for the next pipeline stage instead of in host memory. If the consumer is on the GPU, resident wins outright; if the result must land in host memory, CPU codegen remains the right tool for this cell.
Note
Why the with-mem comparison is asymmetric by construction. A CPU
baseline pays no output copy — its results are born in host memory,
so its with-mem time equals its compute time. The GPU with-mem column
therefore carries the entire PCIe cost on top of compute. The
three-layer reporting contract (compute-only / C++ with-mem /
through-bindings; see test/benchmarks/BENCHMARK_METHODOLOGY.md)
exists precisely to keep that visible: layer 1 is the GPU-resident
design point, layer 2 isolates transfer, layer 3 is what an adopter
pays through the python bindings.
See also#
test/benchmarks/README.md — full bench-harness reference (Pinocchio / MJX / Frax / BARD setup, etc.).
To run a full sweep yourself, use the harness under test/benchmarks/; results are written under
test/benchmarks/results/locally (hardware-specific, e.g. sm_120).CUDA Support Status — currently exercised CUDA paths.
IDSVA / IDSVA-SO (Inverse Dynamics, Second-Order) and FDSVA-SO (Forward Dynamics, Second-Order) — algorithm details.