Benchmarks#

Warning

This page includes dated development measurements, not the release benchmark. Some historical reports compare different timing boundaries or model variants; do not reuse their speedups as headlines. The audited release collection is on the release measurements page.

GRiD ships a benchmark harness that compares GRiD against CPU and GPU baselines (Pinocchio, MJX, Frax, BARD) and against historical GRiD reference points. BARD (PyTorch) is timed on both torch CPU and CUDA, so it contributes bard_cpu / bard_gpu columns to the report. The harness lives under test/benchmarks/.

Quick start#

The full multi-baseline sweep:

.venv/bin/python test/benchmarks/run_benchmarks.py

A single GRiD cell for fast iteration (per-exe path: one TU / exe / process per algorithm, RAM-safe and crash-isolated):

.venv/bin/python test/benchmarks/per_algo_bench.py --robot iiwa14 --base fixed

A multi-version comparison sweep (multiple GRiD configurations + baselines) producing the canonical markdown report:

.venv/bin/python test/benchmarks/run_multi_version.py \
    --columns glass --robots iiwa14 go2 g1 --bases fixed floating \
    --single-call-iters 50000 --batch-iters 500 \
    --output-dir test/benchmarks/results/comparison_<tag> \
    --report test/benchmarks/benchmark_multi_version_<tag>.md

The full sweep takes ~1.5 hours on an RTX 5090. cuBLASDx (glass_nvidia) was removed in v2.0 of the codegen; see cuBLASDx Removal & Any-Thread-Count Library Functions for the rationale and the archive/last-cublasdx git tag for the historical comparison data.

Algorithms measured#

The bench reports 13 rows per cell:

First-order set (10 algorithms):

  • inverse_dynamics — inverse dynamics (RNEA).

  • minv — direct mass-matrix inverse.

  • forward_dynamics — forward dynamics via Minv·(τ − c).

  • aba — articulated body algorithm (independent forward dynamics).

  • crba — composite rigid body algorithm (mass matrix).

  • inverse_dynamics_gradient — ∂(inverse_dynamics)/∂(q, qd).

  • forward_dynamics_gradient — ∂(forward_dynamics)/∂(q, qd).

  • end_effector_pose — end-effector pose.

  • end_effector_pose_gradient — end-effector Jacobian.

  • end_effector_pose_hessian — end-effector Hessian (∂²ee/∂q²).

Second-order set (4 algorithms, 3 rows per cell):

  • idsva_so — second-order inverse dynamics, codegen-time-dispatched to idsva_so_body_frame (fixed-base) or idsva_so_world_frame (floating-base). The dispatcher row reports whichever variant the codegen picked for that cell.

  • idsva_so_body_frame — body-frame propagation, multi-pass.

  • idsva_so_world_frame — world-frame propagation, single-pass.

  • fdsva_so — second-order forward dynamics.

The two IDSVA-SO variants are mathematically equivalent and ship side-by-side so the report shows the body-vs-world crossover; per cell the standalone variant that is bit-identical to the dispatcher’s pick is deduplicated (its row would time the exact same kernel twice), so the four second-order algorithms contribute 3 rows — 13 rows in all. See IDSVA / IDSVA-SO (Inverse Dynamics, Second-Order) for the dispatch rule details.

Beyond the 13 core rows, the per-algo harness’s PER_ALGO_SPECS table (test/benchmarks/baselines/grid/run.py) drives ~42 benchable keys in total — the f_ext gradients, inverse_dynamics_regressor, forward_dynamics_parameter_gradient, the energy regressors, the integrator family, multi_target_position/_gradient, the runtime-EE and frame-Jacobian families, the centroidal set, and the mjx twins — so any of those can be timed the same way.

GRID_BENCH_EXTRA_NVCC_FLAGS is the compile-flag env seam: extra nvcc flags for the bench build (used for compile-flag A/Bs). The value is folded into the bench’s content stamp, so changing the flags correctly invalidates the cached binaries instead of reusing a stale build.

Output schema#

Each cell produces a JSON file at <output-dir>/<robot>_<base>_grid_<column>.json. Each row in the JSON captures three measurements per algorithm:

  • single-call — wall time for one host-level invocation (one timestep). Anti-LICM machinery prevents the compiler from hoisting the call out of the rep-loop, so this is a real per-call cost, not amortized.

  • N=16 (compute) — kernel-only time for a 16-timestep batch (excludes host→device / device→host transfer).

  • N=256 (compute) — kernel-only time for a 256-timestep batch.

Pinocchio numbers are batch with-memory (CPU codegen makes compute/transfer inseparable).

Stability and compile flags#

A flag exists for working around hardware/toolchain quirks; the default is correct on sm_120 (RTX 5090) and likely fine on sm_8x:

  • --ptxas-opt-level 2 — drops ptxas from -O3 to -O2, used when nvcc ptxas itself crashes (rare).

Autotune launch config for your robot / GPU#

GRiD kernels are single-block and thread-count-invariant (same result at any block size), so the optimal (resource_tier, threads_per_block) for each algorithm is a pure performance choice that depends on the robot (DoF / topology) and the GPU. GRiD ships measured-optimal launch configs under config/launch_configs/<robot>/<gpu>.json; codegen bakes the matching file into grid_launch_config.cuh and the host launchers (and therefore the python / jax / torch bindings) default their launch config from it. With no entry for your (robot, GPU), GRiD falls back to a conservative — still correct, just slower — default.

To autotune your robot on your GPU and write the override:

bash config/autotune_robot.sh <robot> [fixed floating]

# examples
bash config/autotune_robot.sh iiwa14             # fixed + floating
bash config/autotune_robot.sh go2 floating        # floating only
GPU_KEY=a40_sm86 bash config/autotune_robot.sh g1  # force the GPU key

This:

  1. Detects your GPU (nvidia-smi name + compute capability) and derives the GPU key <model>_sm<arch> (e.g. rtx5090_sm120). If detection fails, set GPU_KEY=<model>_sm<arch> and re-run.

  2. Runs the GRiD autotune sweep (per_algo_bench.py --mode autotune --stage sweep) for that robot + bases. The build is RAM-safe serial (--compile-jobs 1, one per-algo TU at a time) so the big-robot second-order TUs never OOM the box.

  3. Converts the swept winners into config/launch_configs/<robot>/<gpu>.json in the documented schema (gpu, cuda_arch, gpu_name, autotune_N, source, bases).

Then rebuild to pick it up: re-run codegen for the robot (codegen auto-discovers the new launch_configs file) and rebuild GRiD / the bindings as usual. The host launchers will default to your tuned (tier, threads).

Note

Single-call timing is OFF by default (it is hard to time and needs the -rdc shim). The autotuner tunes on batch timing (N=256 by default; override with AUTOTUNE_N), which is what the host launch config should optimize for. Opt single-call timing back in with --single-timing / GRID_BENCH_SINGLE_TIMING=1 only if you specifically need it.

Run the sweep on a quiet GPU — timing must be isolated, so close other GPU workloads first.

Crowdsource it (please do!): once you have a good config, open a PR adding config/launch_configs/<robot>/<gpu>.json (one file per (robot, GPU)) so everyone defaults to fast launches on your hardware. No code changes are needed — codegen auto-discovers the file. See config/launch_configs/README.md for the full contribution checklist (GPU model, driver / CUDA version, robot DoF / base to include in the PR description).

Cheap re-bake: refresh_launch_configs.sh#

config/refresh_launch_configs.sh re-bakes the live config/launch_configs/<robot>/<gpu>.json picks from an existing tier-sweep without the day-long recompile: it harvests the latest results/tier_sweep_phased_* sweep (or a --sweep-dir you pass), merges the picks, and bakes the per-robot bases (optionally ffi_bases with --with-ffi), printing a git diff of config/launch_configs/ at the end (it never commits).

Which one when: autotune_robot.sh builds a full monolithic per-tier binary, so big robots (g1 / h2_plus) hit the ~hour second-order recompile — use it for a robot/GPU with no prior sweep. refresh_launch_configs.sh reuses the cached split (noSO/SO) binaries from a phased tier sweep (bash test/benchmarks/run_tier_sweep_phased.sh), so re-baking after a sweep costs no recompile at all.

Pre-GLASS regression check#

The multi-version harness supports an automatic regression check against the pre-GLASS reference commit (d2c0d18). Add pre_glass to --columns and the harness will:

  1. Create a worktree at ../GRiD-A2R-pre-glass (or $GRID_PRE_GLASS_WORKTREE).

  2. Initialize submodules pinned to their pre-GLASS revisions.

  3. Run the same cell list using that frozen harness.

The pre_glass column is fixed-base only — the pre-GLASS bench harness predates floating-base support.

GPU-resident pipelines (no-transfer timing)#

GRiD’s jax/torch handles can keep an entire control/rollout loop on the GPU — inputs, dynamics calls, and downstream math never round-trip through host memory. Two example pairs demonstrate and time this (bindings/examples/jax_gpu_resident.py / torch_cuda_graphs.py for the iiwa14, and ..._go2.py twins for the floating-base go2), driven by test/benchmarks/gpu_resident_timing.py.

Measured on the RTX 5090 (sm_120, 2026-08-09, results/overnight_20260809/legC_gpu_resident.json):

  • JAX resident rollout vs host round-trip (a lax.scan rollout calling GRiD’s dynamics each step): iiwa14 24.3x / 14.5x / 7.3x faster at batch 64 / 256 / 1024 (1.5 ms vs 36.6 ms at B=64); go2-floating 6.8x / 6.9x / 3.6x (the floating rollout drives GRiD’s own integrator kernel for the on-manifold base retract inside the scan).

  • Torch CUDA graphs: replay wall-time is roughly break-even with eager (0.91-0.99x) — the win is in CPU submission cost, ~13 us eager vs ~1.9 us replay (~6.6x), which is what matters when the CPU is the bottleneck feeding a real-time control loop. (An earlier “replay slower than eager” reading was a harness artifact — it timed a per-iteration device copy-in; replay-only and copy-in variants are now recorded separately.)

Reproduce with .venv/bin/python test/benchmarks/gpu_resident_timing.py (quiet GPU; results land under test/benchmarks/results/).

Second-order derivatives: the transfer question#

Second-order derivative outputs scale as nv³ per timestep (g1’s idsva_so at N=256 is a 175.6 MB result batch), so for these algorithms the device→host copy — not the kernel — is the dominant with-memory cost. On the g1 idsva_so N=256 cell the measured transfer delta was 13.2 ms of a 15.2 ms with-mem total (87 %) with pageable host buffers, and is still ~80 % of the (much smaller) total after pinning.

Two fixes ship for this:

  • Pinned host staging — grid_host_alloc (cudaMallocHost) replaced pageable malloc for every gridData host buffer, raising effective D2H bandwidth on large payloads from 13.3 to ~21 GB/s (the box’s PCIe-lane ceiling). That alone flipped 8 with-mem cells vs Pinocchio (CPU, 24 threads) (11 losses + 2 ties → 5 losses + 8 wins); the per-cell before/after table is the “W2a re-capture” section of test/benchmarks/results/competitive_render_20260826/competitive_analysis.md.

  • Not copying at all — keep the tensor on the GPU (next section).

Staying resident (jax)#

The grid_rbd jax FFI path is device-in / device-out: under jit / lax.scan composition the nv³ tensor is produced on-device and consumed by the next pipeline stage (a Hessian-vector product, a DDP backward pass) with zero host transfer. Layer-3 timings (jax FFI, medians of 30 iters after warmup, RTX 5090, test/benchmarks/results/night4_20260827/so_resident_results.json), in total-batch µs — resident = device-in/device-out, out_host = outputs pulled to host, roundtrip = numpy-in → numpy-out:

robot

op

B

resident

out_host

roundtrip

iiwa14

idsva_so

16

515

643

790

iiwa14

fdsva_so

16

526

657

799

iiwa14

idsva_so

256

519

807

990

iiwa14

fdsva_so

256

529

831

997

go2

idsva_so

16

784

1081

1140

go2

fdsva_so

16

1018

1333

1474

go2

idsva_so

256

990

4783

4999

go2

fdsva_so

256

1412

4861

5137

g1

idsva_so

16

2003

3529

3801

g1

fdsva_so

16

5893

8310

8417

g1

idsva_so

256

2866

35024

34551

g1

fdsva_so

256

13949

45340

46514

Reading the table:

  • The resident column carries a ~0.4–0.5 ms fixed per-call FFI dispatch overhead (visible as the iiwa14 floor: kernel compute is ~43–72 µs, resident is ~515–530 µs). It is a per-call cost, so it amortizes when the call sits inside a jitted scan/rollout rather than being invoked from python once per step.

  • The resident-vs-out_host gap is pure output transfer, and it grows with nv³·B exactly as predicted: negligible for iiwa14, ~3.8 ms for go2 idsva_so at N=256, ~32 ms for g1.

  • The honest headline for the largest tensor: g1 idsva_so N=256 resident 2.87 ms vs Pinocchio (CPU, 24 threads) 2.38 ms total — near parity on wall time, while the result is already on the GPU for the next pipeline stage instead of in host memory. If the consumer is on the GPU, resident wins outright; if the result must land in host memory, CPU codegen remains the right tool for this cell.

Note

Why the with-mem comparison is asymmetric by construction. A CPU baseline pays no output copy — its results are born in host memory, so its with-mem time equals its compute time. The GPU with-mem column therefore carries the entire PCIe cost on top of compute. The three-layer reporting contract (compute-only / C++ with-mem / through-bindings; see test/benchmarks/BENCHMARK_METHODOLOGY.md) exists precisely to keep that visible: layer 1 is the GPU-resident design point, layer 2 isolates transfer, layer 3 is what an adopter pays through the python bindings.

See also#