GRiD documentation#
GRiD turns a URDF into optimized,
per-robot CUDA C++ for dynamics, kinematics, and collisions, with analytical
derivatives and Hessians for supported numerical operations and a
trajectory-optimization plant layer. Numerical Python interfaces include a numpy handle, a
jax.jit-able FFI surface, or torch.autograd-aware ops, all backed by
one content-addressed .so cache. It is the dynamics layer underneath
GATO, MPCGPU,
HJCD-IK and other A2R Lab GPU
solvers, built on GLASS (block-local linear
algebra), RBDReference (the
Pinocchio-validated numpy oracle every kernel is tested against) and
URDFParser.
Collision routines use the generated CUDA interface; see the collision workflow for geometry coverage and integration. Start with the installation guide and examples below, or visit the project homepage for an overview.
One block per problem, batched. Every algorithm runs as a single CUDA
block per sample with in-block parallelism, so a batch of 16 or 4096 states
is the same kernel design on a Jetson or an RTX 5090 — no multi-block
reductions, no cooperative groups, bit-deterministic and thread-count
invariant. Artifacts are built per target architecture (sm_XX); the
model, the API and the generated code carry over, the binary does not.
load_robot(urdf) → a handle with 20+ batched methods; inputs and
outputs are host arrays, batched on axis 0.
backend="jax" → device-in / device-out FFI targets that compose
under jit, vmap and scan; jax.grad runs the analytical
gradient kernels, never finite differences.
backend="torch" → autograd-aware ops on the current CUDA stream,
plus CUDA-graph capture for fixed-batch replay.
Complete, copy-paste examples for each surface are in the quickstart below.
What you get per robot#
Dynamics: inverse dynamics (RNEA), forward dynamics (Minv-based and ABA), the joint-space inertia matrix (CRBA) and its inverse, gravity, non-linear effects, the Coriolis matrix, external body wrenches and tool loads.
Derivatives: analytical gradients of inverse and forward dynamics (IDSVA / FDSVA), their second-order tensors, inverse-dynamics and forward-dynamics gradients with respect to the inertial parameters, and end-effector pose Jacobians and Hessians.
Kinematics and centroidal quantities: end-effector poses (baked or runtime targets), frame Jacobians and their time derivatives, centroidal momentum matrix, its time variation and the operational-space inertia.
Plant layer: integrators with gradients, quadratic and barrier costs with gradients and Hessians, ready for a trajectory optimizer.
Collisions: self- and environment-collision checks with covering-sphere or native-primitive representations and coarse-to-fine evaluation in CUDA.
Conventions: Pinocchio by default; MuJoCo/mjx-convention twins of values and derivatives on floating-base robots (
handle.mujoco.<op>).
The test suites compare generated CUDA with CPU references and record GPU results in signed receipts. See validation for the test workflow and backend coverage for operation-specific support.
Performance and release measurements#
The release measurements page holds one audited collection — three robots, fifteen operations, batches 16 to 1024, GRiD through every surface beside Pinocchio, MJX, MuJoCo Warp, MuJoCo CPU, BARD and Frax — with the protocol, every cell’s validation status and the cases GRiD loses. The benchmarks page retains dated development experiments and harness guidance; those results are not release evidence.
Portable by construction#
The generated code adapts to the device it runs on rather than to the one it
was tuned on: a per-algorithm resource tier chooses how much of the
working set lives in shared memory (SHARED for peak performance,
LITE and MINIMAL rungs that spill scratch to global memory so the
largest humanoids still fit a 48 KB budget), launch configurations are
autotuned per robot and architecture, and a runtime context carries the device arena, streams
and model tables — several isolated pipelines can share one GPU, one handle
can swap inertial parameters at run time, and autograd refuses to
differentiate a model that changed under it. GRiD targets a single GPU, from
embedded Jetson-class devices to desktop cards, one artifact per
architecture; the tested deployment platform of this release is Linux x86_64
(see compatibility and known limitations).
Quickstart#
Install from a recursive clone (editable install; the extras pin the CPU jax / torch packages, so add the CUDA wheels yourself — see installation):
git clone --recursive https://github.com/A2R-Lab/GRiD && cd GRiD
bash install/base_install.sh && source .venv/bin/activate
pip install -e ".[jax,torch]" "jax[cuda12]" # optional: the JAX and torch surfaces
Each example below runs as written from the repository root. The first
load_robot of a robot generates and compiles its .so — about ten
minutes for the 7-DoF iiwa14 on an RTX 5090, an hour for a humanoid (the dated
cold/warm table and the RAM-safe subset builds are on
fast robot setup).
Every later load is seconds: the artifact is cached by content key and rebuilt
only when the URDF, the options, the GRiD version or the toolchain change.
import numpy as np
import grid_rbd
r = grid_rbd.load_robot("config/robot_assets/iiwa14.urdf")
q, qd, u = (np.zeros((8, r.nq), np.float32) for _ in range(3))
qdd = r.forward_dynamics(q, qd, u) # (8, 7)
dqdd = r.forward_dynamics_gradient(q, qd, u) # (8, 7, 14) = [d/dq | d/dqd]
print(qdd.shape, dqdd.shape)
import jax
import jax.numpy as jnp
import grid_rbd
r = grid_rbd.load_robot("config/robot_assets/go2.urdf", backend="jax",
floating_base=True)
q = jnp.zeros((8, r.nq), jnp.float32).at[:, 6].set(1.0) # unit quaternion (x, y, z, w)
qd = jnp.zeros((8, r.nv), jnp.float32) # nv-wide: the tangent width
u = jnp.zeros((8, r.nv), jnp.float32)
loss = lambda a: r.forward_dynamics(a, qd, u).sum()
g = jax.jit(jax.grad(loss))(q) # analytic VJP, stays on the device
print(g.shape)
import torch
import grid_rbd
r = grid_rbd.load_robot("config/robot_assets/iiwa14.urdf", backend="torch")
q = torch.zeros(8, r.nq, device="cuda", requires_grad=True)
qd = torch.zeros(8, r.nq, device="cuda")
u = torch.zeros(8, r.nq, device="cuda")
qdd = r.forward_dynamics(q, qd, u) # CUDA tensors in and out
qdd.sum().backward() # analytic backward
step = r.capture("forward_dynamics", q.detach(), qd, u) # CUDA-graph replay
print(q.grad.shape, step.replay().shape)
Go deeper#
The task router: calling, generating, testing, benchmarking, debugging — one table.
Design principles, the codegen architecture, resource tiers, runtime contexts, operand validation, the I/O ABI, mjx conventions.
grid-generate robot.urdf emits a self-contained grid.cuh to
#include in your own kernels.
The Python bindings, the URDF parser, the reference algorithms and the code generator.
Competitive and historical sweeps, the with-memory story, and how to re-measure on your hardware.
The CUDA-vs-numpy equivalence suite, the signed GPU receipt and the split test driver.
Citation#
The original ICRA 2022 paper describes the implementation preserved at robot-acceleration/GRiD. Ongoing development lives at A2R-Lab/GRiD. The paper does not describe all current features or establish their performance. If you use GRiD in your research, cite the original paper and record the software commit or release:
@inproceedings{plancher2022grid,
title={GRiD: GPU-Accelerated Rigid Body Dynamics with Analytical Gradients},
author={Brian Plancher and Sabrina M. Neuman and Radhika Ghosal and Scott Kuindersma and Vijay Janapa Reddi},
booktitle={IEEE International Conference on Robotics and Automation (ICRA)},
year={2022},
month={May}
}