GRiD documentation#

GRiD turns a URDF into optimized, per-robot CUDA C++ for dynamics, kinematics, and collisions, with analytical derivatives and Hessians for supported numerical operations and a trajectory-optimization plant layer. Numerical Python interfaces include a numpy handle, a jax.jit-able FFI surface, or torch.autograd-aware ops, all backed by one content-addressed .so cache. It is the dynamics layer underneath GATO, MPCGPU, HJCD-IK and other A2R Lab GPU solvers, built on GLASS (block-local linear algebra), RBDReference (the Pinocchio-validated numpy oracle every kernel is tested against) and URDFParser.

Collision routines use the generated CUDA interface; see the collision workflow for geometry coverage and integration. Start with the installation guide and examples below, or visit the project homepage for an overview.

One block per problem, batched. Every algorithm runs as a single CUDA block per sample with in-block parallelism, so a batch of 16 or 4096 states is the same kernel design on a Jetson or an RTX 5090 — no multi-block reductions, no cooperative groups, bit-deterministic and thread-count invariant. Artifacts are built per target architecture (sm_XX); the model, the API and the generated code carry over, the binary does not.

numpy

load_robot(urdf) → a handle with 20+ batched methods; inputs and outputs are host arrays, batched on axis 0.

Python Wrappers (grid-rbd)
JAX

backend="jax" → device-in / device-out FFI targets that compose under jit, vmap and scan; jax.grad runs the analytical gradient kernels, never finite differences.

Host round trips (numpy): allocate once, reuse
torch

backend="torch" → autograd-aware ops on the current CUDA stream, plus CUDA-graph capture for fixed-batch replay.

Python Wrappers (grid-rbd)

Complete, copy-paste examples for each surface are in the quickstart below.

What you get per robot#

  • Dynamics: inverse dynamics (RNEA), forward dynamics (Minv-based and ABA), the joint-space inertia matrix (CRBA) and its inverse, gravity, non-linear effects, the Coriolis matrix, external body wrenches and tool loads.

  • Derivatives: analytical gradients of inverse and forward dynamics (IDSVA / FDSVA), their second-order tensors, inverse-dynamics and forward-dynamics gradients with respect to the inertial parameters, and end-effector pose Jacobians and Hessians.

  • Kinematics and centroidal quantities: end-effector poses (baked or runtime targets), frame Jacobians and their time derivatives, centroidal momentum matrix, its time variation and the operational-space inertia.

  • Plant layer: integrators with gradients, quadratic and barrier costs with gradients and Hessians, ready for a trajectory optimizer.

  • Collisions: self- and environment-collision checks with covering-sphere or native-primitive representations and coarse-to-fine evaluation in CUDA.

  • Conventions: Pinocchio by default; MuJoCo/mjx-convention twins of values and derivatives on floating-base robots (handle.mujoco.<op>).

The test suites compare generated CUDA with CPU references and record GPU results in signed receipts. See validation for the test workflow and backend coverage for operation-specific support.

Performance and release measurements#

The release measurements page holds one audited collection — three robots, fifteen operations, batches 16 to 1024, GRiD through every surface beside Pinocchio, MJX, MuJoCo Warp, MuJoCo CPU, BARD and Frax — with the protocol, every cell’s validation status and the cases GRiD loses. The benchmarks page retains dated development experiments and harness guidance; those results are not release evidence.

Portable by construction#

The generated code adapts to the device it runs on rather than to the one it was tuned on: a per-algorithm resource tier chooses how much of the working set lives in shared memory (SHARED for peak performance, LITE and MINIMAL rungs that spill scratch to global memory so the largest humanoids still fit a 48 KB budget), launch configurations are autotuned per robot and architecture, and a runtime context carries the device arena, streams and model tables — several isolated pipelines can share one GPU, one handle can swap inertial parameters at run time, and autograd refuses to differentiate a model that changed under it. GRiD targets a single GPU, from embedded Jetson-class devices to desktop cards, one artifact per architecture; the tested deployment platform of this release is Linux x86_64 (see compatibility and known limitations).

Quickstart#

Install from a recursive clone (editable install; the extras pin the CPU jax / torch packages, so add the CUDA wheels yourself — see installation):

git clone --recursive https://github.com/A2R-Lab/GRiD && cd GRiD
bash install/base_install.sh && source .venv/bin/activate
pip install -e ".[jax,torch]" "jax[cuda12]"      # optional: the JAX and torch surfaces

Each example below runs as written from the repository root. The first load_robot of a robot generates and compiles its .so — about ten minutes for the 7-DoF iiwa14 on an RTX 5090, an hour for a humanoid (the dated cold/warm table and the RAM-safe subset builds are on fast robot setup). Every later load is seconds: the artifact is cached by content key and rebuilt only when the URDF, the options, the GRiD version or the toolchain change.

import numpy as np
import grid_rbd

r = grid_rbd.load_robot("config/robot_assets/iiwa14.urdf")
q, qd, u = (np.zeros((8, r.nq), np.float32) for _ in range(3))
qdd = r.forward_dynamics(q, qd, u)                 # (8, 7)
dqdd = r.forward_dynamics_gradient(q, qd, u)       # (8, 7, 14) = [d/dq | d/dqd]
print(qdd.shape, dqdd.shape)
import jax
import jax.numpy as jnp
import grid_rbd

r = grid_rbd.load_robot("config/robot_assets/go2.urdf", backend="jax",
                        floating_base=True)
q = jnp.zeros((8, r.nq), jnp.float32).at[:, 6].set(1.0)   # unit quaternion (x, y, z, w)
qd = jnp.zeros((8, r.nv), jnp.float32)                    # nv-wide: the tangent width
u = jnp.zeros((8, r.nv), jnp.float32)
loss = lambda a: r.forward_dynamics(a, qd, u).sum()
g = jax.jit(jax.grad(loss))(q)                           # analytic VJP, stays on the device
print(g.shape)
import torch
import grid_rbd

r = grid_rbd.load_robot("config/robot_assets/iiwa14.urdf", backend="torch")
q = torch.zeros(8, r.nq, device="cuda", requires_grad=True)
qd = torch.zeros(8, r.nq, device="cuda")
u = torch.zeros(8, r.nq, device="cuda")
qdd = r.forward_dynamics(q, qd, u)                 # CUDA tensors in and out
qdd.sum().backward()                               # analytic backward
step = r.capture("forward_dynamics", q.detach(), qd, u)   # CUDA-graph replay
print(q.grad.shape, step.replay().shape)

Go deeper#

How do I…?

The task router: calling, generating, testing, benchmarking, debugging — one table.

How do I…?
Concepts

Design principles, the codegen architecture, resource tiers, runtime contexts, operand validation, the I/O ABI, mjx conventions.

Concepts
From raw CUDA

grid-generate robot.urdf emits a self-contained grid.cuh to #include in your own kernels.

Generate CUDA for a robot
API Reference

The Python bindings, the URDF parser, the reference algorithms and the code generator.

API Reference
Benchmarks

Competitive and historical sweeps, the with-memory story, and how to re-measure on your hardware.

Benchmarks
Validation and CI

The CUDA-vs-numpy equivalence suite, the signed GPU receipt and the split test driver.

CUDA Validation And Performance Reporting

Citation#

The original ICRA 2022 paper describes the implementation preserved at robot-acceleration/GRiD. Ongoing development lives at A2R-Lab/GRiD. The paper does not describe all current features or establish their performance. If you use GRiD in your research, cite the original paper and record the software commit or release:

@inproceedings{plancher2022grid,
  title={GRiD: GPU-Accelerated Rigid Body Dynamics with Analytical Gradients},
  author={Brian Plancher and Sabrina M. Neuman and Radhika Ghosal and Scott Kuindersma and Vijay Janapa Reddi},
  booktitle={IEEE International Conference on Robotics and Automation (ICRA)},
  year={2022},
  month={May}
}