Fast Robot Setup#
How to get a robot registered on your machine quickly and efficiently — and
how to reuse compiled artifacts. The per-robot .so cache is content-keyed
and persistent, so compatible later loads avoid compilation. Loading and
runtime initialization still take time; sub-second loading is not a guarantee.
This page is the quick canonical
path; the full API tour lives at Python Wrappers (grid-rbd).
The three-line happy path#
import grid_rbd
handle = grid_rbd.load_robot("path/to/robot.urdf")
qdd = handle.forward_dynamics(q, qd, u)
grid_rbd.load_robot() is the frictionless entry point — no name
ceremony, no two-call dance. It derives a stable, content-addressed name
from the URDF bytes ({stem}_{fixed|floating}_{sha256(urdf)[:12]}), so:
the first call generates
grid.cuh, compiles the per-robot.sowithnvcc, and caches it;later calls reuse compatible cached builds. Source, toolchain, GPU architecture, options or cache changes can require regeneration or recompilation even when the URDF bytes are unchanged.
Pass name="my_arm" if you want a human-friendly handle instead, and
floating_base=True for free-base robots (the fixed and floating loads of
one URDF get distinct, stable names and distinct cache entries). Any
register_robot keyword (ee_joint_names, max_batch_size,
algorithm_list, dtype, …) passes straight through.
The three entry points#
All three drive the same persistent cache; pick by ceremony level.
Call |
Does |
Use when |
|---|---|---|
|
build-if-missing and return a handle; the name is auto-derived from the URDF bytes |
the frictionless default — “just give me this robot” |
|
build-if-missing and return a handle, under a name you choose;
exposes the full option surface ( |
you want a stable human-friendly name, or non-default build options |
|
build + populate the cache only (returns manifest entries, no
handle); can prebuild several |
warm the cache out-of-band (CI / Docker / a setup script) so interactive use never compiles |
Once a named robot is in the cache, get_robot(name) looks it up without
touching the URDF (raises RobotNotRegisteredError if absent, and
StaleRobotError when the registered build cannot be loaded by this
grid_rbd — built from another wrapper, torch/jax ABI or GPU arch, or before
the build-identity record existed; re-run register_robot / precompile
and a byte-identical build re-hits the store without nvcc), and
list_registered() shows everything in the manifest. Toolchain provenance
(nvcc, host compiler, GLASS content) is NOT a load blocker: a shipped cache
loads on a box that never had the toolchain that built it.
All of these are idempotent: a robot already in the cache is an instant
no-op — re-registering never re-runs nvcc.
Cache anatomy#
The cache root is grid_rbd.default_cache_dir() — ~/.cache/grid-rbd/
by default, overridden by $GRID_RBD_CACHE_DIR or cache_dir= on any
entry point. Inside it, a manifest.json maps names to keys and each build
lives under store/<cache_key>/ (the generated grid.cuh, the compiled
robot.so, build metadata and log) — see the Cache layout section of
Python Wrappers (grid-rbd) for the full tree.
The store is content-keyed in two stages. A stage-1 input key is a
SHA-256 over the URDF bytes + the canonicalized option dict + the
grid-rbd version + the codegen source tree + the build identity: the
CUDA arch, the nvcc path and version, the host C++ compiler, the content
of the GLASS headers the generator vendors, the compile-flag module, the
torch/jax ABI tags and the generation-time environment knobs
(GRID_CUDA_TARGET_SHARED_MEM_BYTES and friends). bykey/<input_key>
points at a stage-2 content key — a SHA-256 of exactly what nvcc
compiled — under which the .so lives. Each entry keeps a readable
build_inputs.json copy of its build identity, and a stage-1 pointer is
honoured only when that record equals the current identity: a toolkit
upgrade, a dirty GLASS checkout or a changed environment knob can never hand
back a stale .so (the pointer is rejected, the sources are regenerated,
and if the generated bytes are unchanged the same .so is reused without
nvcc). Consequences worth knowing:
Re-registering with identical inputs is a pure lookup — no codegen, no
nvcc, regardless of the name used.An inline
urdf_string=and the equivalent file dedupe to the same.so(the key hashes the bytes).Different option choices (
ee_joint_names,dtype,algorithm_list,floating_base, …) land in separate entries that coexist — nothing is clobbered.CUDA arch in the key means a home directory shared between machines (e.g. NFS) safely keeps separate
.sofiles per GPU; thenvcc/ host compiler identity does the same for machines that differ in toolkit.A codegen edit whose emitted bytes are unchanged (a comment, a refactor) costs the CPU generation half only — never an
nvccrebuild.Diagnosing an unexpected rebuild: compare the entry’s
build_inputs.jsonwithgrid_rbd._cache.build_identity(arch); thegrid_rbdlogger reports the rejected pointer’s reasons at INFO level.Long-lived processes: the GLASS content hash and the compiler versions are memoized behind a
(path, mtime_ns, size)snapshot of the headers and the resolved compiler binaries, so an edit or a toolchain change AFTER an earlier registration in the same notebook is noticed on the next call. That snapshot is metadata: a same-size edit with a preserved timestamp, or a toolchain swapped behind an identical path and stat, is not — callgrid_rbd._cache.refresh_build_identity()in those cases.The cache dir is portable state: build offline once, then ship or keep the cache dir, and every later run on a matching machine starts in well under a second.
Inspect before you build: build_plan#
grid_rbd.build_plan(name, urdf_path, cuda_arch=..., **options) resolves a
registration exactly as register_robot would and stops there: it returns
the effective codegen options, the build identity (toolchain, GLASS content,
environment knobs), the stage-1 input key, and whether the cache already
holds a sound artifact (cached / content_key / stale_reasons /
would). It generates nothing, runs no nvcc, opens no .so and
creates no CUDA context — only the URDF and a cuda_arch (or
nvidia-smi to detect one) are needed. Use it to explain an unexpected
rebuild, to pre-check a deployment’s shipped cache, or to see on a build box
what the build will key on.
precompile(name, urdf, backends=["numpy"], cuda_arch=...) populates the
cache the same way without constructing a handle: a machine with nvcc
but no usable GPU (a CI builder, a container image build) can produce the
.so for the target arch. Only the jax / torch backends need the
device at warm time (they register FFI targets against it).
The baked launch profile is chosen for the architecture you build for:
among config/launch_configs/<robot>/*.json the profile whose name ends
in _sm<arch> matching cuda_arch is used (several profiles for one
architecture resolve to the default, else the first by name — GPU models
sharing an architecture are not told apart yet); if none exists the default
profile is baked with a one-time warning (correct, just not this GPU’s
measured optimum — config/autotune_robot.sh adds one). The choice rides
the cache key and is recorded in meta.json as launch_config_gpu.
What to expect: cold vs. warm#
Measured 2026-09-24 on an RTX 5090 / 24-core host, max_batch_size=1024,
one build per process (/usr/bin/time -v wall and peak RSS). A cold build is
the Python code generation (about a second, even for g1) plus ONE nvcc
translation unit, so a cold build is essentially all nvcc; “generation-only”
is what a codegen edit with byte-identical output costs (the content key re-hits
the compiled .so, no nvcc); a warm start is a fresh process calling
get_robot and one forward_dynamics.
Robot / build |
Cold build |
Peak RSS |
Generation-only |
Warm start |
Host RSS after one call |
|---|---|---|---|---|---|
iiwa14 fixed, all algorithms, pin-only |
10 min 53 s |
2.0 GB |
1.1 s |
1.7 s |
1.1 GB |
go2 floating, all algorithms, pin-only |
14 min 55 s |
3.6 GB |
2.0 s |
1.9 s |
1.8 GB |
go2 floating, all algorithms + mjx twins |
27 min 26 s |
10.1 GB |
— |
— |
— |
g1 floating, all algorithms, pin-only |
53 min 29 s |
8.3 GB |
2.9 s |
3.2 s |
6.9 GB |
g1 floating, all algorithms + mjx twins |
1 h 16 min |
11.0 GB |
— |
3.2 s |
7.1 GB |
g1 floating, |
3 min 0 s |
1.7 GB |
— |
1.7 s |
1.1 GB |
Three things to take from the table. algorithm_list is the biggest lever on
every axis (g1: 53 min → 3 min, 8.3 → 1.7 GB to build, 6.9 → 1.1 GB of host RAM
to run). The mjx twins roughly double a floating-base build and triple its RAM,
so enable_mujoco_kernels=False is the default advice for humanoids. And the
host RSS of a full build is dominated by per-algorithm workspaces sized by
max_batch_size and allocated at init, whether or not the process ever calls
those algorithms — on a shared-memory device (Jetson) size max_batch_size
and algorithm_list to what you run.
The warm path is the whole point: pay the cold cost once — ideally
out-of-band, see below — and interactive sessions, notebooks, tests, and
deployed processes all start in a few seconds (about 0.9 s of that is loading
the .so and initialising the device tables; the first call pays the lazy
allocations; the second call runs at kernel speed).
RAM-safe big-robot builds#
Large floating-base robots are dominated by the MuJoCo-convention (“mjx”) kernel twins that GRiD also emits on floating, non-mimic robots — the Build cost on large floating-base robots section of Python Wrappers (grid-rbd) has the full story. Three levers keep the build tractable:
Skip the mjx twins if you do not need MuJoCo-convention outputs:
grid_rbd.load_robot("g1.urdf", floating_base=True, enable_mujoco_kernels=False)
On g1-floating this is the difference between not building at all and a ~33 min build; fixed-base and mimic robots never get mjx twins, so the flag is a no-op there.
Build only the algorithms you need with
algorithm_list=— the named algorithms plus their auto-expanded dependencies are compiled, dramatically cuttingnvccwall time, peak RAM, and.sosize. Un-built methods raise a clear “add toalgorithm_listand rebuild” error, never a segfault. Viaprecompiletiers:grid_rbd.precompile( "g1", "g1.urdf", floating_base=True, tiers=[{"algorithm_list": ["forward_dynamics", "forward_dynamics_gradient"], "enable_mujoco_kernels": False}], )
Warm out-of-band. Don’t pay a minutes-long compile inside a notebook or an interactive session — run
precompileonce from a script or CI step (see the next two sections), and open the notebook against a warm cache.
Warming the jax / torch backends#
One compiled .so is shared across the numpy, jax, and torch surfaces —
backend= never forces a recompile. precompile can warm each
surface’s artifacts on that shared .so ahead of time:
grid_rbd.precompile("iiwa14", "iiwa.urdf",
backends=("numpy", "jax", "torch"))
Warm the surfaces you will actually use (each backend requires its extra
installed: pip install -e ".[jax]" / ".[torch]" / ".[all]").
Note that the first import jax / import torch and the first
jax.jit trace in a fresh process carry their own one-time framework
costs — those belong to the framework, not to GRiD’s cache.
For agents and CI: the non-interactive warm#
The exact one-liner to warm a robot from a script, CI job, or agent shell (idempotent — safe to run every time):
.venv/bin/python -c "import grid_rbd; grid_rbd.precompile('go2', 'config/robot_assets/go2.urdf', floating_base=True, backends=('numpy', 'jax'))"
Run it once per robot before any interactive or latency-sensitive use; every
subsequent load_robot / get_robot / jax.jit in any process is
then an instant cache hit.
GPU memory: GRiD lives inside the framework allocator’s pool#
The jax and torch surfaces install a device pool automatically at
register_robot / get_robot: GRiD’s entire device arena is carved out
of ONE slab allocated by the framework itself (jax: a buffer from XLA’s
memory pool; torch: a caching-allocator tensor) instead of raw
cudaMalloc. XLA’s default 75% preallocation can therefore stay ON — it no
longer starves GRiD’s allocations (the humanoid “launch failed” class).
handle.capabilities() reports the tuned launch state per method
(including the armed small-batch switch as batch_threshold /
threads_small), and runner.device_pool_used() equals
device_pool_bytes() after a pool-mode init. The numpy-only surface keeps
plain cudaMalloc with its VRAM auto-fit. GRID_WORKSPACE_TIMESTEP_SLOTS
is honored in both modes.
See also#
Python Wrappers (grid-rbd) — the full
grid-rbdAPI tour: method surface, cache layout, jax/torch backends,grid_plant.bindings/examples/AGENT_INTEGRATION_GUIDE.md— GPU-resident usage patterns for agents (stay on device, CUDA graphs, zero-copy interop).examples/notebooks/ — the guided notebook track (start with
01_quickstart_iiwa14).