How do I…?#

Start with source installation, then choose the task below.

I want to…

Start here

Call GRiD from Python (numpy / JAX / torch)

grid_rbd.load_robot("robot.urdf", backend=...) — one call, no name ceremony. Guided tour: Python Wrappers (grid-rbd); API: grid_rbd (Python bindings); agent-facing lifecycle notes: bindings/examples/AGENT_INTEGRATION_GUIDE.md.

Generate CUDA for a new robot

grid-generate path/to/robot.urdf [-f] (ten sample URDFs in config/robot_assets/); walkthrough: Generate CUDA for a robot.

Make a big (humanoid) robot build fit in RAM / finish faster

Fast Robot Setup — subset the build with algorithm_list= and/or skip the mjx twins with enable_mujoco_kernels=False (also grid-generate --algorithm-list ... --no-mujoco-kernels).

Use my own top-level GLASS instead of the copy vendored in grid.cuh

gen_all_code(..., vendor_glass=False): the header #includes glass.cuh from your include path (-I<GLASS root>) and aliases grid::glass to ::glass — one GLASS per translation unit. The default (vendored, self-contained) header is byte-identical. All *_DYNAMIC_SHARED_MEM_BYTES<T[, TIER]>() sizers are constexpr.

Label the GLASS revision when generating from a source archive (no .git)

gen_all_code(..., glass_revision="<sha>") or GRID_GLASS_REVISION=<sha>: the // Pinned commit: line carries the bare revision, so the header is byte-identical to a git checkout’s; a live checkout that disagrees is an error, and gen.glass_revision_source reports git / git-verified / supplied-unverified / unknown. Distribution artifacts record their packaged pin automatically and report bundled.

Embed the generated header in a library or interpreter (no exit() on CUDA errors)

init_robotModel_checked / init_joint_limits_checked / free_robotModel_checked (or robotModel_owner<T>): Library-safe initialization and cleanup.

Add a new algorithm to GRiD

Adding a New Algorithm (numpy oracle in RBDReference first, then the codegen emitter, then equivalence tests).

Run / verify the test suites, or fix a red receipt-verify CI job

CUDA Validation And Performance Reporting — the marker map, the split-suite driver, and the two-tier gpu-proof.json receipt policy (a red verify job after touching fingerprinted tests is BY DESIGN; run the everyday refresh and commit the receipt).

Benchmark GRiD (or compare against Pinocchio / MJX / Warp)

Benchmarks.

Debug a CUDA-vs-numpy mismatch or a weird kernel failure

the agent debugging guide — the accumulated bug-class bible (shared-memory init, output-convention traps, reduction nondeterminism, launch-config pitfalls, …).

Get MuJoCo/mjx-convention inputs & outputs

handle.mujoco.<method>(...) (per-call, thread-safe) or output_convention="mujoco" — values AND derivative/second-order surfaces, floating base; see the conventions section of Python Wrappers (grid-rbd).

Run several pipelines on one GPU without them sharing scratch

handle.context() (numpy, jax and torch handles alike) opens an isolated runtime context on the same artifact — own arena, tables, streams, launch overrides — closed with that handle; handle.device_profile says what it was fitted to. One pipeline per context. Runtime contexts.

Swap inertias / attach a tool at run time and keep autograd honest

set_inertia_params / attach_tool / set_joint_dynamics mutate the context under an exclusive admission lock and bump handle.model_version; a torch/JAX backward whose forward ran under an older version raises — recompute the forward. Same page as above.

Know what this release supports, what changed, and what it does not do

Compatibility and known limitations — platforms and toolchain, runtime contexts and versions, captured graphs, operands, native-interface stability (the wrapper’s C symbols are private; grid.cuh is the supported inline API), differentiation, build cost.

Know what shapes / dtypes / devices a call accepts (and rejects)

One rule set per operand class, enforced natively on every surface: Operand validation.

Understand why something recompiled (or refused to)

Three DIFFERENT “cache keys” exist: (1) the per-robot .so cache key (URDF bytes + codegen options + version + arch — register_robot); (2) the test suites’ content-keyed nvcc compile caches (header/source bytes — byte-identical codegen edits never rebuild); (3) the receipt fingerprints over test/cuda_equivalents + test/python_wrappers (what makes shards stale). Details: Fast Robot Setup and CUDA Validation And Performance Reporting.

See what a registration would build, or why it rebuilt

grid_rbd.build_plan(name, urdf, cuda_arch=...) — options, build identity, keys and cache status without building anything; and precompile(..., backends=["numpy"]) populates the cache without a handle or a CUDA context (build boxes). See Fast Robot Setup.

Tune kernel launch configs for my GPU

config/autotune_robot.sh --help (writes config/launch_configs/<robot>/<gpu>.json, baked at codegen time, overlay-able at runtime via apply_profile_overlay).