CUDA Support Status#
This page summarizes the generated CUDA paths that are currently exercised by the GRiD developer test suite. For the commands that run these checks, see CUDA Validation And Performance Reporting.
Development testing currently targets sm_120 (RTX 5090) for correctness and performance validation. Earlier compute capabilities (sm_8x) remain supported but are not the active development target.
Joint and interface restrictions#
This is a coverage summary, not a guarantee for every combination of joint type, algorithm, backend and GPU resource budget.
Planar and translation joints are decomposed into scalar-joint chains by the parser; mimic joints retain dependent bodies but reduce coordinates.
Mimic and spherical
minvuse a dense inverse of the reduced CRBA matrix rather than the scalar-joint direct-inverse recursion.fk_batchedis a NumPy-only single-leaf pose helper. It is not emitted for spherical models or models with more than 32 joints.Multi-stage integrator gradients are not supported for spherical joints or MuJoCo-output twins.
plant_step_hessiansupports Euler and semi-implicit Euler on fixed and floating bases, via NumPy/CUDA rather than the JAX/PyTorch handles.
See Kinematics (end-effector pose, Jacobian, Hessian) and Integrators and the plant layer for shapes and method scope.
Fixed-Base Robots#
Fixed-base CUDA coverage includes the core dynamics and kinematics paths:
Inverse dynamics / RNEA, direct Minv, forward dynamics, ABA, and CRBA.
Inverse- and forward-dynamics gradients.
End-effector pose, gradient, and Hessian.
Fixed-base forced-fallback coverage for oversized gradient kernels.
IDSVA-SO (body-frame and world-frame variants) and FDSVA-SO. The dispatcher selects body-frame IDSVA-SO for fixed-base models.
Optional per-body external forces (
d_f_ext) on RNEA, forward dynamics, ABA, and the inverse-/forward-dynamics gradients (opt-in;nullptrreproduces the no-force path).The
grid_plantlayer (plant_step, quadratic state/input costs, end-effector position cost, and joint position/velocity/torque log-barriers), emitted as a siblinggrid_plantnamespace.The centroidal family and its derivatives:
com,ccrba,energy,dccrba(∂A/∂q tensor), andcmm_time_variation(Ȧ).The
coriolis_matrixC(q,q̇)and the kinetic / potential inertial-parameter energy regressors.Runtime arbitrary multi-EE:
end_effector_pose_runtimeandend_effector_pose_gradient_runtime(runtime target joint id + per-target offset).Optional runtime-mutable inertial parameters (flag-gated
d_inertia_params+set_inertia_params; the baked default is byte-identical).Arbitrary/skew joint
<axis>(dense 6-vectorS) forinverse_dynamics,crba,abaandminv, including branching models (test_cuda_branching_skew_equivalence.py). This is not a claim of coverage for every skew-joint/backend/derivative combination.
Second-order fixed-base diagnostics are still developer-only. The current
green zero-sample set includes iiwa14, go2, gen3, fr3,
fetch, and rizon4 (healed 2026-09-15: the upstream flexiv xacro
emitted each link’s origin/mass/inertia triple bare, so the local
config/robot_assets/rizon4.urdf now wraps them in proper <inertial>
elements).
Floating-Base Robots#
Floating-base CUDA coverage includes:
Inverse dynamics / RNEA.
Direct Minv, forward dynamics, ABA, and CRBA.
Inverse- and forward-dynamics gradients.
End-effector pose.
Opt-in end-effector pose gradient and Hessian checks.
IDSVA-SO (world-frame variant, dispatcher-selected) and FDSVA-SO.
The centroidal family
com/ccrba/energyand the centroidal derivativesdccrba/cmm_time_variation. The two derivatives also emit on big floating-base robots (e.g.g1/h1_2-floating) via the sweep-pool spill path, so their prior big-floating gap is eliminated.
Floating end-effector Hessian generation uses target-aware spill tiers when needed:
Tier 0 keeps all Hessian scratch in shared memory.
Tier 1 spills chained
d2eeTempscratch tod_workspace.Tier 2 spills both d2XHom and
d2eeTemptod_workspace.
Tier selection is generated from robot dimensions, base mode, topology-derived
transform counts, and GRID_CUDA_TARGET_SHARED_MEM_BYTES. It is not keyed on
robot fixture names.
Known Caveats#
The generated
g1floating-base forward-dynamics inertial-parameter gradient in fp64 requests about 187 KiB of shared memory even at the minimal tier, exceeding the RTX 5090’s per-block limit. An fp64 diagnostic with global scratch is not a supported production launch.Floating IDSVA-SO/FDSVA-SO CUDA generation is now enabled (dispatcher picks
idsva_so_world_framefor floating-base). FDSVA-SO ong1_floatingrequires the selective-spill tier under sm_120’s ~100 KiB per-block dynamic shared-memory cap; the tier selector picks it automatically.Broad nonzero/random floating Hessian coverage is slower than the default smoke suite and remains opt-in.
Compute Sanitizer should be run on a supported GPU/driver setup before treating fallback paths as fully sanitizer-clean.
Performance tier choices can depend on register pressure and occupancy; use ptxas output and timing kernels on the target GPU before saving local baselines. Tiers are named
TIER_SHARED(default) /TIER_LITE/TIER_MINIMAL(the oldTIER_PERFalias has been removed — useTIER_SHARED). See Resource-Tier System (v2.0).Robots with mimic joints: non-gradient algorithms are supported, and every gradient now emits a correct mimic-reduced result on both bases —
inverse_dynamics_gradient/forward_dynamics_gradient,end_effector_pose_gradient/end_effector_pose_hessian, the second-orderidsva_so/fdsva_so, the external-force gradients (f_ext_gradient), and the integrator gradients. No mimic gradient raisesNotImplementedErroranymore. The centroidal kinematics family (com/ccrba/energy) and the centroidal derivatives (dccrba/cmm_time_variation) are now mimic-reduced as well: the per-body world Jacobian and per-unit motion columns carry the mimic multiplier (α), so all five emit + validate against the mimic-aware RBDReference oracle on fixed-base mimic robots.External-force gradients:
f_ext_gradient(∂τ/∂f_ext = −Jᵀ, ∂q̈/∂f_ext = M⁻¹Jᵀ) and the fixed-basef_ext_gradient_dq(−∂Jᵀ/∂q), both with CUDA equivalence tests, and both now fold correctly to the reduced coordinates on mimic robots as well.
Algorithm catalog#
Inverse Dynamics via the Recursive Newton Euler Algorithm (RNEA) from Featherstone
Composite Rigid Body Algorithm (CRBA) for the joint-space mass matrix and the Articulated Body Algorithm (ABA) for forward dynamics, both from Featherstone
The Direct Inverse of Mass Matrix from Carpentier
Forward dynamics:
qdd = Minv @ (u - RNEA(q, qd, 0))Analytical Gradients of Inverse Dynamics from Carpentier
Analytical Gradient of Forward Dynamics from Carpentier
End-effector pose, pose gradient (Jacobian), and pose Hessian
General-frame geometric Jacobian for an arbitrary target frame in any of the three Pinocchio reference frames (
LOCAL,WORLD,LOCAL_WORLD_ALIGNED). The numpy reference additionally provides the Jacobian time-variation J̇ and the operational-space (OSC) inertia Λ = (J·M⁻¹·Jᵀ)⁻¹ — all validated against Pinocchio’sgetFrameJacobian/getJointJacobian,computeJointJacobiansTimeVariation, and(J·M⁻¹·Jᵀ)⁻¹. CUDA codegen emits all three as opt-in keys — J (frame_jacobian), J̇ (frame_jacobian_dot), and Λ (osc_inertia) — each validated on-device against the numpy reference across the three frames (Λ is self-contained: it composes M⁻¹ on-device). All three additionally have the full launchable surface (batched*_kernel+ 3-mode host writing thegridDatad_frame_jacobian/d_frame_jacobian_dot/d_osc_inertiabuffers), so they are benchmarkable + bindable; the launchable surface bakes the leaf-EE target +LOCAL_WORLD_ALIGNEDframe, while the*_devicefunctions stay the arbitrary-target/-frame entry pointsSecond-Order Inverse Dynamics (IDSVA-SO) from Singh, Russell, & Wensing — the dispatcher selects body-frame for fixed-base and world-frame for floating-base models. See Release measurements for measured performance.
Second-Order Forward Dynamics (FDSVA-SO) from Singh, Russell, & Wensing on both fixed and floating bases
A time-integrator family: the discrete step
x_{k+1}plus its gradient∂x_{k+1}/∂(x,u)and a fused value-and-gradient variantOptional per-body external forces (
f_ext), threaded through RNEA, forward dynamics, ABA, and the inverse-/forward-dynamics gradients. Opt-in (anullptr/empty default reproduces the no-force path exactly), supplied in the body-local frame (6*NUM_BODIES, body-major) and subtracted from the per-body force.External-force gradients:
∂tau/∂f_ext = -Jᵀand∂q̈/∂f_ext = M⁻¹Jᵀ, plus the fixed-base∂(inverse_dynamics_gradient)/∂f_ext = -∂Jᵀ/∂qA trajectory-optimization-oriented ``grid_plant`` layer (emitted as a sibling
grid_plantnamespace): aplant_stepintegrator wrapper, quadratic state/input costs, an end-effector position cost (with Gauss-Newton Hessian), and joint position/velocity/torque log-barriers.The Coriolis matrix
C(q,q̇)(withC·q̇ + g(q) = nonlinear_effects)Inertial-parameter energy regressors: kinetic
y_KEand potentialy_PE(each length10·NB, withKE = y_KE·πandPE = y_PE·π)The centroidal derivatives:
dccrba(the ∂A/∂q tensor, 6×NV×NV) andcmm_time_variation(the centroidal-momentum-matrix time variation Ȧ)Runtime arbitrary multi-EE pose / pose-gradient (
end_effector_pose_runtime+_gradient): the end-effector target joint id and a per-target offset become runtime arguments instead of codegen-baked, so one compiled robot serves any leaf/target frame.Runtime-mutable inertial parameters (flag-gated): a
set_inertia_paramsdevice entry mutates an on-device parameter table (sysID / domain randomization) with no recompile; the baked default path is byte-identical.The joint-torque inertial-parameter regressor (
inverse_dynamics_regressor): the classicY(q, q̇, q̈)withtau = Y·πThe analytic regressor gradient (
inverse_dynamics_regressor_gradient): ∂Y/∂q and ∂Y/∂q̇, satisfying the contraction identitydY_dx[c]·π == ∂τ/∂x[:,c]The forward-dynamics parameter gradient (
forward_dynamics_parameter_gradient):∂q̈/∂π = −M⁻¹YContact-frame wrench mapping (
f_ext_contact): maps a contact-frame wrench to the joint-localf_extlayout, with the∂/∂f_cand∂/∂qderivativesRuntime multi-target positions (
multi_target_position/multi_target_position_gradient): batched runtime-target position queries and their gradientsThe integrator Hessian (
integrator_hessian): the second-order sensitivity of the discrete integrator stepThe collision family: two-tier
config_freecollision checking over spherized collision geometry