Input / Output ABI (h_q_qd_u)#
Read this before you pack a buffer by hand. GRiD is built for power users who call the generated kernels directly from their own CUDA (see Algorithm Design Principles & Best Practices), which means you own the layout of the input buffer. This page is the contract for that raw kernel buffer. Getting it wrong on a floating base does not crash and does not warn: it silently returns wrong dynamics.
The Python handles and the C ABI of the compiled grid-rbd wrapper do not
expose this buffer. They take q at the configuration width nq and every
velocity-like input (qd, qdd, u) at the tangent width nv, and
return dynamics vectors at nv, exactly as Pinocchio and MuJoCo do. The
padded layout below is staged inside the wrapper. See Python Wrappers (grid-rbd).
The per-timestep input block#
Every algorithm that takes (q, qd, u) — forward dynamics, its gradients, the
second-order kernels, the integrators, the regressor — reads one contiguous block
per timestep out of h_q_qd_u / d_q_qd_u. That block is three
``NUM_POS``-wide slots:
stride = Q_QD_U_STRIDE = 3 * NUM_POS # NUM_JOINTS == NUM_POS
timestep k occupies [k*stride, (k+1)*stride)
block + 0 q
block + NUM_POS qd
block + 2*NUM_POS u (or qdd, for the q|qd|qdd kernels)
The generated kernels slice it exactly that way:
T *s_q = s_q_qd_u; T *s_qd = &s_q_qd_u[NUM_POS]; T *s_u = &s_q_qd_u[2*NUM_POS];
The header also emits the slot offsets as named constants, right next to the contract stated as a comment block:
grid::GRID_Q_OFFSET // 0
grid::GRID_QD_OFFSET // NUM_POS
grid::GRID_U_OFFSET // 2*NUM_POS
grid::GRID_QDD_OFFSET // == GRID_U_OFFSET (qdd shares slot 2 with u)
Always derive your offsets from these emitted constants (or NUM_POS /
Q_QD_U_STRIDE) rather than hardcoding integers.
Warning
Do not pack ``q|qd|u`` tightly. The slots are NUM_POS wide, not
NUM_VEL wide. On a fixed base with independent scalar joints nq == nv, so a tight packing happens to
produce identical bytes and the mistake is invisible. On a quaternion floating
base without spherical joints nq == nv + 1; each independent spherical
joint adds another quaternion offset. A tight packing writes u at NUM_POS + NUM_VEL
while every kernel reads it from 2*NUM_POS. That read is in bounds, so
nothing traps: you get plausible, wrong numbers. Fixed-base code that is
“known good” therefore proves nothing about your floating-base packing.
Floating base#
For a floating base the root contributes 7 positions but only 6 velocities:
q (NUM_POS = 7 + n_joints) [ translation(3) | quaternion xyzw(4) | joint positions ]
qd (NUM_VEL = 6 + n_joints) [ linear vel(3) | angular vel(3) | joint velocities ]
In the raw kernel buffer qd and u occupy ``NUM_POS``-wide slots: their
NUM_VEL meaningful values fill the leading entries of the slot and the
trailing entry is a pad (write zero). The kernels index velocities by tangent
index, so this “leading entries, then padding” rule holds for every model,
floating or spherical. The quaternion is xyzw; identity is [0, 0, 0, 1].
The user-facing root velocity is ordered [linear; angular]; GRiD permutes it
into the internal Featherstone [angular; linear] spatial ordering for you.
Note
Kernel buffer slots are nq-wide; every public width is physical. Mass
matrices, Minv, the dynamics gradients and, on the wrapper surfaces, the
dynamics vectors and the velocity-like inputs are all NUM_VEL wide. Only a
hand-packed h_q_qd_u carries the padding.
Which surface protects you#
Python bindings and the C ABI (**``grid-rbd``)** — protected. You pass
qd/qdd/uatnvand the wrapper stages the padded buffer itself; annq-wide (padded) array on a floating-base robot raises a preciseValueErrornaming the tangent width. Nothing is auto-padded or sliced.Direct CUDA consumers — unprotected by construction. You pack
h_q_qd_uyourself, so this page is the only contract. Size the buffer as3 * NUM_JOINTS * NUM_TIMESTEPS(whatinit_gridDataallocates) and take the offsets from the constants above.
Stability#
Q_QD_U_STRIDE == 3 * NUM_POS and the three slot offsets are a published
ABI. They are pinned by test/cuda_equivalents/test_cuda_input_abi.py, which
checks the emitted offsets themselves (not merely the stride constant) on both
floating and fixed robots, and asserts that the device allocation, the host
memcpy, and the kernel stride argument all still agree. A change here breaks every
consumer that packs the buffer, so it is a deliberate, announced change — not a
refactor.