CUDA Validation And Performance Reporting#

This page summarizes the developer CUDA validation flow for generated GRiD headers. The checks are intentionally staged so long GPU runs provide useful progress and can be resumed with cached generated artifacts.

Everyday refresh entry point#

Plan without collecting tests or starting any compilation:

.venv/bin/python test/run_validation.py --out docs/open-tasks/validation-next

Add --execute only on a clean committed candidate in a coordinated quiet window. The launcher holds /tmp/a2rlab-timing.lock throughout the run, records logs, validates fresh/carried header-key provenance, checks the reviewed minimum test-node set, and stages the generated receipt and header manifest. Its awaiting-evidence-commit state is successful recording, not a failed GPU run. Commit only those generated evidence files, then verify the clean descendant without rerunning any tests:

git add gpu-proof.json test/gpu-proof-header-keys.json
git commit -m "Record GPU correctness evidence"
.venv/bin/python test/run_validation.py --finalize docs/open-tasks/validation-next

Finalization rejects altered evidence, dirty trees and source changes since the recorded candidate. It does not commit files or start GPU work. --resume <split-directory> reuses completed shards; choose a new --out directory for the resumed launch’s logs.

For a formal release, add --full to both the dry plan and execution. This disables carried shards and selects the strict release policy at finalization. Resuming a full run also requires --full. Freeze implementation and package inputs before starting this long run; installed-artifact checks complement, but do not replace, the source correctness receipt.

Main requires test/gpu-proof-scope.json’s node IDs. A smoke receipt cannot replace full-scope evidence. Carried shards remain permitted by the everyday policy; this does not require a fresh release-policy run. After a reviewed full collection adds tests, explicitly extend the scope with test/receipt_integrity.py --receipt gpu-proof.json --scope test/gpu-proof-scope.json --extend-scope. Removing or renaming required tests requires a reviewed scope edit.

Local Docs Build#

Developer installs include the Sphinx documentation dependencies. To install them manually in an existing environment:

.venv/bin/python -m pip install -r docs/requirements.txt

Build the HTML docs from the repository root:

.venv/bin/python -m sphinx -W --keep-going -b html docs/source docs/build/html

CUDA Correctness Checks (crash-isolated split driver)#

The full GPU pass runs through the crash-isolated split driver, which partitions the cuda_equivalence and python_wrappers suites into bounded shards (one module’s abort can’t eat the rest) and is pausable/resumable:

# everything, sharded + receipt-signed (the standard full pass):
SPLIT=1 test/run_gpu_proof.sh

# only the shards whose inputs changed vs the committed receipt:
SPLIT=1 SPLIT_REFRESH=1 test/run_gpu_proof.sh

# the driver directly (no receipt), e.g. wrappers only:
.venv/bin/python test/run_split_suite.py --domains wrappers
.venv/bin/python test/run_split_suite.py --domains wrappers,cuda --changed-only

touch <out>/PAUSE stops cleanly between shards; --resume <out> continues an interrupted run without re-running completed shards. See CLAUDE.md and test/run_split_suite.py --help for the full contract (RAM-aware compile pool, shard bin-packing, receipt merge/carry).

For a quick targeted check, plain pytest still works:

.venv/bin/python -m pytest -m cuda_equivalence -q   # CUDA vs numpy oracle
.venv/bin/python -m pytest -m python_wrappers -q    # jax/torch handles

(The retired run_staged_cuda_checks.py stage runner predates the split driver; its stages map onto the shard partition + --changed-only.)

GPU-Proof Signed Receipts#

Because the CUDA equivalence suite needs a real GPU (and a full cold run is hours), GRiD records correctness in a signed receipt so a merge can be gated CPU-only, without re-running GPU tests in CI. This uses the pytest-gpu-proof plugin, installed from PyPI via requirements-dev.txt (it was previously a vendored test/pytest-gpu-proof submodule):

.venv/bin/python -m pip install -r requirements-dev.txt   # brings in pytest-gpu-proof

Generate the receipt on a quiet GPU box. The scope is tiered so it is never an all-or-nothing barrier — the signature, code fingerprint, and commit-SHA proof are identical regardless of how many tests the receipt attests:

SCOPE=smoke   test/run_gpu_proof.sh   # ~2 robots, cached cells — minutes; proves the plumbing
SCOPE=curated test/run_gpu_proof.sh   # representative robot set — tens of minutes
SCOPE=full    test/run_gpu_proof.sh   # every gpu_proof test — hours cold, the nightly job (default)

The gpu_proof marker is auto-applied to every cuda_equivalence and python_wrappers item by test/conftest.py (no per-test annotation), so receipt membership tracks the existing marker taxonomy. Running the script signs gpu-proof.json in place with your local SSH key and refuses a dirty tree (allow_dirty:false in test/gpu-proof-policy.yaml): the fingerprint cannot descend into the codegen/GLASS submodules, so a clean tree is what makes the receipt’s commit SHA an honest pin of the code under test.

CI (.github/workflows/verify-gpu-proof.yml) verifies whatever receipt is committed — signature (via github.com/{signer}.keys), fingerprint, commit SHA, freshness — with no GPU and no secrets, and skips gracefully when no receipt is present so code can ship before the long GPU run lands. A second always-on CPU lane runs the no-GPU tests (descriptor parity, kernel-attr manifest, plant launch hygiene).

The two-tier receipt policy (when CI goes red, and the fix)#

The committed gpu-proof.json goes stale — and CI’s verify-receipt job goes RED — the moment a push touches a fingerprinted test file (anything under test/cuda_equivalents/ or test/python_wrappers/). That red is by design; the fix is a refresh, not a revert:

  • Everyday (test/gpu-proof-policy.yaml, allow_carried: true): SPLIT=1 SPLIT_REFRESH=1 test/run_gpu_proof.sh re-runs ONLY the shards whose narrow fingerprints changed vs the committed receipt and CARRIES the rest; commit the refreshed receipt and CI goes green — minutes-to-hours, not a full pass. A carried shard attests “these tests, whose files are unchanged, passed at an ancestor commit ≤30 days old”; cross-cutting codegen/bindings changes are NOT re-proven by carry. GRID_SPLIT_REFRESH_DRY=1 previews the stale/carried plan (clean tree required).

  • Release (test/gpu-proof-policy-release.yaml, allow_carried: false): refuses carried shards outright — a release receipt requires ONE fresh full SPLIT=1 pass at the release tip. Verify with --policy test/gpu-proof-policy-release.yaml.

Maintainer internals: stale shard NAMES are recycled into the fresh partition automatically (the carry contract has no “superseded” state); a deleted test module forces a full pass; and both compile caches are CONTENT-keyed (header/source bytes), so byte-identical codegen edits cost seconds of regeneration, never an nvcc rebuild. Two newer carry gates: the per-shard header-key replay (test/header_key_replay.py, records in test/gpu-proof-header-keys.json) byte-precisely re-checks that each carried shard’s generated headers regenerate identically; and the robot-asset gate (codegen_neutrality.changed_robot_assets) demotes the covering shards when an oracle-side robot asset changes. Related lesson: every emission-shaping env var must be folded into the header cache key — GRID_ENABLE_MUJOCO_KERNELS once changed emission without changing the key and served stale headers.

The full test-pyramid rationale (which suite exists for what, and why the receipt is trustworthy) lives in test/TESTING_STRATEGY.md; a bucket map of everything in test/’s root is test/README.md.

CUDA Artifact Cache#

The CUDA equivalence tests cache both generated headers and compiled runners. This avoids regenerating and recompiling the same robot/configuration between test stages.

Environment variables:

  • GRID_CUDA_CACHE_DIR=.grid_build_cache/cuda controls the cache root.

  • GRID_CUDA_DISABLE_CACHE=1 disables header and runner reuse.

  • GRID_CUDA_VERBOSE_CACHE=1 prints cache hit/miss details.

  • GRID_CUDA_ARCH=120 overrides architecture detection when needed.

The header cache key includes robot identity, base mode, codegen source hash, target shared-memory bytes, profile/algorithm selection, and cache schema version. The runner cache key also includes CUDA architecture, nvcc version, L2 mode, compile flags, and runner source hash.

Progress Output#

CUDA equivalence tests print live progress through pytest’s terminal reporter. This keeps long GPU runs from looking stuck under normal pytest capture.

Environment variables:

  • GRID_CUDA_PROGRESS=0 disables progress output.

  • GRID_CUDA_VERBOSE_PROGRESS=1 enables per-algorithm detail.

  • GRID_CUDA_VERBOSE_CACHE=1 includes cache details in progress output.

Sample Selection#

Use GRID_CUDA_SAMPLE_NAMES to choose deterministic samples:

GRID_CUDA_RANDOM_SAMPLES=0 \
GRID_CUDA_SAMPLE_NAMES=zero,positive,negative,mixed_sign,tiny \
.venv/bin/python -m pytest test/cuda_equivalents/test_cuda_executable_equivalence.py -q -s

Recognized deterministic sample names include:

  • zero

  • positive

  • negative

  • mixed_sign

  • tiny

  • velocity_only

  • accel_or_torque_only

  • near_limit

  • floating_quat_identity

  • floating_quat_positive

  • floating_quat_mixed

Set GRID_CUDA_SAMPLE_NAMES=all to include all deterministic corner samples. Set GRID_CUDA_RANDOM_SAMPLES=N to append N seeded random samples.

Shared-Memory And L2 Controls#

Generated CUDA kernels use target-aware shared-memory tiers. The generator keeps the most shared-memory-heavy, all-shared implementation when it fits, and only selects spill tiers when the generated dynamic-shared arena exceeds the target byte budget. The current default target is 98304 bytes (96 KiB), which matches the opt-in dynamic shared-memory size available on many CUDA GPUs.

Use GRID_CUDA_TARGET_SHARED_MEM_BYTES at code-generation time to choose a different target:

# More portable target; useful for testing low-shared-memory fallback paths.
GRID_CUDA_TARGET_SHARED_MEM_BYTES=49152 grid-generate path/to/robot.urdf

# Current default: prefer all-shared kernels up to 96 KiB.
GRID_CUDA_TARGET_SHARED_MEM_BYTES=98304 grid-generate path/to/robot.urdf

# Experimental only: use a larger target if the deployment GPU supports it.
GRID_CUDA_TARGET_SHARED_MEM_BYTES=120000 grid-generate path/to/robot.urdf

At runtime, generated host initialization checks the actual device’s per-block shared-memory limit, including CUDA’s opt-in limit when available, and reports a clear launch/configuration error if the generated request is too large for the device.

Useful environment variables for fallback testing:

  • GRID_CUDA_TARGET_SHARED_MEM_BYTES=10000 forces low-target fallback paths.

  • GRID_CUDA_ENABLE_L2_PERSISTING=1 opts back in to the persisting-L2 hint. The generated default is now 0: a controlled A/B measured the pin never helping and hurting by up to 23%. The harness forwards the env var to the compile as -DGRID_CUDA_ENABLE_L2_PERSISTING=<n>.

  • GRID_CUDA_FORCE_SHARED_TIER=GRID_SPILL_DA_DF_OUTPUT forces a gradient shared-memory tier when generated support is available.

The shared-memory target is a legality and first-pass performance heuristic, not a full occupancy model. It does not currently account for register pressure, instruction-cache effects, or topology-specific compiler behavior. For example, on one robot a 96 KiB all-shared tier may be fastest, while on another robot a selective-spill tier can win because it uses fewer registers or allows better occupancy.

Register-Pressure And Tier Analysis#

For performance work, compare a small set of generated variants on the target GPU instead of assuming one target is best for every robot/topology. A practical local sweep is:

  1. Generate the same robot/profile with several shared-memory targets, usually 49152, 98304, and optionally the device’s opt-in maximum.

  2. Compile with ptxas verbosity enabled, for example by adding -Xptxas=-v to the compile command used for the timing binary.

  3. Record ptxas output for each kernel: registers per thread, local memory, and any compiler-reported spills.

  4. Run the generated timing kernels with GPU warmup and internal repeats.

  5. Prefer the fastest correct tier on that GPU, while keeping the 96 KiB default for generated code that has not been locally tuned.

The non-failing performance reporter can capture ptxas lines and timing summaries around an existing timing command:

.venv/bin/python test/benchmarks/perf_regression_report.py \
  --robot g1 \
  --base-mode floating \
  --profile all \
  --fallback-tier auto \
  --precision float \
  -- ./path/to/generated_timing_binary

If the compile command emits ptxas lines, the report prints a ptxas summary section. Use this to look for tier choices that trade shared memory for substantially higher register count or local-memory spills. Save baselines only from stable, representative GPU machines:

.venv/bin/python test/benchmarks/perf_regression_report.py \
  --robot g1 --base-mode floating --profile all \
  --fallback-tier GRID_SPILL_DA_DF_OUTPUT --precision float --save \
  -- ./path/to/generated_timing_binary

Second-order diagnostics are opt-in:

GRID_CUDA_RUN_SECOND_ORDER_FALLBACK_SMOKE=1 \
GRID_CUDA_SECOND_ORDER_SMOKE_ROBOT=iiwa14 \
.venv/bin/python -m pytest test/cuda_equivalents/test_cuda_second_order_fallback.py -q -s

Use GRID_CUDA_SECOND_ORDER_TEST_THREADS to override the diagnostic runner’s thread count.

Linear Algebra Backend#

Generated headers include a GRiD-owned linear algebra backend adapter that vendors the SIMT subset of GLASS (L1/dot, L2/gemv, L3/gemm) at codegen time. Generated headers are self-contained — no external SDK dependencies. The cuBLASDx-backed glass-nvidia path was removed in v2.0; see cuBLASDx Removal & Any-Thread-Count Library Functions.

Performance Reporting#

Performance reports are informational and non-failing. They are intended for a stable, fast GPU machine after correctness is green.

Run the standard benchmark suite with:

.venv/bin/python test/benchmarks/run_benchmarks.py

or run one GRiD benchmark slice directly:

.venv/bin/python test/benchmarks/baselines/grid/run.py \
  --robot g1 --base floating

Use test/benchmarks/perf_regression_report.py when you want a non-failing delta report around a timing command or a local JSON baseline. The reporter records:

  • GPU name

  • compute capability

  • CUDA version

  • robot

  • fixed/floating base mode

  • codegen profile

  • fallback tier

  • precision

It reports min, median or mean, max, spread, and percent delta from matching JSON baselines. Use --save only when the run should become the local baseline for that GPU class.