CUDA Validation And Performance Reporting#
This page summarizes the developer CUDA validation flow for generated GRiD headers. The checks are intentionally staged so long GPU runs provide useful progress and can be resumed with cached generated artifacts.
Everyday refresh entry point#
Plan without collecting tests or starting any compilation:
.venv/bin/python test/run_validation.py --out docs/open-tasks/validation-next
Add --execute only on a clean committed candidate in a coordinated quiet
window. The launcher holds /tmp/a2rlab-timing.lock throughout the run,
records logs, validates fresh/carried header-key provenance, checks the
reviewed minimum test-node set, and stages the generated receipt and header
manifest. Its awaiting-evidence-commit state is successful recording, not
a failed GPU run. Commit only those generated evidence files, then verify the
clean descendant without rerunning any tests:
git add gpu-proof.json test/gpu-proof-header-keys.json
git commit -m "Record GPU correctness evidence"
.venv/bin/python test/run_validation.py --finalize docs/open-tasks/validation-next
Finalization rejects altered evidence, dirty trees and source changes since
the recorded candidate. It does not commit files or start GPU work.
--resume <split-directory> reuses completed shards;
choose a new --out directory for the resumed launch’s logs.
For a formal release, add --full to both the dry plan and execution. This
disables carried shards and selects the strict release policy at finalization.
Resuming a full run also requires --full. Freeze implementation and package
inputs before starting this long run; installed-artifact checks complement,
but do not replace, the source correctness receipt.
Main requires test/gpu-proof-scope.json’s node IDs. A smoke receipt cannot
replace full-scope evidence. Carried shards remain permitted by the everyday
policy; this does not require a fresh release-policy run. After a reviewed
full collection adds tests, explicitly extend the scope with
test/receipt_integrity.py --receipt gpu-proof.json --scope test/gpu-proof-scope.json --extend-scope.
Removing or renaming required tests requires a reviewed scope edit.
Local Docs Build#
Developer installs include the Sphinx documentation dependencies. To install them manually in an existing environment:
.venv/bin/python -m pip install -r docs/requirements.txt
Build the HTML docs from the repository root:
.venv/bin/python -m sphinx -W --keep-going -b html docs/source docs/build/html
CUDA Correctness Checks (crash-isolated split driver)#
The full GPU pass runs through the crash-isolated split driver, which partitions
the cuda_equivalence and python_wrappers suites into bounded shards (one
module’s abort can’t eat the rest) and is pausable/resumable:
# everything, sharded + receipt-signed (the standard full pass):
SPLIT=1 test/run_gpu_proof.sh
# only the shards whose inputs changed vs the committed receipt:
SPLIT=1 SPLIT_REFRESH=1 test/run_gpu_proof.sh
# the driver directly (no receipt), e.g. wrappers only:
.venv/bin/python test/run_split_suite.py --domains wrappers
.venv/bin/python test/run_split_suite.py --domains wrappers,cuda --changed-only
touch <out>/PAUSE stops cleanly between shards; --resume <out> continues
an interrupted run without re-running completed shards. See CLAUDE.md and
test/run_split_suite.py --help for the full contract (RAM-aware compile
pool, shard bin-packing, receipt merge/carry).
For a quick targeted check, plain pytest still works:
.venv/bin/python -m pytest -m cuda_equivalence -q # CUDA vs numpy oracle
.venv/bin/python -m pytest -m python_wrappers -q # jax/torch handles
(The retired run_staged_cuda_checks.py stage runner predates the split
driver; its stages map onto the shard partition + --changed-only.)
GPU-Proof Signed Receipts#
Because the CUDA equivalence suite needs a real GPU (and a full cold run is
hours), GRiD records correctness in a signed receipt so a merge can be gated
CPU-only, without re-running GPU tests in CI. This uses the pytest-gpu-proof
plugin, installed from PyPI via requirements-dev.txt (it was previously a
vendored test/pytest-gpu-proof submodule):
.venv/bin/python -m pip install -r requirements-dev.txt # brings in pytest-gpu-proof
Generate the receipt on a quiet GPU box. The scope is tiered so it is never an all-or-nothing barrier — the signature, code fingerprint, and commit-SHA proof are identical regardless of how many tests the receipt attests:
SCOPE=smoke test/run_gpu_proof.sh # ~2 robots, cached cells — minutes; proves the plumbing
SCOPE=curated test/run_gpu_proof.sh # representative robot set — tens of minutes
SCOPE=full test/run_gpu_proof.sh # every gpu_proof test — hours cold, the nightly job (default)
The gpu_proof marker is auto-applied to every cuda_equivalence and
python_wrappers item by test/conftest.py (no per-test annotation), so
receipt membership tracks the existing marker taxonomy. Running the script signs
gpu-proof.json in place with your local SSH key and refuses a dirty tree
(allow_dirty:false in test/gpu-proof-policy.yaml): the fingerprint cannot
descend into the codegen/GLASS submodules, so a clean tree is what makes the
receipt’s commit SHA an honest pin of the code under test.
CI (.github/workflows/verify-gpu-proof.yml) verifies whatever receipt is
committed — signature (via github.com/{signer}.keys), fingerprint, commit
SHA, freshness — with no GPU and no secrets, and skips gracefully when no
receipt is present so code can ship before the long GPU run lands. A second
always-on CPU lane runs the no-GPU tests (descriptor parity, kernel-attr
manifest, plant launch hygiene).
The two-tier receipt policy (when CI goes red, and the fix)#
The committed gpu-proof.json goes stale — and CI’s verify-receipt job goes
RED — the moment a push touches a fingerprinted test file (anything under
test/cuda_equivalents/ or test/python_wrappers/). That red is by
design; the fix is a refresh, not a revert:
Everyday (
test/gpu-proof-policy.yaml,allow_carried: true):SPLIT=1 SPLIT_REFRESH=1 test/run_gpu_proof.shre-runs ONLY the shards whose narrow fingerprints changed vs the committed receipt and CARRIES the rest; commit the refreshed receipt and CI goes green — minutes-to-hours, not a full pass. A carried shard attests “these tests, whose files are unchanged, passed at an ancestor commit ≤30 days old”; cross-cutting codegen/bindings changes are NOT re-proven by carry.GRID_SPLIT_REFRESH_DRY=1previews the stale/carried plan (clean tree required).Release (
test/gpu-proof-policy-release.yaml,allow_carried: false): refuses carried shards outright — a release receipt requires ONE fresh fullSPLIT=1pass at the release tip. Verify with--policy test/gpu-proof-policy-release.yaml.
Maintainer internals: stale shard NAMES are recycled into the fresh partition
automatically (the carry contract has no “superseded” state); a deleted test
module forces a full pass; and both compile caches are CONTENT-keyed
(header/source bytes), so byte-identical codegen edits cost seconds of
regeneration, never an nvcc rebuild. Two newer carry gates: the per-shard
header-key replay (test/header_key_replay.py, records in
test/gpu-proof-header-keys.json) byte-precisely re-checks that each
carried shard’s generated headers regenerate identically; and the
robot-asset gate (codegen_neutrality.changed_robot_assets) demotes
the covering shards when an oracle-side robot asset changes. Related lesson:
every emission-shaping env var must be folded into the header cache key —
GRID_ENABLE_MUJOCO_KERNELS once changed emission without changing the
key and served stale headers.
The full test-pyramid rationale (which suite exists for what, and why the
receipt is trustworthy) lives in test/TESTING_STRATEGY.md; a bucket map of
everything in test/’s root is test/README.md.
CUDA Artifact Cache#
The CUDA equivalence tests cache both generated headers and compiled runners. This avoids regenerating and recompiling the same robot/configuration between test stages.
Environment variables:
GRID_CUDA_CACHE_DIR=.grid_build_cache/cudacontrols the cache root.GRID_CUDA_DISABLE_CACHE=1disables header and runner reuse.GRID_CUDA_VERBOSE_CACHE=1prints cache hit/miss details.GRID_CUDA_ARCH=120overrides architecture detection when needed.
The header cache key includes robot identity, base mode, codegen source hash,
target shared-memory bytes, profile/algorithm selection, and cache schema
version. The runner cache key also includes CUDA architecture, nvcc version,
L2 mode, compile flags, and runner source hash.
Progress Output#
CUDA equivalence tests print live progress through pytest’s terminal reporter. This keeps long GPU runs from looking stuck under normal pytest capture.
Environment variables:
GRID_CUDA_PROGRESS=0disables progress output.GRID_CUDA_VERBOSE_PROGRESS=1enables per-algorithm detail.GRID_CUDA_VERBOSE_CACHE=1includes cache details in progress output.
Sample Selection#
Use GRID_CUDA_SAMPLE_NAMES to choose deterministic samples:
GRID_CUDA_RANDOM_SAMPLES=0 \
GRID_CUDA_SAMPLE_NAMES=zero,positive,negative,mixed_sign,tiny \
.venv/bin/python -m pytest test/cuda_equivalents/test_cuda_executable_equivalence.py -q -s
Recognized deterministic sample names include:
zeropositivenegativemixed_signtinyvelocity_onlyaccel_or_torque_onlynear_limitfloating_quat_identityfloating_quat_positivefloating_quat_mixed
Set GRID_CUDA_SAMPLE_NAMES=all to include all deterministic corner samples.
Set GRID_CUDA_RANDOM_SAMPLES=N to append N seeded random samples.
Register-Pressure And Tier Analysis#
For performance work, compare a small set of generated variants on the target GPU instead of assuming one target is best for every robot/topology. A practical local sweep is:
Generate the same robot/profile with several shared-memory targets, usually
49152,98304, and optionally the device’s opt-in maximum.Compile with ptxas verbosity enabled, for example by adding
-Xptxas=-vto the compile command used for the timing binary.Record ptxas output for each kernel: registers per thread, local memory, and any compiler-reported spills.
Run the generated timing kernels with GPU warmup and internal repeats.
Prefer the fastest correct tier on that GPU, while keeping the 96 KiB default for generated code that has not been locally tuned.
The non-failing performance reporter can capture ptxas lines and timing summaries around an existing timing command:
.venv/bin/python test/benchmarks/perf_regression_report.py \
--robot g1 \
--base-mode floating \
--profile all \
--fallback-tier auto \
--precision float \
-- ./path/to/generated_timing_binary
If the compile command emits ptxas lines, the report prints a ptxas summary
section. Use this to look for tier choices that trade shared memory for
substantially higher register count or local-memory spills. Save baselines only
from stable, representative GPU machines:
.venv/bin/python test/benchmarks/perf_regression_report.py \
--robot g1 --base-mode floating --profile all \
--fallback-tier GRID_SPILL_DA_DF_OUTPUT --precision float --save \
-- ./path/to/generated_timing_binary
Second-order diagnostics are opt-in:
GRID_CUDA_RUN_SECOND_ORDER_FALLBACK_SMOKE=1 \
GRID_CUDA_SECOND_ORDER_SMOKE_ROBOT=iiwa14 \
.venv/bin/python -m pytest test/cuda_equivalents/test_cuda_second_order_fallback.py -q -s
Use GRID_CUDA_SECOND_ORDER_TEST_THREADS to override the diagnostic runner’s
thread count.
Linear Algebra Backend#
Generated headers include a GRiD-owned linear algebra backend adapter
that vendors the SIMT subset of GLASS (L1/dot, L2/gemv,
L3/gemm) at codegen time. Generated headers are self-contained — no
external SDK dependencies. The cuBLASDx-backed glass-nvidia path
was removed in v2.0; see cuBLASDx Removal & Any-Thread-Count Library Functions.
Performance Reporting#
Performance reports are informational and non-failing. They are intended for a stable, fast GPU machine after correctness is green.
Run the standard benchmark suite with:
.venv/bin/python test/benchmarks/run_benchmarks.py
or run one GRiD benchmark slice directly:
.venv/bin/python test/benchmarks/baselines/grid/run.py \
--robot g1 --base floating
Use test/benchmarks/perf_regression_report.py when you want a non-failing
delta report around a timing command or a local JSON baseline. The reporter
records:
GPU name
compute capability
CUDA version
robot
fixed/floating base mode
codegen profile
fallback tier
precision
It reports min, median or mean, max, spread, and percent delta from matching
JSON baselines. Use --save only when the run should become the local
baseline for that GPU class.