Running Tests#
GLASS ships a pytest suite that compiles small CUDA binaries and checks their output against NumPy/SciPy references.
Running the suite#
cd test
pip install -r requirements.txt
pytest -v
The first run compiles the CUDA test binaries with nvcc. Subsequent runs
skip recompilation unless the source changed (see the compile cache below).
Compiled binaries land in test/build/.
Selecting tests#
pytest test_l1.py -v
pytest test_l1.py -k "simple" # glass:: threadIdx variants only
pytest test_l1.py -k "cg" # glass::cgrps:: cooperative-groups variants
pytest test_l1.py -k "simple_hs" # high_speed warp-shuffle variants
The compile cache#
Compilation is driven by test/conftest.py. The session-scoped bins
mapping compiles a binary lazily, only when a selected test first accesses it,
and caches it by a SHA-256 hash of its source set.
For each test binary, _hash_sources() hashes its .cu file, shared helper,
umbrella headers, and implementation headers for the binary’s operation family
(vector, dense/factor, solver, robotics, or Nvidia). The hash is written to
test/build/<name>.hash. A local change therefore rebuilds the selected
family without rebuilding unrelated binaries. New binaries default to a
whole-source-tree hash until assigned to a family.
GPU-arch detection#
detect_arch() queries nvidia-smi --query-gpu=compute_cap and turns the
reported compute capability into an nvcc arch flag (e.g. 8.6 → sm_86).
If nvidia-smi is unavailable, it falls back to sm_75. Every binary is
compiled with -std=c++17 -arch=<detected>.
Optional binaries (graceful skips)#
Some test binaries need extra dependencies, and the fixtures skip rather than fail when those aren’t present:
Binary |
Requirement / behavior |
|---|---|
|
Exercises the SIMT-only batched APIs ( |
|
Round-2 compile-time auto-dispatch features. Needs |
|
The |
When MATHDX_ROOT is set and cublasdx.hpp is found, the cuBLASDx-gated
binaries are compiled with --expt-relaxed-constexpr -DGLASS_BENCH_CUBLASDX
plus the MathDx include paths automatically.