.. _glass-gpu-linear-algebra-simple-subroutines: GLASS Code Documentation ======================== For the research overview, paper, figures, and robotics case studies, visit `the GLASS project page `_. These pages document installation, numerical contracts, tuning, and the API. `GLASS `_ is a header-only CUDA C++ library of composable ``__device__`` primitives for small, block-local linear algebra and robotics math, architecture-tuned for edge robotics and beyond. It is the foundational linear-algebra layer underneath `GRiD `_, `MPCGPU `_, `GATO `_, `HJCD-IK `_, and other A2R Lab GPU solvers. The library covers BLAS, factorizations and solvers, structured systems, spatial algebra, Lie groups, projections, and small estimation kernels. Block-scoped operations compose inside a caller-owned kernel; selected warp- and thread-scoped forms pack smaller independent problems more densely. See the :doc:`API reference ` for the exact surface and :doc:`robotics conventions ` for the load-bearing layout choices. Interfaces ---------- GLASS separates **execution scope** (thread, warp, or block) from **implementation family** (dependency-free GLASS code or optional NVIDIA libraries). The spellings intentionally overlap, but do not all expose every operation. .. grid:: 2 :gutter: 3 .. grid-item-card:: Block — ``glass::block::`` :link: user_guide/getting_started/library_overview :link-type: doc One **block** per problem; the block's threads cooperate over shared/global data. Pure SIMT, **no dependencies** (``#include "glass.cuh"``). This explicit implementation is never re-dispatched. .. grid-item-card:: Warp — ``glass::warp::`` :link: api_reference/warp :link-type: doc One **warp** per problem (``__shfl_*_sync``, no ``__syncthreads``), so warps run independently. Choose this to pack **many small independent problems** into one block for intra-block parallelism. Requires a full 32-lane warp. .. grid-item-card:: Thread — ``glass::thread::`` :link: api_reference/thread :link-type: doc One problem per **thread**, 32 packed per warp. This is a compile-time, branch-free subset. It is usually register-resident around ``N≤7``; larger measured sizes remain correct and can still win despite spills. .. grid-item-card:: Nvidia — ``glass::nvidia::block::`` :link: user_guide/concepts/backend_dispatch :link-type: doc CUB / cuBLASDx / cuSOLVERDx at block scope. Choose this when the measured vendor implementation wins at your size (needs NVIDIA MathDx). .. grid-item-card:: Nvidia warp — ``glass::nvidia::warp::`` :link: api_reference/nvidia :link-type: doc CUB ``WarpReduce`` L1 reductions, one full 32-lane warp per problem and explicit per-warp scratch. .. grid-item-card:: Nvidia thread — ``glass::nvidia::thread::`` :link: api_reference/nvidia :link-type: doc cuSOLVERDx 0.4+ LAPACK, one independent packed problem per CUDA thread; no dynamic shared scratch or block-wide synchronization. .. note:: ``glass::cgrps::`` is a cooperative-groups *adapter* for the **Block** interface — identical numerics (the same SIMT loop, indexed via a ``thread_group`` handle), for callers already in a cooperative-groups context or tiling arbitrary sub-block groups. It is **not** a separately-tuned backend. ``#include "glass-cgrps.cuh"``. .. note:: **Bare** ``glass::op`` is the **measured-default face**: the same block-scope calling contract, with the implementation body chosen per (op, size, dtype) by ``glass::dispatch_body()`` (``glass-dispatch.cuh``, regenerated by the measured ``tune.py --legs body`` sweep). The selection is a ``constexpr`` decision **inside the device function** — resolved per call site at compile time, with no host-side dispatcher and no runtime branch. Cells with a robust measured win route to a warp- or thread-body inside the block; every other name is the *same entity* as ``glass::block::``. Determinism-sensitive callers pin ``glass::block::`` explicitly; see :doc:`user_guide/concepts/namespaces`. NVIDIA calls always name ``block``, ``warp``, or ``thread``; there is no bare ``glass::nvidia::op`` re-export. Measured defaults ----------------- ``glass::recommend()`` exposes architecture-specific measured execution plans for operations included in the tuning ladder. Measurements are not a promise that interfaces have identical coverage or reduction order. See :doc:`user_guide/concepts/tuning` for the selection policy and :doc:`user_guide/tutorials/sweep_results` for dated, configuration-specific results. .. grid:: 2 :gutter: 3 .. grid-item-card:: Get started :link: user_guide/getting_started/installation :link-type: doc Header-only install, the single-block execution model, and an optional MathDx setup for the ``glass::nvidia::`` backend. .. grid-item-card:: API reference :link: api_reference/index :link-type: doc The L1 / L2 / L3 and NVIDIA device functions, generated from the header doc-comments via Doxygen + Breathe. Quick start ----------- .. code-block:: cpp #include "glass.cuh" // One block solves one problem; threads stride over the data. __global__ void saxpy_kernel(uint32_t n, float a, float *x, float *y) { glass::axpy(n, a, x, y); // y = a*x + y } saxpy_kernel<<<1, 256>>>(n, 2.0f, d_x, d_y); See :doc:`user_guide/tutorials/quickstart` for a complete, compilable example, and :doc:`user_guide/tutorials/examples` for a worked program per concept. .. _measured-performance: Paper results ------------- The release evaluates GLASS on Jetson AGX Orin, Jetson AGX Xavier, and RTX 5090. See :doc:`user_guide/tutorials/paper_results` for the current measurements, comparison populations, and release figures. The dated :doc:`user_guide/tutorials/sweep_results` archive preserves earlier measurements. .. toctree:: :hidden: :caption: Getting Started user_guide/getting_started/index .. toctree:: :hidden: :caption: Concepts user_guide/concepts/index .. toctree:: :hidden: :caption: Tutorials user_guide/tutorials/index .. toctree:: :hidden: :caption: API Reference api_reference/index .. toctree:: :hidden: :caption: Developer Guide contribution_guidelines sphinx_edit_guide