A2R LAB · ICRA 2024

Parallel computation.
Real-time control.

MPCGPU accelerates nonlinear model predictive control with a GPU-optimized preconditioned conjugate gradient solver.

Three snapshots of a simulated Kuka iiwa arm moving through a five-goal pick-and-place circuit.
THE TASK A simulated iiwa manipulator tracks a five-goal circuit. Figure 3 from our ICRA 2024 paper.
01 / THE METHOD

Structure is the advantage.

Exploit the block-tridiagonal structure of trajectory optimization, from system assembly to the final update.

MPCGPU forms the Schur system in parallel, solves it with GBD-PCG, then uses a parallel line search to update the robot trajectory.
Algorithm overview · Figure 2, our ICRA 2024 paper
01

Assemble in parallel

Build the Schur system and its structured preconditioner across the trajectory horizon.

02

Solve cooperatively

GBD-PCG distributes block rows across a cooperative CUDA grid, sharing work through synchronized iterations.

03

Update the trajectory

Evaluate candidate steps in parallel, apply the selected update, and warm-start the next control step.

02 / RESULTS

A solver designed
for longer horizons.

MPCGPU solves the structured linear systems in nonlinear MPC with GPU-accelerated PCG. On the five-goal pick-and-place task from our paper, the current code compares GBD-PCG with CPU QDLDL on an NVIDIA RTX 5090.

GBD-PCG averages 50–64 µs per linear-system solve at every horizon from 32 to 512 knot points, 1.8× to 17× faster than QDLDL. It sustains kilohertz control rates with horizons as long as 512 knot points.

The gain over the published results comes mainly from fewer PCG iterations per solve. A Gauss-Newton cost Hessian and a corrected horizon update make the controller's linear systems easier, and GLASS makes each iteration slightly cheaper. How the results changed ↗

Read the protocol and complete results ↗
Current average linear-system solve times for CPU QDLDL and GPU GBD-PCG from 32 to 512 knot points; GBD-PCG stays between 50 and 64 microseconds while QDLDL grows to 1100 microseconds.
Average linear-system solve time on the pick-and-place task at 500 Hz, current code. Grey labels give the paper's speedups.
Average SQP iterations per control step at 250 Hz, 500 Hz and 1 kHz for QDLDL and GBD-PCG across horizons, with the paper's values beneath each cell.
SQP iterations completed per control step within the control period. QDLDL cannot keep up at 512 knot points and 1 kHz.
03 / THE SOFTWARE

One focused stack.

MPCGPU combines generated robot dynamics, GPU linear algebra and cooperative PCG in a single CUDA/C++ solver.

DYNAMICS

GRiD

Generated rigid-body dynamics and analytical derivatives from the repository’s iiwa model.

GRiD repository ↗
BLOCK PRIMITIVES

GLASS

GPU linear algebra for the block operations shared by MPCGPU and GBD-PCG.

GLASS repository ↗
COOPERATIVE SOLVER

GBD-PCG

A grid-wide block-tridiagonal PCG solver, maintained and tested inside MPCGPU.

API and examples ↗

CUDA/C++ implementation: fixed-base Kuka iiwa examples, PCG and QDLDL backends, and configurable trajectory horizons. Explore the documentation ↗

04 / GET STARTED

From checkout
to a first solve.

Build the solver and run an end-effector tracking example. The quickstart covers dependencies, CUDA architecture selection and validation.

Open the quickstart ↗
MPCGPU / checkout
git clone --recurse-submodules \
  https://github.com/A2R-Lab/MPCGPU.git
cd MPCGPU
python3 -m venv .venv
.venv/bin/pip install -r requirements-dev.txt
make build_qdldl
make examples ARCH=sm_120
LD_LIBRARY_PATH="$PWD/qdldl/build/out" \
  ./examples/pcg.exe

Choose the architecture for your GPU. The PCG backend uses CUDA cooperative launches.

CITE THIS WORK

MPCGPU

Real-Time Nonlinear Model Predictive Control through Preconditioned Conjugate Gradient on the GPU

Emre Adabag · Miloni Atal · William Gerard · Brian Plancher
IEEE International Conference on Robotics and Automation, 2024

BibTeX

@inproceedings{adabag2024mpcgpu,
  title={MPCGPU: Real-Time Nonlinear Model Predictive
    Control through Preconditioned Conjugate Gradient on the GPU},
  author={Adabag, Emre and Atal, Miloni and Gerard, William
    and Plancher, Brian},
  booktitle={IEEE International Conference on Robotics and Automation},
  year={2024}
}

This work was supported by the National Science Foundation under Award 2246022. Opinions and conclusions are those of the authors and do not necessarily reflect those of the funding organization.