A GPU-accelerated library for robot dynamics, kinematics, and collisions, with analytical derivatives and Hessians for supported numerical operations.
GRiD turns a URDF into optimized, per-robot CUDA C++ for rigid-body dynamics, kinematics, their analytical
first- and second-order derivatives and a trajectory-optimization plant layer, then hands you that code three
ways — a numpy handle, a jax.jit-able FFI surface, or torch.autograd-aware ops — from one content-addressed
.so cache. One CUDA block per problem, batched, bit-deterministic and thread-count invariant; the same
model and API rebuild for each target architecture (one artifact per sm_XX), with runtime memory
adaptation from embedded Jetson class devices to desktop GPUs. Website: https://a2r-lab.github.io/GRiD/.
GRiD builds on our URDFParser, RBDReference, and GLASS packages (URDF parsing, Pinocchio-validated reference dynamics, and GPU linear algebra), together with its own bundled code generator. Using its scripts, users can easily generate and test optimized rigid body dynamics CUDA C++ code for their URDF files.
Ongoing development and the upcoming rerelease live in A2R-Lab/GRiD. The original ICRA 2022 paper describes the implementation preserved in the archival robot-acceleration/GRiD repository, not the full feature set or performance of the upcoming release. See the project website for the overview. Collision routines use the generated CUDA interface; numerical Python interface coverage is documented separately.
| Task | Start here |
|---|---|
| Call GRiD from Python (numpy/JAX/torch) | grid_rbd.load_robot("robot.urdf", backend=...) — Python wrappers docs · agent guide |
| Generate CUDA for a new robot | grid-generate config/robot_assets/iiwa14.urdf — see Quick Start below |
| Fit a humanoid build in RAM | fast robot setup (algorithm_list=, enable_mujoco_kernels=False) |
| Add an algorithm | adding an algorithm |
| Run tests / fix a red receipt CI job | CUDA validation + test/run_gpu_proof.sh --help |
| Benchmark | benchmarks |
| Debug a CUDA-vs-numpy mismatch | docs/agent_debugging_guide.md — the bug-class bible |
| Get MuJoCo/mjx-convention I/O | handle.mujoco.<method>(...) — values AND derivatives/second-order |
| Everything else | How do I…? on the docs site |
Start-here track: examples/README.md routes the four usage tracks — the examples/notebooks/ Python-bindings tour (01-quickstart → 07-inline-cuda), the runnable bindings/examples/ scripts, codegen scripts, and hand-written-CUDA walkthroughs.
This package contains submodules make sure to run git submodule update --init --recursive after cloning!
Install (creates a local venv and registers the grid-generate CLI):
bash install/base_install.sh
source .venv/bin/activateGenerate CUDA code for your robot (ten ready-to-use URDFs ship in
config/robot_assets/ — iiwa14, go2, fr3, g1, h1_2, …):
# Via the installed CLI (works on a clean base install):
grid-generate config/robot_assets/iiwa14.urdf # arm, fixed base
grid-generate config/robot_assets/go2.urdf -f # quadruped, floating base
grid-generate path/to/robot.urdf [-t EE_JOINT_NAME] [-n NAMESPACE] [-f] [--algorithm-list LIST] [-o OUT.cuh]
# Or via a hardcoded zero-config example (these two pull their URDFs from the
# robot_descriptions package — a DEV dependency; install install/requirements-dev.txt first):
python examples/codegen/generate_iiwa14.py # iiwa14 fixed base
python examples/codegen/generate_go2_floating.py # Go2 floating baseValidate and debug:
# Print CPU reference values for all algorithms:
python examples/codegen/print_reference_values.py path/to/robot.urdf
# Compile and run the CUDA print kernel (requires nvcc):
python examples/codegen/print_grid.py path/to/robot.urdfWrite your own CUDA kernel against the generated header:
# Step-by-step walkthrough + compiling/validated example kernels:
# examples/cuda/README.md (and examples/cuda/wrapper_types.md)
bash examples/cuda/build_and_validate.sh # generate → nvcc → run → validateRequires a C++17-capable host compiler (e.g., g++ ≥ 7, clang++ ≥ 5). The benchmark and codegen runtime compile with
-std=c++17— needed for inline variables in the bench common header. With the[torch]extra the per-robot.sofollows torch's ATen requirement (-std=c++20from torch 2.14; needs CUDA 12+ and g++ ≥ 10).
grid-generate PATH_TO_URDF— generategrid.cuh; add-dfor full debug mode,-ffor floating base,-t JOINT_NAMEto target a specific end-effector jointpython examples/codegen/print_reference_values.py PATH_TO_URDF— print CPU reference values for all algorithms to validate CUDA outputpython examples/codegen/print_grid.py PATH_TO_URDF— compile and run the CUDA print kernel against the generated header
Floating-base parsing and the Python reference path now accept a public floating-base convention flag. The default is Pinocchio-compatible:
floating_base_convention="pinocchio":q = [x, y, z, qx, qy, qz, qw],v = [vx, vy, vz, wx, wy, wz]floating_base_convention="legacy":q = [x, y, z, qw, qx, qy, qz],v = [wx, wy, wz, vx, vy, vz]
GRiD normalizes both public conventions into one shared internal floating-base
representation, so the code generator and RBDReference stay consistent under
the hood while callers can choose the input/output ordering they need.
Contributor-facing test workflows (floating-convention regression suite, CUDA equivalence env overrides, shared-memory targets) moved to CONTRIBUTING.md; the receipt/verification policy lives in the CUDA validation guide.
GRiD currently fully supports any robot model consisting of revolute, prismatic, and fixed joints that does not have closed kinematic loops. Arbitrary/skew joint axes (a non-cardinal <axis>) are also supported via a dense 6-vector motion subspace — currently for inverse_dynamics and crba only (cardinal-axis robots stay byte-identical; other algorithms and the helical/planar/spherical joint types are later stages).
GRiD implements the full modern rigid-body-dynamics stack: RNEA / CRBA / ABA /
Minv / forward dynamics; analytical first-order gradients (ID + FD, incl.
external-force gradients); the second-order derivatives (IDSVA-SO both frames
with a codegen-time dispatcher, FDSVA-SO); the kinematics family (EE pose /
Jacobian / Hessian, general-frame frame_jacobian/J̇/OSC inertia, runtime
multi-EE targets); integrators + integrator gradients; the centroidal family
(CoM, CCRBA, dccrba, CMM time-variation, Coriolis matrix, energy/ID
regressors); the inertial-parameter (π) family (the joint-torque regressor
Y with its analytic gradient ∂Y/∂(q,v) and the FD parameter gradient
∂q̈/∂π); contact-frame wrench mapping (contact_fext /
register_robot(contact_frames=...)); runtime tool/payload welding
(attach_tool/tool_fext); runtime multi-target positions; the collision
family (two-tier config_free); a trajectory-optimization grid_plant
cost/step layer; and
runtime-mutable inertia/transform/joint-dynamics tables. The complete
per-algorithm catalog with citations and per-feature detail lives in the
CUDA support status page.
RBDReference additionally provides numpy reference oracles — validated against Pinocchio — for generalized gravity, nonlinear effects, kinetic/potential/mechanical energy, the Coriolis matrix, the centroidal quantities (CoM, CoM Jacobian, CCRBA, centroidal momentum) and their derivatives (the analytic dccrba ∂A/∂q tensor — replacing the prior finite-difference oracle — and cmm_time_variation Ȧ), the inverse-dynamics and kinetic/potential-energy regressors, the general-frame Jacobian / J̇ / OSC inertia described above, and the plant/cost/barrier layer above.
Dual-surface equivalence. Every algorithm exists on two surfaces that are tested for numerical agreement: the RBDReference numpy implementation (the oracle, checked against Pinocchio) and the generated CUDA C++ kernels (checked against that same numpy reference). This keeps the GPU codegen honest against an independent, Pinocchio-validated baseline.
Mimic-joint support: per-robot gating is now essentially eliminated. Non-gradient algorithms (RNEA, forward dynamics, ABA, CRBA, …) work for robots with mimic joints, and every gradient emits a correct mimic-reduced result on both the fixed and floating base: inverse_dynamics_gradient/forward_dynamics_gradient, end_effector_pose_gradient/end_effector_pose_hessian, the second-order idsva_so/fdsva_so, the external-force gradients (f_ext_gradient), and the integrator gradients. The centroidal family — com, ccrba, energy, and the centroidal derivatives dccrba/cmm_time_variation — now also runs on mimic robots (the per-body Jacobian and per-unit motion columns carry the mimic multiplier α, validated against the mimic-aware reference). dccrba/cmm_time_variation additionally run on big floating-base robots (e.g. g1/h1_2-floating) via the sweep-pool spill path. No algorithm raises NotImplementedError for mimic robots anymore.
Additional algorithms and features are in development. If you have a particular algorithm or feature in mind please let us know by posting a GitHub issue. We'd also love your collaboration in implementing the Python reference implementation of any algorithm you'd like implemented!
| Directory | Owns | Entry doc |
|---|---|---|
grid_codegen/ |
the code-generation engine: emits grid.cuh AND the checked-in generated binding regions, all driven by the abi_specs.py table |
codegen architecture |
bindings/ |
the grid-rbd Python package (numpy/jax/torch handles over a cached per-robot .so) |
bindings/README.md · agent guide |
external/ |
the peer-product submodules: GLASS (GPU linear algebra), RBDReference (Pinocchio-validated numpy oracle), URDFParser |
each submodule's README |
examples/ |
the start-here track: notebooks/ (Python tour), codegen/, cuda/ |
examples/README.md |
test/ |
pytest suites + the split-suite/receipt machinery (run_split_suite.py, run_gpu_proof.sh, compile_sched.py) |
CUDA validation |
config/ |
ten sample URDFs (robot_assets/) + tuned per-GPU launch configs (launch_configs/) + autotune_robot.sh |
config/robot_assets/URDF_SOURCES.md |
docs/ |
the Sphinx site (source/) + agent_debugging_guide.md (the bug-class bible) |
docs site |
install/ |
install scripts (base_install.sh, developer_install.sh) + requirements files |
installation guide |
For each algorithm GRiD emits four layers: *_inner (core math on
shared-mem inputs), *_device (allocates scratch + calls _inner),
*_kernel (global entry point with batched timestep loop), and the
host wrapper (CPU launcher with H↔D copies). See the
codegen architecture docs
for the rationale and concrete signatures.
For Python users the grid-rbd package (in bindings/) wraps
the per-robot codegen behind a register-then-run UX with numpy, jax,
and torch backends. It ships as part of the single repo distribution — a
pip install -e . (what install/base_install.sh runs) installs the codegen
toolkit and the grid_rbd wrapper together. The base install is minimal;
pick a backend extra for the surface you want:
pip install -e "." # base: numpy backend only
pip install -e ".[jax]" # + JAX FFI surface
pip install -e ".[torch]" # + torch backend (CUDA wheel matching your GPU arch)
pip install -e ".[all]" # jax + torchSee the install matrix in bindings/README.md
for what each extra unlocks (and the torch CUDA-wheel note).
import grid_rbd
# numpy (default), jax, or torch; urdf_string= also accepted instead of urdf_path
handle = grid_rbd.register_robot("iiwa14", urdf_path="iiwa.urdf", backend="torch")
qdd = handle.forward_dynamics(q, qd, u) # autograd-aware torch.Tensor
qdd.sum().backward() # gradients flow to q, qd, uThe torch backend exposes autograd-aware inverse_dynamics / forward_dynamics /
aba / integrator (analytic backward passes) plus CUDA-Graphs capture,
and the handle also surfaces the grid_plant cost/barrier methods. inverse_dynamics
(alias rnea) / forward_dynamics (alias fd) take an optional qdd= (the
autograd gradient is qdd-aware, returning the correct ∂τ/∂(q,q̇) including the
∂(M·q̈)/∂q term), and all three backends expose the value ops coriolis_matrix,
kinetic_energy_regressor, potential_energy_regressor, dccrba, and
cmm_time_variation (forward-only on jax/torch). The π-regressor family
(inverse_dynamics_regressor, the differentiable
inverse_dynamics_wrt_params/forward_dynamics_wrt_params, and the
forward_dynamics_parameter_gradient ∂q̈/∂π), runtime tool welding
(attach_tool/tool_fext, via enable_tool=True), and multi-contact
wrench mapping (contact_fext, via register_robot(contact_frames=[...]))
are bound as well. For true fp64 compute build with
register_robot(..., dtype="float64") (its own cache entry); allow_fp64=True
is only the numpy handle's fp64-in/fp64-out convenience cast on an fp32 build
(ignored when dtype="float64"). See
bindings/README.md and the
Python wrappers docs.
To cite GRiD in your research, please use the following bibtex for our paper "GRiD: GPU-Accelerated Rigid Body Dynamics with Analytical Gradients":
@inproceedings{plancher2022grid,
title={GRiD: GPU-Accelerated Rigid Body Dynamics with Analytical Gradients},
author={Brian Plancher and Sabrina M. Neuman and Radhika Ghosal and Scott Kuindersma and Vijay Janapa Reddi},
booktitle={IEEE International Conference on Robotics and Automation (ICRA)},
year={2022},
month={May}
}
Release measurements from the 27 September 2026 run on one NVIDIA RTX 5090 with an Intel Core Ultra 9 285K cover RNEA, its analytical gradient (∇RNEA), and its analytical Hessian (∇²RNEA) on iiwa14 (fixed base, 7 velocities), go2 (floating base, 18), and G1 (floating base, 35) at batch sizes 16–1024. The release measurements page gives the method, every timing boundary, and the caveats; the benchmark harness reproduces the collection.
Ratios are baseline time divided by GRiD time; above 1× favors GRiD. Each column names its timing boundary: GRiD host
calls including copies against the CPU libraries, and GRiD compute-only calls against the GPU libraries' resident
calls. * marks cells where the evaluated baseline path required fp64 and ~ a side whose run means span more than
1.5×. Colors are clipped at 100×.
Microseconds per complete batch on a log axis. GRiD's bar splits into its CUDA compute-only call, the GPU–CPU I/O increment, and the JAX wrapper increment. These are differences of measured call times, not isolated measurements of each component.
Call wall times through each API boundary: native CUDA, the C++ host call, NumPy, PyTorch, and JAX. Pick the boundary your application uses. For the Python surfaces the solid bar is the call with its buffers allocated once and reused (measured 2 October 2026), and the hatched cap reaches the default call, which allocates its output every time. With reused buffers, NumPy and PyTorch land within a few percent of the C++ host call on large outputs.
The Quick Start above covers the common-case install. For CUDA Toolkit setup, developer dependencies (Pinocchio, robot_descriptions, benchmarks), and Docker, see the full installation guide.
On Ampere (sm_86 / CUDA 12.6) the bench harness can wedge nvcc /
ptxas at 100 % CPU when compiling heavy floating-base GRiD harnesses.
Pass --ptxas-opt-level 2 to test/benchmarks/run_multi_version.py — it
forwards -Xptxas -O2 to floating-base compiles only. Blackwell (sm_120) does
not hit this. Typical user code that includes grid.cuh and calls the
batch host wrappers (e.g. grid::forward_dynamics<T>(...)) does not
trigger the hang — it's specific to the timing-bench template surface.
Contributions welcome — see CONTRIBUTING.md for the workflow (and CLAUDE.md for the repo conventions AI agents and humans both follow).
Brian Plancher |
Zachary Pestrikov |
Kwamena A |
Danelle Tuchman |
Ann Li |
= |
EmreAdabag |
Cael Yasutake |
Naren Loganathan |
emilyburnett2003 |
pruyontrarakk |
Kimiya Shahamat |
Kimiya Shahamat |



