Performance
Hardware-conscious software extracts more inference performance from workstations and local multi-GPU systems.
Independent nonprofit Open source
We develop open kernels, tools, benchmarks, and research for AI inference on hardware you control.
01 Why local
Local AI improves privacy, ownership, access, and room to experiment. We build software and publish research for demanding inference workloads.
Hardware-conscious software extracts more inference performance from workstations and local multi-GPU systems.
Reproducible benchmarks and practical research support configuration, tuning, and hardware decisions.
Public code and shared artifacts make results reproducible and available for independent development.
02 Areas of focus
We focus on NVIDIA Blackwell platforms available outside traditional datacenters, with an emphasis on inference performance, efficiency, and usability.
GeForce RTX 50 Series and RTX PRO Blackwell systems.
GB10 Grace Blackwell systems for compact local inference.
Workstation-class Blackwell systems for large local workloads.
03 Flagship project
b12x is an open SM120/SM121 CuTe DSL and Triton kernel library for local LLM inference, targeting DGX Spark and Blackwell-based RTX systems.
Explore b12x on GitHub$ pip install b12x
04 Join the lab
Run our projects on your hardware, report results, and contribute fixes through the public repositories.