Hanzo AI

Research Papers

All papers

Refutation-Driven Performance Engineering

An Empirical Study of a Multi-Week GPU Kernel Campaign

Hanzo AI Research

EngineGPU KernelsBenchmarksMethodology

Abstract

An N=1 observational study of a multi-week campaign to close the llama.cpp inference gap on three accelerators (AMD gfx1151 RDNA3.5, NVIDIA GB10 Blackwell, Apple M4 Max) inside a one-source GPU kernel DSL. The primary artifact is not a kernel but a refutation log: twenty-five numbered hypotheses, most plausible, each killed by a cheap decisive experiment. Reports a Vulkan-prefill win ladder (212→1734 tok/s, 8.2× of the gap closed in one ~36-hour cadence), a CUDA decode gap closed to parity by a single model-eligibility entry with zero kernel work, and an instrument fix that collapsed prefill benchmark noise from ±42% to ±0.8% (52×). Formalizes three recurring algorithms — a multi-fidelity Gate Ladder, a decisive-experiment selection rule, and refutation-log-as-prior — and a measurement doctrine of same-run, sustained-only ratios. Honest about standing: two wins, one parity, three open gaps.

Preview unavailable in this browser. Download the PDF.