MaxKernel: Agentic Kernel Generation for TPUsExplained for Beginners
Shangkun Wang, Nina Cai, Charles Hoong +7 more
Abstract
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-time compiler feedback to build agentic systems for kernel generation. In this work, we present MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: (1) a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; (2) an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and (3) a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space. All three paradigms leverage a shared pool of specialized sub-agents to handle planning, implementation, self-debugging, testing, and hardware profiling. We evaluate MaxKernel on JaxBench, a comprehensive suite of 50 diverse kernel tasks for TPUs, alongside complex, real-world workloads from state-of-the-art open-source models. We demonstrate that MaxKernel consistently generates highly optimized implementations, matching expert hand-tuned baselines and delivering significant performance across the benchmark. Our agent is open-sourced and available https://github.com/AI-Hypercomputer/accelerator-agents/tree/main/MaxKernel.
MaxKernel: Agentic Kernel Generation for TPUs
If you’ve ever tried to squeeze every last drop of performance out of a specialized chip like a TPU, you know the drill: you dive into arcane documentation, manually tweak tiling sizes, and battle cryptic compiler errors that seem to appear just when you think you’re done. It’s deep, meticulous work that usually requires a PhD-level specialist. Enter MaxKernel, a framework from Google that aims to hand off much of this drudgery to software agents. In this explanation, we’ll break down what MaxKernel is, how its multi-agent system operates, and why the results matter—even if you’re not planning to write a TPU kernel tomorrow.
The Problem
Designing custom kernels for accelerators has long been the domain of hardware gurus. Whether you’re working with CUDA for NVIDIA GPUs, Triton for AMD, or JAX/Pallas for Google’s TPUs, the task involves manually managing the movement of data between different types of memory (like the high-bandwidth memory, or HBM, and the much smaller, faster scratchpad memory), orchestrating Direct Memory Access (DMA) transfers to hide latency, and devising multi-dimensional tiling strategies that keep the compute units busy without starving for data.
Traditionally, engineers either write these by hand—or rely on heavy compiler infrastructure like TVM or Halide, which still require significant domain knowledge to guide. The problem is twofold: first, there’s a severe talent bottleneck. Second, the space of possible configurations is astronomically large. Even for a single operation, you might have dozens of tile sizes, loop orderings, and memory allocation strategies to consider. Exploring all of them exhaustively would take a human operator forever.
Large Language Models (LLMs) have shown promise in code generation, but accelerator programming is notoriously brittle. A functionally correct kernel might still crash the hardware or run slowly because of a subtle memory violation that the LLM isn’t aware of. To make LLMs reliable here, you need real-time feedback from the compiler and profiling tools. MaxKernel was built precisely to bridge that gap: it marries the creative code-generation power of LLMs with the rigid, empirical feedback loops that real hardware demands.
How It Works (The Technical Mechanics)
At its heart, MaxKernel is a multi-agent system. Rather than asking one monolithic LLM to “write me a fast TPU kernel,” the framework decomposes the problem into a suite of specialized sub-agents, each with a clearly defined role. Think of it less as a single AI writer and more as a small team of engineers, each expert in a different phase of kernel development.
The Sub-Agents
- Planning Agent: Before any code is written, this agent sketches a high-level optimization plan. It looks at the reference implementation and the TPU hardware specs, then decides, “Okay, we should try tiling the matrix this way, and we’ll use a specific DMA pattern to overlap computation and data transfer.”
- Implementation Agent: Takes that plan and translates it into valid Pallas code—the domain-specific language for TPU kernels. It’s responsible for the syntax, making sure the loops are structured correctly for the TPU’s systolic array.
- Validation & Fix Loop: This is where the “agentic” part gets real. The system tries to compile the code. If the compiler spits out an error—maybe a VMEM allocation failure or a dimension mismatch—the Feedback loop captures that error, feeds it back to the Planning or Implementation agent, and they iterate. It’s like a compiler that doesn’t just say “error,” but suggests a fix or a different approach.
- Testing & Verification Agents: Once the code compiles, we need to make sure it actually does the math correctly. A Test Synthesis Agent builds a suite of test cases (inputs and expected outputs), and an Execution Agent runs those tests on actual TPU hardware. It compares the results against the standard JAX reference, allowing for tiny floating-point tolerances (because bfloat16 math on hardware can vary slightly).
- Autotuning Agent: This agent engages in hyperparameter search. It systematically varies block sizes, tile dimensions, and other configuration knobs to find the sweet spot where the kernel runs fastest.
- Profiling Agent: After a kernel runs, this agent captures traces via a tool called XProf. It measures actual latency, memory bandwidth utilization, and “compute density” (how much of the hardware’s theoretical peak compute is actually being used). This data becomes the signal that guides the next iteration of optimization.
The Three Orchestration Paradigms
MaxKernel doesn’t just have these agents; it has three ways of orchestrating them, depending on how much human control you want.
1. Human-in-the-Loop (HITL): This is the “collaborative” mode. The system walks the user through discrete phases—plan, generate, validate, test. After each phase, the agent pauses and returns control to the human developer. You can review the optimization plan, look at a draft of the code, and say “no, try a different tile size” or “that test case is wrong.” It’s ideal for complex, novel kernels where a human’s architectural intuition is needed to steer the LLM away from a dead end.
2. Autonomous Loop (Auto): This is the “set it and watch it iterate” mode. The agents are chained together in a closed loop: Plan → Generate → Compile → Test → Tune → Profile → (feedback to Plan). The system runs this loop automatically, often for a fixed number of iterations (e.g., 5). It’s essentially performing a localized hill-climbing search. If it gets stuck in a local optimum (a configuration that’s okay but not great), it can backtrack and try a new plan. At the end, it rolls back to the best version it found. This mode is great for benchmarking many kernels quickly, where you want consistent, repeatable results without holding someone’s hand.
3. Graph-Based Autonomous Search: This scales the Auto agent for broad exploration. Instead of a single linear loop, the system models the kernel design space as a graph. Each node is a candidate kernel configuration. The orchestrator then uses search algorithms to decide which nodes to explore next.
- Parallel Search: Imagine launching five independent Auto agents simultaneously, each starting from the same baseline but taking different random paths. Because they don’t interfere with each other, each gets a “deep” budget of iterations to mature its code. The system then picks the single best result from all five. This is excellent for stabilizing performance—giving complex code enough time to debug itself.
- Beam Search: Here, the system maintains a “beam” of the top-k (e.g., top-3) best candidates at any given time. It explores a little bit with each, but if a candidate starts performing poorly, it gets pruned away aggressively. This is faster and explores more diverse solutions but might miss solutions that require a long sequence of small improvements to bear fruit.
A key technical detail across all these modes is the Retrieval-Augmented Generation (RAG) knowledge base. Loading the entire TPU instruction set into an LLM’s context window is wasteful and often noisy. Instead, MaxKernel dynamically retrieves relevant snippets—like memory layout guides or Pallas documentation—just-in-time. Crucially, it excludes hand-tuned kernel code from this knowledge base, ensuring the agents are forced to discover optimizations themselves rather than just memorizing human-written solutions.
Key Results & Benchmarks
The proof is in the pudding, and MaxKernel has some impressive numbers. The paper evaluates the system on JaxBench, a suite of 50 diverse kernel tasks for TPUs, plus real-world workloads from state-of-the-art open-source models.
When comparing the different orchestration modes against a baseline of simple zero-shot LLM generation (which managed a pitiful 1.08× speedup and only got 20% of kernels to compile correctly), the differences are stark.
- MaxKernel Auto (single trajectory): Shows high stability in producing compilable, correct kernels. However, because a single run can get stuck in a suboptimal state, the speedup numbers have wide variance (represented by median and bracket values).
- MaxKernel Parallel Search: This mode selects the best outcome from five independent Auto runs. It achieves a 1.58× geometric mean speedup across all 50 JaxBench tasks. It also achieves a 100% compilation and correctness rate—every single kernel compiled and ran correctly. Furthermore, 34 out of 50 tasks achieved a “fast-1” speedup (meaning they were at least 1x faster, and often much more).
- MaxKernel Beam Search: This structured search achieves a geometric mean speedup of 1.49× and a 100% correctness rate. While slightly lower than Parallel Search’s top-line speed, it successfully navigates the design space, avoiding local minima and finding diverse high-performing solutions.
Perhaps the most striking results come when comparing against human-written, hand-tuned baselines. On eight specific production kernels where experts have already done the hard work, MaxKernel’s Parallel Search achieved a 2.32× geometric mean speedup over the reference code. The Beam Search variant close behind at 1.78×. Notably, on seven out of eight workloads, the agents surpassed the human-tuned references. For example, on Paged Attention, Beam Search discovered a kernel that achieved a 6.74× speedup, more than doubling the human expert’s 2.41×. On MLA Attention, where human experts struggled to improve upon the baseline, the agents successfully found implementations that beat the standard JAX/XLA compiler.
Beyond the aggregate numbers, MaxKernel demonstrated real-world impact on complex models. For Qwen3-Next Gated DeltaNet, generating kernels for both forward and backward passes reduced forward pass latency by 1.63× and accelerated the overall training step by up to 4.70×. On DeepSeek-V4 Sparse Attention, speedups ranged from 2.36× for small decoding shapes to a massive 7.85× for user-defined prefill operations. Even for memory-bound algorithms like Mamba v2 SSD, MaxKernel kept intermediate matrices snugly in fast Vector Memory, yielding a 1.10× speedup.
Why It Matters (Key Takeaways)
So, why should you, as a reader—whether you’re a product manager, a software engineer, or just a curious AI enthusiast—care about this?
1. It Lowers the Barrier to High-Performance Computing. Historically, getting the most out of accelerators required rare expertise. MaxKernel democratizes this. By automating the iterative compile-test-tune loop, it means that teams can explore performance optimizations without needing a dedicated kernel engineer on staff. It shifts the bottleneck from “who knows the hardware?” to “who can frame the problem for the agent?”
2. It Shows a Path to Faster Model Iteration. In the benchmarks on models like Qwen3 and DeepSeek-V4, we see speedups of up to 5× on training steps. In a production AI setting, a 5× reduction in training step time translates directly to cheaper compute bills and faster iteration cycles for researchers. Faster kernels mean you can experiment with more architectures in the same amount of time.
3. It Highlights the Power of “Agentic” Workflows. MaxKernel is a concrete example of how LLMs can be made reliable for complex engineering tasks. The secret sauce isn’t just the model’s ability to write code, but the surrounding infrastructure: the real-time compiler feedback, the structured sub-agents, the search graphs. This pattern—LLM + tool use + iterative feedback—is likely to appear in many other domains beyond kernel generation, such as hardware design, system administration, or complex scientific modeling.
4. There Are Trade-offs to Watch. The paper notes that different search strategies excel in different scenarios. Parallel Search is great for stability and reaching the absolute best performance given enough time, but it can be computationally expensive because it runs multiple independent trajectories. Beam Search is more efficient and explores a broader frontier, but might miss solutions that require long, winding sequences of tweaks to compile correctly. As the field matures, we’ll likely see hybrid approaches that dynamically switch strategies based on the kernel being optimized.
5. The Open-Source Availability. MaxKernel is open-source. This is significant. It means the community can inspect, extend, and adapt the framework. Whether you want to plug in a different LLM, experiment with new search algorithms, or integrate it with a different accelerator’s compiler, having the code publicly available accelerates the entire ecosystem’s ability to iterate on these ideas.
In summary, MaxKernel represents a meaningful step toward automating the drudgery of low-level accelerator programming. It doesn’t eliminate the need for hardware understanding—you still need to know what you’re optimizing—but it drastically reduces the human effort required to get there. For the broader AI community, it promises faster models and lower costs; for the engineer, it offers a powerful new tool in the quest for performance.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →