Agents & InferencearXiv

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

With 50 optimization iterations per kernel, Kernel Forge generated CUDA replacements that beat PyTorch eager on 14 kernels, including 2.83× softmax speedup on Gemma 4 E2B and 1.70× group_norm on Stable Diffusion 3.5 Medium. The practical shift is that agent-written kernels can now be tested against whole, unmodified PyTorch models rather than toy isolated ops, making this closer to a drop-in latency/cost optimization loop for production inference stacks.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →