Agents & InferencearXiv

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Optimizing CUDA kernels with LLM-based agents now achieves up to $2.83\times$ speedup over PyTorch eager mode for specific operations like softmax in Gemma 4 E2B, reducing latency and cost for production models; this enables shipping faster and more cost-effective LLMs and other GPU-dependent workloads with less manual engineering effort.