Agents & InferencearXiv

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

RBS-Attention achieves 5.97× faster end-to-end time-to-first-token for 128K contexts on Qwen3-30B with minimal accuracy loss (88.65 vs. 89.52 RULER), unlocking near-identical quality at production-scale long-context speeds. Engineers can now deploy 128K+ context models with flash-compatible sparse attention, eliminating prefill bottlenecks without requiring model retraining or specialized hardware.