Agents & InferencearXiv

AgentKVShift cuts agentic memory prefill latency 2-3.5x on A100

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

AgentKVShift gets near full-recompute quality for agentic memory retrieval while refreshing only 10–30% of KV cache, producing 2–3.5x prefill speedups on a single A100 versus no KV reuse. The key production implication is that structured, metadata-heavy agent memories can be cached and reused without the quality collapse seen in RAG-oriented KV reuse, including under 2- and 4-bit KV quantization.