Agents & InferencearXiv

RENDER benchmark shows memory formats boost LLM accuracy up to 72.6 points

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Matched-budget “resolved packets” beat recency-truncated raw dialogue by 42.4–72.6 points on 500 LongMemEval questions across nine models, and deployed-style memory templates varied by 24.6–48.8 points within the same model. For production RAG/memory systems, the format you feed the model—summary, typed record, ChatGPT-style memory entry, or raw transcript—is not a neutral implementation detail; it can dominate measured quality, so evaluations must lock or report the reader-facing artifact before comparing retrievers, memory stores, or models.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →