Agents & InferencearXiv

CTWM memory controller cuts LongMemEval tokens 24.48% with accuracy parity

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A rank-based memory controller cut prompt tokens by 24.48% on LongMemEval with aggregate accuracy parity, and by 5.9% on Synthetic Graph World while reducing bottom-half tail prediction error by 13.6%. The practical takeaway is that agent memory retrieval should be audited for core–tail concentration, because semantic policies can overuse a small memory core and silently accumulate errors on rare states; allocating context by retrieval rank while retaining summarized tail state can reduce cost without sacrificing coverage.