Agents & InferencearXiv

MERIT finds memory lifts tool-agent success from 0.00 to 0.55-1.00

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Memory implementations in tool-using LLM agents can shift task success rates by up to 60 points, with structured fact stores outperforming embedding retrieval on updated facts (70-100% success vs. 30-95%). This means choosing the wrong memory strategy can drastically reduce agent performance, and full replay is never cost-effective—optimal memory use delivers 2.7-3.9x more utility per dollar. Engineers must carefully select memory architectures to maximize both task success and cost efficiency in production systems.