MERIT finds memory lifts tool-agent success from 0.00 to 0.55-1.00
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
Memory implementations in tool-using LLM agents can shift task success rates by up to 60 points, with structured fact stores outperforming embedding retrieval on updated facts (70-100% success vs. 30-95%). This means choosing the wrong memory strategy can drastically reduce agent performance, and full replay is never cost-effective—optimal memory use delivers 2.7-3.9x more utility per dollar. Engineers must carefully select memory architectures to maximize both task success and cost efficiency in production systems.
The marginal utility of long-term memory in tool-using LLM agents can increase task success rates from 0.00 to 0.55-1.00, but the choice of memory implementation can move task success by up to 60 points and affect costs by a factor of 2.7-3.9x; this variability directly impacts the cost-effectiveness and reliability of production LLM agents.