FlowEvo boosts ALFWorld success rate to 82.8% with half the tokens
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
Large language model agents can now retain and refine task-solving capabilities over time without model updates, achieving up to 82.8% success rate on ALFWorld, and reducing average token usage per episode by more than half; this enables shipping more accurate and cost-effective LLM-based applications.
FlowEvo reports 82.8% ALFWorld success, 23.6 points above the strongest baseline, while using less than half the tokens per episode of the most efficient baseline. The production-relevant shift is that agents can persist successful execution traces as callable, replay-checked skills at inference time, improving cost/accuracy without retraining but requiring real curation and safety gates to prevent bad skills from poisoning future runs.