Agents & InferencearXiv

FlowEvo boosts ALFWorld success rate to 82.8% with half the tokens

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Large language model agents can now retain and refine task-solving capabilities over time without model updates, achieving up to 82.8% success rate on ALFWorld, and reducing average token usage per episode by more than half; this enables shipping more accurate and cost-effective LLM-based applications.