Agents & InferencearXiv

CITA improves Tool F1 and task success in long-horizon tool-use agents

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A new method, CITA, improves tool-use agents by estimating the likelihood of a tool invocation leading to task success, and it consistently improves Tool F1 and task success across three benchmarks. This enables more accurate decision-making in long-horizon tool use for LLMs, directly impacting production agents' ability to choose the right tool invocations. This improvement matters for shipping reliable LLM-based applications that rely on complex tool invocation sequences.