Agents & InferencearXiv

ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

"ToolSense" is a new diagnostic framework designed to evaluate how well large language models (LLMs) understand and retrieve tools from large catalogs. It introduces benchmarks to test retrieval accuracy under realistic query ambiguity and probes models' factual knowledge about tools, revealing gaps in performance compared to traditional benchmarks. The framework highlights a dissociation between retrieval success and actual tool knowledge in some models.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →