Agents & InferencearXiv

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

AI-generated research papers from leading autonomous scientist systems score 2.14–2.47 (on a 1–5 scale) versus 1.00–1.87 for others, with FARS benchmark papers consistently outperforming by 2x—validated by strong inter-model agreement ($\rho$ = 0.907). This proves LLM-based automated peer review can reliably rank research quality, enabling scalable evaluation of AI-generated science without human reviewers. If you deploy autonomous research agents, you should adopt multi-model (Gemini/Claude) scoring to audit output quality at scale.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →