Agents & InferencearXiv

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

FARS-generated papers score 2.1–2.5 on a 1–5 scale, more than 2× higher than the next-best AI scientist system. This means you can now use multi-LLM review (Gemini + Claude) as a drop-in replacement for human peer review to validate autonomous research agents in production, cutting evaluation time from weeks to hours while maintaining >0.9 correlation with expert judgment.