Agents & InferencearXiv

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A new method cuts false-pass rates in LLM-as-judge evaluations by ~50% (0.115 vs. 0.173 on tau-bench) without touching model weights—just by evolving a human-readable rubric from a small set of labeled trajectories. This matters because every false pass ships a broken agent to production, so you can now deploy with half the risk of silent failures while keeping the same inference cost and latency. The rubric is frozen text, so you can audit or tweak criteria without retraining.