Agents & InferencearXiv

When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

MARCH declared both code solutions equally good in 78–95% of comparisons and fell to 4.4% accuracy where direct judging hit 43.7%, because its evidence was not candidate-distinguishing in code tasks. For production code-eval agents, multi-agent verification is not automatically safer than direct LLM judging; you need label-free grounding checks and an abstain path, which can lift accuracy from 20.7% to 36.9% while only answering about half of comparisons.