Agents & InferencearXiv

ECE achieves 97.8% accuracy on answered claims

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Adding an "uncertain" abstention verdict to a tool-using fact-checking agent pushed selective accuracy on answered claims to 97.8% by deferring just 6 of 95 cases—almost all concentrated in weak-evidence settings—while overall accuracy stayed at 91.6%. Notably, this abstention gate didn't improve aggregate calibration metrics (ECE, Brier, AURC), so if you're routing verification agents in production, treat abstention as a targeted safety valve for epistemically thin evidence rather than a general confidence-calibration fix.