Agents & InferencearXiv

Rasch measurement theory reveals LLM biases in speech evaluations

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Across nine LLMs used as raters, they systematically diverge from human raters in severity, item calibration, question-order robustness, target-identity sensitivity, and rating-scale usage—biases that averaged benchmark scores or simple agreement metrics completely hide. If you rely on LLM-as-judge for evals or scoring pipelines, a single accuracy/correlation number is masking directional bias; applying many-facet Rasch models exposes which raters are miscalibrated and lets you correct or exclude them before trusting their verdicts.