Rasch measurement theory reveals LLM biases in speech evaluations
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
LLMs as raters/judges have **systematic biases**—severity, item miscalibration, and identity sensitivity—that standard benchmarks hide. Rasch measurement theory (RMT) exposes these flaws by decomposing ratings into comparable, debuggable facets. Adopting RMT means your evals will catch hidden biases before they skew rankings, break fairness in production, or silently degrade downstream tasks like content moderation or model selection.
Across nine LLMs used as raters, they systematically diverge from human raters in severity, item calibration, question-order robustness, target-identity sensitivity, and rating-scale usage—biases that averaged benchmark scores or simple agreement metrics completely hide. If you rely on LLM-as-judge for evals or scoring pipelines, a single accuracy/correlation number is masking directional bias; applying many-facet Rasch models exposes which raters are miscalibrated and lets you correct or exclude them before trusting their verdicts.