GAVEL reduces discrepancies from 7.63 to 0.85 per report
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
The GAVEL judge protocol identifies extraction errors between structured timelines and raw text with up to 89.4% accuracy, enabling automated merging that slashes clinical timeline discrepancies from 7.63 to 0.85 per report. For production pipelines extracting complex chronological data, this eliminates the need for human gold-standard references by establishing a reliable self-correction loop that adjudicates and merges multi-agent outputs. Implementing this grounded verification system allows you to ship highly reliable chronological extractions with a 77% preference rate over single-model outputs while drastically cutting manual validation overhead.
GAVEL cut evaluated-timeline discrepancies from 7.63 to 0.85 per clinical case report after LLM-guided merging, with merged timelines preferred in 77.0% of comparisons. The practical takeaway is that timeline extraction evals can move away from brittle “gold” annotations toward grounded pairwise adjudication against source text, but event matching remains a constraint: true match rates were only 60% just below and 48% just above the 0.10 cutoff, so alignment logic still needs careful validation before production use.