Agents & InferencearXiv

GAVEL reduces discrepancies from 7.63 to 0.85 per report

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

GAVEL cut evaluated-timeline discrepancies from 7.63 to 0.85 per clinical case report after LLM-guided merging, with merged timelines preferred in 77.0% of comparisons. The practical takeaway is that timeline extraction evals can move away from brittle “gold” annotations toward grounded pairwise adjudication against source text, but event matching remains a constraint: true match rates were only 60% just below and 48% just above the 0.10 cutoff, so alignment logic still needs careful validation before production use.