Agents & InferencearXiv

GuideSkill boosts clinical LLM accuracy by 18.49% without model updates

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

GuideSkill-Evo raises gold-label diagnostic skill coverage from 56.5% to 99.5% and beats direct LLM inference by 18.49% relative without updating the backbone. For production clinical agents, the useful shift is from retrieving guideline text to executing guideline-derived rules as an external, inspectable scoring layer, which can improve accuracy and coverage while avoiding model fine-tuning and making reasoning paths easier to audit.