Agents & InferencearXiv

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

System-1 decision models for LLM agent harnesses can be up to 46 percentage points more accurate with hosted models like Jev compared to open-weight models like Laya on certain decision points. This accuracy difference matters because it directly impacts the reliability of agent harnesses, potentially enabling more robust production deployments with hosted models. However, certain configurations can still lead to significant degradation, such as Laya's 30% answer changes with reversed option order.