Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
System-1 decision models for LLM agent harnesses can be up to 46 percentage points more accurate with hosted models like Jev compared to open-weight models like Laya on certain decision points. This accuracy difference matters because it directly impacts the reliability of agent harnesses, potentially enabling more robust production deployments with hosted models. However, certain configurations can still lead to significant degradation, such as Laya's 30% answer changes with reversed option order.
Deploying open-weight System-1 models for agent routing yields an actual cost savings of just 4.3%—far below the claimed 23.9%—once mandatory pre-screening overhead is factored in. Furthermore, these cheaper classifiers are highly brittle, with the open-weight Laya model changing 30% of its answers when option order is reversed and dropping to 31% accuracy on 50 nearest-neighbor tools. For production agents, this means relying on lightweight classifiers for tool gating will break your routing logic at scale while failing to deliver any meaningful cost reductions.