Agents & InferencearXiv

Native harnesses don't always solve more coding tasks

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Same-model harness swaps on 256 private coding tasks showed no reliable overall win for vendor-native harnesses: Opus 4.8 was 48.8% vs 50.0%, and GPT-5.5 was 55.6% vs 54.4%, both within wide confidence intervals. For production agent teams, the harness choice should be treated as workload- and cost-dependent rather than assumed native-best: repository vs contest tasks diverged sharply for Opus, timeouts sometimes contained passing patches, and the neutral harness cost about 1.2–1.6× more per solved task on observed usage.