Agents & InferenceTechCrunch

Nvidia just showed that the harness, not the AI model, is now the real hero

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Nvidia researchers achieved a 100% score on the ARC-AGI-3 reasoning benchmark using Claude Opus 5 wrapped in a custom memory-managed supervisor harness, up from the raw model's baseline of 30%. For engineers shipping production agents, this proves that solving complex, long-horizon tasks depends far less on upgrading to the latest raw frontier model and far more on building sophisticated runtime scaffolding, memory management, and supervisor layers.