Unreleased Anthropic model tested 650 ideas on the Riemann hypothesis
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
An unreleased Anthropic model autonomously ran a 1.5-day, 60-subagent orchestration burning 31 million output tokens across 650 candidate approaches, and produced a genuine, Lean-formalized advance on the Riemann hypothesis—directed by a non-mathematician. The signal for you isn't the math result but the operational proof point: long-horizon, self-coordinating multi-agent runs at massive token spend can now generate expert-verified novel output, which means your agent architectures and cost/token budgeting should plan for extended autonomous swarms with dedicated validator agents rather than single-shot calls.
An unreleased Anthropic model tested 650 different ideas and spent 31 million output tokens to make significant progress on the Riemann hypothesis, a longstanding math problem, with minimal human guidance. This demonstrates that large language models can achieve substantial mathematical breakthroughs autonomously, potentially changing how mathematicians approach research and raising questions about authorship and responsibility. This capability shift may significantly impact the development and deployment of LLMs in scientific and mathematical applications.