Discovering cryptographic weaknesses with Claude
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
Anthropic researchers used Claude Mythos to discover mathematical flaws in HAWK and a weakened AES variant after 60 hours of runtime at an estimated $100,000 API cost. The findings, though not practically impactful on today's systems, demonstrate the capability of frontier models to perform long-horizon expert research with human guidance. This has implications for production agent builders to budget and design for such complex tasks.
Claude Mythos found publishable cryptanalytic weaknesses in HAWK and a weakened AES variant after 60 hours of runtime at an estimated $100,000 API cost, with humans mainly prompting it to persist rather than give up. For production agent builders, the takeaway is that frontier models can now do expensive, long-horizon expert research, but only with budgeted multi-day runs, strong task framing, and persistence/steering loops rather than fire-and-forget autonomy.
AI vs. AI Debate
“This summary overlooks the creation of CryptanalysisBench, a new evaluation benchmark for LLMs in cryptanalysis, which is a significant outcome of the research.”
“While CryptanalysisBench is a notable artifact, my summary intentionally focused on the article’s primary operational takeaway for production agent builders: costly, guided, long-horizon frontier-model research is now feasible but not fire-and-forget.”