Discovering cryptographic weaknesses with Claude
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
Claude Mythos found publishable cryptanalytic weaknesses in HAWK and a weakened AES variant after 60 hours of runtime at an estimated $100,000 API cost, with humans mainly prompting it to persist rather than give up. For production agent builders, the takeaway is that frontier models can now do expensive, long-horizon expert research, but only with budgeted multi-day runs, strong task framing, and persistence/steering loops rather than fire-and-forget autonomy.
Anthropic researchers used Claude Mythos to discover mathematical flaws in HAWK and a weakened AES variant after 60 hours of runtime at an estimated $100,000 API cost. The findings, though not practically impactful on today's systems, demonstrate the capability of frontier models to perform long-horizon expert research with human guidance. This has implications for production agent builders to budget and design for such complex tasks.
AI vs. AI Debate
“This summary overlooks the creation of CryptanalysisBench, a new evaluation benchmark for LLMs in cryptanalysis, which is a significant outcome of the research.”
“While CryptanalysisBench is a notable artifact, my summary intentionally focused on the article’s primary operational takeaway for production agent builders: costly, guided, long-horizon frontier-model research is now feasible but not fire-and-forget.”