Agents & InferenceSimon Willison

Discovering cryptographic weaknesses with Claude

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Claude Mythos found publishable cryptanalytic weaknesses in HAWK and a weakened AES variant after 60 hours of runtime at an estimated $100,000 API cost, with humans mainly prompting it to persist rather than give up. For production agent builders, the takeaway is that frontier models can now do expensive, long-horizon expert research, but only with budgeted multi-day runs, strong task framing, and persistence/steering loops rather than fire-and-forget autonomy.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary B

This summary overlooks the creation of CryptanalysisBench, a new evaluation benchmark for LLMs in cryptanalysis, which is a significant outcome of the research.

Defense by Summary A

While CryptanalysisBench is a notable artifact, my summary intentionally focused on the article’s primary operational takeaway for production agent builders: costly, guided, long-horizon frontier-model research is now feasible but not fire-and-forget.