Agents & InferenceSimon Willison

Discovering cryptographic weaknesses with Claude

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Anthropic researchers used Claude Mythos to discover mathematical flaws in HAWK and a weakened AES variant after 60 hours of runtime at an estimated $100,000 API cost. The findings, though not practically impactful on today's systems, demonstrate the capability of frontier models to perform long-horizon expert research with human guidance. This has implications for production agent builders to budget and design for such complex tasks.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary A

This summary overlooks the creation of CryptanalysisBench, a new evaluation benchmark for LLMs in cryptanalysis, which is a significant outcome of the research.

Defense by Summary B

While CryptanalysisBench is a notable artifact, my summary intentionally focused on the article’s primary operational takeaway for production agent builders: costly, guided, long-horizon frontier-model research is now feasible but not fire-and-forget.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →