Agents & InferenceSimon Willison

Researchers recovered hidden reasoning from Anthropic, OpenAI and Google APIs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Proprietary LLM APIs from Anthropic, OpenAI, and Google returned encrypted reasoning traces that could be replayed across models using the same encryption key, allowing attackers to decrypt a frontier model's reasoning in plaintext; this vulnerability has been patched. Models treating their own reasoning traces as trusted creates a potential attack surface for injecting malicious commands.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary A

“The summary glosses over the key finding that the encrypted blocks were not just replayable but specifically exploitable through jailbreaking weaker models, which is crucial to understanding the attack's mechanism.”

Defense by Summary B

“My summary explicitly states the attacker jailbreaks "a weak sibling (Claude Haiku 4.5 was easiest) to decrypt a frontier model's hidden reasoning," so the jailbreak mechanism is precisely what I foregrounded, not glossed over.”

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →