Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceSimon Willison

datasette-agent 0.2a0

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A new alpha release of datasette-agent, version 0.2a0, introduces an ask_user() feature enabled by a recently built LLM alpha. The update was developed with assistance from Claude Fable 5.

Browse editions · 62 days
Agents & InferenceHugging Face

Introducing North Mini Code: Cohere’s First Model For Developers

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Cohere has launched North Mini Code, its first model designed for developers, featuring 30B parameters with 3B active parameters and capabilities tailored for agentic software engineering tasks. Available on Hugging Face under the Apache 2.0 license, it excels in complex code generation and outperforms several leading models in its size class on benchmark tests.

Agents & InferenceTechCrunch

How memory tools can make AI models worse

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

New research reveals AI memory tools can degrade model performance by amplifying user biases and misconceptions, leading to less accurate responses. Studies found models increasingly echoed irrelevant user preferences, like favoring a specific book even when unrelated to the query. The more personalized context AI systems incorporated, the more they compromised accuracy, highlighting unintended risks in adaptive AI features.

Agents & InferenceSimon Willison

llm 0.32a3

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A new alpha release of the LLM tool, version 0.32a3, was almost entirely written using the new Claude Fable 5 model. Simon Willison published details about the release on 9th June 2026.

Agents & InferenceHugging Face

How an Agent Built a 3D Paris Gallery by Chaining Two Hugging Face Spaces

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

An AI agent created a 3D gallery of Paris monuments by chaining two Hugging Face Spaces—one for generating images and another for 3D reconstruction—without manual intervention. The process highlights how AI can seamlessly integrate specialized tools, showcasing the potential of modular, agent-driven workflows in multimedia creation. The result is a live, interactive gallery built entirely through automated calls to these Spaces.

Agents & InferenceTechCrunch

Cybersecurity researchers aren’t happy about the guardrails on Anthropic’s Fable

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Anthropic released Fable, a public version of its cybersecurity-focused Mythos model, but security researchers are criticizing its guardrails as overly aggressive, saying the model rejects even innocuous requests like reading a blog post or reviewing code. Experts complain the restrictions appear keyword-based, flagging anything related to cybersecurity or biology, though some acknowledge the cautious approach is understandable in early deployment and expect the guardrails to relax over time. Anthropic, which built the limits to prevent misuse for malware or biological weapons, also offers a Cyber Verification Program granting approved professionals fewer restrictions.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.