Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHugging Face

Adding MCP Tools to Reachy Mini

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Reachy Mini's conversation app now supports adding tools hosted on Hugging Face Spaces via MCP, allowing users to expand the robot's capabilities without modifying the app directly. These remote tools, such as weather checks or web searches, run in the cloud rather than locally on the user's machine. Users can also publish their own tools for others to utilize through the platform.

Browse editions · 101 days
Agents & InferenceSimon Willison

datasette-agent-micropython 0.1a0

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Simon Willison announced the alpha release of Datasette Agent Micropython 0.1a0, aimed at safely generating and executing Python code. Early testing shows promise, with GPT-5.5 unable to break the sandbox security measures. The update was shared on June 2, 2026, alongside a sponsorship offer for exclusive LLM development insights.

Agents & InferenceTechCrunch

Meta’s AI agent for WhatsApp Business is now available globally

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Meta has rolled out its customer support AI bot, now called Meta Business Agent, globally within WhatsApp and Instagram DMs after roughly two years of testing in markets like India and Mexico. The agent can answer customer questions, recommend products, book appointments, qualify sales leads, and route queries to humans, with planned features including daily chat briefings and integrations with tools like Shopify and Zendesk. Meta intends to monetize the tool through WhatsApp Business Premium subscription tiers and token-based pricing for large businesses.

Agents & InferenceHugging Face

Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

JetBrains has released Mellum2, an open 12-billion-parameter Mixture-of-Experts model optimized for low-latency text and code tasks. Building on the original Mellum code-completion model, it activates only a subset of parameters per token to deliver more than twice the inference speed of similarly sized open models while remaining competitive on coding, reasoning, science, and math benchmarks. JetBrains positions it as a "focal" model for high-frequency operations such as routing, RAG pipelines, sub-agent tasks, and private self-hosted deployments.

Agents & InferenceSimon Willison

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Uber has imposed a $1,500 monthly spending limit per employee for AI coding tools like Claude Code to control costs after exceeding its AI budget. The cap applies separately to each tool, allowing engineers to spend up to $3,000 monthly if using two tools. This policy reflects Uber's effort to balance AI tool benefits with cost management as usage surged unexpectedly.

Agents & InferenceTechCrunch

New Microsoft tool lets devs spin up AI behavior tests using text descriptions

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Microsoft has released ASSERT, an open source framework that lets developers test whether their AI systems behave as intended by turning plain-language descriptions of goals and policies into scored, structured tests. The tool generates problem scenarios, runs them against the target system, and records the AI's actions and tool calls so developers can pinpoint failures. It can be used during development, after deployment, or for continuous monitoring, addressing the need for application-specific evaluations that broader benchmarks miss.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.