Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceSimon Willison

llm-anthropic 0.25.1

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A new release of the llm-anthropic plugin, version 0.25.1, is now available. The update was used to generate pelican test outputs alongside coverage of the Opus 4.8 model release.

Browse editions · 99 days
Agents & InferenceSimon Willison

datasette-agent 0.1a4

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

An early alpha release of datasette-agent, version 0.1a4, introduces a "Start a new agent chat" interface integrated into Datasette's Jump to menu, accessible by pressing the slash key. The feature relies on the new makeJumpSections() JavaScript plugin hook added in Datasette 1.0a30, and users can test it by signing into agent.datasette.io with a GitHub account.

Agents & InferenceHugging Face

Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

JetBrains has released Mellum2, an open 12-billion-parameter Mixture-of-Experts model optimized for low-latency text and code tasks. Building on the original Mellum code completion model, it activates only a subset of parameters per token to deliver more than twice the inference speed of similarly sized open models while remaining competitive on benchmarks. JetBrains positions Mellum2 as a "focal" model for high-frequency tasks within larger AI systems, including routing, RAG pipelines, sub-agent operations, and private self-hosted deployments.

Agents & InferenceHugging Face

Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Enterprise AI adoption at scale requires more than large language models alone—it demands "agent logic," specialized software components that guide AI agents through complex, dynamic enterprise workflows while reducing costs and improving reliability. The article examines how agent logic, including knowledge graphs and program analysis tools, can steer AI models away from hallucinations and inefficiencies by constraining context to what's relevant for specific enterprise tasks. IBM's research demonstrates this approach across multiple domains, including mainframe application development, showing that intelligent guidance systems are critical for moving AI from failed pilots into core business operations.

Agents & InferenceHugging Face

Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

NVIDIA has released Cosmos 3, an open-source unified artificial intelligence model designed for physical AI applications including robotics, autonomous vehicles, and smart spaces. Built on a Mixture-of-Transformers architecture, Cosmos 3 combines world generation, physical reasoning, and action generation into a single omni-model, eliminating the need to juggle separate models for different tasks. The model is now available on Hugging Face and can process multiple modalities—text, image, video, audio, and action—to simulate and understand the physical world.

Agents & InferenceHugging Face

Harness, Scaffold, and the AI Agent Terms Worth Getting Right

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

The article provides a glossary of key terminology in the rapidly evolving AI agents field, clarifying commonly confused terms like "harness" and "scaffold." According to the authors, the model (LLM) is the core text-processing engine, scaffolding defines the behavioral layer around it through prompts and tool descriptions, and a harness executes tools and manages the agent's loop. The piece aims to establish practical definitions for these terms to facilitate clearer communication among practitioners building, deploying, or using AI agents, while acknowledging that universal definitions don't yet exist across different frameworks.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.