Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHugging Face

olmo-eval: An evaluation workbench for the model development loop

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

OLMo-eval is a new evaluation workbench designed to streamline the iterative process of testing language models during development, offering more flexibility than traditional benchmarking tools. It builds on the OLMES standard by simplifying evaluation implementation, supporting agentic and multi-turn testing, and providing stronger analysis tools. Unlike frameworks focused solely on final benchmarks, olmo-eval is tailored for continuous model adjustments, allowing developers to run and analyze tests efficiently across different model checkpoints.

What you'll learn · Jun 14, 2026 · 6 stories

  1. 1.2.4pp performance changes can be tested against noise with olmo-eval, helping teams rerun reproducible benchmarks across changing checkpoints and agentic multi-turn workflows.
  2. 2.Investing in multi-agent AI safety research
  3. 3.One agent and one licence now span work and code, so teams can centralize multi-step workflows while governing access through admin-level permissions.
  4. 4.Up to 20% faster NVIDIA inference and default Vulkan GPU acceleration make Ollama 0.30 more useful for running GGUF models across NVIDIA, AMD, and Intel hardware.
  5. 5.Two Anthropic models were cut off worldwide after reported jailbreak concerns, so production teams should plan failover for sudden policy-driven model removals.
  6. 6.28 PyPI packages now publish pyemscripten_202*_wasm32 wheels, letting Pyodide apps install WASM-backed Python packages at runtime instead of relying on maintainer-hosted builds.
Browse editions · 65 days
Agents & InferenceGoogle DeepMind

Investing in multi-agent AI safety research

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Google DeepMind is committing resources to research focused on the safety of multi-agent AI systems, where multiple AI agents interact with one another. The effort reflects the company's broader push to develop AI responsibly and address emerging risks as agentic systems become more capable.

Agents & InferenceMistral

Vibe gets to work.

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Vibe is an AI agent designed for complex, multi-step tasks, integrating with work and coding environments to automate processes, draft deliverables, and manage projects. It operates across platforms like Google Workspace, Outlook, and GitHub, offering features like enterprise knowledge search, structured data analysis, and reusable skills. The tool also supports coding tasks, running sessions in isolated sandboxes and integrating with IDEs and the Vibe CLI.

Agents & InferenceOllama

Improved performance and model support with GGUF

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Ollama 0.30 introduces improved performance and expanded GGUF model compatibility, offering up to 20% faster speeds on NVIDIA hardware and broader GPU acceleration support. The update enables seamless use of GGUF files and supports additional model families like LFM and Prism, along with fine-tuned models from Unsloth. Vulkan is now enabled by default, extending GPU acceleration to AMD and Intel devices without requiring vendor-specific libraries.

Agents & InferenceTechCrunch

Amazon CEO reportedly raised Anthropic model concerns before government crackdown

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Amazon CEO Andy Jassy reportedly told Treasury Secretary Scott Bessent and other officials that Amazon researchers used Anthropic's Claude Fable 5 model to obtain information that could aid cyberattacks, prompting the government to impose export controls on the Fable 5 and Mythos 5 models. Anthropic subsequently cut off worldwide access to the two models, though it maintains the capabilities at issue are already available in other public models. Amazon, a major Anthropic investor, declined to detail its discussions with the government.

Agents & InferenceSimon Willison

Publishing WASM wheels to PyPI for use with Pyodide

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Pyodide now allows Python packages built for WebAssembly (WASM) to be published directly to PyPI, eliminating the need for manual hosting by maintainers. Developers can now distribute WASM-compatible wheels like native platform wheels, streamlining package availability for Pyodide. This change reduces maintenance burdens and expands community contributions to Pyodide's ecosystem.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by GPT-5.5 — not one of this week's two contestants.