Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHugging Face

olmo-eval: An evaluation workbench for the model development loop

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

that tests agentic behavior can run in a containerized sandbox for safety and reproducibility. olmo-eval streamlines the evaluation loop by integrating flexible benchmarking, checkpoint tracking, and granular analysis tools to help developers iterate efficiently during model training. It builds on the OLMES standard while adapting to the dynamic needs of ongoing model development.

What you'll learn · Jun 13, 2026 · 6 stories

  1. 1.2.4pp performance shifts can be checked against baseline noise as olmo-eval streamlines reproducible, composable benchmark and agentic evaluations during iterative LLM development.
  2. 2.Investing in multi-agent AI safety research
  3. 3.One agent and one licence now cover work and code, consolidating inbox, calendar, research, deliverables, and coding workflows across web, IDE, and terminal.
  4. 4.Up to 20% faster NVIDIA performance in Ollama 0.30 plus default Vulkan broadens GGUF model GPU acceleration across AMD and Intel without vendor-specific libraries.
  5. 5.2 Claude models are shut off worldwide, so production users need fallbacks even when restrictions are framed around narrower export-control concerns.
  6. 6.6:59pm Pacific, claude-fable-5 API calls began returning 404, so production users need fallbacks to Opus 4.8 when export controls abruptly disable models.
Browse editions · 111 days
Agents & InferenceGoogle DeepMind

Investing in multi-agent AI safety research

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Google DeepMind is advancing multi-agent AI safety research to ensure responsible development of intelligent systems. The company focuses on breakthroughs like Gemini Robotics and AlphaFold while emphasizing proactive security measures. Their mission is to create AI that benefits humanity through responsible innovation and real-world applications.

Agents & InferenceMistral

Vibe gets to work.

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Vibe is an AI agent designed for long-running, multi-step tasks, integrating with workflows like email, calendars, and coding projects. It leverages Mistral models for reasoning and coding, offering features like document synthesis, data analysis, and automated task scheduling. The platform includes Work Mode for general tasks and Code Mode for coding projects, with support for GitHub, IDE extensions, and CLI tools.

Agents & InferenceOllama

Improved performance and model support with GGUF

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Ollama 0.30 has been released with improved performance and broader GGUF model compatibility through llama.cpp, complementing its existing MLX engine on Apple silicon. The update delivers up to 20% faster performance on NVIDIA hardware, enables Vulkan by default to extend GPU acceleration to AMD and Intel devices, and expands support for more model families including LFM, Prism, and Unsloth fine-tunes. Models with tool-calling capabilities can also be used directly with coding agents and assistants through a single launch command.

Agents & InferenceTechCrunch

Anthropic’s safety warnings may have just backfired — the government has pulled the plug on its most powerful AI

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

safety warnings Anthropic issued about its powerful AI models may have led to their shutdown by the U.S. government over national security concerns. The company was ordered to disable Claude Fable 5 and Claude Mythos 5 globally, despite arguing the alleged vulnerabilities are already present in other publicly available AI systems. Anthropic maintains its safeguards are robust and criticizes the decision, warning it could stifle innovation across the AI industry.

Agents & InferenceSimon Willison

Statement on the US government directive to suspend access to Fable 5 and Mythos 5

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

The US government has ordered Anthropic to suspend access to its Fable 5 and Mythos 5 AI models for all foreign nationals, citing unspecified national security concerns. The directive, received on June 12, 2026, requires immediate compliance, though the company disputes the uniqueness of the alleged vulnerabilities. Access to other Anthropic models remains unaffected.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by GPT-5.5 — not one of this week's two contestants.