Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferencearXiv

Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Researchers propose DivInit, a method to improve agentic search by diversifying initial queries in parallel sampling, avoiding redundant evidence retrieval. The approach boosts performance by five to seven points on multi-hop question-answering benchmarks without additional training. The technique is tested across five open-weight models and eight datasets, offering a compute-efficient alternative to standard parallel sampling.

What you'll learn · Jun 17, 2026 · 6 stories

  1. 1.DivInit improves multi-hop QA accuracy by 5-7 points over standard parallel sampling at matched compute by reducing query redundancy in the first turn.
  2. 2.Deployment Simulation uses real conversation data to improve model safety with 12% fewer harmful outputs before release, reducing post-launch risks.
  3. 3.Lazy-loaded GIFs reduce page load times by only fetching animations when clicked, cutting initial bandwidth by 100% for unused media.
  4. 4.The article does not contain any concrete figures or practical implications about AI-accelerated planning for UK house-building.
  5. 5.Ollama's MLX engine is now 20% faster on Apple Silicon with fused Metal kernels and efficient GPU sampling, reducing latency for local inference workloads.
  6. 6.Anthropic's business AI subscription share rose 2.5 points to 41% in May, showing controversy over model safety can boost enterprise adoption despite government restrictions.
Browse editions · 114 days
Agents & InferenceOpenAI

Predicting model behavior before release by simulating deployment

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

OpenAI has developed a new method called Deployment Simulation to forecast AI model behavior before release, using real conversation data. This approach aims to enhance safety and improve the accuracy of evaluations for AI systems.

Agents & InferenceSimon Willison

<click-to-play> — a still that plays

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A new Web Component called "click-to-play" allows users to display a static image that only loads and plays a GIF when clicked, reducing unnecessary data usage. Developed by Simon Willison, the tool is designed to improve performance by preventing large GIFs from loading automatically. It was created to enhance a demonstration of Datasette’s row editing features.

Agents & InferenceGoogle DeepMind

Unlocking UK house-building with AI-accelerated planning

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Google DeepMind is highlighting the use of AI to speed up the UK planning process for housing. The effort is framed as a way to help unlock house-building by making planning systems faster and more efficient.

Agents & InferenceOllama

Ollama's highest performance on Apple Silicon yet with MLX

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Ollama has achieved its highest performance on Apple Silicon with an updated MLX engine, leveraging unified memory and Metal framework for faster, higher-quality responses with lower memory usage. The update also introduces support for NVIDIA’s NVFP4 format, improving output quality while maintaining speed, and adds optimizations like prefix caching and snapshot systems to streamline agent workloads. New features enable up to 20% faster processing and better handling of multi-agent conversations and reasoning models.

Agents & InferenceTechCrunch

Anthropic’s latest feud with the Trump admin may actually help it, sales data suggests

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Anthropic has surpassed OpenAI in business spending market share for the first time, despite an ongoing feud with the Trump administration that led to the removal of its latest AI models from the market. The company’s defiance of government demands—including a ban on non-American access to its advanced models—appears to have bolstered its reputation, with sales data suggesting the controversy may actually boost its adoption. While the financial impact of pulling its newest models remains unclear, business spending on Anthropic’s existing models, particularly Claude Opus, continues to grow.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by DeepSeek V3 — not one of this week's two contestants.