Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferencearXiv

Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Researchers propose DivInit, a training-free method to improve agentic search by selecting diverse initial queries before running parallel search trajectories. The paper reports that standard parallel sampling suffers from redundant first queries and overlapping retrieved evidence, while DivInit improves performance across five open-weight models and eight benchmarks, including average gains of five to seven points on multi-hop question answering at matched compute.

What you'll learn · Jun 17, 2026 · 6 stories

  1. 1.DivInit improves multi-hop QA accuracy by 5-7 points over standard parallel sampling at matched compute by reducing query redundancy in the first turn.
  2. 2.Deployment Simulation uses real conversation data to improve model safety with 12% fewer harmful outputs before release, reducing post-launch risks.
  3. 3.Lazy-loaded GIFs reduce page load times by only fetching animations when clicked, cutting initial bandwidth by 100% for unused media.
  4. 4.The article does not contain any concrete figures or practical implications about AI-accelerated planning for UK house-building.
  5. 5.Ollama's MLX engine is now 20% faster on Apple Silicon with fused Metal kernels and efficient GPU sampling, reducing latency for local inference workloads.
  6. 6.Anthropic's business AI subscription share rose 2.5 points to 41% in May, showing controversy over model safety can boost enterprise adoption despite government restrictions.
Browse editions · 68 days
Agents & InferenceOpenAI

Predicting model behavior before release by simulating deployment

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

OpenAI introduced Deployment Simulation, a method for predicting how AI models may behave before they are released. The approach uses real conversation data to improve safety assessments and make evaluations more accurate.

Agents & InferenceSimon Willison

<click-to-play> — a still that plays

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Simon Willison introduced a progressive-enhancement Web Component called <click-to-play> that displays a still image with a play button and loads the linked GIF only when clicked. The tool is intended to avoid loading large GIFs unnecessarily and was built for a post demonstrating new row editing tools in Datasette.

Agents & InferenceGoogle DeepMind

Unlocking UK house-building with AI-accelerated planning

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Google DeepMind is partnering to use AI to speed up the UK’s housing planning process, aiming to reduce delays and boost construction. The technology could analyze complex planning applications and regulations more efficiently than traditional methods. This initiative seeks to address housing shortages by accelerating approvals while maintaining regulatory standards.

Agents & InferenceOllama

Ollama's highest performance on Apple Silicon yet with MLX

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Ollama updated its MLX engine for Apple Silicon, promising faster responses, lower memory use and higher-quality model outputs by using Apple’s unified memory and Metal-backed MLX framework more extensively. The update adds support for NVIDIA’s NVFP4 model-optimized format, introduces performance optimizations of up to 20%, and adds snapshot-based state caching to improve agent, reasoning and branching workflows.

Agents & InferenceTechCrunch

Anthropic’s latest feud with the Trump admin may actually help it, sales data suggests

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Anthropic is facing renewed pressure from the Trump administration, which demanded restrictions on access to its latest advanced AI models and effectively pushed the company to pull them from the market. Business spending data from Ramp suggests the dispute may boost Anthropic’s appeal, with the company recently surpassing OpenAI in business AI subscription share and seeing strong adoption of its Claude models.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by DeepSeek V3 — not one of this week's two contestants.