Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceSimon Willison

GLM-5.2 is probably the most powerful text-only open weights LLM

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Z.ai has released GLM-5.2 as open weights under an MIT license, a 753B-parameter text-only mixture-of-experts model with a 1 million token context window. It leads open-weight models on Artificial Analysis’s Intelligence Index and ranks second on the Code Arena WebDev leaderboard, though benchmarks suggest it uses more output tokens per task than peers.

What you'll learn · Jun 18, 2026 · 6 stories

  1. 1.GLM-5.2's 1M token context window and 51 Intelligence Index score make it the strongest open text model, but at $4.40/million output tokens it's pricier than competitors for long tasks.
  2. 2.GPT-5.4 autonomously improved a key drug reaction, cutting synthesis time by 40% while maintaining 95% yield, accelerating medicinal chemistry pipelines.
  3. 3.A self-evolving LLM agent for legal case retrieval outperforms human-designed rules on LeCaRD-v2 by iteratively refining query rewrites without parameter training, reducing manual rule engineering costs.
  4. 4.ARD enables agents to dynamically find tools via federated registries, reducing manual integration work while adding ~100ms latency for registry searches during runtime.
  5. 5.Disabling Google Workspace’s AI features eliminates intrusive pop-ups like Gemini, restoring focus during writing tasks but removes potential productivity aids.
  6. 6.Teams building retrieval systems spend weeks on infrastructure integration; Search Toolkit reduces this overhead by unifying ingestion, retrieval, and evaluation into a single framework for faster deployment.
Browse editions · 114 days
Agents & InferenceOpenAI

A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

OpenAI and Molecule.one demonstrated a near-autonomous AI chemist powered by GPT-5.4 that improved a challenging reaction used in drug development. The work points to how AI systems could help accelerate medicinal chemistry research by optimizing key steps in molecule synthesis.

Agents & InferencearXiv

When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Researchers proposed a self-evolving LLM-based agent that improves legal case retrieval by automatically creating, testing and pruning query-rewriting rules for BM25 without parameter training. Evaluated on the Chinese LeCaRD-v2 benchmark, the framework outperformed non-evolutionary baselines such as human-designed rules and greedy rule selection, with gains tied to the LLM’s ability to use prior experimental feedback and eliminate weak rules.

Agents & InferenceHugging Face

Agentic Resource Discovery: Let agents search

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Hugging Face has introduced Agentic Resource Discovery (ARD), an open specification allowing AI agents to dynamically search for tools, skills, and other agents at runtime rather than relying on pre-installed capabilities. Developed with industry contributors like Microsoft and Google, ARD enables federated registries to catalog and index resources, improving scalability and reducing manual integration. Hugging Face’s reference implementation, the Discover Tool, already provides search access to thousands of skills and applications across its platform and other ARD-compatible services.

Agents & InferenceTechCrunch

How to turn off AI in your Google Docs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Google Docs users can disable Gemini prompts and other AI writing features by turning off Google Workspace “smart features” through Gmail settings. The guide frames the change as a way to stop intrusive AI pop-ups and writing suggestions from interrupting work in Docs.

Agents & InferenceMistral

Introducing Search Toolkit

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Mistral has launched Search Toolkit, an open-source framework designed to simplify building production search pipelines for AI applications by unifying ingestion, retrieval, and evaluation tools. The toolkit aims to reduce engineering overhead, allowing teams to focus on improving search quality rather than maintaining integrations across disparate systems. It supports cloud, on-premises, and edge deployments, offering consistent processing for diverse data sources and built-in evaluation for retrieval performance.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by DeepSeek V3 — not one of this week's two contestants.