Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHacker News

GLM-5.2 – How to Run Locally

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

GLM-5.2 is an AI model that can be operated on local machines. The process for running GLM-5.2 locally is explained in the associated post. Details on setup and implementation are likely covered.

What you'll learn · Jun 23, 2026 · 6 stories

  1. 1.GLM-5.2 can be run locally, shifting inference control on-prem while requiring teams to verify hardware fit before production.
  2. 2.0.2B parameters can run in-browser via WebGPU, making lightweight inpainting feasible without a PyTorch and NVIDIA CUDA deployment.
  3. 3.3k–10k token budgets show DPTS needs enough exploration while SSDP can deplete frontiers, so production ToT should adapt search and pruning to available compute.
  4. 4.Four reasoning benchmarks and six LLM backbones suggest PEAR can improve multi-agent debate accuracy by adaptively rerouting roles, while reducing routing complexity and positional bias.
  5. 5.1-prompt limits make Codex’s context preservation useful for complex projects that need work to continue across longer-running tasks.
  6. 6.GPT-5.5-Cyber and Codex Security could help security teams find, validate, and patch vulnerabilities at scale across organizations.
Browse editions · 134 days
Agents & InferenceSimon Willison

Porting the Moebius 0.2B image inpainting model to run in the browser with Claude Code

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

The Moebius 0.2B image inpainting model, originally requiring PyTorch and NVIDIA CUDA, was successfully ported to run in a browser using WebGPU. The ported model allows users to mark regions of an image to remove and fill the space with a generated alternative. A demo of the browser-based model is available at simonw.github.io/moebius-web/.

Agents & InferencearXiv

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A new research study examines how different Tree of Thought reasoning strategies perform under varying computational budgets, model sizes, and problem difficulties. The analysis reveals that existing search methods suffer from opposing limitations, with one approach struggling at low budgets due to high exploration requirements and another failing to scale because of aggressive path pruning. To overcome these constraints, the researchers suggest that future artificial intelligence reasoning agents must employ adaptive search strategies that adjust based on available resources and real-time search progress.

Agents & InferencearXiv

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Researchers have introduced PEAR, a protocol for multi-agent debate that dynamically reconfigures communication roles to improve the reliability of large language models. PEAR prevents persistent positional biases and uneven influence across debates, and has been shown to improve average accuracy across four reasoning benchmarks and six diverse large language model backbones. The protocol is designed to be permutation-equivariant and sparse, reducing routing complexity and improving generalization.

Agents & InferenceOpenAI

Codex-maxxing for long-running work

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Jason Liu utilizes Codex to manage and preserve context in long-running projects, enabling work to continue beyond a single prompt. Codex helps Liu handle complex tasks and maintain continuity. This approach supports more efficient project management.

Agents & InferenceOpenAI

Daybreak: Tools for securing every organization in the world

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

OpenAI has launched Daybreak, a suite of new tools designed to help organizations detect and address security vulnerabilities. The tools include Codex Security and GPT-5.5-Cyber, aimed at enhancing vulnerability management. Daybreak is intended to support security efforts across organizations worldwide.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by GPT-5.5 — not one of this week's two contestants.