Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHacker News

ZCode 3.0 optimizes GLM-5.2 for multi-agent collaboration

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

ZCode is an agentic coding harness tuned for GLM-5.2, positioning itself as a Claude Code/Cursor-style multi-agent workspace where task-scoped agents plan, edit files, and run shell/git commands autonomously (the demo shows an agent bootstrapping a full Gomoku app from an empty repo in ~3 minutes). If you're evaluating GLM-5.2 as a cheaper alternative to Anthropic/OpenAI models for coding agents, this gives you a ready-made loop rather than building tool orchestration yourself—worth benchmarking against your current harness on real repos, since the value hinges on whether GLM-5.2's tool-use reliability holds up beyond greenfield demos.

What you'll learn · Jul 2, 2026 · 6 stories

  1. 1.GLM-5.2 gains faster multi-agent workflows in ZCode 3.0; teams can now plan, code, and deploy with lower coordination overhead.
  2. 2.GLM-5.2 integration in ZCode 3.0 speeds up multi-agent coding workflows for planning, review, and deployment without toolchain changes.
  3. 3.31B Gemma 4 on Cerebras hardware reduces P95 latency to under 300ms, enabling real-time voice agents in 9,000+ robots.
  4. 4.Sonnet 5 cuts inference costs by ~20% vs Opus 4.8 while matching accuracy, but English tokens cost 1.4x more than Sonnet 3.5.
  5. 5.Default blocks on ad-heavy pages may cut AI training data by half unless crawlers split search from agent use.
  6. 6.Gemini Spark beta on Mac lets Ultra subscribers automate file tasks and app workflows, competing with Claude and Copilot at no added cost.
Browse editions · 88 days
Agents & InferenceHacker News

ZCode 3.0 adds multi-agent support for GLM-5.2 coding tasks

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Zhipu's GLM team shipped a Claude Code clone (ZCode) with native desktop clients across macOS/Windows/Linux, tuned specifically for GLM-5.2 with multi-agent orchestration, 20+ tool integrations, and remote task triggering via WeChat/Feishu/Telegram. If you're paying Anthropic per-token for agentic coding, this is a drop-in alternative on a cheaper China-hosted model stack—worth evaluating for cost, but weigh the data-residency and latency implications of routing your codebase through z.ai infrastructure.

Agents & InferenceHugging Face

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A fully open cascaded speech-to-speech stack—Parakeet STT → Gemma 4 31B on Cerebras → Qwen3TTS—now runs fast enough to hit conversational latency, with the key win being tail stability (P95) rather than just median, which is what actually kills voice UX. If you're building voice agents or embodied AI, this is a swappable, self-hostable alternative to closed real-time APIs, and the same pipeline is already in production on 9,000+ Reachy Mini robots. The bet is that Cerebras's inference speed makes multi-turn tool-calling and multimodal steps feel real-time, so the constraint you're designing around shifts from LLM response time to your STT/TTS choices.

Agents & InferenceSimon Willison

What's new in Claude Sonnet 5

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Claude Sonnet 5’s new tokenizer can make the same English documents about 1.4x more tokens, with smaller increases for Spanish and Python and little change for Simplified Mandarin. Even if the model is priced below Opus 4.8 while offering similar performance, production cost and context-budget assumptions need to be remeasured on your actual payloads before switching.

Agents & InferenceTechCrunch

Cloudflare’s new policy pushes AI companies to pay for publishers’ content

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Starting September 15, 2026, Cloudflare will block by default any mixed-use crawler—one that blends search, training, and agent fetching—from ad-hosting pages, hitting all free customers and new sites unless owners opt out. If your crawlers don't cleanly separate search from training/agent use, you risk losing default access to a large slice of the web, and Cloudflare's Pay Per Use model means you may soon be billed when content drives value, not just when fetched—so plan for segmented, transparent bot identities and a metered content-cost line item.

Agents & InferenceTechCrunch

Gemini Spark, Google’s agentic assistant, is now available on Mac

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Gemini Spark is now a macOS desktop agent in beta, but only for U.S. Google AI Ultra subscribers, and it can already read/use local Mac files to organize them or generate Google Workspace docs and spreadsheets. For teams shipping agent workflows, this makes Google’s stack a more direct Claude Desktop/Copilot competitor, with the key implication that desktop file permissions, Workspace data access, and upcoming MCP/app connectors need to be treated as production security and governance surfaces rather than chat-only integrations.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by Mistral Large — not one of this week's two contestants.