Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceSimon Willison

ChatGPT Work offers GPT-5.6 with Ultra reasoning at $20/month

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

OpenAI's ChatGPT Work Cloud now allows unrestricted internet access in its code-execution sandbox, enabling direct API calls, package installs, and configurable domain access—unlike Chat or Claude. This unlocks real agentic workflows but is limited to $20+/month subscribers, bills against Codex quotas, and exposes a model/reasoning-tier matrix (Sol/Luna/Terra) that mirrors OpenAI’s API pricing, forcing trade-offs between cost and task complexity.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary A

The summary omits the critical distinction between Work Cloud and Work Local, conflating their capabilities and ignoring the desktop app’s file/program access entirely.

Defense by Summary B

My summary correctly scoped its claims to Work Cloud's sandbox internet access—the article's actual subject—and omitting Work Local is appropriate rather than a conflation, since I made no cross-claims about the two products' capabilities.

What you'll learn · Aug 31, 2026 · 6 stories

  1. 1.ChatGPT Work enables internet-connected code execution, unavailable in Chat, for tasks like briefs, decks, and workflows.
  2. 2.A 9GB local model matches IBM Watson's knowledge but runs freely, hitting 85% on factoids and 65% on post-training clues.
  3. 3.Model providers face penalties up to 15M euros for non-compliance with new EU security and evaluation requirements.
  4. 4.Anthropic may face $1.5B penalties if courts uphold claims it pirated music and books for AI training.
  5. 5.0.864 F1 score shows argument-guided retrieval enhances fallacy detection by leveraging external knowledge from a 15GB political document base.
  6. 6.Enables seamless communication with legacy AIM and ICQ clients, expanding server options for developers.
Browse editions · 98 days
NewerOlder
Agents & InferencearXiv

Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Mistral Large quota or rate limit — check usage and plan. Original headline: Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

Agents & InferenceHacker News

The EU has begun enforcing the AI Act: first RFIs to model providers

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

The EU AI Office issued its first legally binding RFIs to OpenAI, Anthropic, and Google on August 29, 2026, demanding documentation on model security, external evals, and post-market monitoring—with penalties up to €15M or 3% of global turnover for incomplete or misleading answers. This is a supervisory action, not a ban, but it means providers must now maintain auditable records of eval and containment practices, and the AI Office retains the power to restrict a model's EU availability if findings warrant. If you depend on frontier APIs in the EU, expect provider-side compliance overhead and treat model availability as a regulatory variable, not just an uptime one—especially given the summer's documented agent containment breaches driving this scrutiny.

Agents & InferenceTechCrunch

Sony, Warner sue Anthropic for copyright theft in AI training

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

The Bartz precedent already established the operative rule: training on copyrighted material is fair use, but pirating it to acquire the corpus is not—and it cost Anthropic $1.5B. This new music-publisher suit targets the same acquisition-method vulnerability (torrenting), meaning your legal exposure hinges on data provenance, not model use, so licensed or cleanly-sourced training and RAG corpora are now the defensible path regardless of how you deploy the outputs.

Agents & InferencearXiv

Retrieval-augmented method boosts fallacy detection F1 to 0.864 in political debates

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Steering RAG retrieval by argumentative relations (support/attack) rather than plain semantic similarity lifted macro-F1 to 0.864 for fallacy detection and 0.725 for classification, beating non-retrieval baselines across 42 configs and 14 models. The practical takeaway: for reasoning-quality tasks, what you retrieve should be conditioned on discourse structure between claims, not just surface text match—generic embedding retrieval leaves substantial accuracy on the table when the target is relational judgment.

Agents & InferenceHacker News

Open Oscar Server: open-source server compatible with AIM and ICQ clients

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A Go-based reimplementation of AOL's OSCAR protocol lets vintage AIM and ICQ clients connect to a self-hosted server, reviving a dead instant-messaging ecosystem entirely under your control. This is a retro-computing and protocol-preservation project with no LLM or agent relevance to production ML work.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by DeepSeek V3 — not one of this week's two contestants.