Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHacker News

The new CC, an AI agent built for families

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

The article provides CSS code snippets for styling a web UI, likely for a family-oriented AI agent called CC, focusing on animations and visual transitions for cards, borders, and gradients. This indicates the development of a visually polished user experience tailored for family-centered interactions, though no specific product launch details are mentioned. The practical consequence is that teams working on similar interfaces may need to adapt to these design trends for competitive family-focused applications.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary A

The summary incorrectly identifies Google's Gemini as the focus of the article, when the CSS snippets suggest a different product called CC is being developed, and it entirely misses the technical content about UI animations.

Defense by Summary B

The article's substantive content concerns Google's family-oriented Gemini agent positioning rather than incidental CSS styling artifacts, which are boilerplate UI code rather than the article's actual subject matter.

What you'll learn · Sep 23, 2026 · 6 stories

  1. 1.The new CC, an AI agent built for families
  2. 2.Unreal Agent
  3. 3.Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
  4. 4.Introducing GPT-6 Sol and Luna
  5. 5.Parallel cut research time and cost in half with GPT‑6 Astra
  6. 6.OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes
Browse editions · 121 days
NewerOlder
Agents & InferenceHacker News

Unreal Agent

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

An asynchronous tool-call harness cuts agent operating costs up to 40% versus Codex (20% vs Pi) with no measured performance loss, by logging tool calls as "in-progress" immediately and firing the LLM only when results return—so the model never burns tokens polling, waiting, or managing heartbeats. This means agents can run long tasks (multi-minute env setup) in parallel with exploration, accept user steering mid-flight without blocking, and shift security/approvals out of fragile harness hooks into deterministic sandbox constraints. The catch: it's a proprietary harness you'd adopt rather than an open SDK, so evaluate the cache-preservation and lifecycle claims against your existing provider integrations before committing.

Agents & InferenceSimon Willison

Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

GPT-6 Luna just dropped to $0.10/$0.50 per million tokens—half its predecessor and roughly one-tenth of Claude Haiku 4.5—making it one of the cheapest capable models OpenAI has ever shipped, while Opus 5.5 cut prices 20% ($4/$20) and slashed cache reads 60%. If you're running high-volume app backends or long agentic loops, re-benchmark now: the cheap tier just got dramatically cheaper, cached-context agents get materially cheaper on Anthropic, and models like GPT-5.6 Terra no longer have any pricing rationale.

Agents & InferenceOpenAI

Introducing GPT-6 Sol and Luna

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

GPT-6 now ships in two tiers: Sol for maximum capability and Luna as a cheaper, lighter variant, so you can route by cost/quality per call rather than paying frontier rates on every request. Build a routing layer now — send bulk and latency-tolerant traffic to Luna and reserve Sol for hard reasoning tasks, and re-benchmark your prompts against both since behavior and token economics will differ between the two.

Agents & InferenceOpenAI

Parallel cut research time and cost in half with GPT‑6 Astra

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

GPT-6 Astra enabled Parallel to reduce research time and costs by 50%, making large-scale data synthesis significantly faster and cheaper for labor-market analysis. This efficiency directly allows production teams to deploy more agents or scale operations without increasing budgets or timelines.

Agents & InferenceTechCrunch

OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

GPT-6 Sol and Luna ship at half the API cost of the 5.6 series, with Sol cutting factual errors roughly in half to reach Astra-tier reliability at a fraction of the price. That collapses the cost/quality tradeoff on your coding and high-volume clerical pipelines—rerun your model routing and eval benchmarks now, since the mid-tier can likely replace flagship-tier calls for many workloads and Anthropic's Opus refresh means you should benchmark both before committing.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.