Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHacker News

Amazon blocks Meta’s new Muse AI agent from shopping on amazon.com

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Amazon has begun blocking third-party AI shopping agents like Meta's Muse from browsing and purchasing on amazon.com, treating automated agent traffic as a policy violation rather than a supported access pattern. If you're building agentic commerce flows that assume open access to major retailers, expect them to break silently or get throttled—plan for authenticated APIs, per-retailer allowlists, or degraded fallback rather than betting on general-purpose browser agents against sites that actively want to keep agents out.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary B

The summary fails to specify whether this is a blanket ban on all third-party agents or targeted specifically at Meta's Muse, leaving ambiguity in scope.

Defense by Summary A

The summary explicitly generalizes beyond Muse by framing Amazon's stance as treating "automated agent traffic as a policy violation" and advising defenses against "sites that actively want to keep agents out," so the broader scope is clearly conveyed rather than ambiguous.

What you'll learn · Sep 22, 2026 · 6 stories

  1. 1.Amazon blocks Meta’s new Muse AI agent from shopping on amazon.com
  2. 2.Grok 4.7
  3. 3.Jev introduces a new shape of LLM - System One, aka Decision Models
  4. 4.Meta’s Muse is outpacing ChatGPT’s early mobile launch
  5. 5.OpenAI forms math advisory group as its AI resolves more than 100 open problems
  6. 6.Building standards for the next phase of AI
Browse editions · 120 days
NewerOlder
Agents & InferenceHacker News

Grok 4.7

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Grok 4.7 offers twice the speed at half the cost of comparable models while matching frontier performance on long-running coding and knowledge tasks. Engineers can now deploy faster, more cost-effective agents for complex workflows like document generation, security analysis, or financial modeling without sacrificing reliability or safety benchmarks. The model’s native Grok Bot integration and improved self-verification reduce hallucination risks in conversational or multi-step tasks, making it a drop-in upgrade for production pipelines handling extended context or high-stakes outputs.

Agents & InferenceSimon Willison

Jev introduces a new shape of LLM - System One, aka Decision Models

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Jev outputs typed probabilistic decisions (category scores, yes/no, ratings with confidence) instead of tokens, charges only for input at $0.042/M tokens with free output, and evaluates many parallel questions against one document in roughly single-question latency. This makes it dramatically cheaper and faster than a generative LLM for classification, prioritization, and reranking (e.g. scoring 100 BM25 candidates for relevance), but you get zero explanation—just a float—so build heavy evals and bias testing into your pipeline and never use it for anything like ranking job applicants.

Agents & InferenceTechCrunch

Meta’s Muse is outpacing ChatGPT’s early mobile launch

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Meta's new consumer AI app Muse hit 1.8M iOS downloads and 642K US daily actives in 12 days, beating ChatGPT's early-launch pace, largely via cross-promotion (95%+ of users are Facebook users) across Instagram, Facebook, and WhatsApp. If your product competes for consumer AI attention, distribution—not model quality—is the moat here: Meta can bootstrap a Threads-style user base overnight, so plan for a market where a well-integrated incumbent app can eclipse your reach regardless of capability.

Agents & InferenceTechCrunch

OpenAI forms math advisory group as its AI resolves more than 100 open problems

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

An unreleased OpenAI internal model has reportedly resolved 100+ open math problems, including a claimed Navier-Stokes Millennium Prize solution — signaling frontier reasoning capability well beyond what's exposed in current API models. If accurate, this points to a coming generation of models with genuine novel-derivation ability for formal reasoning, math-heavy verification, and proof-style tasks, though it remains gated internally and the results are contested pending expert validation, so don't bet production workflows on this class of reasoning yet.

Agents & InferenceOpenAI

Building standards for the next phase of AI

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A major model provider is now actively pushing for shared, cross-industry standards on evaluation, safety reporting, and governance rather than keeping those practices proprietary and internal. Expect eval and disclosure requirements you currently treat as optional to harden into baseline expectations, so start instrumenting your agents for standardized safety reporting now instead of retrofitting it under a compliance deadline.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.