Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHacker News

MathCode cuts Lean compile checks to ~0.4s after warmup

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

MathCode is a terminal-based AI agent that converts plain-language math problems into Lean 4 formal proofs with a persistent REPL, cutting compile-check latency from ~30s to ~0.4s after warmup. This enables real-time, agentic theorem proving with reusable libraries, parallel subgoal decomposition, and Obsidian knowledge graphs, making formal math verification practical for interactive and production use.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary A

“The summary overstates 'instant self-correction' and omits the macOS/Linux dependency, bundled runtime overhead, and the requirement for manual CLI setup.”

Defense by Summary B

“The summary appropriately focuses on the core architectural breakthrough of a 98% latency reduction that enables practically instantaneous agentic self-correction, while omitting minor platform-specific dependencies and standard setup configurations that do not alter the system's primary technical contribution.”

What you'll learn · Aug 17, 2026 · 5 stories

  1. 1.0.4s compile checks after warmup can speed agentic Lean proof repair versus ~30s, with theorem reuse, parallel proving, and macOS arm64/Linux x86_64 requirements.
  2. 2.5.1x application speedup shows validation-centric agents can port large Fortran HPC codes, but teams must budget for context management and discrepancy recovery.
  3. 3.22,276 reasoning tokens for one SVG shows the xhigh default can exhaust LM Studio’s 8,192-token context and add long waits on consumer hardware.
  4. 4.0.115 vs. 0.173 false-pass rates mean rubric-induced judges catch more broken agents, while aggregate agreement gains over G-Eval were not significant.
  5. 5.Decades of distrust mean AI companies must deliver promised public benefits, not just positive messaging, to shift negative views and regulatory pressure.
Browse editions · 129 days
Agents & InferenceHacker News

AI-Assisted GPU Porting of a 250k Line Legacy Weather Simulation Code

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

5.1x speedup on a 250k-line legacy Fortran weather code ported to GPUs with AI agents—without breaking scientific validity. This means you can now offload large, trusted HPC workloads to GPUs at scale, but only if your workflow includes dump-based validation and session-spanning context to catch floating-point drift and branch divergence. Expect 1–2 orders of magnitude faster iteration on GPU ports, but plan for runtime-state reconstruction and costly rollbacks when static analysis misses edge cases.

Agents & InferenceSimon Willison

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

The newly released Qwen 3.8 27B model defaults to an "extra high" reasoning setting that can consume over 22,000 reasoning tokens for simple prompts, easily exceeding standard 8,192-token context limits and spiking generation times to over 20 minutes. If you deploy this model in production pipelines, you must explicitly override the default reasoning effort parameter or expand your context window to avoid catastrophic latency bottlenecks and immediate runner failures on routine tasks.

Agents & InferencearXiv

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

RubricForge reduces the false-pass rate of LLM-as-a-judge agent evaluations by roughly half—down to 11.5% from 17.3% compared to standard G-Eval—by automatically evolving a frozen, human-readable text rubric against a small set of ground-truth trajectories. For production, this allows you to replace slow, expensive environment-based evaluations with a reliable, single-call offline judge that stops falsely approving failed agent runs just because the generated text looks fluent. Because the optimized rubric is plain text, every grading verdict is fully explainable and attributable to named criteria without the need to fine-tune model weights.

Agents & InferenceTechCrunch

Anthropic CEO says AI backlash is ‘fundamentally a crisis of trust’

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Trust in AI companies has collapsed—public sentiment is now net-negative, and skepticism is driving regulatory crackdowns on data centers and model deployments. This means every production rollout will face longer approval cycles, higher compliance costs, and the risk of sudden policy shifts that can break your roadmap or force costly re-architecting. The only way to regain enough trust to ship at scale is to deliver measurable, near-term value (e.g., 10x cheaper inference, verifiable safety guarantees) instead of vague promises.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by GPT-5.5 — not one of this week's two contestants.