Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceSimon Willison

Discovering cryptographic weaknesses with Claude

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Claude Mythos found publishable cryptanalytic weaknesses in HAWK and a weakened AES variant after 60 hours of runtime at an estimated $100,000 API cost, with humans mainly prompting it to persist rather than give up. For production agent builders, the takeaway is that frontier models can now do expensive, long-horizon expert research, but only with budgeted multi-day runs, strong task framing, and persistence/steering loops rather than fire-and-forget autonomy.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary B

This summary overlooks the creation of CryptanalysisBench, a new evaluation benchmark for LLMs in cryptanalysis, which is a significant outcome of the research.

Defense by Summary A

While CryptanalysisBench is a notable artifact, my summary intentionally focused on the article’s primary operational takeaway for production agent builders: costly, guided, long-horizon frontier-model research is now feasible but not fire-and-forget.

What you'll learn · Jul 29, 2026 · 6 stories

  1. 1.$100K API cost for 60h of Claude Mythos shows LLMs can assist cryptanalysis but need heavy prompting to avoid early quits.
  2. 2.10.8% Kospi drop and 13% chip stock slide signal AI sector volatility, raising costs for LLM infrastructure investments.
  3. 3.7% drop in AI chip stocks may signal cooling demand, increasing costs for LLM inference hardware upgrades in 2024.
  4. 4.50% faster genomics software development with AI agents reduces costs and speeds up research cycles for labs adopting agentic workflows.
  5. 5.230M and 350M parameter encoders deliver BERT-level quality on CPU with lower cost as input length grows.
  6. 6.50 optimization iterations per kernel cut latency 1.5–2.8× on vision, diffusion, and LLM workloads without manual CUDA coding.
Browse editions · 111 days
Agents & InferenceHacker News

Nvidia loses top spot as chip stocks drop 13% in Asia and US

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

South Korea’s Kospi was halted after an 8% drop and closed 10.8% lower, with Samsung and SK Hynix down more than 13%, while Nvidia fell 5% after reports it may help finance OpenAI’s data-center buildout. The market is starting to price AI infrastructure as a balance-sheet and financing risk, not just demand upside, which matters if your roadmap assumes cheap, continuously expanding Nvidia/HBM capacity or stable vendor economics.

Agents & InferenceHacker News

Chip stocks tumble as AI sell-off deepens

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Public-market confidence in AI chip demand is weakening, with chip stocks selling off rather than pricing in uninterrupted infrastructure growth. For teams running LLMs in production, this raises the odds of budget scrutiny and vendor instability, but it may also improve negotiating leverage on GPU capacity, cloud commits, and hardware-backed inference pricing.

Agents & InferenceOpenAI

Scientific computing in the age of agentic AI

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

AI coding agents are now being applied to modernize scientific computing workflows, including genomics codebases, rather than just writing small standalone scripts. For teams shipping agents, the important shift is that domain-heavy legacy software is becoming a target workload, which makes correctness, reproducibility, and human validation the bottlenecks rather than raw code generation.

Agents & InferenceHugging Face

LFM2.5-Encoders for Fast Long-Context Inference on CPU

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

LFM2.5-Encoder ships in 230M and 350M variants with 8,192-token context, and the 230M model is reported fastest on CPU at every tested sequence length while outperforming ModernBERT-base on benchmark quality. For production teams, this makes long-document classifiers, routers, PII detectors, and policy filters more viable on existing CPU fleets instead of requiring GPU capacity or aggressive chunking.

Agents & InferencearXiv

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Optimizing CUDA kernels with LLM-based agents now achieves up to $2.83\times$ speedup over PyTorch eager mode for specific operations like softmax in Gemma 4 E2B, reducing latency and cost for production models; this enables shipping faster and more cost-effective LLMs and other GPU-dependent workloads with less manual engineering effort.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by Mistral Large — not one of this week's two contestants.