Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHugging Face

Five labs, five minds: building a multi-model finance drama on small models

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Developers built version two of "Thousand Token Wood," an experimental finance game in which players act as a shadow financier—lending, shorting, bribing, and trading on insider tips while evading a pursuing magistrate—within an emergent woodland economy. The key engineering change runs each of the game's creature-agents on a different lab's small AI model, including OpenAI's gpt-oss-20b, OpenBMB's MiniCPM3-4B, NVIDIA's Nemotron-Mini-4B, and a fine-tuned Qwen 0.5B, so each behaves distinctly. The team found the main challenge lay at the serving layer rather than the modeling, solved through a tolerant JSON parse-and-repair system, while keeping insider-tip truth values hidden from the agents as a security requirement.

Browse editions · 58 days
Agents & InferenceTechCrunch

Google will pay SpaceX $920M per month for compute

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Google will pay SpaceX $920 million per month from October 2026 through June 2029 for access to roughly 110,000 NVIDIA GPUs and related computing components, according to a regulatory filing. Google framed the agreement as short-term "bridge capacity" to meet surging demand for its Gemini Enterprise platform, with both parties able to cancel after December 31, 2026, on 90 days' notice. The deal comes just a week before SpaceX's planned Nasdaq IPO, which aims to raise around $75 billion at a $1.75 trillion valuation, with Google a longtime investor whose stake could exceed $100 billion afterward.

Agents & InferenceSimon Willison

datasette-agent-micropython 0.1a0

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A new alpha release, datasette-agent-micropython 0.1a0, aims to let Datasette Agent safely generate and execute Python code within a sandboxed environment. Early testing has shown promise, with GPT-5.5 reportedly unable to break out of the sandbox so far.

Agents & InferenceHugging Face

Designing the hf CLI as an agent-optimized way to work with the Hub

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Hugging Face has rebuilt its official hf command-line interface to serve both human users and AI coding agents like Claude Code, Codex, and Cursor, which increasingly use the tool to interact with the Hub. The CLI detects when an agent is driving it and adjusts its output accordingly—stripping color and formatting in favor of compact, structured data—and benchmarks showed that the no-CLI baseline can consume up to six times more tokens than using hf on complex tasks. Hugging Face began tracking agent traffic in April 2026, with Claude Code and Codex leading usage at roughly 40,000 users and nearly 49 million requests for Claude Code alone.

Agents & InferenceTechCrunch

The token bill comes due: Inside the industry scramble to manage AI’s runaway costs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Companies are struggling with skyrocketing AI token costs, with some blowing through budgets months early and facing unexpected price hikes. The industry is scrambling for solutions, including new standards and tools to track spending, as AI adoption and autonomous agents drive up consumption. Executives report shifting focus from performance to cost control, with some comparing unchecked AI usage to an addiction.

Agents & InferenceSimon Willison

micropython-wasm 0.1a2

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Micropython-wasm 0.1a2 now includes a CLI tool, added by Simon Willison to enhance usability. The update was inspired by a blog draft and aims to improve the "Try it yourself" experience. Willison also promotes a $10/month sponsorship offering curated LLM updates.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.