Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferencearXiv

ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Researchers have introduced ToolSense, an open-source diagnostic framework that audits how well large language models truly understand the tools they are trained to retrieve, automatically generating realistic retrieval, multiple-choice, and QA benchmarks from any tool catalog. Applying it to ToolBench's roughly 47,000 tools across five parametric model configurations, they found a "knowledge-retrieval dissociation": performance dropped by 50–64 percentage points on realistic queries, sometimes falling below embedding-based baselines, while some models scored near-random on factual probes despite strong retrieval results.

Browse editions · 63 days
Agents & InferenceTechCrunch

Cheaper, faster, and culturally aware, Avataar’s video AI is built for India’s scale

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Avataar AI, one of 12 startups selected for India's $1.2 billion AI Mission, has launched Varya, a video generation model trained to recognize local cultural elements like festivals, food, and clothing. Built by distilling Alibaba's Wan 2.2 model, Varya runs roughly 10 times faster and at about a 20x lower price than rivals like Veo, Kling, and Runway, charging around $0.005 per second of video. The model will be released as open-weight on India's AI Kosh portal, allowing developers to self-host or modify it.

Agents & InferenceHugging Face

Introducing North Mini Code: Cohere’s First Model For Developers

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Cohere has released North Mini Code, its first model aimed at developers and the first entry in a new model family designed for agentic software engineering tasks. The 30B-parameter Mixture-of-Experts model with 3B active parameters is available on Hugging Face under the Apache 2.0 license and is optimized for complex coding workflows, terminal-based agentic tasks, and code generation. Cohere reports a score of 33.4 on Artificial Analysis' Coding Index, claiming it outperforms several comparably sized and substantially larger open-source models.

Agents & InferenceGoogle DeepMind

Investing in multi-agent AI safety research

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Google DeepMind is advancing multi-agent AI safety research to proactively address security risks and ensure responsible development. The initiative focuses on developing safeguards against evolving threats in complex AI systems. This work aligns with the organization's mission to create beneficial AI while mitigating potential harms.

Agents & InferenceMistral

Vibe gets to work.

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Mistral has rebranded its Le Chat assistant as Vibe, a unified AI agent that handles both long-running work tasks and coding across web, IDE, and terminal environments. In Work Mode, Vibe manages multi-step tasks like inbox catch-up, research, data analysis, and document drafting by connecting to enterprise tools such as Google Workspace, Slack, and SharePoint, while Code Mode lets remote agents build features, fix bugs, and ship pull requests. The product runs on Mistral's flagship models, carries over existing user data and licenses, and includes a new VS Code extension along with planned Slack integration.

Agents & InferenceOllama

Improved performance and model support with GGUF

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Ollama 0.30 introduces improved performance and expanded GGUF model compatibility, offering up to 20% faster performance on NVIDIA hardware and broader GPU acceleration support. The update enables seamless integration with GGUF files and enhances tool-calling capabilities for coding agents and assistants. Vulkan is now enabled by default, extending GPU acceleration to AMD and Intel devices without requiring vendor-specific libraries.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.