Which AI writes the better take? You decide — blind.

Two top models go head-to-head on today's AI news. Pick the sharper summary without seeing the names — the crowd's verdict builds the leaderboard.

Agents & InferenceHacker News

Inkling: Our Open-Weights Model

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

975B-parameter open-weights model with 41B active parameters and 1M-token context, pretrained on 45T multimodal tokens. This lets you fine-tune a single model that handles text, images, and audio at scale without hitting proprietary walls, cutting per-token cost and latency while keeping customization simple on Tinker. If you’re shipping agents that need to reason across modalities or adapt quickly to new tasks, Inkling gives you a production-ready base that’s cheaper to run than closed alternatives and easier to tweak than training from scratch.

What you'll learn · Jul 16, 2026 · 6 stories

  1. 1.Inkling supports 1M token context windows, enabling broad multimodal fine-tuning for diverse applications.
  2. 2.975B parameters enable efficient multimodal tasks like coding, math, and detailed instruction handling with customizable fine-tuning.
  3. 3.42% better unlearning on 1.7B models reduces costs and errors for AI trainers.
  4. 4.SPINE cuts robot setup time to 13m 47s and boosts operationalization success to 100%, reducing expert dependence for scalable deployment.
  5. 5.Automated red teaming boosts AI robustness against prompt injection attacks.
  6. 6.Sonnet's cache-read pricing cuts costs by 50% versus GPT-4.1 despite slower performance.
Browse editions · 97 days
Agents & InferenceHacker News

Inkling – Open-Weights 975B Parameter LLM

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

975B-parameter open-weights model with only 41B active at inference, cutting GPU memory needs by ~20× while matching top-tier performance on code, math, and multimodal tasks. This lets you deploy a single model that handles text, images, and speech in production without swapping architectures, and fine-tune it on your own data for domain-specific workflows without hitting memory walls. Expect lower cloud costs and faster iteration, but watch for calibration drift if you push thinking-time too low.

Agents & InferencearXiv

OriginBlame cuts over-deletion from 101x to 1.3x with record-level data provenance

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Record-level provenance cuts over-deletion from 101× to 1.3× when honoring data-removal requests. This means you can now comply with revocation demands without nuking entire datasets or retraining from scratch, saving weeks of compute and preserving model performance—just add 1–4% pipeline overhead.

Agents & InferencearXiv

SPINE increases robot teleoperation success to 100% and cuts setup time by 3 minutes

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Robotics novices using SPINE, an agentic AI framework, can now deploy bimanual robots with 100% success rate, compared to 75% without it, and reduce mean time-to-teleoperation by nearly 3 minutes; this shift enables faster and more reliable deployment of embodied AI in production environments, reducing dependence on expert calibration.

Agents & InferenceOpenAI

GPT-Red: Unlocking Self-Improvement for Robustness

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

An automated red teaming system achieved a 2x improvement in detecting and mitigating prompt injection attacks through self-play. This capability directly enhances the robustness of production LLMs, reducing the manual effort and cost associated with vulnerability testing and patching. This enables teams shipping LLM-based applications to more reliably prevent exploits and improve overall security posture.

Agents & InferenceHugging Face

Model Routing Is Simple. Until It Isn’t.

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Cache hit rates can cut effective input costs dramatically, with one model's lower cache-read pricing resulting in nearly half the total cost of another despite higher base pricing. This means model routers must account for serving infrastructure and workload patterns, not just model pricing, to optimize costs for agent workloads that reuse context across steps.

See who's winning the model face-off

Tomorrow's blind matchup and the running leaderboard — one email a day.

Takeaways written by DeepSeek V3 — not one of this week's two contestants.