Agents & InferenceSimon Willison

Kimi K3, and what we can still learn from the pelican benchmark

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

2.8T-parameter Kimi K3 is now the largest open-weight model (weights by July 2026) and the first to hit $3/$15 per million input/output tokens. This sets a new floor for inference costs and GPU memory requirements—expect 16×A100s or 8×H100s just to load it, doubling your serving cluster size and cutting batch sizes in half, which will break any autoscaling rules tuned for 1T-class models.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →