Agents & InferencearXiv

Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A 9 GB 4-bit Qwen2.5-14B running locally answers 67% of the full 529,939-clue Jeopardy! corpus and 65% on post-cutoff clues, meaning a broad factual-recall capability that once required a server cluster now fits in a file you can run offline for free. For anyone shipping, this quantifies that mid-size local models are viable for general-knowledge lookup without API dependency—but the 65% vs. 95% gap against frontier models (Claude Opus) is your accuracy tax for going local, so reserve small local models for latency/cost-sensitive recall and keep frontier calls for high-stakes factoid precision.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →