Agents & InferencearXiv

Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Gemma3:1B and Llama3.2:1B hit 0.56–0.65 J/token on a single RTX 4060Ti—4.4× more efficient than 7B-Mistral—while pushing >170 tok/s. This means you can slash cloud GPU spend or run 4× more local inference on the same power budget, but only if you swap out larger models for these smaller, quantized architectures.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →