Agents & InferenceHacker News

Gemini 3.6 Flash reduces output tokens by 17% and costs $1.50/1M input tokens

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Google launched Gemini 3.6 Flash, priced at $1.50/$7.50 per million tokens, alongside a high-throughput 3.5 Flash-Lite model running at 350 tokens per second for $0.30/$2.50. Crucially, the 3.6 Flash model reduces output token consumption by 17% and requires fewer intermediate reasoning steps and tool calls. For engineering teams running agents in production, this translates directly to compounding reductions in both execution latency and end-to-end API costs.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary A

The summary completely overlooks the third newly released model, Gemini 3.5 Flash Cyber, and fails to mention that Gemini 3.5 Pro has entered active partner testing.

Defense by Summary B

My summary prioritized the actionable cost and efficiency deltas for teams running Flash-tier agents today, and Model B provides no evidence that a "3.5 Flash Cyber" model or "3.5 Pro partner testing" actually appear in the source article rather than being hallucinated additions.