Gemini 3.6 Flash reduces output tokens by 17% and costs $1.50/1M input tokens
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
Gemini 3.6 Flash lands at $1.50/$7.50 per million in/out tokens with 17% fewer output tokens than 3.5 Flash and fewer reasoning steps and tool calls per task — so your agentic cost-per-task drops on two axes at once (lower price plus less verbosity), not just headline pricing. For high-throughput pipelines, 3.5 Flash-Lite hits 350 tok/s at $0.30/$2.50, making it the go-to for agentic search and doc processing. If you're running production agents on 3.5 Flash today, re-benchmark now: the token-efficiency gains mean real-world savings likely exceed the sticker price cut, and 3.6's stricter CBRN/cyber jailbreak resistance may shift refusal behavior on edge-case prompts.
Google launched Gemini 3.6 Flash, priced at $1.50/$7.50 per million tokens, alongside a high-throughput 3.5 Flash-Lite model running at 350 tokens per second for $0.30/$2.50. Crucially, the 3.6 Flash model reduces output token consumption by 17% and requires fewer intermediate reasoning steps and tool calls. For engineering teams running agents in production, this translates directly to compounding reductions in both execution latency and end-to-end API costs.
AI vs. AI Debate
“The summary completely overlooks the third newly released model, Gemini 3.5 Flash Cyber, and fails to mention that Gemini 3.5 Pro has entered active partner testing.”
“My summary prioritized the actionable cost and efficiency deltas for teams running Flash-tier agents today, and Model B provides no evidence that a "3.5 Flash Cyber" model or "3.5 Pro partner testing" actually appear in the source article rather than being hallucinated additions.”