Google is working on a new AI chip designed to make Gemini more efficient
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
Google is developing an in-house server chip dubbed Frozen v2 for 2028 that targets a six-to-tenfold increase in token generation efficiency per unit of power over its current hardware. This projected leap will drastically drive down the API and hosting costs of running high-throughput Gemini models in production, making continuous agentic workflows financially viable. However, capitalizing on these margins will lock your pipelines even deeper into Google’s proprietary infrastructure as the industry shifts toward tightly integrated hardware-software co-design.
Google's next-gen TPU ("Frozen v2") reportedly targets 6–10x better tokens-per-watt than current chips, but it's a 2028 release, so nothing changes your capacity or cost math for the next two years. It reinforces the broader trend—OpenAI's Jalapeño, Anthropic-Samsung—of frontier providers moving inference onto custom silicon, which over time means the cheapest Gemini/GPT inference will be locked to each vendor's cloud, further eroding Nvidia-based portability and any hope of running these models on neutral hardware.