Google reportedly developing ‘Frozen v2’ chip with Gemini’s architecture etched into the silicon — engineers project 6 to 10 times more tokens per watt than latest TPUs

Google is developing a server chip, informally dubbed “Frozen v2,” that would etch part of its Gemini model’s architecture directly into the silicon, according to a report published Monday by The Information, citing two people with direct knowledge of the matter. Engineers on the project have projected that the chip could serve six to ten times more tokens per unit of power than the newest generation of Google’s TPUs, with deployment targeted for as soon as 2028. The two sources said the project is partly a response to an AI compute shortage severe enough that Google Cloud has turned down deals with outside customers.

Go deeper with TH Premium: AI and data centers

Microsoft data center in Mount Pleasant, Wisconsin

(Image credit: Microsoft)

A TPU, like a GPU, runs whatever model is loaded onto it, which means the hardware makes time-consuming runtime decisions as it interacts with each one. Frozen v2 would have some of those decisions for Gemini fixed in the transistors, reducing the number of steps the chip takes and the amount of data it shuttles around per query. That could cut response latency enough to enable new applications, one of the sources said.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *