Google is developing “Frozen v2,” an AI chip designed to hardwire Gemini models into silicon, delivering up to 10x power efficiency by 2028.
In a bold strategic bid to overcome severe data center capacity bottlenecks and slash astronomical hardware power demands, Alphabet has set its sights on a fundamental rethink of silicon architecture. Officially reported by technology media outlets on Monday, July 20, 2026, Google is actively developing a custom server chip designed to hardwire core components of its flagship Gemini artificial intelligence models directly into physical circuitry. Internally codenamed “Frozen v2,” the proprietary accelerator aims to dramatically alter the cost-performance equation of running large language models at scale. By deeply integrating model architecture into the silicon layer itself rather than relying entirely on general-purpose matrix math units, Google engineers estimate the new processor could generate between six and ten times more AI tokens per unit of electricity compared to the company’s existing custom chips.
The technical breakthrough that makes Frozen v2 distinct from conventional processors lies in its co-design methodology. Traditional Tensor Processing Units (TPUs) and graphics processing units (GPUs) act as flexible mathematical engines, executing software code layer by layer. In contrast, Frozen v2 hardwires static parameters and structural weights of the Gemini architecture directly onto the chip die. While this reduces software flexibility, it unlocks unprecedented energy savings during inference operations, the process of serving real-time AI responses to hundreds of millions of search users and enterprise clients. In a statement provided to TechCrunch, a Google spokesperson emphasized that while not every experimental project enters mass production, co-designing hardware and software from the ground up ensures systems are tightly optimized for demanding, real-world workloads.
From an operational standpoint, these chips will be deployed across Google Cloud’s hyper-scale data center footprint across North America, Europe, and Asia. Rather than replacing Google’s flagship family of TPUs, the Frozen project is structured as a complementary tier of specialized inference accelerators operating alongside existing clusters. This dedicated silicon layer is designed to handle high-volume background tasks and conversational traffic, freeing up traditional TPU nodes for intensive model pre-training.
Regarding the strategic timeline of when this technology will hit live production, engineering sources indicate that Google plans to deploy Frozen v2 across its data center network as early as 2028. Design teams are currently finalizing the chip’s internal logic gates and determining exactly what percentage of Gemini model weights should be permanently burned into the silicon. The project comes at a delicate moment for Google’s broader AI roadmap, arriving shortly after reports that the public release of its Gemini 3.5 Pro model experienced minor scheduling adjustments while engineers refined its complex coding capabilities.
See Also: Head of U.S. AI Safety Agency Resigns After Three Months
The underlying why driving this multi-billion-dollar hardware effort reflects both compute scarcity and intense financial pressures. A severe shortage of AI computing capacity has created internal friction at Google Cloud, occasionally forcing the division to turn down lucrative contracts with enterprise customers due to server availability constraints. With Alphabet forecasting annual capital expenditures between $180 billion and $190 billion to expand its technical infrastructure, dramatically improving token efficiency is essential to protecting profit margins. Furthermore, by building custom inference hardware, Google joins rivals like OpenAI, which unveiled its custom “Jalapeño” inference chip in June, in aggressively seeking to reduce reliance on third-party silicon suppliers like Nvidia.

