Alphabet-owned Google is reportedly developing a new custom Google AI chip for its servers, internally codenamed “Frozen v2,” that could make its Gemini models six to ten times more power-efficient than the company’s latest Tensor Processing Units (TPUs). Rather than replacing Google’s existing TPU lineup, the chip is expected to form a new class of highly specialized AI accelerators designed specifically for Gemini inference by embedding parts of the model’s architecture directly into the hardware.
The project reflects Google’s growing focus on optimizing AI infrastructure as demand for Gemini-powered services continues to surge. Reports suggest the company is facing internal compute constraints that have, at times, forced Google Cloud to decline external AI infrastructure deals. By creating hardware tailored to Gemini, Google hopes to dramatically improve inference efficiency while lowering energy consumption and operating costs.
What Is “Frozen v2”?
Unlike general-purpose AI accelerators that are designed to run a wide variety of models, Frozen v2 is reportedly being engineered specifically around the Gemini family of AI models.
The chip would:
- Hardwire portions of Gemini’s architecture into silicon.
- Reduce unnecessary computation during inference.
- Increase the number of AI tokens generated per watt of power.
- Lower latency for Gemini-powered applications.
This approach resembles the concept of model-specific AI hardware, where the processor is optimized for a particular neural network architecture instead of supporting a broad range of models.
Frozen v2 at a Glance
| Feature | Details |
|---|---|
| Chip codename | Frozen v2 |
| Company | Google (Alphabet) |
| Primary purpose | Run Gemini AI models more efficiently |
| Expected efficiency gain | 6–10× more tokens per watt than latest TPUs |
| Deployment target | As early as 2028 (reported) |
| Status | Under development |
Why Google Is Building a Gemini-Specific Chip
AI inference has become one of the largest operating expenses for companies deploying large language models at scale.
Every response generated by Gemini consumes:
- Compute resources.
- Memory bandwidth.
- Electricity.
- Data center capacity.
By embedding parts of Gemini directly into the hardware, Google could eliminate redundant processing steps, allowing the chip to generate more AI output while consuming significantly less power. This would improve the economics of serving billions of AI queries across products such as Search, Workspace, Android, YouTube, and Google Cloud — a scale that matters as Gemini gains ground on ChatGPT in everyday usage.
Potential Benefits
| Advantage | Impact |
|---|---|
| Higher power efficiency | Lower electricity consumption |
| Faster inference | Reduced response times |
| Lower operating costs | Cheaper AI deployment |
| Greater compute capacity | More users served with existing infrastructure |
Complementing, Not Replacing, Google’s TPUs
Reports indicate that Frozen v2 is not intended to replace Google’s Tensor Processing Units (TPUs).
Instead, Google is expected to maintain two hardware strategies:
- TPUs for flexible AI training and diverse inference workloads.
- Frozen v2 for highly optimized Gemini inference.
This dual-chip approach would allow Google to continue training future Gemini models on programmable TPUs while deploying specialized Frozen chips to serve mature production models at much lower cost.
TPU vs Frozen v2
| TPU | Frozen v2 |
|---|---|
| General-purpose AI accelerator | Gemini-specific AI accelerator |
| Supports multiple AI workloads | Optimized primarily for Gemini |
| Flexible architecture | Hardware tailored to model architecture |
| Training and inference | Primarily inference |
Addressing Google’s AI Compute Crunch
The reported development comes as Google faces rapidly growing demand for AI infrastructure.
According to reports, shortages in AI computing capacity have created internal tensions and even limited Google Cloud’s ability to accept certain external customer workloads. Frozen v2 is intended to alleviate these bottlenecks by dramatically increasing inference efficiency without requiring proportional growth in data center infrastructure.
Improving inference efficiency could also reduce Google’s dependence on third-party AI hardware while maximizing the value of its in-house silicon strategy. The same cost pressure is pushing rivals in a different direction: Microsoft is testing cheaper open-weight models for Copilot instead of building model-specific silicon.
Risks of Specialized AI Hardware
Although specialized chips can deliver significant performance gains, they also carry important trade-offs.
Potential challenges include:
- Gemini’s architecture may evolve before the chip reaches production.
- Less flexibility than general-purpose accelerators.
- High design and manufacturing costs.
- Long development timelines.
Because Frozen v2 reportedly targets deployment around 2028, engineers must ensure that the hardware remains compatible with future Gemini model architectures.
Opportunities and Risks
| Opportunities | Risks |
|---|---|
| Lower AI serving costs | AI models may evolve rapidly |
| Better energy efficiency | Specialized hardware is less flexible |
| Reduced latency | Expensive chip development |
| Increased infrastructure capacity | Longer deployment timelines |
Why This Matters for India
India is among the largest user bases for Google Search, Android, YouTube and Workspace, so a large share of Gemini inference ultimately serves Indian queries. Cheaper tokens per watt translate directly into cheaper AI features for Indian consumers and lower cloud bills for Indian startups and IT services firms that build on Google Cloud. Power efficiency also matters locally because data centre capacity in India is constrained by electricity availability and cooling costs.
There is a supply-chain angle too. Custom AI silicon depends on advanced foundry capacity and export rules that are increasingly politicised, as seen in Beijing’s proposed export controls on AI models and chips. Google has not announced any India-specific deployment plans for Frozen v2, and the chip remains an unconfirmed internal project rather than a launched product.
Looking Ahead
Google’s reported Frozen v2 project highlights the next phase of the AI infrastructure race, where companies are moving beyond designing general-purpose AI accelerators toward building chips optimized for specific foundation models. If Frozen v2 delivers its reported six-to-tenfold improvement in tokens per watt, it could significantly reduce the cost of running Gemini while expanding Google’s AI capacity across consumer products and Google Cloud.
Although the project remains under development and is reportedly targeted for deployment as early as 2028, it underscores Google’s long-term commitment to vertically integrating AI software and hardware. As competition intensifies among Google, OpenAI, Microsoft, Meta, and Anthropic, specialized AI silicon may become an increasingly important differentiator in delivering faster, more efficient, and more cost-effective generative AI services.
Frequently Asked Questions
What is Google’s Frozen v2 AI chip?
Frozen v2 is the internal codename for a custom Google AI chip reportedly designed to run Gemini models. It hardwires parts of Gemini’s architecture into silicon so inference uses far less computation and power than a general-purpose accelerator.
What is a Google TPU and how is it different?
A TPU, or Tensor Processing Unit, is Google’s general-purpose AI accelerator used for both training and inference across many models. Frozen v2 would be narrower by design — optimised primarily for Gemini inference rather than flexible across workloads.
When will the Frozen v2 chip be deployed?
Reports point to deployment as early as 2028, but the chip is still under development and Google has not officially confirmed the project, its specifications, or any launch timeline. Those details should be treated as unconfirmed.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



