First reported by The Information, the unannounced chip would reportedly hardwire parts of Gemini’s architecture while leaving its weights updatable. That compromise could give Google much of the efficiency of model-specific silicon without making the hardware obsolete every time Gemini changes.
A spokesperson for Google tells The New Stack, "Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers. While not every project moves into production, this rigorous exploration is central to our full stack approach. By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads."
According to the reporting, Google hopes the chip will help relieve the AI compute crunch that’s made it harder for cloud providers to keep up with demand, while also making Gemini much cheaper and more efficient to serve. Internal projections reportedly estimate the design could deliver six to ten times more tokens per watt than Google’s current generation of AI chips.
If the project moves forward, it would represent a different approach to AI infrastructure because Google would be building hardware specifically for Gemini. For developers, that’s an early indication that future AI systems may be designed with much tighter integration between the model and the hardware beneath it.
Specialized Silicon Replaces Flexibility
Currently, the vast majority of AI inferences run on Nvidia GPUs or Google’s own Tensor Processing Units (TPUs). Because these chips are intended to accommodate a wide variety of AI models, they carry a high degree of processing load. This inherent flexibility, while beneficial for training diverse models, introduces overhead and limits peak efficiency when deployed for inference on a specific, optimized model.
There’s precedent for this kind of shift in the technological landscape. The cryptocurrency sector, for instance, followed a similar trajectory. Initial Bitcoin mining operations relied on general-purpose Central Processing Units (CPUs). As the computational demands increased, miners transitioned to Graphics Processing Units (GPUs), which offered greater parallelism. Ultimately, the pursuit of maximum efficiency led to the development of Application-Specific Integrated Circuits (ASICs), hardware designed exclusively for Bitcoin mining. AI inference may be headed in a comparable direction. While training still benefits from the flexibility of GPUs to explore and refine various model architectures, once a model reaches production and its architecture is finalized, the priority shifts dramatically. The objective becomes serving as many inference requests as possible with the lowest possible power consumption and latency.
The economic implications of this shift are significant. The cost of running AI models, particularly large language models like Gemini, is a substantial operational expense for cloud providers and enterprises. Estimates suggest that inference costs can account for a significant portion, potentially over 80%, of the total cost of operating AI models in production environments. By developing specialized hardware, companies like Google aim to drastically reduce these per-inference costs, making AI services more accessible and scalable.
Competitors Hardwire Their Own
Google isn’t the only entity looking beyond general-purpose GPUs for AI inference. As inference becomes a larger share of AI workloads—often comprising the bulk of operational costs—more companies are experimenting with specialized hardware designed to improve performance while consuming less power. This trend is not confined to startups; established players are also making significant investments.
Even Nvidia, whose GPUs currently dominate the AI market, has recognized the growing importance of inference optimization. The company has reportedly invested heavily in this area, striking a significant deal with Groq, a startup specializing in AI inference acceleration, to license its technology. This move by Nvidia signals an acknowledgment that while their general-purpose GPUs are powerful, dedicated inference solutions offer a distinct advantage in terms of efficiency and cost-effectiveness for specific workloads.
Among the more ambitious efforts in the specialized silicon space is Canadian startup Taalas. This company has demonstrated a chip that embeds an entire 8-billion-parameter Llama model directly into the silicon. By keeping the model’s parameters on the chip itself, rather than requiring constant data transfer between processing units and external memory, Taalas claims it can dramatically speed up inference. Their reported throughput of roughly 17,000 tokens per second represents a significant leap in efficiency compared to conventional hardware.
Meanwhile, other companies are focusing on optimizing the spatial relationship between memory and compute to circumvent issues like hardware lock-in and memory bottlenecks. For example, d-Matrix’s new Corsair platform utilizes an SRAM-based in-memory compute architecture. This approach eschews reliance on standard High-Bandwidth Memory (HBM) packaging, which can become a bottleneck for data-intensive AI workloads. By performing computations closer to or directly within the memory, d-Matrix aims to reduce data movement and improve energy efficiency.
Similarly, SambaNova is deploying custom dataflow technology featuring a three-tier memory architecture. This hierarchical memory system is designed to maximize tokens per watt for advanced AI workflows by intelligently managing data flow and access patterns.
Google’s reported “Frozen v2” project appears to occupy a unique position within this evolving landscape. It seems to borrow the hyper-efficiency characteristic of hardwired architectures, akin to Taalas’s approach, but crucially retains a degree of flexibility. This flexibility is designed to ensure the hardware remains viable across multiple product cycles, a critical consideration for long-term investment in AI infrastructure.
Freezing Architecture, Not Weights
The name “Frozen” reportedly stems from the concept of permanently etching specific components of Gemini’s design into the chip’s silicon fabric. According to The Information, Google’s engineers have dedicated years to navigating the complex trade-off between achieving peak efficiency and maintaining essential flexibility in AI hardware.
An earlier conceptualization, reportedly spearheaded by Jeff Dean, Google DeepMind’s Chief Scientist, explored embedding Gemini’s model weights directly into the silicon. This approach, while promising extreme speed, was ultimately deemed impractical. The primary drawback was that it would have inextricably tied the hardware to a single, static version of the Gemini model. As AI models are continuously updated and refined, this would have dramatically limited the useful lifespan of the hardware, rendering it obsolete with each significant model iteration.
The “Frozen v2” approach, as reported, represents a strategic pivot. Instead of locking in the model’s weights—the learned parameters that define the model’s behavior—this new chip design would hardwire parts of Gemini’s underlying architectural components. These architectural elements are more stable and evolve at a slower pace than the model weights themselves. Crucially, this design would still allow the model weights to be updated over time. This means that Google can continue to enhance and improve the Gemini model with new data and algorithmic advancements without necessitating a complete hardware replacement each time a new version of Gemini is released. This hybrid approach aims to capture the performance benefits of hardware specialization without sacrificing the adaptability required for a rapidly evolving AI landscape.
Cheaper Inference Reaches Developers
By integrating a significant portion of Gemini’s execution pipeline directly into the specialized silicon, Google aims to drastically reduce the overhead associated with running complex models on more general-purpose hardware. This optimization can translate into tangible benefits for end-users and developers, most notably in terms of reduced latency. For applications that demand near real-time responses, such as interactive chatbots, real-time translation services, or dynamic content generation, lower latency is paramount.
Google’s reported internal projection of delivering six to ten times more tokens per watt directly addresses the core challenge of inference efficiency. This substantial improvement in energy efficiency has a direct impact on operational costs. For cloud providers, it means being able to serve more inference requests with the same amount of energy, thereby increasing capacity and potentially reducing the need for costly infrastructure expansions.
For enterprise teams building applications on top of Gemini, these efficiency gains could eventually manifest in several ways. One likely outcome is a reduction in API costs. If Google can significantly lower the cost of serving Gemini inferences, they may pass these savings on to their customers, making AI-powered features more economically viable for a wider range of businesses. Alternatively, the increased efficiency could lead to greater availability of compute resources, allowing developers to run more complex or more frequent inferences without hitting capacity limits. While Google has not yet detailed how these gains would be distributed, the broader industry consensus is clear: reducing the cost of AI inference has become a critical priority. This is driven by the exponential growth in AI adoption and the increasing realization that inference, not training, represents the bulk of ongoing AI expenditure. The success of projects like “Frozen v2” could therefore have a ripple effect across the entire AI ecosystem, democratizing access to powerful AI capabilities by making them more affordable and accessible.
The development of specialized inference hardware represents a significant evolution in AI infrastructure. It signals a move away from a one-size-fits-all approach toward a more tailored and efficient ecosystem. As AI models become more ubiquitous and computationally demanding, the drive for optimized hardware solutions will undoubtedly intensify, shaping the future of how artificial intelligence is developed, deployed, and consumed.
