Now Reading
Google Developing New Chip to Bake Gemini AI into Silicon

Google Developing New Chip to Bake Gemini AI into Silicon

Google Gemini AI chip development

Google is developing a new custom server chip that could integrate its Gemini AI model directly into silicon, according to a recent report. The move could improve power efficiency significantly as the company faces rising demand for AI computing resources. However, the approach marks a major shift from traditional AI processing methods.

The new chip would include a “frozen” version of Gemini’s neural network weights inside custom hardware. Therefore, instead of running Gemini as software on general-purpose processors, the chip would operate as an application-specific integrated circuit designed for a fixed AI model.

As a result, the technology could deliver much higher efficiency per AI token compared with Google’s current custom chips. The company reportedly aims to deploy the technology as early as 2028, although the approach comes with limitations because frozen models cannot receive updates without creating new chips.

A New Direction for AI Hardware

Meanwhile, the idea of embedding AI models directly into silicon is gaining attention across the technology industry. Finnish startup Taalas previously demonstrated a similar approach with its HC1 chip, which places Meta’s Llama model directly into hardware.

The company claimed that its chip could achieve speeds up to ten times faster than competing inference platforms. Additionally, it reported that the system required only 12–15 kilowatts per rack, compared with 120–600 kilowatts for traditional GPU-based racks.

However, the main challenge remains flexibility. Once an AI model is permanently built into silicon, developers cannot easily modify or upgrade it. Instead, they must create new hardware when major model changes are required.

AI Demand Increases Pressure on Google Infrastructure

Furthermore, Google’s new chip project comes as the company experiences growing pressure on its AI computing infrastructure. During Alphabet’s first-quarter 2026 earnings call, executives revealed that Google Cloud faced capacity limitations as demand continued to increase.

At the same time, Google Cloud exceeded $20 billion in quarterly revenue for the first time, while its backlog reportedly doubled to $462 billion. In addition, reports indicated that Google informed Meta in March 2026 that it could not provide all the computing capacity requested for AI workloads.

Therefore, Google has expanded its AI hardware strategy in several areas. In April, the company introduced its eighth-generation Tensor Processing Units (TPUs), separating training and inference operations. Moreover, Broadcom signed a long-term agreement to help develop future generations of Google’s custom AI chips through 2031.

See Also
Samsung ChatGPT Enterprise deployment

Google has also started offering TPU capacity to external customers, including OpenAI. Consequently, the company is building a wider AI infrastructure ecosystem to meet increasing demand.

Frozen Gemini Chip to Work Alongside TPUs

The upcoming Gemini-focused chip is expected to complement Google’s existing TPU lineup rather than replace it. While TPUs remain programmable and support updated AI models, the frozen chip could handle large volumes of repeated inference tasks more efficiently.

As a result, Google could reduce energy costs for common AI requests while reserving flexible hardware for training new models and developing future AI systems.

Meanwhile, investor confidence in Google’s AI infrastructure plans has continued to grow. Alphabet shares increased following reports about the chip development, with the company’s stock already gaining momentum ahead of its upcoming earnings announcement.

View Comments (0)

Leave a Reply

Your email address will not be published.

© 2024 The Technology Express. All Rights Reserved.