Geoff Tate believes that graphics processing units (GPUs) still account for about 70% of the world’s installed AI computing capacity by manufacturer, compared with 30% for custom chips, which are primarily dominated by Google and Amazon. However, this advantage is no longer guaranteed, as cloud companies and advanced-model developers have begun designing their own hardware to reduce the operating costs of specific workloads and improve energy consumption.
Epoch.ai data cited by the author indicates that global installed AI computing capacity grew 17-fold over three years. In 2023, Nvidia held a 90% share compared with 10% for Google, while custom computing accounted for 20% of growth over the past three years. Tate notes that the data does not include Meta, Microsoft, or Cerebras, whose capacities he describes as relatively small within this measurement.
Nvidia’s Advantages Go Beyond the Processor
According to the author, Nvidia’s strength does not depend on an individual GPU, but on an integrated ecosystem that includes interconnect and expansion units, racks, central processing units, and networking processors. The company offers the NVL72 platform, the Vera processor aimed at AI-agent workloads, the BlueField unit, and Spectrum-6 networking, in addition to the Groq 3 LPX option for low-latency inference.
Nvidia says that the Vera processor delivers 1.8 times higher performance on agent workloads compared with a conventional processor. The company also presents Vera Rubin systems as achieving improvements ranging from twofold to 30-fold compared with GB300 in the DeepSeek-v4-PRO benchmark, depending on interactivity. The article also states that Vera Rubin systems have begun operating in racks at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud, and Nebius.
Nvidia expects 70% growth in its next fiscal year, while noting that demand is constrained by supply. Its commitments to secure supplies also increased, according to the earnings call, from $119 billion to $279 billion. Tate believes that securing manufacturing capacity from TSMC and supplying HBM memory will remain decisive factors in the competition.
Custom Chips Move Closer to the Heart of the Competition
Designing chips internally gives cloud companies and model developers direct knowledge of their workloads, the distinction between training and inference requirements, and infrastructure bottlenecks. As computing scales expand, designing two specialized chips—one for training and another for inference—has become more economically viable, particularly if this improves throughput per watt by 10% to 30% and reduces silicon area by a similar percentage.
OpenAI stands out in the article through its Jalapeño chip, an inference accelerator designed by the company for its workloads, with support for both the prefill and decode stages while keeping the KV Cache local. The system uses Broadcom Tomahawk 6 switches to build a domain containing 128 chips and a global domain reaching 2,048 chips. According to OpenAI’s initial tests, Jalapeño outperformed Nvidia GB300 on GPT-OSS, DeepSeek R1, and Kimi K2.5 workloads. However, the tests relied on single-token prediction and have not yet included multi-token prediction, which the company expects to increase performance by threefold to fivefold.
Tate describes the comparison between Jalapeño and Vera Rubin as an approximate estimate based on projecting published performance curves, rather than a complete direct test. He therefore believes the results are promising but require verification in production deployment and across broader scales and workloads. Jalapeño is also dedicated to inference and does not address training requirements, which may account for one-third or more of OpenAI’s computing needs.
Google, Meta, and AMD Expand the Alternatives
In its eighth-generation TPU, Google is launching two synchronized versions: TPU 8t for training and TPU 8i for inference. The article notes that TPU 8i is approximately 1.4 times larger than 8t because it requires additional HBM memory, while the interconnect topology differs between the two versions according to the differing requirements of training and inference. Google says that the new generation delivers approximately twice the performance of previous generations, while Morgan Stanley estimated TPU sales at about $9 billion in the second half of 2026, $84 billion in 2027, and $108 billion in 2028.
Meta is developing the MTIA series for recommendation-system workloads, followed by training and generative AI, while Cerebras introduced the CS-4 chip, which it says doubles tokens per user compared with CS-3 and achieves up to ten times the throughput per watt. AMD also presented the MI455x unit and the Helios system. According to the article, its share of global installed capacity rose from zero to 6%, with no performance chart for a specific AI benchmark included in the referenced presentation.
What Does This Mean for GPU Manufacturers?
certi.news’s reading: The data does not point to an imminent end for GPUs, but rather to a shift in competition from general-purpose chip performance toward full-system efficiency, the availability of power, memory, and manufacturing, and the hardware’s ability to serve specific workloads. GPU flexibility will remain an important advantage for customers running diverse models and workloads, while custom chips can outperform them when a company’s workloads are sufficiently known and stable to justify a dedicated design.
Tate suggests that Nvidia consider developing separate chips for training and inference instead of relying on a single general-purpose design, and that Nvidia and AMD adopt optical interconnect technologies such as optical circuit switching and co-packaged optics as copper approaches its practical limits. However, he clarifies that some of these proposals are analytical opinions, and that he serves on the boards of two companies working on optical-interconnect technologies, which means they should be treated as recommendations from the article’s author rather than independent facts.
The author’s conclusion is that Nvidia has significant defensive barriers, but faces customers with substantial resources and a direct incentive to run their models on their own hardware. If Jalapeño demonstrates performance in production deployment comparable to or better than Vera Rubin, that would be a strong indicator that GPU dominance has become the subject of real competition, not merely a theoretical possibility.