Cloud Computing and Data Centers

AI Data Centers Face a Stranded Power Problem That Could Be Converted Into Computing Capacity

AI data centers consume part of their available electrical capacity because of backup power systems and reliability requirements, leaving capacity that is not actually used. The analysis examines voltage-reduction and load-management techniques that could enable this capacity to be utilized while allowing rapid fallback when one power source fails, although scheduling, security, and fault-response challenges remain.

2026-09-15
5 min read
3 views
فريق تحرير certi.news
AI Data Centers Face a Stranded Power Problem That Could Be Converted Into Computing Capacity

The power problem in AI data centers is not only how much electricity is consumed, but also how much capacity is paid for and reserved to handle failures and then remains unused. An analysis published by Semiconductor Engineering on September 15, 2026, presents a basic paradox: reducing chip and server consumption improves efficiency, but it may increase the gap between available capacity and capacity actually used, while backup-power architecture requires operators to leave a substantial portion of capacity outside normal use.

This problem is becoming more important as AI workloads expand, because obtaining new electrical capacity may require transmission lines, permits, and additional on-site generation. According to Marisa Hummon, chief technology officer at Utilidata, adding equipment to capacity that is already available may be easier than building a new data center or waiting for increased supply.

Why Does Capacity Remain Unused?

Data centers are typically designed to withstand the loss of one power source without interrupting service. In a “3+1” arrangement, the facility needs three full-capacity sources and adds a fourth as a backup, so that the four sources operate at about 75% of their capacity and the remaining three can cover the load if one fails. In a 2N architecture, power is supplied through two separate paths, and each path is maintained at about 50% load in anticipation of losing the other.

Hummon explains that a site configured for 2 gigawatts may effectively be treated as a 1-gigawatt site when 2N is adopted, and only about 70% to 80% of that capacity is typically used. As a result, the capacity actually reaching the computing infrastructure, when the site is considered as a whole, may amount to about one-third of the nominal 2-gigawatt capacity. This is not waste caused by equipment malfunction, but a direct result of reliability requirements and operating margins.

Reducing Chip Consumption Does Not Solve the Problem on Its Own

Energy optimization begins with circuit design and management of the intellectual property integrated into the chip. Arif Khan of Cadence pointed to the importance of shutting down unused interfaces and channels, using multiple standby states, and adjusting clock frequency and voltage according to the work in progress. Low-power AMBA interfaces, such as the Q and P channels, provide mechanisms for controlling clock gating, power domains, and multiple-state management.

The on-chip interconnect also plays a central role, as SignatureIP relies on clock-gating and independent power states for network-on-chip nodes. However, reducing power requires balancing the amount of savings against the time needed to return to operation. Chips also do not all behave the same way: voltage is typically set according to the worst load, temperature, and aging conditions, even though those conditions may not always occur together.

According to Noam Brossard of proteanTecs, on-die monitoring can measure the actual safety margin and reduce voltage when maximum margins are not required, then raise it when the load or conditions change. A fast protection mechanism is still needed, because reducing voltage too far could threaten circuit timing when the load changes suddenly or a dynamic voltage drop occurs. Protection may involve lowering the clock frequency or temporarily limiting performance until the voltage returns to a safe level.

What Changes in Practice at the Data Center?

The broader idea is to use backup capacity during normal operation, while reducing loads before or when a power source is lost. Utilidata proposes two control loops: the first is slower and connects to workload scheduling, providing forecasts of available capacity for each rack, row, or hall so that the scheduler can assign GPU tasks according to changing capacity. The second is faster, using interfaces such as DVFS and server management to control power over a millisecond timescale.

The technology does not schedule or cancel work directly; instead, it influences the scheduler through power data. It uses the NVIDIA management library and the baseboard management controller, or BMC, to access power behavior and measurements, with direct rack-level measurement rather than aggregating estimates from individual servers. According to Hummon, this approach could allow the four lines in a 3+1 arrangement to operate at about 95% to 98% of their capacity under normal conditions, then reduce the load in an orderly manner when necessary.

Limitations and Open Questions

Using backup capacity does not mean that adding servers becomes free or unrestricted. Additional equipment requires space, cooling, and connectivity, while moving AI workloads may require draining active requests and loading model weights onto another instance, processes that may take seconds. Conversely, reducing chip voltage alone is not enough to address a failure in a power source; monitoring and response must operate at the rack and facility levels.

This architecture also adds a sensitive security surface. Every server is equipped with a BMC, while GPU hosts expose interfaces for setting power limits for processes with sufficient privileges. Utilidata therefore says its system relies on secure boot, a hardware root of trust, signed firmware, and mutual authentication on interfaces, while running the control loop locally without relying on the internet. Editorially, the importance of this development lies in transforming backup power from a rigid margin into a manageable resource, but its success depends on the operator’s ability to ensure rapid response, avoid harming workloads, and secure the control tools that could affect thousands of servers at once.

News source
Semiconductor Engineering
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news