Chips and Semiconductors

1-Kilowatt AI Accelerators Redefine the Meaning of Reliability Verification

Operation at the 1-kilowatt level does not represent a single test condition, as risks vary according to the workload, heat location, operating duration, and power-delivery path. Semiconductor Engineering’s analysis argues that verification must cover the chip, package, test system, and field data—not merely the detection of electrical failures.

2026-09-10
6 min read
8 views
فريق تحرير certi.news
1-Kilowatt AI Accelerators Redefine the Meaning of Reliability Verification

AI accelerators approaching 1 kilowatt are driving semiconductor engineers to reconsider the concept of test coverage. Reaching the same power figure does not prove that the chip was tested under the conditions it will actually face; two different workloads can consume the same amount of power while activating different computing blocks, memory interfaces, and power-distribution paths, and generating different hot spots inside the package.

According to an analysis published by Semiconductor Engineering on September 10, 2026, verification at these levels extends beyond the device under test (DUT) to include the package, cooling system, load board, sockets and contact points, and automated test equipment. The article quoted Christopher Rand of Nordson Test & Inspection as saying that these system tools have become part of the ecosystem that should be qualified and verified, rather than merely neutral means of measuring the chip.

Total Power Does Not Reveal Where the Risk Lies

Chip testing has long relied on fault coverage and test patterns that expose structural defects and push the device toward its operating limits. But Brent Bullock of Advantest believes that verifying AI accelerators requires linking fault coverage to thermal coverage. A particular pattern may test the required logic without placing heat in the same locations exposed during a real functional workload, or without keeping the chip there long enough for its electrical and thermal effects to appear.

Simply operating all units at maximum intensity is not necessarily sufficient. Scan tests, even when power is taken into account, may create simultaneous switching that does not represent functional operation. Mission-mode and built-in self-test can more closely approach usage conditions, but they require workloads that reproduce the condition the engineer wants to examine.

The problem is that the figure of 1,000 watts describes power at an overall level, but does not specify how it is distributed across the die. The package surface temperature may appear acceptable while a smaller region becomes hot enough to affect timing or long-term reliability. Therefore, Lang Lin of Synopsys advocates running the real design with the actual workload and measuring temperatures throughout the device, rather than relying solely on tables or general maximum limits.

Time Is Part of Coverage

Not all failure mechanisms appear on the same time scale. Bullock noted that the chip may enter thermal runaway within a few hundred microseconds, requiring rapid monitoring of voltage droop or current indicators and test shutdown before significant damage occurs. Other problems, however, require seconds of sustained current for contact points to heat up, increasing their resistance and reducing their current-carrying capacity.

Glenn Cunningham of Modus Test explained that a contact point may pass the initial continuity test and then fail when realistic current is applied. Differences in resistance among similar contact points can also cause current to concentrate in a single path, raising the temperature and potentially leading to electrical overstress (EOS) that damages the contact point, socket, and device.

This creates a practical trade-off: qualification or burn-in tests cannot simulate years of operation, while outsourced semiconductor assembly and test (OSAT) companies are required to maintain production speed. According to Brad Booth of NLM Photonics, these systems cannot be operated indefinitely to confirm that they will last forever. The test duration must therefore be selected according to the targeted failure mechanism, while recognizing that a single period will not simultaneously cover rapid runaway, contact heating, and degradation caused by repeated thermal cycles.

From the Chip to the Test System

At these power levels, the source of a fault may lie in the test path rather than in the silicon. The load board, socket, handler or prober, cooling system, and automated test equipment all change electrically and thermally during operation. Bullock believes that separating these components as was historically done is no longer sufficient, and that the “full stack” should be examined to determine the condition that actually reaches the device.

The package itself adds another layer of complexity. Heat can alter charge-carrier movement, transistor performance, and timing, creating an interconnected loop involving power, temperature, stress, and timing. Voids, early delamination between layers, and cracking at interfaces can also weaken thermal conduction or develop after the device has passed the initial electrical test. Phenomena such as package warpage may disappear when heat is removed, even though the damage caused by the stress remains.

This is why nondestructive inspection before and after loading is becoming important, using techniques such as acoustic inspection and X-ray imaging, with destructive analysis used when a specific hypothesis exists about the defect’s location or cause.

certi.news’s Reading: Verification Moves from Measurement to Correlation

The actual change is not simply that tests require higher current or stronger cooling, but that the meaning of “coverage” has become multidimensional. Electrical coverage should be linked to thermal, workload, spatial, and temporal coverage. Timing margin should also be linked to resistance-induced voltage droop, temperature, and component degradation, rather than treating each measurement in isolation.

The article indicates that identifying these relationships does not end at the laboratory gate. Some interactions between the chip, system, and data center are difficult to reproduce during qualification and may appear only after field deployment. Monitoring timing margin, voltage droop, voltage fluctuation, and cycle-to-cycle variation therefore becomes an important source of understanding reliability after the device is deployed.

The fundamental limitation is that this perspective does not provide a single test that conclusively determines reliability. It requires better modeling, faster measurement equipment, characterization of the power and cooling paths, and correlation between test data and actual workload conditions. The open question remains how much can be collected before production and what will still need to be learned from thousands of devices after they enter service.

News source
Semiconductor Engineering
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news