Artificial intelligence

Microsoft: The Measure of Progress in AI Should Be “Useful Yield,” Not Infrastructure Scale

Rani Borkar of Microsoft believes that the success of AI infrastructure should not be measured by the number of chips or tokens processed, but by the amount of useful intelligence produced from each unit of energy, memory, and investment. She calls for an integrated design extending from data centers and chips to models, software, and agents.

2026-09-02
6 min read
7 views
فريق تحرير certi.news
Microsoft: The Measure of Progress in AI Should Be “Useful Yield,” Not Infrastructure Scale

Rani Borkar, who at Microsoft leads the organizations responsible for planning, engineering, developing, and deploying the hardware and infrastructure for the cloud computing platform, presents a different vision for measuring progress in the age of AI. Instead of focusing on the number of chips, the size of data centers, or the number of tokens a system can generate, she proposes adopting the concept of “yield”: the amount of useful output that can be produced from the available resources.

Borkar borrows the concept from the semiconductor industry, where yield refers to the number of usable chips that can be extracted from each wafer. In her view, this principle should be expanded to include capital, energy, memory, communications, models, and software, ultimately extending to the value AI delivers in work and everyday life.

Why Is Scaling Infrastructure No Longer Enough?

The article notes that AI adoption has accelerated faster than what occurred with the internet, personal computers, and smartphones, but its global use still reaches only about 18% of the workforce, while most usage is concentrated in conversations. As systems move toward inference, planning, tool use, and the execution of longer agentic workflows, infrastructure requirements are changing fundamentally.

Borkar says that a single agentic task may use more than 3,400 times the number of tokens consumed by typical conversational interactions. This escalation puts pressure on energy, rack density, package size, and memory capacity. She believes that the traditional response of adding more silicon, memory, power, and fiber cannot continue indefinitely, because each new improvement may require greater inputs than the previous generation.

The vision proposes two parallel paths: incremental improvements to current architectures to increase efficiency, utilization, and economic viability, and deeper transformations that reshape the performance curve through new architectures, materials, and methods for designing systems and models. Borkar cites the transition to multicore processor designs after increases in clock frequency encountered a power barrier, as well as the transition of NAND memory from planar to vertical designs.

Yield Is the Result of Collaboration Across All Layers

The article emphasizes that a bottleneck is not necessarily solved in the layer where it appears. Measuring the performance of an individual component is insufficient, because gains and losses accumulate across the data center, silicon, models, and the tools that coordinate agent tasks. Borkar therefore calls for “co-design” that first defines the desired output and then optimizes the entire system rather than maximizing the performance of a single element.

In memory, Microsoft sees the problem as more than merely a shortage of components; it is a system-level challenge. Inference requires accommodating larger models and longer contexts while delivering data quickly, whereas agents add generation, retrieval, tool use, and persistent memory operations within loops that may extend for minutes or hours. The experience of the Azure Maia platform shows that reducing KV-cache memory consumption can combine model engineering, data science, and compression, along with software management of memory hierarchies, silicon optimization, data movement, and compilers.

The idea here is not simply to add more bytes, but to extract more useful intelligence from every available byte.

From Networks to Energy

At the cluster level, intelligence does not come from a single chip, but from thousands of chips operating as one system. As a result, output depends not only on link speeds, but also on congestion management, fault recovery, workload distribution, programming complexity, and the boundaries between silicon, systems, and software.

Borkar says that the design of the Maia platform began with the intended outcome—efficient inference at fleet scale—rather than with a preexisting network design. This included building a two-level scale-out network, integrating network-card functions into the chip, and developing a dedicated transport layer. According to the article, this approach delivered scalable performance in dense inference clusters, simplified programming, increased workload flexibility, and reduced the networking hardware required.

Energy, meanwhile, has shifted from being a resource consumed by the system to a constraint that must be incorporated into the design from the electrical grid to the chip. The article points to rack power rising from tens of kilowatts to hundreds, and to data-center campuses operating at gigawatt scale. It also mentions solutions such as solid-state transformers and 800-volt direct-current distribution to reduce distribution losses.

Microsoft’s Arm-based Azure Cobalt 200 server processor provides an example of this co-design. It allows each core to have dedicated voltage and frequency control, while enabling the power consumption of each virtual machine to be specified through software. The company says that precise control allows power to be adjusted while protecting the performance of critical workloads, making it possible to run more servers within the same power limits.

What Matters in Practice?

The importance of this argument lies in shifting the discussion from a race for raw capacity to the efficiency of converting resources into results. For cloud and data-center operators, this means that hardware evaluation cannot be separated from networking, software, cooling, and workload management. For model and agent developers, memory efficiency, task length, and tool integration become factors affecting scalability just as much as model size.

However, this conclusion remains a vision from Microsoft, based in part on its experience building and operating AI infrastructure at large scale, rather than an independent standard established for all environments. The article also does not provide detailed comparative figures to measure the gains of Maia or Cobalt 200, nor does it specify the cost or timeline for implementing the solutions mentioned outside the Azure context.

Nevertheless, the “need for yield” poses a practical question to the industry: Do growing investments in energy, chips, and data centers produce intelligence that can be accessed at an appropriate cost, as well as tangible productivity and value? In Borkar’s vision, yield is not complete when tokens are produced, but when those outputs become faster scientific discovery, an earlier-detected medical signal, better learning, or new opportunities for small businesses.

News source
Microsoft Official Blog
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news