Artificial intelligence

High Bandwidth Flash Memory Could Reduce Large Language Model Serving Time by up to 87%

A paper from UC Berkeley and FuriosaAI examines the use of High Bandwidth Flash to expand the memory of AI accelerators when running large language models. The simulations indicate a reduction in completion time of up to 87%, with potential improvements in energy consumption and write endurance through cache-aware scheduling.

2026-10-02
3 min read
12 views
certi.news Editorial Team
High Bandwidth Flash Memory Could Reduce Large Language Model Serving Time by up to 87%

A research paper from the University of California, Berkeley, and FuriosaAI proposes using High Bandwidth Flash, or High Bandwidth Flash (HBF), as an additional layer alongside High Bandwidth Memory (HBM) and host memory to meet the serving needs of large language models. The paper addresses the problem posed by growing model sizes and longer conversation contexts, in which memory capacity and bandwidth become bottlenecks for inference performance.

Large language model serving systems store model weights and KV cache data in memory. Keeping this data available is particularly important in agentic workloads, which conduct repeated interactions within contexts that expand over time. However, expanding memory with HBF involves trade-offs: flash is slower than HBM to access, and its write endurance is limited compared with volatile memory.

What did the paper test?

The researchers presented a hierarchical storage architecture combining HBM, HBF, and host memory, along with a cache-aware scheduling mechanism that takes cache contents and data locations into account. They used trace-driven simulations to analyze the impact of data placement and scheduling decisions on completion time, energy consumption, and the expected write lifetime of HBF.

According to the reported results, the fastest HBF-supported systems reduced completion time by between 36.1% and 87.0% compared with systems relying on HBM alone. The calculated energy savings also reached 55.8% in the evaluated scenarios.

The gain is not guaranteed for every workload

The results do not present HBF as a direct replacement for HBM. The paper noted that using HBF increased energy consumption in some lightweight workloads, meaning that the benefit depends on the nature of the load, the amount of retained data, and how read and write operations are scheduled.

The importance of write management is clear in the lifetime estimates. Cache-aware scheduling increased the expected HBF write lifetime from 4.77 to 14.82 years in the evaluated configuration.

Why does this research matter?

The results show that adding a larger memory layer alone is not enough to address model-serving bottlenecks. The gains are tied to coordinating data placement with the scheduling policy, allowing inference workloads with long contexts to benefit from the additional capacity without exhausting flash endurance or incurring unnecessary energy costs.

For AI accelerator designers and infrastructure operators, the paper proposes a research path for expanding memory capacity beyond HBM, but it does not yet establish the feasibility of broad commercial deployment. The results are based on simulations and operational traces, while actual performance and endurance will depend on hardware design, workload characteristics, and memory-management policies in real-world systems.

The paper is titled “Characterizing High Bandwidth Flash for LLM Serving,” and was prepared by Yu, Zack, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. It was published as a research preprint on arXiv in September 2026 under arXiv:2609.39131.

News source
Semiconductor Engineering
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news