Follow the latest coverage, related explainers and connected technology stories.
A paper from UC Berkeley and FuriosaAI examines the use of High Bandwidth Flash to expand the memory of AI accelerators when running large language models. The simulations indicate a reduction in completion time of up to 87%, with potential improvements in energy consumption and write endurance through cache-aware scheduling.