Follow the latest coverage, related explainers and connected technology stories.
PrismML has launched the Bonsai 2 27B model, which compresses the Qwen3.8 27B model to 5.9 gigabytes while retaining 98% of its benchmark results, bringing the operation of advanced reasoning models closer to personal computers and potentially high-end phones. The company is betting on ternary weights to reduce memory requirements without significant performance loss.
ITmedia argues that growing interest in running large language models locally is explained not only by concerns about reliance on cloud services, but also by the maturation of three elements: open-weight models, hardware, and software. The first part of the series focuses on the evolution of models, with DeepSeek, Qwen, and other Chinese models emerging prominently.