Artificial intelligence

PrismML Compresses a Reasoning Model to 5.9 Gigabytes for Local Operation

PrismML has launched the Bonsai 2 27B model, which compresses the Qwen3.8 27B model to 5.9 gigabytes while retaining 98% of its benchmark results, bringing the operation of advanced reasoning models closer to personal computers and potentially high-end phones. The company is betting on ternary weights to reduce memory requirements without significant performance loss.

2026-09-17
4 min read
2 views
فريق تحرير certi.news
PrismML Compresses a Reasoning Model to 5.9 Gigabytes for Local Operation

Artificial intelligence startup PrismML has launched its Bonsai 2 27B model, which compresses Alibaba’s open-source Qwen3.8 27B model into a file measuring only 5.9 gigabytes. The company says the compressed model achieves 98% of the original model’s aggregate scores on benchmarks, while reducing memory requirements by roughly ninefold to tenfold.

The new size makes it possible to run the model on a personal computer, and potentially on a high-end smartphone, rather than relying entirely on cloud servers. This does not mean that every phone will be able to run it in practice, since the source presents phone compatibility as a possibility, and actual performance depends on the hardware and software surrounding the model’s operation.

How Does the Compression Technology Work?

PrismML relies on what it calls ternary weights. Model weights are values that store some of the information learned during training and typically require 16 bits per weight, according to the company’s explanation. The new approach represents each weight using only one of three values: plus one, minus one, or zero. Reducing the number of possible values significantly decreases the space required to store the model.

The company is led by Babak Hassibi, a Caltech professor specializing in compression techniques, while Ion Stoica serves as its adviser. Its backers include Khosla Ventures, Cerberus Capital, and Caltech. PrismML raised a $22.25 million seed round. It is not the only company working on compressing large language models, as the source also identified Multiverse Computing as a competitor in the field.

Better Results Than the Previous Release

Bonsai 2 represents an improvement over the first version of Bonsai, which the company launched in March and which achieved 95% of the original model’s scores, according to what TechCrunch reported from the company. PrismML says the first version was downloaded more than 11 million times, while its smaller models recorded an additional 2.6 million downloads.

However, 98% does not mean complete equivalence with the original model, nor does it by itself prove that the remaining gap will not appear in practical uses. Hassibi acknowledges that compression will likely have some effect on performance. Benchmark tests also do not reflect every real-world task, and the software environment in which the model operates also affects its accuracy.

Why Does This Development Matter?

If PrismML continues to reduce model sizes while preserving most of their capabilities, running reasoning models locally could become more practical. This would allow data to be processed on the user’s device instead of being sent to the cloud, which could support privacy and reduce reliance on external connectivity, according to Stoica’s argument. However, it does not yet resolve questions about power consumption, response speed, and device compatibility, nor does it prove that the model will deliver the same level of performance in every scenario.

The company plans to apply its compression technology to much larger models, and Hassibi said upcoming releases could fall within the range of several hundred billion parameters over the next few months. He believes that larger models may give the compression process more room to preserve their capabilities, but this goal remains a future plan rather than an available result. The report also pointed to rumors of discussions with Apple, but Hassibi declined to comment on them, so they cannot be considered an announced partnership or agreement.

News source
TechCrunch AI
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news