PrismML’s Bonsai 2 27B shrinks AI model to 5.9GB for local devices
The new ternary model retains 98.2% of Qwen3.8 27B benchmark performance while reducing memory requirements by a factor of nine to ten.
AI startup PrismML has released Bonsai 2 27B, a compressed large language model designed to run on personal computers and high-end smartphones. The release marks a significant step in the trend towards local AI processing, allowing complex models to operate outside of data centres.
The new model utilises a ternary weight compression technique to reduce the footprint of Alibaba’s Qwen3.8 27B. This process results in a file size of just 5.9GB, representing a nine- to ten-fold reduction in memory requirements compared with the original model.
According to PrismML, Bonsai 2 27B retains 98.2% of the benchmark performance of the base Qwen3.8 27B model. While the company describes the compression as near-lossless, the 1.8% performance drop may still be relevant in specific edge cases, though the overall retention remains high for a compressed model.
Beyond standard language processing, the model supports multimodal and agentic capabilities. These features suggest the model can handle various data types and perform task-oriented actions, although the specific scope and limitations of these agentic functions are not detailed in the initial release materials.
The launch highlights the growing interest in optimising large language models for consumer hardware. By drastically lowering memory demands, PrismML aims to make advanced AI accessible on devices that previously lacked the resources to run such models.
The release was announced on 17 September 2026. As a recent entry in the AI compression space, the model’s performance relative to other 27B-class competitors remains to be seen, with current comparisons strictly limited to the Qwen3.8 27B baseline.


