PrismML releases Bonsai 2 27B, which compresses Alibaba's Qwen3.8 27B to 5.9 GB, small enough for smartphones, while retaining 98.2% of Qwen's benchmark scores
If AI lab PrismML isn't on your radar yet, it should be — not because it's raised gobs of money (it hasn't yet …
TechCrunchJulie Bort
Context & Ripple Effects
PrismML emerged from stealth in March with a claim that 1-bit LLMs could sharply reduce model size without a comparable performance hit. By July, it had turned that thesis into a demonstration of a 27B Qwen model on an iPhone 17 Pro and said its earlier Bonsai release ran natively on Apple devices via MLX.
Bonsai 2 moves the comparison to Qwen3.8 and puts a concrete footprint behind PrismML’s compression pitch. The broad same-day pickup, including developer communities, indicates that near-lossless local inference is being treated as a deployment question rather than solely a model-training result.
First-order effects
PrismML gives developers a 5.9 GB derivative of Alibaba’s Qwen3.8 27B that it says retains 98.2% of the source model’s aggregate benchmark scores, widening the set of consumer hardware that can host a 27B-class model locally.
Apple’s reported evaluation of PrismML’s approach gains a more mature reference point after the earlier Bonsai 27B Apple-device launch, while Alibaba gains another route for Qwen-derived capability to reach devices outside a hosted service.
Second-order effects
Device makers and app developers can weigh local Qwen-derived inference against cloud calls with a smaller storage and memory burden, making on-device latency and offline operation more central product trade-offs.
Compression specialists gain a clearer competitive benchmark: PrismML’s claimed retention rate shifts attention from parameter counts toward whether reduced-weight models preserve task quality on broad evaluations.
Third-order effects
If comparable compression holds across successive open-weight releases, distribution advantage may move toward companies that control device integration and application surfaces rather than only the operators of large inference clusters.
The pattern strengthens a split AI stack in which base-model developers supply weights while compression layers and device platforms determine where inference runs and who owns the user relationship.
The trend: Large open-weight models are becoming inputs to a local-inference stack, with compression determining whether frontier-scale capability stays in the cloud or reaches consumer devices.
Very excited about today's release of Bonsai 2 27B. It is remarkable how quickly the quality gap between ternary and full-precision models has been narrowed. Across a suite of 20 benchmarks, Ternary Bonsai 2 27B retains 98.2% of Qwen3.8 27B's aggregate performance, despite being …
everyone is obsessed with massive models but prismml gets it. tiny llms will win because latency is the only feature that actually matters for ai video. the era of waiting for a token stream to finish is dead. https://techcrunch.com/...
Today, we're announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footp…
it's so nice to have a capable local model i can recommend to people in my life who only have normal laptops, super valuable drop. check out the new prism-ml quantization of Qwen3.8-27b. 6GB weights! — blog: prismml.com/news/bonsai-... webgpu (browser-native!) demo: huggingf…
Today's my birthday, and I could not have asked for a better gift. Bonsai 2 fits on most phones shipping today (anything >8GB). And it outperforms Opus 4.6 and 5.6 Luna — on tasks that would have sounded absurd to attempt anywhere two years ago. I try to be disciplined about time…