/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

PrismML launches Bonsai 27B, a model based on Qwen3.6 27B that it says runs natively on Apple devices via MLX; its CEO says Apple is evaluating the tech

Apple is in talks with a small Silicon Valley company that says it can shrink powerful artificial intelligence models enough to run directly …

CNBC MacKenzie Sigalos

Context & Ripple Effects

Apple’s on-device AI work has previously centered on a roughly 3B-parameter model, alongside larger models served through its Apple-silicon Private Cloud Compute infrastructure. Its internal testing has also reportedly spanned models up to 150B parameters.

PrismML’s claim that a Qwen3.6-derived 27B model runs through Apple’s MLX framework follows reports that it demonstrated the approach on an iPhone 17 Pro. Apple’s reported evaluation makes the launch relevant as a potential extension of the software stack Apple has built around MLX.

First-order effects

  • PrismML gains a concrete product and demonstration point for its claim that substantially larger language models can operate natively on Apple hardware via MLX.
  • Apple has another externally developed approach to assess as it weighs the boundary between on-device inference and models handled through Private Cloud Compute.

Second-order effects

  • If the approach proves practical, it could increase pressure on Apple’s model teams and MLX ecosystem to prioritize memory-efficient deployment of larger models rather than relying solely on smaller local models.
  • Developers targeting Apple silicon would have a clearer incentive to test MLX-based model optimization, because the value proposition shifts from framework access to running more capable models locally.

Third-order effects

  • The more capable models that can be deployed locally, the more AI product design may move toward hybrid systems in which private-cloud capacity is reserved for workloads that cannot fit or perform adequately on-device.
  • This is not yet evidence of an Apple product decision: the durable implication depends on whether PrismML’s reported result holds across real-world latency, memory, power, and quality requirements.

The trend: This is one data point in the push to make larger generative-AI models viable on consumer devices, reducing the portion of inference that must be sent to centralized cloud systems.

Discussion

  • @prismml @prismml on x
    Today, we're announcing Bonsai 27B: the first 27B-class model to run on a phone. Bonsai 27B is the new multimodal flagship of the Bonsai family. Based on Qwen3.6 27B, it brings a new capability tier to local AI: multi-step reasoning, structured tool use, long-context workflows, […
  • @prismml @prismml on x
    Here is Ternary Bonsai 27B running an end-to-end agentic workflow locally with Hermes on an NVIDIA GeForce RTX 5090 GPU. The model reasons, calls tools, reads outputs, modifies files, and surfaces insights - all on consumer hardware, while all private files, intermediate states, …
  • @scaling01 @scaling01 on x
    A 27B model using a mere 3.9GBs [image]
  • @deryatr_ Derya Unutmaz on x
    This is unbelievable! 27B class model running on your phone! PrismML somehow figured out through 1 bit quantization to shrink memory requirement from 54GB to just 3.8GB (-93%), while retaining its intelligence! This is a big deal!
  • @evaninwords Evan Walters on x
    We're releasing 27B 1-bit and ternary models based on Qwen 27B! Super excited about this release, they run easily on your macbook and maintain vision capabilities. Yesterday I was controlling my macbook with the ternary using a simple script with pyautogui. Excited to see what
  • @prismml @prismml on x
    The footprint reduction does not come at the expense of the capabilities that matter. Across 15 benchmarks spanning knowledge, reasoning, math, coding, instruction following, tool use, and vision, Ternary Bonsai 27B retains 95% of the full-precision model's performance. The [imag…
  • @prismml @prismml on x
    The phone threshold is even harder than the storage number suggests. A phone exposes only part of its memory to an application, and the model must share that budget with its KV cache, activations, runtime, and the rest of the product. At 3.9 GB, 1-bit Bonsai 27B clears that [vide…
  • @babakhassibi Babak Hassibi on x
    Today PrismML is bringing the maximal possible intelligence that can run locally on an iPhone. Based on Qwen 3.6 27B, our 1-bit Bonsai 27B reduces the memory footprint from 54GB to 3.9GB and retains 90% of the benchmark scores across knowledge, reasoning, math, coding,
  • @prismml @prismml on x
    Raw capability determines what a model can do. Intelligence density determines where it can do it. Bonsai 27B moves the Pareto frontier left again: 27B-class capability in a footprint smaller than many full-precision 2B models. By intelligence density, 1-bit Bonsai 27B delivers […
  • @omarsar0 Elvis on x
    Huge if true! We are talking about a 27B multimodal model that runs locally on a phone. That's wild! Bonsai 27B reaches up to 163 tok/s in 1-bit and 134 tok/s in Ternary on an NVIDIA GeForce RTX 5090. On an M5 Max, it reaches up to 87 tok/s in 1-bit and 58 tok/s in Ternary.
  • @prismml @prismml on x
    Why does this matter? Because modern AI workflows are no longer single prompts. They are sustained loops. A capable agent may take hundreds of steps: reasoning, calling tools, reading outputs, updating its state, and iterating toward a result. When every step is remote, [image]
  • @togethercompute @togethercompute on x
    Bonsai 27B is a big step for local AI: 27B-class multimodal capability in a phone-class footprint. Congrats to @PrismML on the launch. Try Ternary Bonsai 27B on Together AI: https://www.together.ai/...
  • @nvidiartxspark @nvidiartxspark on x
    Congratulations to @PrismML on the launch of Bonsai 27B family of models! Try them out today on NVIDIA RTX GPUs & DGX Spark.
  • @timkellogg.me Mr. Tim on bluesky
    it fits into 3.9 GiB!!  —  this is part of the new trend — tiny LLM that's part of a larger multi-agent system (involving frontier models, software & more)  —  an LLM small enough to keep in your pocket, to use obsessively for long-running agent tasks [embedded post]
  • @jessefelder.com Jesse Felder on bluesky
    ‘The technology, according to PrismML, could ultimately extend well beyond phones and laptops to robotics, autonomous systems and other products that need to make decisions quickly without relying on a cloud connection.’ www.cnbc.com/2026/07/14/a...
  • r/LocalLLaMA r on reddit
    Bonsai 27B: The First 27B-Class Model to Run on a Phone
  • r/LocalLLaMA r on reddit
    Apple in talks with startup PrismML that shrinks AI models to run on an iPhone