Sources: at WWDC, Apple is likely to showcase how 15 years of designing chips gives it an advantage in running AI locally, and will use a distilled Gemini model
Aaron Tilley /The Information:
Context & Ripple Effects
Apple’s AI effort has been building across two tracks: a dedicated on-device Neural Engine reported in 2017, and later work on server-side AI hardware, including use of M2 Ultra systems and a reported dedicated server-chip project.
The reported WWDC plan connects those tracks to a product narrative: Apple can emphasize hardware it controls for local inference while drawing on an external model provider for at least some demonstrations.
First-order effects
- Apple can frame local AI execution as a differentiator for its device lineup and chip-design stack at WWDC.
- Gemini gains a visible role in Apple’s AI presentation, while Apple’s own models and hardware remain central to the stated local-computing message.
Second-order effects
- Developers building for Apple platforms would have stronger incentive to design AI features around on-device constraints and Apple hardware capabilities, rather than assume every task is cloud-served.
- The combination of local processing and an external distilled model raises the importance of how Apple divides workloads between device models, its own server infrastructure, and partner technology.
Third-order effects
- If Apple sustains this approach, AI competition in consumer devices may increasingly turn on vertically integrated silicon and runtime efficiency, not only the scale of a provider’s frontier model.
- The earlier reports of both device and server AI chips point to a hybrid architecture becoming a durable strategic layer: local inference where hardware permits it, backed by server capacity for heavier tasks.
The trend: This is one data point in the shift toward hybrid AI products, where proprietary device silicon is used to make local inference a product and platform advantage while outside model providers can fill capability gaps.