Qualcomm’s 45-TOPS NPU cleared the 40-TOPS threshold for a next-generation AI PC, yet it did not create a separate “edge AI” purchase order. The compute arrived inside a laptop refresh. Cloud compute arrives with a contract, a power requirement, and several billion dollars of punctuation.
Key takeaways
- Edge AI capacity is often hidden inside ordinary device purchases: Qualcomm’s 45-TOPS Snapdragon X Elite NPU cleared the 40-TOPS AI-PC threshold, but buyers record the investment as a PC refresh rather than a separate AI deployment.
- The durable architecture is hybrid inference, with each task routed between device and cloud according to privacy, latency, memory, power, model quality, reliability, and cost.
- Qualcomm’s planned Modular acquisition targets the software bottleneck: compilers and runtimes are needed to make models portable and allocate work across endpoint and data-center chips.
- Data-center processors give Qualcomm a visible, separately priced growth business that complements endpoint distribution; the company projects $15 billion in data-center chip sales by 2029.
- Local AI creates value only when it improves device demand or reduces serving costs without sacrificing task quality; raw TOPS do not guarantee useful applications.
NPUs disappear into the PC refresh cycle
In 2024, Intel executives said next-generation AI PCs would require 40 TOPS of NPU performance. Qualcomm then distributed models optimized for the 45-TOPS Hexagon NPU in Snapdragon X Elite. The hardware cleared the threshold, developers gained models prepared for it, and PC makers gained another reason to replace existing machines.
For a corporate IT buyer, the budget records a PC refresh while developers inherit a local accelerator. Each order adds inference capacity under the accounting label for a computer. Phone, vehicle, camera, and industrial-system makers use the same route through existing procurement cycles.
Apple and Perplexity route work task by task
Apple made the structure explicit when it described Apple Intelligence as a roughly 3-billion-parameter on-device model paired with a larger model on Apple silicon servers through Private Cloud Compute. The system places work according to privacy, memory, latency, capability, and elasticity.
Two years later, Perplexity split Computer tasks between local and cloud models, emphasizing private data and token efficiency. Application teams lose margin on unnecessary remote calls and user value when underpowered local models fail.
Teams should optimize cost per useful task while meeting requirements for privacy, latency, reliability, memory, power, and model quality. A small local model saves little when poor results erase user value.
Chip benchmarks measure available capacity, not whether an application uses it well. Developers capture the value of local prediction only after redesigning the workflow so routine operations stay near the user and demanding workloads escalate automatically.
Modular turns TOPS into usable capacity
A 40- or 45-TOPS NPU sits idle until developers can build useful applications for it. Qualcomm’s optimized model release gave them a path from silicon capability to product capacity.
Qualcomm agreed to acquire Modular for nearly $4 billion. Modular builds a chip software platform and a proprietary programming language. Its platform translates models across chips, optimizes execution, and helps developers deploy software across Qualcomm’s endpoint business and planned data-center products.
Developers then face portability problems across generations of endpoint silicon, cloud hardware, and different memory and power envelopes. Compilers and runtimes can inspect the task, available models, device state, privacy requirements, and escalation cost before allocating the work.
Dragonfly adds a separately priced revenue line
Qualcomm says Meta will use its Dragonfly C1000 data-center CPU when production starts in 2028. The company also projects $15 billion in data-center chip sales by 2029 and raised its non-handset chip-revenue forecast from $22 billion to $40 billion.
If the Modular deal closes, Qualcomm would sell device silicon, data-center processors, and the software layer that maps models across them. Qualcomm could collect silicon or software revenue as application teams change workload placement.
Qualcomm’s device business earns through unit volumes and silicon content, limiting how directly it can monetize each additional inference task. Data-center processors give the company a separately forecast revenue line tied to centralized build-outs.
Customers still decide whether local AI pays
Customers reward local AI when it changes the device they buy or lowers a provider’s serving bill.
Phones impose hard limits on memory, processing power, and energy. Local models must fit, run reliably, and improve the experience enough to influence a replacement decision.
Cloud providers can point to visible demand. Google said it needed to double AI compute capacity every six months to keep up. Centralized infrastructure translates that demand into power, accelerators, reserved capacity, and revenue forecasts.
An enterprise buyer can put privacy-sensitive or latency-critical work on NPU-equipped PCs and reserve cloud capacity for models that exceed the device. A common toolchain spares application teams from hand-building placement logic across chips, models, and device classes.
The 45-TOPS NPU is easy to miss because it is purchased as part of a PC; the data center is easy to see because it is sold as capacity. They are two destinations within the same workflow. The strategic unit is the stack that reads memory, privacy, latency, power, model quality, and price, then places each task where it belongs.
Qualcomm’s planned path from software to data-center revenue
- June 25, 2026 — Qualcomm announced a nearly $4 billion acquisition of Modular, expected to close in H2 2026.
- 2028 — Production of the Dragonfly C1000 is scheduled to start, with Meta expected to use the data-center CPU.
- 2029 — Qualcomm projects $15 billion in data-center chip sales.
- Fiscal 2029 — Qualcomm forecasts $40 billion in non-handset revenue, up from its prior $22 billion forecast.
Frequently asked questions
Is edge AI disappearing as cloud AI spending grows?
No. Edge capacity is increasingly bundled into PCs, phones, vehicles, cameras, and industrial systems, making it less financially visible than contracted data-center capacity.
What determines whether an AI task runs locally or in the cloud?
Applications must balance privacy, latency, memory, power, reliability, model quality, and price. Routine or sensitive work can remain local, while tasks exceeding device capabilities escalate to larger cloud models.
Why does Qualcomm want to acquire Modular?
Modular’s chip software platform and programming language could help developers deploy and optimize models across different processors. That software layer would connect Qualcomm’s endpoint silicon with its planned data-center products.
How does Qualcomm expect to make money from hybrid inference?
It could earn device-silicon revenue at the endpoint, processor revenue in data centers, and potentially software revenue from the tooling that maps models across both environments.
When does local AI deliver economic value?
It pays when local execution improves the device experience enough to influence purchases or lowers a provider’s remote-inference bill. A local model that produces inadequate results can destroy more user value than it saves in compute costs.