/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Tech companies are racing to run generative AI natively on mobile devices to reduce computing costs, but face hurdles like limited memory and processing power

[running in both] the data centre and locally — otherwise it will cost too much... https://twitter.com/...

Financial Times Richard Waters

Context & Ripple Effects

The economics problem this story sits inside has been visible for years: research showing that AI increasingly requires datacenter-scale computation raised concerns that only a handful of companies could afford frontier work at all. Running every user query through those data centres compounds the bill, which is why the industry's answer has been shrinking the models themselves — Microsoft, Meta, Google and others pitching smaller, cheaper-to-train language models as a way to cut costs and hardware requirements.

First-order effects

  • Pushing inference onto the handset directly cuts the per-query compute bill for Microsoft, Meta and Google, whose generative AI services otherwise scale costs linearly with usage.
  • Phone makers and their chip suppliers now face hard requirements — more memory and faster on-device processing — because the article notes models must run partly locally or the cost 'will be too much'.

Second-order effects

  • The small-model push gets a second engine: models sized for phones serve the same frugal-AI demand already driving startups and researchers without access to top chips to build smaller open-weight models, widening the market for compact architectures beyond cost-cutting alone.
  • Hybrid deployment splits the value chain — whoever controls the on-device runtime captures part of the inference relationship that cloud providers currently own outright.

Third-order effects

  • If native mobile inference proves viable at scale, it loosens the centralization dynamic flagged since 2019, when datacenter-scale compute threatened to concentrate advances among a few large companies — distribution of capable models no longer requires distribution of datacenter capacity.
  • The constraint flips from training-scale bragging rights to efficiency engineering, rewarding players who can compress models rather than those who can buy the most silicon — though whether on-device quality can match hosted frontier models stays genuinely unresolved.

The trend: Generative AI is splitting into hybrid cloud-plus-device architectures, with inference economics — not raw model capability — increasingly dictating where computation physically runs.

Discussion

  • @martijnrasser Martijn Rasser on x
    “Google...managed to run a version of...its latest large language model, on a handset...the move is the latest sign that a form of AI that has required computing resources only found in a data centre is quickly starting to find its way into many more places.” https://www.ft.com/.…
  • @mukulneetika Mukul Kumar on x
    Interesting article on moving AI to mobile devices, which would reduce computing costs and improve speed of AI chatbots. Could fuel the next wave of innovations at chip companies such as Qualcomm and others. https://twitter.com/...
  • @brianroemmele Brian Roemmele on x
    This will be interesting. We already have LLM AI models running on Raspberry Pi systems and some Android systems. It is really a matter of the will to do it and accepting the days of cloud based private personal AI is over. It is never coming back. https://twitter.com/...
  • @tprstly Theo on x
    Artificial Intelligence on the edge, on all your personal devices, so what's the true cost?.... this will be Apple's ‘privacy’ play for future iPhones no doubt. “You need to make the AI hybrid — [running in both] the data centre and locally — otherwise it will cost too much... ht…