Due to concerns over cloud computing costs, some companies are shifting their data and ML models to on-premises servers and investing in hybrid infrastructure
A quick-service restaurant chain is running its AI models on machines inside its stores to localize delivery logistics. Tweets: @tariqkrim and @counternotions Tweets: @tariqkrim : Why AI and machine learning are drifting away from the cloud - TLDR :cheaper faster and more control https://www.protocol.com/... Kontra / @counternotions : The eternal wheel of compute deployment: on-prem → cloud → colocation → hybrid...machine/deep learning edition. https://www.protocol.com/...
Context & Ripple Effects
This 2022 piece was an early marker of the repatriation wave that later coverage has confirmed rather than reversed. A 2020 analysis of AI-centric businesses had already flagged the structural problem — high compute usage compressing gross margins versus traditional software — and a 2021 study of 50 public software companies put a number on the escape route, finding workload repatriation could halve cloud spend. The quick-service chain running delivery-logistics models on in-store machines shows the pattern reaching beyond software firms into physical operations.
Since then the drift has widened rather than corrected: hardware vendors like Dell and Qualcomm spotted an opening as hyperscalers strained under AI demand (WSJ coverage of the on-prem opportunity), banks and airlines are reviving mainframes to run AI locally (the mainframe revival), and GPU-focused 'neoclouds' like CoreWeave and Lambda Labs emerged as a middle path (SemiAnalysis's neocloud economics deep dive) — evidence the pendulum Kontra described is swinging again.
First-order effects
- Companies running large ML workloads face an immediate cost decision: the restaurant chain's in-store deployment trades cloud elasticity for lower per-inference spend and latency gains at the store level.
- Hardware vendors gain a direct sales channel into enterprises that were previously pure cloud buyers, as on-prem servers become a budget line item again.
Second-order effects
- Hyperscalers are pushed to defend share not just on price but through hybrid offerings, while a new intermediary tier — GPU rental providers like Crusoe, Lambda Labs, and CoreWeave — captures demand from buyers who want cloud-like flexibility without hyperscaler pricing.
- Cost pressure splits AI strategy down the middle, as later coverage shows: some firms standardize on cheaper models to cut spend while others like Shopify mandate frontier models outright, making infrastructure placement a board-level choice rather than an engineering default.
Third-order effects
- If the pattern holds, AI infrastructure settles into a hybrid equilibrium where training stays centralized but inference migrates to the cheapest viable location — store shelves, mainframes, or rented GPUs — reversing two decades of default centralization.
- Compute deployment becomes cyclical infrastructure management, as Kontra's framing suggests: each new workload class (now machine learning) re-litigates the on-prem-versus-cloud question rather than ending it.
The trend: AI workloads are repeating the historical compute cycle — on-prem to cloud to hybrid — with inference economics, not ideology, deciding where models physically run.