Wafer, which makes AI agents that optimize open-source models for a business's workload, raised a $40M Series A, a source says at a $200M+ valuation
Context & Ripple Effects
Wafer's financing sits alongside a 2026 wave of agent-led infrastructure companies: Factory raised a Series C for coding agents that select models by task complexity, while ChipAgents raised an expanded round for agents that accelerate chip design. Earlier, Granulate raised funding for AI-based computing-infrastructure optimization.
The common thread is software that makes expensive AI and compute capacity more productive, rather than selling a single underlying model. Wafer applies that layer to enterprises using open-source models for their own workloads.
First-order effects
- Wafer gains $40 million to build and commercialize its agent-driven optimization layer for enterprise open-source-model workloads, at a reported valuation above $200 million.
- The round gives Wafer a larger financial base for competing on inference speed and GPU efficiency rather than on ownership of a proprietary foundation model.
Second-order effects
- Model-routing and infrastructure-optimization vendors, including Factory and Granulate, face a sharper contest to become the software layer enterprises use to translate workload requirements into model and compute choices.
- For enterprise buyers, optimization vendors create another route to lower the operating cost of open-source models, making performance-per-dollar a more central procurement criterion.
Third-order effects
- If enterprise AI spending continues to be constrained by inference and GPU economics, value may concentrate in orchestration software that selects, tunes, and operates models across heterogeneous workloads.
- The pattern points toward AI infrastructure platformization: the durable control point is increasingly the layer that manages models and compute, not necessarily the model developer or chip supplier alone.
The trend: Enterprise AI infrastructure is shifting toward agentic software layers that optimize model selection and compute use for specific workloads.