Sales data: AI evaluation startups have modest revenue; Weights & Biases hit $50M ARR by Dec. 2024, but only 2% of that was from its AI evaluation product Weave
Startups whose products evaluate artificial intelligence to make sure it isn't making things up or delivering unsafe results … Bluesky: @edzitron.com . X: @srimuppidi Bluesky: Ed Zitron / @edzitron.com : lmfao, absolutely nobody is making money in AI [embedded post] X: Sri Muppidi / @srimuppidi : NEW: In today's AI Agenda, we break down revenue for AI evaluation startups like Braintrust, Galileo, and Patronus. These startups evaluate AI to make sure it isn't making things up or delivering unsafe results. Read more in @theinformation: https://www.theinformation.com/ ... [image]
Context & Ripple Effects
The report puts a revenue test to the AI-evaluation pitch: Weights & Biases' broader platform had reached $50M in ARR, while Weave contributed only 2% of that total. That gap matters because AI adoption has spread faster than clearly demonstrated productivity gains, as evidence of productivity impact remained thin.
It also sharpens the contrast between capital flowing into AI startups and the ability of specialized tooling to monetize. Later coverage of Arena's rapid revenue run rate for evaluation analytics suggests the category's commercial outcomes may vary sharply by product positioning and distribution.
First-order effects
- Braintrust, Galileo, Patronus, and Weights & Biases face immediate evidence that standalone evaluation products have not yet translated into large revenue streams.
- For Weights & Biases, Weave appears to be a small add-on to its core business rather than a material driver of the company's reported ARR.
Second-order effects
- Evaluation vendors will face greater pressure to prove that their products reduce costly failures or become embedded in customers' existing development workflows, rather than selling safety checks as a separate budget line.
- The modest revenue signal makes distribution more consequential: vendors with an established developer platform or broader analytics suite can attach evaluation capabilities more readily than point-solution entrants.
Third-order effects
- If this pattern persists, AI evaluation may consolidate into broader model-development, observability, and analytics platforms instead of supporting many large standalone software businesses.
- The category's long-term economics will depend on whether buyers treat evaluation as a required operating layer for production AI, or as an intermittent feature purchased only when model risk becomes visible.
The trend: AI infrastructure is moving from funding-led category creation toward a harder test of whether specialized tools can secure recurring budgets and durable distribution.