Anthropic's Claude 3.7 Sonnet reportedly cost “a few tens of millions of dollars” to train, similar to Claude 3.5 and cheaper than GPT-4, which cost over $100M
“Assuming Claude 3.7 Sonnet indeed cost just ‘a few tens of millions of dollars’ to train, not factoring in related expenses, it's a sign of how relatively cheap it's becoming to release state-of-the-art models.” — https://techcrunch.com/... X: Ethan Mollick / @emollick : After publishing the post, I was contacted by Anthropic who told me that Sonnet 3.7 would not be considered a 10^26 FLOP model and cost a few tens of millions of dollars, though future models will be much bigger. I updated the post to reflect this, though it doesn't change much. Piotr Cieluchowski / @thisguyoftheai : So, it turns out training the latest AI wonder from Anthropic was a budget affair—just a “few tens of millions” and a sprinkle of computing power. Who knew revolutionary tech could be so affordable? Dive into the genius of Claude 3.7 Sonnet here: https://techcrunch.com/... Peter Wildeford / @peterwildeford : Pretty interesting revelations here: - Claude 3.7 is much smaller model size than Grok yet better? - Trained for “few tens of millions of dollars” (single model training run size, this would be apples-to-apples comparison with DeepSeek's $6M) - Future models will be much bigger Tim Crawford / @tcrawford : Lowering the cost will speed up innovation and potentially open the field to new entrants. Anthropic's latest flagship AI might not have been incredibly costly to train. https://techcrunch.com/... #CIO #AI
Context & Ripple Effects
Anthropic has moved the Claude line from the three-tier Claude 3 launch to Claude 3.5 Sonnet, which it positioned ahead of its prior flagship, making repeated capability upgrades central to the product’s arc. The reported training-cost figure adds an economics datapoint to that release cadence.
The comparison with GPT-4 matters because it separates the cost of one training run from the broader expense of building and operating a model business. Anthropic also said future models will be much bigger, limiting what this figure alone can establish about the direction of total spending.
First-order effects
- If accurate, the reported spend gives Anthropic a lower training-cost benchmark for Claude 3.7 Sonnet than the reported cost of GPT-4, while remaining broadly in line with Claude 3.5.
- The disclosure shifts attention from model scale alone to the cost required for a competitive release; it does not account for related expenses or ongoing serving costs.
Second-order effects
- Rival model developers face more pressure to show that their larger training budgets translate into differentiated performance or commercially viable pricing.
- For enterprise buyers and developers, lower reported training costs strengthen the case for comparing models on useful output and deployment economics rather than treating training scale as a standalone quality signal.
Third-order effects
- If successive capable models can be trained at lower-than-earlier frontier-model cost levels, competition may widen beyond the few firms able to fund the largest individual training runs.
- That shift would not eliminate compute concentration: Anthropic’s indication that future models will be bigger suggests training-cost reductions and escalating scale ambitions can coexist.
The trend: This is one datapoint in the shift from headline model scale toward the economics of delivering useful AI capability at sustainable cost.