A study of 14 open-source LLMs: reasoning models use exponentially more energy, likely producing more greenhouse gas emissions, but don't answer more accurately
When it comes to artificial intelligence, more intensive computing uses more energy, producing more greenhouse gases.
Context & Ripple Effects
Open-source LLMs have expanded the set of models available for independent scrutiny since the earlier wave of open-model releases. This study adds an inference-side comparison to longstanding calls for transparent AI carbon accounting.
It also complicates the assumption that more compute necessarily buys more useful answers: prior coverage found larger models can be more likely to answer incorrectly than admit uncertainty when they do not know.
First-order effects
- Developers evaluating the 14 studied open-source models now have evidence that reasoning-oriented variants can carry far higher energy use without a corresponding accuracy gain.
- For deployments where electricity use or emissions matter, the result makes model-selection trade-offs more explicit: extra reasoning compute is not, on this evidence, an automatic quality upgrade.
Second-order effects
- Model builders and evaluators face pressure to report energy use alongside accuracy benchmarks, rather than presenting capability scores alone.
- Teams operating open models may favor lighter models or restrict reasoning modes to tasks where their value can be demonstrated, reducing avoidable inference demand.
Third-order effects
- If comparable results persist across evaluations, AI competition will increasingly turn on reasoning economics—quality per unit of inference energy—not simply on maximum test performance.
- The finding strengthens the case for standardized inference-energy and emissions disclosure, though broader conclusions depend on replication across workloads, hardware, and electricity sources.
The trend: AI is moving toward an efficiency-first phase in which the economic and environmental cost of inference is judged alongside model capability.