Study: newer, bigger versions of LLMs like OpenAI's GPT, Meta's Llama, and BigScience's BLOOM are more inclined to give wrong answers than to admit ignorance
Nicola Jones / Nature :
Context & Ripple Effects
This finding extends early evidence that ChatGPT could deliver confident errors on basic arithmetic, documented in testing of ChatGPT's arithmetic failures, and concerns from scientific users about misleading technical responses raised as chatbots entered research workflows. It shifts the reliability question from isolated mistakes to how model scaling affects willingness to acknowledge uncertainty.
Related coverage also shows that model behavior is being examined across several dimensions, from political bias to susceptibility to persuasion. That makes calibrated uncertainty a product and evaluation issue alongside raw capability, rather than a narrow quality-control concern.
First-order effects
- OpenAI, Meta and BigScience face a sharper reliability trade-off: newer and larger versions of the cited model families may require stronger uncertainty calibration and clearer limits on answers they cannot support.
- Teams deploying these models in factual or technical workflows have reason to test abstention behavior, not just answer accuracy, because a wrong answer presented instead of an admission of uncertainty changes the review burden.
Second-order effects
- Model benchmarks and procurement evaluations may place more weight on calibrated refusal or deferral rates; a model that answers more often is not necessarily the safer choice for high-consequence uses.
- The result reinforces demand for surrounding safeguards—retrieval, verification and human review—rather than treating a larger base model as a standalone reliability upgrade.
Third-order effects
- If replicated across model generations, scaling competition will increasingly be judged on epistemic calibration as well as capability: providers will need to demonstrate when systems know enough to answer and when they do not.
- This points toward a more mature AI evaluation regime in which behavioral properties such as bias, persuasion susceptibility, and uncertainty handling are measured separately instead of being inferred from model size or benchmark scores.
The trend: As frontier models become more capable, the central deployment challenge is shifting from generating answers to reliably calibrating when an answer should not be given.