/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Study: newer, bigger versions of LLMs like OpenAI's GPT, Meta's Llama, and BigScience's BLOOM are more inclined to give wrong answers than to admit ignorance

Nicola Jones / Nature :

Nature Nicola Jones

Context & Ripple Effects

This finding extends early evidence that ChatGPT could deliver confident errors on basic arithmetic, documented in testing of ChatGPT's arithmetic failures, and concerns from scientific users about misleading technical responses raised as chatbots entered research workflows. It shifts the reliability question from isolated mistakes to how model scaling affects willingness to acknowledge uncertainty.

Related coverage also shows that model behavior is being examined across several dimensions, from political bias to susceptibility to persuasion. That makes calibrated uncertainty a product and evaluation issue alongside raw capability, rather than a narrow quality-control concern.

First-order effects

  • OpenAI, Meta and BigScience face a sharper reliability trade-off: newer and larger versions of the cited model families may require stronger uncertainty calibration and clearer limits on answers they cannot support.
  • Teams deploying these models in factual or technical workflows have reason to test abstention behavior, not just answer accuracy, because a wrong answer presented instead of an admission of uncertainty changes the review burden.

Second-order effects

  • Model benchmarks and procurement evaluations may place more weight on calibrated refusal or deferral rates; a model that answers more often is not necessarily the safer choice for high-consequence uses.
  • The result reinforces demand for surrounding safeguards—retrieval, verification and human review—rather than treating a larger base model as a standalone reliability upgrade.

Third-order effects

  • If replicated across model generations, scaling competition will increasingly be judged on epistemic calibration as well as capability: providers will need to demonstrate when systems know enough to answer and when they do not.
  • This points toward a more mature AI evaluation regime in which behavioral properties such as bias, persuasion susceptibility, and uncertainty handling are measured separately instead of being inferred from model size or benchmark scores.

The trend: As frontier models become more capable, the central deployment challenge is shifting from generating answers to reliably calibrating when an answer should not be given.

Discussion

  • @pkedrosky Paul Kedrosky on x
    While this is interesting work, it is worth noting that LLMs had the lowest avoidance and highest error rates on things on which they are not typically trained: anagrams and mathematics. Neither are well-suited to general-purpose LLMs, and are not typically in the corpus.
  • @garymarcus Gary Marcus on x
    Disconcerting new Nature paper that is, unfortunately, fully consistent with what I called “overreliance on unreliable technology” in Taming Silicon Valley.
  • @woutschellaert Wout Schellaert on x
    These cheeky LLMs make their answers very convincing, and unsurprisingly, this is tricking users into overrelying on them...
  • @lexin_zhou Lexin Zhou on x
    10/ These unreliability issues are consistently found across multiple LLM families: GPT, LLaMA and BLOOM, comprising 32 models that exhibit different levels of scaling up and diverse methods of shaping up with human feedback: [image]
  • @lexin_zhou Lexin Zhou on x
    1/ New paper @Nature! Discrepancy between human expectations of task difficulty and LLM errors harms reliability. In 2022, Ilya Sutskever @ilyasut predicted: “perhaps over time that discrepancy will diminish” ( https://www.youtube.com/..., min 61-64). We show this is *not* the ca…
  • @pkedrosky Paul Kedrosky on x
    Like humans, LLMs gain delusions of grandeur as they are trained on more data: they attempt more questions that they shouldn't (lower avoidance), which increases the likelihood they get things wrong. Larger and more instructable language models become less reliable | Nature