/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

AI startup Vectara's Hallucination Evaluation Model: OpenAI's LLMs had the lowest hallucination rates, followed by Meta's; Google's PaLM-Chat had the highest

When summarizing facts, ChatGPT technology makes things up about 3 percent of the time, according to research from a new start-up.

New York Times Cade Metz

Context & Ripple Effects

Hallucination risk had already emerged as a central limitation of chatbot-style AI: earlier coverage described systems filling gaps with plausible-sounding output rather than grounded facts when requested information is absent from training data. OpenAI researchers had also pointed to process supervision as a possible route to reducing hallucinations, while Anthropic positioned Claude around a similar reliability claim when it introduced Claude as less prone to hallucination.

Vectara’s comparative evaluation turns those broad concerns and vendor claims into a concrete ranking across named models. That matters because factual summarization is a use case where even a relatively low error rate can determine whether a model is suitable without added review.

First-order effects

  • OpenAI gains an independent comparative signal for factual-summarization reliability in this evaluation, with Meta next and Google’s PaLM-Chat at the bottom of the reported ranking.
  • Teams assessing these models now have a specific benchmark and reported ChatGPT error rate to weigh alongside capability and access when selecting a summarization system.

Second-order effects

  • Google and Meta face greater pressure to demonstrate hallucination reductions with comparable evaluations, rather than relying only on general model-performance claims.
  • The result reinforces the value of safeguards and human review in factual workflows: a lower hallucination rate does not eliminate the need to check generated summaries.

Third-order effects

  • If independent, task-specific evaluations become common, model competition may shift from broad capability narratives toward measurable reliability on production tasks.
  • Reliability could become part of AI’s cost-per-useful-task equation: models that require less correction may be more valuable even when raw model access is not the only consideration.

The trend: Generative-AI competition is moving toward task-level reliability benchmarks as enterprises judge whether model output can be used in factual workflows.

Discussion

  • @ahmedrezat Ahmed Reza on x
    There's a lot more here than meets the eye 😎 everyone is super excited about @OpenAI releases. I am too, but this here is something fellow Geeks 🤓 are really excited about. Thanks @awadallah @amin3141 and the whole team @B_BaderH will have his hands full 😅
  • @mccannatron Chris McCann on x
    Check out the new release of @vectara's hallucination evaluation model: https://www.nytimes.com/... TLDR: On average LLM models hallucinate (make up facts) between 3-27% of the time.
  • @hsu_steve Steve Hsu on x
    Hallucination rate >3% for ChatGPT/GOOG AIs. The other day I was preparing a lecture and found Bard mixed up light years and AUs. Still need human-in-the-loop unless you have some specialized tech like SuperFocus to help your AI 🤖🧠 https://www.nytimes.com/...
  • @vectara @vectara on x
    Breaking news! In today's @nytimes @CadeMetz showcases Vectara's work to solve AI hallucinations. Read the article to learn about our new Hallucination Evaluation Model, plus expert perspectives from @Awadallah and Simon Hughes, PhD ⬇️ https://www.nytimes.com/... #AI #ML
  • @henkvaness @henkvaness on x
    Google says that there are no African countries starting with a K. It starts accepting #chatgpt hallucinations as a truth. Read @mimbsy excellent “AI Search Is Turning Into the Problem Everyone Worried About” https://www.theatlantic.com/ ... (1/3) [image]
  • @ofermend Ofer Mendelevitch on x
    Excited about our new Hallucination Evaluation Model, which measures how much LLMs hallucinate. You can use it right from Huggingface: https://huggingface.co/...
  • @nigelesl Nigel Caplan on x
    Chatbots are like self-driving cars is the textbook definition of a “self-own” ("a bit less dangerous than human drivers" turned out to be a great slogan for Cruise, right?). Read to the end for a spectacular kicker about the methodology of this “research” https://www.nytimes.com…
  • @dlberes Damon Beres on x
    “The bot told me that I'm wedded to my own uncle, linking to my grandfather's obituary as evidence—which, for the record, does not state that I am married to my uncle.” Read @mimbsy on what's happening to Google: https://www.theatlantic.com/ ... https://www.theatlantic.com/ ...
  • @vectara @vectara on x
    @vectara's work on quantifying #hallucinations featured in today's @nytimes article by @CadeMetz highlights our ongoing work not only as a model builder but as a steward committed to responsible #ai and inspired to partner with #foundationmodel builders, by releasing our...
  • @vectara @vectara on x
    To bring transparency to the GenAI marketplace, Vectara is launching the Hallucination Evaluation Model. This release features an open-source model on Hugging Face together with a publicly accessible LLM Leaderboard. Learn more: 🔗 https://www.globenewswire.com/ ... #AI #LLM
  • @alphawatchai @alphawatchai on x
    We all know LLMs hallucinate. Do you agree with the results from @Vectara's new Hallucination Evaluation Model? GPT-4 tops the charts for now, and results will be updated monthly on Github. [image]
  • r/ChatGPT r on reddit
    Chatbots May ‘Hallucinate’ More Often Than Many Realize