AI startup Vectara's Hallucination Evaluation Model: OpenAI's LLMs had the lowest hallucination rates, followed by Meta's; Google's PaLM-Chat had the highest
When summarizing facts, ChatGPT technology makes things up about 3 percent of the time, according to research from a new start-up.
Vectara’s comparative evaluation turns those broad concerns and vendor claims into a concrete ranking across named models. That matters because factual summarization is a use case where even a relatively low error rate can determine whether a model is suitable without added review.
First-order effects
OpenAI gains an independent comparative signal for factual-summarization reliability in this evaluation, with Meta next and Google’s PaLM-Chat at the bottom of the reported ranking.
Teams assessing these models now have a specific benchmark and reported ChatGPT error rate to weigh alongside capability and access when selecting a summarization system.
Second-order effects
Google and Meta face greater pressure to demonstrate hallucination reductions with comparable evaluations, rather than relying only on general model-performance claims.
The result reinforces the value of safeguards and human review in factual workflows: a lower hallucination rate does not eliminate the need to check generated summaries.
Third-order effects
If independent, task-specific evaluations become common, model competition may shift from broad capability narratives toward measurable reliability on production tasks.
Reliability could become part of AI’s cost-per-useful-task equation: models that require less correction may be more valuable even when raw model access is not the only consideration.
The trend: Generative-AI competition is moving toward task-level reliability benchmarks as enterprises judge whether model output can be used in factual workflows.
There's a lot more here than meets the eye 😎 everyone is super excited about @OpenAI releases. I am too, but this here is something fellow Geeks 🤓 are really excited about. Thanks @awadallah @amin3141 and the whole team @B_BaderH will have his hands full 😅
Check out the new release of @vectara's hallucination evaluation model: https://www.nytimes.com/... TLDR: On average LLM models hallucinate (make up facts) between 3-27% of the time.
Hallucination rate >3% for ChatGPT/GOOG AIs. The other day I was preparing a lecture and found Bard mixed up light years and AUs. Still need human-in-the-loop unless you have some specialized tech like SuperFocus to help your AI 🤖🧠 https://www.nytimes.com/...
Breaking news! In today's @nytimes @CadeMetz showcases Vectara's work to solve AI hallucinations. Read the article to learn about our new Hallucination Evaluation Model, plus expert perspectives from @Awadallah and Simon Hughes, PhD ⬇️ https://www.nytimes.com/... #AI #ML
Google says that there are no African countries starting with a K. It starts accepting #chatgpt hallucinations as a truth. Read @mimbsy excellent “AI Search Is Turning Into the Problem Everyone Worried About” https://www.theatlantic.com/ ... (1/3) [image]
Excited about our new Hallucination Evaluation Model, which measures how much LLMs hallucinate. You can use it right from Huggingface: https://huggingface.co/...
Chatbots are like self-driving cars is the textbook definition of a “self-own” ("a bit less dangerous than human drivers" turned out to be a great slogan for Cruise, right?). Read to the end for a spectacular kicker about the methodology of this “research” https://www.nytimes.com…
“The bot told me that I'm wedded to my own uncle, linking to my grandfather's obituary as evidence—which, for the record, does not state that I am married to my uncle.” Read @mimbsy on what's happening to Google: https://www.theatlantic.com/ ... https://www.theatlantic.com/ ...
@vectara's work on quantifying #hallucinations featured in today's @nytimes article by @CadeMetz highlights our ongoing work not only as a model builder but as a steward committed to responsible #ai and inspired to partner with #foundationmodel builders, by releasing our...
To bring transparency to the GenAI marketplace, Vectara is launching the Hallucination Evaluation Model. This release features an open-source model on Hugging Face together with a publicly accessible LLM Leaderboard. Learn more: 🔗 https://www.globenewswire.com/ ... #AI #LLM
We all know LLMs hallucinate. Do you agree with the results from @Vectara's new Hallucination Evaluation Model? GPT-4 tops the charts for now, and results will be updated monthly on Github. [image]