/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A review of Claude, Gemini, GPT-4, Llama 2, and Mixtral found the models often gave inaccurate or misleading election information; GPT-4 was least inaccurate

Proof

Context & Ripple Effects

This review arrives as model competition was increasingly framed through broad capability benchmarks, with several systems reported to be nearing or surpassing GPT-4 on those measures. It highlights that benchmark progress and dependable answers on high-stakes civic questions are not interchangeable.

Later audits reinforced the concern: leading chatbots were found to repeat claims from a pro-Kremlin disinformation network in a NewsGuard audit of chatbot misinformation. The issue is therefore less a single model ranking than a recurring reliability test for widely deployed assistants.

First-order effects

  • Users seeking election information from Claude, Gemini, GPT-4, Llama 2, or Mixtral face a meaningful risk of receiving inaccurate or misleading guidance; GPT-4's comparatively better result does not make it a verified election-information source.
  • The evaluated model providers gain a concrete safety and trust benchmark beyond general capability comparisons, putting pressure on them to identify and correct election-specific failure modes.

Second-order effects

  • Model comparisons will need to weigh factual reliability in sensitive domains alongside performance claims; this matters when providers promote benchmark gains such as Llama 3's claimed advantages over similarly sized rivals.
  • Organizations that place chatbots in consumer-facing information flows may need stronger sourcing, escalation, or retrieval safeguards, because an incorrect answer can create downstream editorial and reputational costs.

Third-order effects

  • If repeated audits continue to find election-related errors across leading models, political evenhandedness and factual grounding are likely to become durable evaluation categories, not optional safety add-ons—as suggested by the later open method for scoring political evenhandedness.
  • The industry may increasingly separate general-purpose chat from high-stakes information services, with trust depending on evidence, monitoring, and correction processes rather than a model's overall ranking alone.

The trend: Generative AI competition is shifting from headline benchmark scores toward demonstrable reliability, provenance, and governance in high-consequence information settings.

Discussion

  • @jeffjarvis @jeffjarvis on x
    The more interesting question is how accurate LLMs are when restricted to a limited corpus (e.g., an interview transcript, a meeting transcript, a court filing, a book). We KNOW raw models cannot general questions (it ain't AGI, folks!). Can they summarize well?
  • @jeffjarvis @jeffjarvis on x
    We KNOW that generative AI models without an unconstrained corpus of data have NO sense of meaning or fact & thus should not be used by news sites or search engines to write news or answer factual queries. I'm not sure what the point is of “testing” them. https://www.proofnews.or…
  • @geomblog Suresh Venkatasubramanian on x
    One thing that was very heartening though was seeing election officials approach AI initially with some trepidation, and then have them realize that in fact they WERE the best authority on election related information and that the AI systems were not even that good. 2/2
  • @geomblog Suresh Venkatasubramanian on x
    A fascinating report out from @JuliaAngwin @alondra and Proof News about the use of chatbots to get reliable information about elections and voting. The short answer - DON'T! 1/2 https://www.proofnews.org/...
  • @emilybell Emily Bell on x
    And here is the chaser - independent research showing *just how unreliable AI is for truth-critical tasks like informing people about their right to vote etc https://www.proofnews.org/...