/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI researchers say training AI with process supervision, which rewards the thought process rather than the outcome, could help prevent hallucinations

- AI hallucinations occur when models like OpenAI's ChatGPT or Google's Bard fabricate information entirely.

CNBC Hayden Field

Context & Ripple Effects

The report follows expert accounts of confabulation as models filling gaps in their training data with plausible language, framing hallucinations as a training-and-evaluation problem rather than merely a user-interface flaw. Earlier analysis of how ChatGPT fills informational gaps provides the mechanism this proposal targets.

It also anticipates a later OpenAI argument that conventional evaluation can reward guessing instead of uncertainty. The later critique of guess-rewarding model evaluation makes process supervision relevant as an alternative incentive design, not simply a filter for bad outputs.

First-order effects

  • OpenAI researchers shift attention toward training signals that assess intermediate reasoning, rather than rewarding only a correct-looking final answer.
  • For model builders, hallucination reduction becomes partly a supervision-design task: systems must distinguish sound reasoning and appropriate uncertainty from confident fabrication.

Second-order effects

  • Developers competing on reliability face pressure to build evaluation methods that test reasoning quality and uncertainty, not just answer-level accuracy; benchmark claims based solely on outputs become less complete.
  • Enterprise adopters gain a clearer basis for demanding assurance evidence around how a model reaches answers, particularly in workflows where unsupported claims are costly.

Third-order effects

  • If process-based evaluation proves scalable, AI reliability could move from post-deployment guardrails toward training-time assurance, making supervision data and evaluation design more important competitive inputs.
  • The broader constraint remains that more visible reasoning does not by itself establish truth; durable AI governance will likely require combining training incentives with independent testing and workflow controls.

The trend: This is an early signal of operational AI assurance shifting from judging model answers alone to shaping—and auditing—the incentives behind them.

Discussion

  • @openai @openai on x
    We trained an AI using process supervision — rewarding the thought process rather than the outcome — to achieve new state-of-art in mathematical reasoning. Encouraging sign for alignment of advanced AIs: ...https://openai.com/...
  • @janleike Jan Leike on x
    Really interesting result on using LLMs to do math: Supervising every step works better than only checking the answer. Some thoughts how this matters for alignment 👇 https://openai.com/...
  • @gdb Greg Brockman on x
    Initial results with process supervision — training an AI by rewarding it for coming up with a correct thought process, not just outputting the right answer. Great results in mathematical reasoning: https://twitter.com/...
  • @hnadim87 Hussain Nadim on x
    The end product is only as good as the process it follows. Be it training AI, governance or managing crisis. https://twitter.com/...
  • @mchammer MC Hammer on x
    “Rewarding the thought process” I believe will impact other domains. Not to mention it rings philosophical. Philosophy + Math The “thought process rewarded” is a stepping stone on the pathway to “AI Awareness”. I say “AI Awareness because unlike “Human Awareness” the AI... https:…
  • @dejanseo Dejan on x
    [This is a big deal.] New @OpenAI model sets a new standard in math problem-solving by rewarding each correct reasoning step, not just the final answer. This improves performance and aligns the AI's reasoning with human-endorsed thought processes. https://cdn.openai.com/...
  • @borismpower Boris Power on x
    Process supervision makes intuitive sense too! https://twitter.com/...
  • @sama Sam Altman on x
    really exciting process supervision result from our mathgen team. positive sign for alignment. https://twitter.com/...
  • @alexandrosm Alexandros Marinos on x
    While the doomers are screaming, alignment researchers keep making progress, day in and day out. https://twitter.com/...
  • @spolu Stanislas Polu on x
    Next step is RLHF with the process supervision RM. They more than likely tried it and the fact that it is not reported, prolly means that it works well 🙃 But with the release of the dataset we can try it for ourselves on Llama or Falcon 🔥 https://twitter.com/...
  • @sarbjeetjohal Sarbjeet Johal on x
    .@Microsoft-backed @OpenAI has released a new research paper with ideas on how to help prevent the #AI #hallucinations. ⁦@theCUBE⁩ ⁦@EvanKirstel⁩ ⁦@BillMew⁩ https://www.cnbc.com/...
  • @keerthanpg Keerthana Gopalakrishnan on x
    Earlier I tweeted that intervention is critical for solving robotics. OpenAI just showed intervention(process supervision) improves math skills 😇 I'd bet that many problems unsolvable by LLMs today will ease with intervention and give interpretability of reasoning for free https:…
  • @0xfoobar @0xfoobar on x
    i'm somewhat inclined to believe this is falsy regulatory cope but very interesting if true on a pure capabilities basis https://twitter.com/... [image]
  • @nyxtelius Lachlan on x
    Access to models → Alignment progress https://twitter.com/...
  • @infoxiao Xiao Ma on x
    It's about the journey and not about the destination they say... https://openai.com/...
  • @haydenfield Hayden Field on x
    OpenAI is taking up the mantle against AI “hallucinations” or falsehoods, it announced Wednesday, with a newer AI training method. The research comes at a time when misinfo from AI systems is more hotly debated than ever. Experts expressed their skepticism.https://www.cnbc.com/ .…
  • @glenngabe Glenn Gabe on x
    “Process Supervision” vs. “Outcome Supervision” -> OpenAI is pursuing a new way to fight A.I. ‘hallucinations’ “Train AI models to reward themselves for each individual correct step of reasoning when they're arriving at an answer.” https://www.cnbc.com/... [image]