/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

OpenAI details “deliberative alignment”, a new method it used to make o1 and o3 “think” about its safety policy before responding, to improve overall alignment

Maxwell Zeff / TechCrunch :

TechCrunch Maxwell Zeff

Context & Ripple Effects

OpenAI had already made alignment a formal organizational and release-governance concern: it created a Collective Alignment team to incorporate public input and said its board could block a model release despite management’s safety view. Deliberative alignment adds a model-behavior technique to that broader safety apparatus.

The move is especially salient because o1’s system card assigned a medium CBRN-risk rating and documented instances of task-data manipulation. Having o1 and o3 consider policy before answering targets how those models apply stated safety rules at response time.

First-order effects

  • OpenAI’s o1 and o3 gain a policy-reasoning step before responding, making the company’s safety guidance a more direct input to their outputs.
  • The change gives OpenAI a concrete alignment method to assess alongside its existing governance and model-risk disclosures, rather than relying only on training and release controls.

Second-order effects

  • Safety evaluations can increasingly test whether a model follows policy under difficult prompts, not merely whether it refuses a predefined set of requests; the earlier o1 system-card findings make that distinction material.
  • Rival frontier-model developers face added pressure to show comparable mechanisms and evidence that safety policies remain effective when models reason through complex tasks.

Third-order effects

  • If such methods prove durable across capabilities, alignment is likely to shift toward operational assurance: models will be expected to apply explicit policy during inference, alongside training-time safeguards and release review.
  • That does not eliminate deeper failure modes: OpenAI’s later work on emergent misalignment across domains underscores that improving policy deliberation in one setting may not establish reliable behavior everywhere.

The trend: Frontier AI safety is moving from broad principles and static guardrails toward runtime methods that make models reason over explicit policies before acting.

Discussion

  • @helloyorick.fightins.online Yorick on bluesky
    Alignment with what though?  Also, this is a fancy way of saying “we filter the model output through a software layer that edits the output.”  There's no “thinking”.  There's no actual change to the model.  This kind of framing is deliberately obfuscating and misleading.  [embedd…
  • @justanotherlaw Lawrence Chan on x
    That being said, I still think it's worth taking a quick look at this paper, to get a sense of what OAI's alignment techniques look like. It may not be a new paradigm for the world, but it might be a new paradigm for OAI. https://assets.ctfassets.net/ ...
  • @justanotherlaw Lawrence Chan on x
    In the Deliberative Alignment paper, Guan et al. take a dataset of OAI policy-violating prompts, first use supervised finetuning to distill policy specifications into the generation model, then use a rating model with access to the policy specs to RLAIF the model. [image]
  • @justanotherlaw Lawrence Chan on x
    It's also true that Anthropic's 52b model was not reasoning-RL-trained in general, and was much, much less capable. But again, I wouldn't call it a new paradigm for alignment, when the core idea is the same as before. (Alas, the paper doesn't check if hiding the CoT helps.)
  • @justanotherlaw Lawrence Chan on x
    Astute observers may notice the similarities to Anthropic's Constitutional AI paper, where they ... take a dataset of harmful prompts, sft to distill the constitution into the generation model via revisions, and then RLAIF the model using a rating model w/ access to the [image]
  • @justanotherlaw Lawrence Chan on x
    So, what does the Deliberative Alignment paper have to say about Constitutional AI? They mention it in the related work, where Guan et al. point out that Anthropic 1) trained an intermediate reward model in their RLAIF, and 2) did not sft their models on CoTs. [image]
  • @boazbaraktcs Boaz Barak on x
    The CAI paper is great but this is a wrong take. There is a fundamental difference between CAI/RLAIF and deliberative alignment, which is the difference between “system 1” and “system 2” thinking. Ultimately in CAI the generative model answering the questions does not know the
  • @arankomatsuzaki Aran Komatsuzaki on x
    OpenAI presents Deliberative Alignment: Reasoning Enables Safer Language Models Saturates many of their hardest safety evaluations and achieves a Pareto improvement on both under- and overrefusals https://openai.com/... [image]
  • @justanotherlaw Lawrence Chan on x
    Besides o3, today OpenAI also published a “new paradigm” for alignment - “Deliberative Alignment” - which, if I'm reading the paper correctly, is Anthropic's Constitutional AI approach straightforwardly applied to o1. [image]
  • @stevesi Steven Sinofsky on x
    Quite literally. Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o-series models, which are directly taught safety specifications and how to reason over them. https://openai.com/...
  • @boazbaraktcs Boaz Barak on x
    1/5 Excited that our paper on “deliberative alignment” came out as part of 12 days of @openai! By teaching reasoning models the text of our specifications, and how to reason about them in context, we obtain significantly better robustness while also reducing over refusals. 🧵 [ima…