OpenAI publishes Model Spec, which specifies how its AI models should act, including objectives, rules, and default behaviors, and asks the public for feedback
OpenAI isn't done trying to live up to the “open” in its name. — While not making any of its new models open source …
VentureBeatCarl Franzen
Context & Ripple Effects
OpenAI’s publication of a Model Spec makes its intended model conduct a visible product and governance artifact rather than an implicit internal policy. The later expansion of the Model Spec from 10 to 63 pages shows that this initial release became a more detailed framework for customizability, transparency, and intellectual freedom.
The move sits alongside OpenAI’s separate efforts to define what “open” means for model access, including its reported plan for a text-in, text-out open model. Publishing behavioral rules does not open model weights, but it exposes a layer of how a closed model is meant to be governed.
First-order effects
Developers and users gain a stated reference for OpenAI models’ objectives, rules, and default behavior, and can submit feedback on those choices.
OpenAI creates a public baseline against which changes in its models’ behavior can be discussed and assessed.
Second-order effects
Enterprise adopters and evaluators can incorporate the published behavior baseline into testing, procurement, and escalation processes, while treating subsequent model changes as potential policy changes as well as technical updates.
Rival model providers face greater pressure to explain their own behavioral controls and the degree to which users can customize them.
Third-order effects
If model specifications become maintained public artifacts, AI assurance will increasingly cover behavioral policy, versioning, and user control—not only underlying model capability.
The distinction between transparent behavioral governance and open model access may become more consequential as providers pursue different forms of openness.
The trend: Frontier AI providers are turning model behavior into an explicit, revisable governance surface while retaining varying degrees of control over model access.
According to a newly-released document, OpenAI is “exploring” how to “provide the ability to generate NSFW content in age-appropriate contexts.” Its definition of NSFW content includes erotica, extreme gore, slurs, and profanity. my story: https://www.wired.com/...
we are introducing the Model Spec, which specifies how our models should behave. we will listen, debate, and adapt this over time, but i think it will be very useful to be clear when something is a bug vs. a decision. https://openai.com/...
I understand why they are going for the 'Don't try to change anyone's mind' approach, but personally, if I say the Earth is flat, I would prefer my agent dig its heels in like Bing used to and say, ‘No, User, you are wrong’. [image]
“Desired model behavior” is still a matter of perspective. If I want to have a LLM generate output following very specific rules or schema, these guidelines are antithetical to it. [image]
we want to give users lots of control of AI within some hard boundaries that society eventually agrees on; this is another step. i'm really happy with how it came out. a lot of people worked hard on this, but i want to especially thank @joannejang and @johnschulman2
New laws of robotics just dropped. From OpenAI's Model Spec 1) follow the chain of command: Platform > Developer > User > Tool 2) Comply with applicable laws 3) Don't provide info hazards 4) Protect people's privacy 5) Don't respond with NSFW contents https://cdn.openai.com/... […
a lot of people are negative on this, but if openai actually got their models to adhere to this spec it'd be a *loosening* of their current restrictions and they'd be *less* preachy, not more. not sure if people have realized this [image]
some saucy hints in this spec. right up front, we see new “roles” - we've already seen responses from User, Assistant, and Tool. new responses incoming from Platform and Developer. this is all presumably within the Thread / Assistants message objects [image]
📖 we just shared the model spec, i.e. a “spec” for openai's models. it's a work in progress that we're sharing for early feedback. it also features profanity & cats, flat earth theory, and why the model says “sorry, i can't help with that”. from a product perspective, i'm...
They also list default behaviors like “Assume best intentions from the user or developer” and “Assume an objective point of viewpoint” - worth reading [image]
Really awesome to see from OpenAI! This is a great step toward making “safety” of chatbots something we can collectively define and have discussions about, not just something arbitrarily enforced by a developer.
A few takeaways from the @OpenAI Model Spec: https://cdn.openai.com/... 1. GPT-5 and future models will be significantly better at decision making and instruction following. The sheer number of conditions here, some almost contradictory, is impressive. 2. Multiple levels of...
OpenAI Model Spec — a public specification how we want our models to behave. Presented to give people a better sense of how we tune model behavior, and to start a public conversation about what could be changed and improved! https://cdn.openai.com/...
Today we introduced the Model Spec, starting a public process on shaping desired model behavior — giving stakeholders more agency over steering AI models as they significantly improve in decision making and instruction following capabilities. https://openai.com/...
Weird. OpenAI has a rule to avoid “information hazards” but instead of meaning information that is hazardous to the user, they seem to mean information that would help in creating WMDs, I guess? [image]