Due to safety concerns like facial recognition abuse, OpenAI hasn't widely shipped GPT-4's “multimodal” capability that can respond to images and text prompts
An advanced version of ChatGPT can analyze images and is already helping the blind.
New York TimesKashmir Hill
Context & Ripple Effects
When GPT-4 launched, OpenAI held back its image-and-text mode rather than shipping it alongside the text model, citing facial-recognition abuse among the risks. The related coverage shows how the gap closed: a hands-on test of ChatGPT's image features found the shipped version simply refuses to discuss faces — the guardrail that made release possible.
Blind users are the immediate beneficiaries: an advanced ChatGPT version that can analyze images is already in use as a visual aid, making accessibility the beachhead market for a capability held back from everyone else.
Paying ChatGPT subscribers gain image analysis ahead of the general public, with face-related queries blocked by design rather than left to user discretion.
Second-order effects
Rival labs shipping their own vision features now have to match not just the capability but the refusal policy — 'won't discuss faces' becomes a competitive baseline, not a differentiator.
Accessibility tools built on ChatGPT's vision get a capability moat competitors can only close by accepting the same facial-recognition restrictions, shaping what assistive products can promise.
Third-order effects
If the pattern holds, frontier capabilities ship in gated stages — restricted preview, paid tiers, general release — with safety committees deciding the cadence, turning release timing itself into a governance instrument.
Refusal-based guardrails like the face ban become the template regulators and enterprises expect when evaluating whether a multimodal model is safe to deploy.
The trend: Frontier labs are shifting from shipping capabilities all at once to staged, safety-gated rollouts where the guardrail policy — not just the model — defines the product.
“What OpenAI doesn't want ChatGPT to become is a facial recognition machine.” — “Google offers an opt-out for well-known people who don't want to be recognized, and OpenAI is considering that approach.” — Tech companies love playing the “well, you can always opt-out” game bec…
Blind users have had access for months to a version of ChatGPT that analyzes images. One user described it as extraordinary. But it recently started blurring people's faces. I talked to OpenAI about why: https://www.nytimes.com/...
And in case you are wondering, you can upload images to Bard and use Lens, but if it's a person, that image is removed and Bard explains it cannot provide info about people yet. So Google is in alignment with OpenAI here. [image]
Blind users have had access for months to a version of ChatGPT that analyzes images. One user described it as extraordinary. But it recently started blurring people's faces so it can't give information about them. I talked to OpenAI about why: https://www.nytimes.com/...
OpenAI worries a tool intended for blind people would share things incl gender or emotional state. OpenAI is ‘figuring out how to address these and other safety concerns before releasing the image analysis feature’> These shouldn't be internal decisions ↘️ https://www.nytimes.com…
Yet again confirming my 2001 hypothesis that neural networks would hallucinate until they bridged gap with symbols, new reports show that GPT-4's unreleased visual system hallucinates objects (eg remote control buttons that don't exist; absent faces, etc). https://www.nytimes.com…
OpenAI says it will follow public opinion on whether GPT4 should include face, emotion and gender recognition. Which public? In which nations? Whose opinions will count? https://www.nytimes.com/...
It's surprising that safety concerns have delayed the release by 4+ months, considering that Bing has already rolled it out with what appears to be a simple fix (face blur). A different approach than the much more cavalier release of Code Interpreter. https://twitter.com/...
The reason I predicted this is that facial recognition is “merely” extreme multiclass image classification, so if you train a model to describe images there isn't a strong reason it won't also learn faces. Still, facial recognition being emergent behavior feels creepy and sci-fi.
When @kashhill asked me a while back why OpenAI hadn't released multimodal GPT-4 yet, I wildly speculated that it's because it can do facial recognition and OpenAI doesn't want that out in the wild. Well... https://www.nytimes.com/...