Hands-on with ChatGPT's image recognition and voice features: image recognition isn't perfect and refuses to discuss faces, voice is fluid and natural, and more
Kevin Roose / New York Times : X: @kevinroose X: Kevin Roose / @kevinroose : Really like the new ChatGPT image tool for stuff like writing descriptions for FB Marketplace. OpenAI has trained it to refuse face requests, tho, so I wasn't able to get empirical proof that I am better looking than @CaseyNewton [image]
Context & Ripple Effects
OpenAI had kept GPT-4’s image-and-text capability from broad release because of concerns including facial-recognition abuse; this test shows that restriction surfacing as a product-level refusal rather than an unrestricted visual-analysis feature. The earlier safety-limited multimodal rollout provides the immediate backdrop.
The hands-on follows ChatGPT Plus and Enterprise gaining voice prompts and image uploads. That multimodal expansion shifts ChatGPT from a text-only interface toward an assistant that can take in everyday visual and spoken inputs, while retaining explicit boundaries on sensitive image use.
First-order effects
- Users can use voice interaction and image understanding for practical tasks such as drafting marketplace descriptions, but face-related requests are declined and image results remain uneven.
- OpenAI gains a more natural-feeling interaction layer while making its facial-image safety boundary visible in normal use.
Second-order effects
- Use cases that depend on identifying, evaluating, or discussing people from images cannot rely on ChatGPT, directing users toward workflows that do not require facial analysis.
- Competing assistants face a paired product decision: match low-friction voice and vision inputs while defining equally legible safeguards for sensitive visual requests.
Third-order effects
- If multimodal assistants continue to improve, conversational voice and camera inputs could become a standard assistant interface rather than a distinct feature category.
- The durable differentiation may rest not only on model capability but on whether safety limits are predictable enough for people and businesses to build workflows around them.
The trend: This is an early instance of assistants becoming multimodal work surfaces, with convenience expanding alongside embedded safety constraints.