OpenAI introduces two new GPT-3 models: CLIP, which classifies images into categories from arbitrary text, and DALL·E, which can generate images from text
With GPT-3, OpenAI showed that a single deep-learning model could be trained to use language in a variety of ways simply by throwing it vast amounts of text.
MIT Technology ReviewWill Heaven
Context & Ripple Effects
OpenAI’s GPT-3 had already drawn attention as a step toward broader language capability, even as high benchmark performance did not settle questions of real-world understanding. CLIP and DALL·E extend that work from text into image recognition and image generation.
The announcement becomes the starting point for a DALL·E product line: DALL-E 2 was built using CLIP, and later versions were positioned inside ChatGPT rather than as standalone model demonstrations.
First-order effects
OpenAI adds two image-focused capabilities to its GPT-3-era research portfolio: CLIP maps images to arbitrary text categories, while DALL·E maps text prompts to generated images.
Researchers and developers gain a foundation for building language-directed image classification and generation workflows, alongside GPT-3’s text capabilities.
Second-order effects
OpenAI’s subsequent image models can compound the value of CLIP and DALL·E rather than begin from scratch, as reflected in the CLIP-based DALL-E 2 successor.
Making DALL·E controllable through ChatGPT shifts image generation toward a conversational workflow, a direction OpenAI outlined with DALL-E 3’s planned ChatGPT integration.
Third-order effects
If image generation is distributed through general-purpose assistants, competition shifts from standalone creative models toward the assistant, access tier, and safety controls surrounding them; OpenAI later paired wider DALL-E 3 access with a safety mitigation stack.
The arc points to multimodal models becoming commercial product capabilities, with generation increasingly governed and embedded in existing user workflows rather than presented solely as research demos.
The trend:Generative AI is moving from separate text and image models toward governed, workflow-native multimodal assistants.
I've gently suggested that this sort of technology will make many of the current startup ideas I see being pitched completely obsolete. Follow this closely. https://twitter.com/...
Tag yourself. I'm Gamer Avo Chair second row, second from left. I'm staying the hell away from Giant Pit Avo Chair bottom row middle https://www.technologyreview.com/ ... https://twitter.com/...
How long until twitter's little GIF widget actually generates the GIFs? I'm sick of no good results for “Exasperated and tired dachshund wearing red bowtie and pointing emphatically at text of 47 U.S.C. § 230.” https://twitter.com/...
CLIP — our new neural network for classifying images into categories based on arbitrary text. Works like GPT-3's semantic search, but for images: https://andrewmayneblog.wordpress.com/ ... Maybe a path towards vision nets that don't make silly mistakes: https://openai.com/... htt…
“DALL·E is the kind of system that Riedl imagined submitting to the Lovelace 2.0 test, a thought experiment... for measuring artificial intelligence. It assumes that one mark of intelligence is the ability to blend concepts in creative ways.” https://twitter.com/...
An NN takes a list of category names, and outputs (in a zero-shot manner) a visual classifier. It beats RN50 on ImageNet zero-shot, while being far more robust to unusual images: https://openai.com/... https://twitter.com/...
They limited the possible inputs to this tool out of fear of what it could show you. Its generative capacity is limitless, infinite, obscene. You can ask it for anything Abby Shapiro's nudes. The true face of any anon. A picture of your future wife. The girl reading this. https:/…
This is absolutely mind melting. Open AI can generate nearly unlimited images of a chair that looks like an avocado. https://www.technologyreview.com/ ...
Wow - this AI can now generate unique images *entirely* from text prompts. Impressive, although somewhat disappointed the available prompts are so... vanilla. Also GG, illustrators https://twitter.com/...
Researchers asked an AI model to draw “an avocado armchair.” The results show how we might make artificial intelligence smarter. https://www.technologyreview.com/ ...
I was going to show a few more examples but honestly it's just worth 5 minutes of your time to go play around with different combinations: https://openai.com/...
And, for fun, generated images of “an illustration of a baby shark in a wizard hat wielding a blue light saber”. (You can try other combinations here: https://openai.com/...) https://twitter.com/...
... i wanna get into this and see if it'll write assessment reports for me. but im gonna start with using it to do all the slides for my next con talk, which will be “how to vet your security vendors” https://twitter.com/...
Consider this your notice, if you're a manga artist: you have N years left before you're out of a job. I wish that I had any grasp whatsoever how to relate N to announcements like these. My initial sense is N=2, wisely adjusted upwards to “actually after the end of the world”. ht…
I, for one, am looking forward to the copyright lawsuits over who holds the copyright for these images (in many cases the answer should be “no one, they're public domain"). https://twitter.com/...