Two papers, one publish date
OpenAI published separate blog posts for DALL-E and CLIP on the same day, January 5, 2021. DALL-E was described as a 12-billion-parameter version of the GPT-3 architecture, trained to generate images from a text caption. It ingested up to 1,280 tokens of text and image data as a single stream, the same next-token approach GPT-3 used for language. [2]
The training set for DALL-E combined roughly 250 million text-image pairs pulled from the internet and from existing datasets including Conceptual Captions and YFCC100M. CLIP, described in its own paper, was trained on a separate set of 400 million image-text pairs collected from the web, using a contrastive objective: predicting which caption belongs with which image rather than predicting pixels or words directly. [3]
The armchair that carried the coverage
DALL-E's demonstration images, including an armchair shaped like an avocado, spread quickly because the model appeared to combine unrelated concepts in a plausible way rather than retrieving something close to an existing photo. Coverage at the time noted the model also struggled once a caption named too many objects, gave inconsistent results when a caption was reworded to mean the same thing, and sometimes produced outputs that missed part of the description, such as rendering a window in the wrong material. [2]
Atlas interpretation: OpenAI did not release DALL-E's weights or an API alongside the announcement. What the public got was a curated set of sample outputs and a description of the method, which made the demonstration easy to share and hard to independently test. [2]
CLIP measured itself differently, and that mattered later
CLIP's headline result was zero-shot: without training on any of ImageNet's 1.28 million labeled examples, it matched the accuracy of the original ResNet-50, a model that had been trained directly on that data. That framing measured transfer to a task the model was never tuned for, not a new state-of-the-art score on a fixed benchmark. [3]
Atlas interpretation: A model that scores images against arbitrary text captions is a general-purpose piece of plumbing, not just a classifier. Later text-to-image systems, including OpenAI's own DALL-E 2, used CLIP or CLIP-like models to judge how well a generated image matched its prompt. The armchair image is what people remembered from that day; the scoring function is what kept shipping. [3][2]
Sources
- DALL·E: Creating images from text
OpenAI · Jan 5, 2021
- This avocado armchair could be the future of AI
MIT Technology Review · Jan 5, 2021
- Learning Transferable Visual Models From Natural Language Supervision
arXiv · Mar 4, 2021