DALL-E and CLIP: OpenAI's Two-Model 2021 Release

OpenAI introduced DALL-E for text-to-image generation and CLIP for matching images with text on January 5, 2021. See their training data, methods and access limits.

Two papers, one publish date

OpenAI published separate blog posts for DALL-E and CLIP on the same day, January 5, 2021. DALL-E was described as a 12-billion-parameter version of the GPT-3 architecture, trained to generate images from a text caption. It ingested up to 1,280 tokens of text and image data as a single stream, the same next-token approach GPT-3 used for language. [2]

The training set for DALL-E combined roughly 250 million text-image pairs pulled from the internet and from existing datasets including Conceptual Captions and YFCC100M. CLIP, described in its own paper, was trained on a separate set of 400 million image-text pairs collected from the web, using a contrastive objective: predicting which caption belongs with which image rather than predicting pixels or words directly. [3]

The armchair that carried the coverage

DALL-E's demonstration images, including an armchair shaped like an avocado, spread quickly because the model appeared to combine unrelated concepts in a plausible way rather than retrieving something close to an existing photo. Coverage at the time noted the model also struggled once a caption named too many objects, gave inconsistent results when a caption was reworded to mean the same thing, and sometimes produced outputs that missed part of the description, such as rendering a window in the wrong material. [2]

Atlas interpretation: OpenAI did not release DALL-E's weights or an API alongside the announcement. What the public got was a curated set of sample outputs and a description of the method, which made the demonstration easy to share and hard to independently test. [2]

CLIP measured itself differently, and that mattered later

CLIP's headline result was zero-shot: without training on any of ImageNet's 1.28 million labeled examples, it matched the accuracy of the original ResNet-50, a model that had been trained directly on that data. That framing measured transfer to a task the model was never tuned for, not a new state-of-the-art score on a fixed benchmark. [3]

Atlas interpretation: A model that scores images against arbitrary text captions is a general-purpose piece of plumbing, not just a classifier. Later text-to-image systems, including OpenAI's own DALL-E 2, used CLIP or CLIP-like models to judge how well a generated image matched its prompt. The armchair image is what people remembered from that day; the scoring function is what kept shipping. [3][2]

Sources

  1. DALL·E: Creating images from text

    OpenAI · Jan 5, 2021

  2. This avocado armchair could be the future of AI

    MIT Technology Review · Jan 5, 2021

  3. Learning Transferable Visual Models From Natural Language Supervision

    arXiv · Mar 4, 2021