A generator developers could run themselves
On August 22, 2022, Stability AI announced Stable Diffusion's public release with links to weights, code and a model card. The recommended weights were v1.4. Downloading a trained model let developers build around the generator itself instead of relying only on a hosted image service. [2]
Stability AI expected the release to use 6.9 GB of GPU memory and recommended NVIDIA hardware. Optimizations for AMD and Apple's M1/M2 chips were still forthcoming. Downloadable weights did not remove the need for compatible hardware and software. [2]
Doing the repeated work in a smaller space
Diffusion models learn to remove noise. At generation time, they start with random noise and repeatedly refine it toward an image. Running those steps directly over every pixel is expensive. The latent-diffusion paper moved the repeated denoising into a smaller, learned representation of an image. [3]
An autoencoder learns to compress images into that representation and reconstruct them. The denoiser works in the compressed space; the decoder converts its final output back into pixels. Compression discards some information, but avoids carrying full-resolution pixels through every denoising step. That tradeoff helped make high-resolution generation less demanding. [3]
The v1.4 model card specifies an eightfold reduction in each spatial dimension: a 512-by-512 image corresponds to a 64-by-64 latent grid with four channels. This describes the size of the representation, not an equivalent multiplier for speed. [4]
CLIP supplied the connection to words
CLIP, announced alongside DALL·E in 2021, learned associations between images and their accompanying text. Stable Diffusion v1.4 reused a pretrained CLIP text encoder to represent a prompt numerically. Cross-attention let the image denoiser use that text representation while constructing the image. [5][4]
Atlas interpretation: The text encoder supplies guidance; it does not draw the picture itself. This separation made an earlier text-and-image research result useful as one component of a generator with a different architecture. [5][3]
Downloadable did not mean unrestricted or reliable
The release used CreativeML OpenRAIL-M. The announcement allowed commercial and non-commercial use, required the license to accompany distributions and be available to service users, and assigned ethical and legal responsibility to users. The package also included an adjustable safety classifier. Operating the generator meant taking responsibility for its deployment. [2]
Atlas interpretation: The v1.4 model card documents failures with legible text, faces and compositions involving several objects, along with an English-caption bias. A prompt can name the right objects without producing the right relationships between them. Generating an attractive image and following a detailed visual specification are separate tests of usefulness. [4]
Sources
- Stable Diffusion Public Release
Stability AI · Aug 22, 2022
- Stable Diffusion Public Release
Stability AI · Aug 22, 2022
- High-Resolution Image Synthesis with Latent Diffusion Models
arXiv · Dec 20, 2021
- Stable Diffusion v1-4 Model Card
CompVis · Sep 8, 2026
- CLIP: Connecting text and images
OpenAI · Jan 5, 2021