GPT Image 2: Reasoning, Multilingual Text, 2K Output, API

OpenAI’s gpt-image-2 reasons before generating or editing images, demonstrates non-Latin text, supports 2K output, and uses tiered API access limits.

A showcase, not a spec sheet

OpenAI's own announcement is built almost entirely around example images rather than benchmark tables: dozens of generated pieces spanning photography, illustration, manga, pixel art and editorial layout, framed as evidence of "greater precision and control" in turning a prompt into a finished image. [1]

The multilingual claim is demonstrated the same way: examples set in Japanese, Korean, Arabic, Chinese, Hindi, Cyrillic, Bengali and Greek scripts, including manga pages, South Asian typography and Korean advertising layouts, rather than a stated accuracy figure for text rendering. [1]

Atlas interpretation: That is a deliberate choice of evidence, not a gap. A claim like "renders Devanagari correctly" is hard to verify from a press release regardless of format, and a gallery of examples in scripts a Latin-only model would visibly mangle is a more legible demonstration to a general reader than a benchmark score would be. It does mean the summary's specific figures, such as the 2K output ceiling and the eight-image set limit, are not confirmed on the announcement page itself; the model's own documentation is the source for those. [1][2]

gpt-image-2, gated by tier

The model behind the announcement is called gpt-image-2 in OpenAI's API, reachable through the image generation and image editing endpoints. It accepts both text and image inputs and returns image outputs; it does not support streaming, function calling, structured outputs, fine-tuning or predicted outputs. [2]

OpenAI's documentation rates the model in a higher performance tier and a medium speed tier, and caps image requests per minute by account tier, from 5 up to 250 depending on where a developer's account sits. Access is account-gated rather than open to any API key on sign-up. [2]

Atlas interpretation: Rate limits scaled by account tier, rather than a flat cap, is the same access model OpenAI has used for its text models since GPT-4: capability ships to everyone with API access at once, but the volume anyone can actually draw on is rationed by account history. An image model is more compute-hungry per request than a chat completion, which is likely why the ceiling here starts as low as five requests a minute at the bottom tier. [2]

The reasoning stack, not a new one

Atlas interpretation: OpenAI's image line runs from DALL·E through DALL·E 2, both diffusion models with no text-understanding step in front of them. The summary's claim, that OpenAI put its reasoning stack in front of image generation, describes a lineage change rather than a new research direction: the same chain-of-thought reasoning OpenAI built for GPT-5 and its text-based successors is what now runs before a pixel is drawn, which is also why OpenAI describes the model as researching, planning and self-checking a prompt rather than only rendering it. [1]

Atlas interpretation: No independent, dated coverage of the April 21 announcement was available within this page's coverage window. What is cited here is limited to OpenAI's own materials: the announcement itself and the model's documentation. [1][2]

Sources

  1. Introducing ChatGPT Images 2.0

    OpenAI · Apr 21, 2026

  2. gpt-image-2

    OpenAI · Sep 8, 2026