DeepSeek-OCR: Vision-Token Compression & Benchmarks

DeepSeek-OCR encodes page images into 64 to 400 vision tokens; the page examines its reported compression-to-precision curve and OmniDocBench comparisons.

Pages as pixels, not as tokens

DeepSeek posted DeepSeek-OCR to GitHub on October 20, 2025 under an MIT license, with the accompanying paper, "DeepSeek-OCR: Contexts Optical Compression," appearing on arXiv the next day. Rather than running OCR to hand a downstream language model a stream of text tokens, the model encodes a rendered page image into a fixed, much smaller pool of vision tokens and treats that pool as the context the decoder reasons over. [1][2]

Four preset resolutions set the token budget directly: a 512x512 render costs 64 vision tokens, 640x640 costs 100, 1024x1024 costs 256, and 1280x1280 costs 400. A "Gundam" mode tiles several 640x640 crops alongside one 1024x1024 view for dense or oversized pages, trading a larger token count for more detail where it is needed. [1]

How much a page can lose before OCR notices

The paper's own headline result is a tradeoff curve, not a fixed accuracy number. When the text being reconstructed runs to within 10 times the number of vision tokens feeding the decoder, a compression ratio under 10x, decoding precision on the paper's OCR evaluation reaches 97 percent. Pushed to a 20x ratio, precision falls to about 60 percent. The paper reports that range as evidence for further work on long-context compression and on modeling how memory degrades in language models, not as a finished production number. [2]

Fewer tokens than the competition, on the paper's own benchmark

On OmniDocBench, the paper reports beating GOT-OCR2.0, which spends 256 tokens per page, using only 100 vision tokens, and outperforming MinerU2.0, which averages more than 6,000 tokens per page, while staying under 800 vision tokens. DeepSeek also reports that a single A100-40G GPU running the pipeline can generate labeled OCR training data at a rate of more than 200,000 pages a day. [2]

Atlas interpretation: Those OmniDocBench comparisons and the throughput figure come from DeepSeek's own paper and repository rather than a third-party reproduction, which matters more here than it would for a plain accuracy leaderboard entry: the paper is using the benchmark wins to support a claim about token economics, and that claim has not yet been checked by anyone running the comparison independently. [2][1]

The pitch is about context windows, not scanners

DeepSeek's own release note describes the project as investigating "the role of vision encoders from an LLM-centric viewpoint," and outside commentary that followed the release largely read past the OCR framing to the same question: whether rendering old conversation history or documents as images and re-encoding them as a small block of vision tokens is a cheaper way to keep that content in a language model's context than carrying it as text tokens. [1][3]

Atlas interpretation: Read that way, the compression-ratio-to-precision curve matters more than the OmniDocBench wins. A model that can recover most of a page's text from a tenth as many tokens, and something usable from a twentieth as many, is a proposal for how a language model might discard and later approximately recall old context cheaply, an idea the paper connects explicitly to memory forgetting rather than to document digitization. [2]

Sources

  1. deepseek-ai/DeepSeek-OCR

    DeepSeek · Aug 20, 2026

  2. DeepSeek-OCR: Contexts Optical Compression

    arXiv · Oct 21, 2025

  3. DeepSeek-OCR Explained: How Contexts Optical Compression Works

    BentoML · Oct 24, 2025