The vision encoder is gone, not just smaller
Google released Gemma 4 12B on June 3, 2026, describing it as a unified model with no separate multimodal encoders: vision and audio inputs are projected directly into the language model backbone. Google's announcement put it plainly: no multimodal encoders, with image and audio inputs flowing directly into the LLM. Vision uses what the company calls a lightweight embedding module built from a single matrix multiplication, positional embeddings and normalization; raw audio is projected into the same token space. [1][3]
The accompanying technical report is more specific about scope: the encoder-free design is a property of the 12B model within the Gemma 4 family, which spans 2.3B to 31B parameters across dense and mixture-of-experts variants. The other sizes in the family kept improved vision and audio encoders rather than dropping them. [3]
Atlas interpretation: That distinction matters more than the headline. This is not Google removing encoders across a model line; it is one 12B variant used as the vehicle for a specific bet, that a single set of transformer weights can absorb what an encoder used to do, in exchange for one model instead of an encoder plus a backbone to train, align and serve separately. Whether that bet pays off for larger models remains untested by anything in this release. [1][3]
16GB VRAM, with a footnote
Google described the model as small enough to run locally with 16GB of VRAM or unified memory, which covers a wide range of consumer laptops and single-GPU desktop cards rather than requiring a data-center accelerator. Google's own model documentation gives the actual memory figures behind that claim: about 26.7GB at full 16-bit precision, 13.4GB at 8-bit quantization, and 6.7GB at 4-bit quantization. [1][4]
Atlas interpretation: The 16GB figure only holds once the model is quantized below full precision; run at 16-bit, it needs more memory than that. This is a normal and disclosed tradeoff, not a discrepancy Google hid, but the marketing framing states the quantized number without saying so, and a reader taking the blog post's number at face value and loading full-precision weights on a 16GB card would run out of memory. [4]
Google positioned the 12B model as reaching performance close to its own larger 26B mixture-of-experts model on standard benchmarks while using under half the memory. The Hugging Face model card lists instruction-tuned scores including 77.2 percent on MMLU Pro, 78.8 percent on GPQA Diamond and 69.1 percent on MMMU Pro, the multimodal benchmark most relevant to the encoder-free redesign. [1][2]
Atlas interpretation: All of these numbers come from Google's own reporting on Google's own benchmark suite, run by the lab that built the model. That is the normal state of affairs for a same-day model release, since independent evaluation takes time the announcement does not wait for, but it means the comparison to the 26B model is a claim to watch for third-party confirmation rather than a settled result. [1][2]
Open weights kept coming from Google, not from the labs chasing frontier scale
The release carries an Apache 2.0 license, the same permissive terms Google has used for prior Gemma releases, and the model is distributed through Hugging Face and Kaggle rather than gated behind an API. [2]
Atlas interpretation: By mid-2026, open-weight releases from a US frontier lab were notable mostly for still happening at all. Meta had pulled back from its earlier Llama openness, and OpenAI and Anthropic keep their frontier weights closed. Google is the exception among the US labs still shipping openly licensed weights at meaningful scale, alongside a wave of open Chinese releases from labs like Zhipu and DeepSeek, which is the specific gap this event is worth noting for. [2]
Google kept building on the Gemma line after this release. By August 20, 2026, the company reported the Gemma family had passed a billion cumulative downloads and more than 100,000 community-built variants, which is the kind of adoption an openly licensed model can accumulate that a closed API model cannot. [5]
Sources
- Gemma 4 12B: A unified, encoder-free multimodal model
Google · Jun 3, 2026
- google/gemma-4-12b
Hugging Face · Jun 3, 2026
- Gemma 4
Google · Jun 3, 2026
- Gemma 4 Technical Report
Google DeepMind · Jun 3, 2026
- Inside the Gemmaverse: Celebrating one billion Gemma downloads
Google · Aug 20, 2026