Four sizes, one hardware target
Google released Gemma 3 in four sizes, 1B, 4B, 12B and 27B parameters, each with a 128,000-token context window. The 4B, 12B and 27B models add vision, letting the model analyze images and short videos; the 1B model is text only. Google listed out-of-the-box support for more than 35 languages and pretrained support for over 140, and shipped official quantized versions to cut memory and compute further. [1]
Google's own framing for the release was "the most capable model you can run on a single GPU or TPU," and the launch post put a chart next to that claim showing rival models needing as many as 32 accelerators to match Gemma 3 27B's benchmark position. Ars Technica's coverage the same day described the range as running on "almost anything, from powerful GPUs to a smartphone," tying the four sizes to that same one-GPU pitch rather than treating each size as a separate story. [1][2]
Atlas interpretation: The interesting number in this release was never a benchmark score. It was that a 27B model with vision and a 128,000-token window could sit on one card at all, at a time when comparably ranked models were built around multi-GPU serving. That is a deployment argument, not a capability argument, and it is the one Google chose to lead with. [1]
The benchmark was a leaderboard, not a test suite
Google's evidence for Gemma 3 27B's ranking was a preliminary human-preference evaluation on LMArena's leaderboard, where it scored a listed Elo of 1338, ahead of Llama3-405B, DeepSeek-V3 and o3-mini in that setting. Google's own post described the result as preliminary. [1]
Atlas interpretation: An Elo score from a crowdsourced preference arena measures which output people click as better in a blind pairwise vote, not accuracy on a fixed task with a checkable answer. It says Gemma 3 27B's answers read well to arena voters against those specific competitors on that specific day. It does not establish that the model matches Llama 3 405B or DeepSeek-V3 on reasoning, coding or factual benchmarks, and Google's release did not claim that it did. [1]
The catch was in the license, not the weights
Gemma 3 shipped under Google's custom Gemma Terms of Use rather than a standard open source license such as Apache 2.0. Reporting two days after launch found the terms gave Google the right to "restrict (remotely or otherwise) usage" of Gemma that it judged to violate its prohibited use policy, a restriction that also reached derivative models, including ones trained on Gemma-generated synthetic data. [3]
Developers raised the concern publicly that this made commercial use of the models a legal risk rather than a settled question: a company could build a product on Gemma 3, and Google's terms left room to restrict that usage later. One applied scientist quoted in the coverage flagged the possibility of a "clawback" reaching downstream customers and fine-tuning businesses built on top of the model, not just Google's direct relationship with the company that trained on it. [3]
Atlas interpretation: Open weights and an open license are not the same commitment, and Gemma 3 is a clean example of the gap. The weights were downloadable and ran on consumer hardware; the terms attached to them kept a lever in Google's hands that Apache 2.0 does not give a vendor. Google switched Gemma 4 to Apache 2.0 more than a year later, after this exact complaint had circulated since Gemma 3's launch, which is the kind of ordinary walk-back that a licensing complaint from launch week turns into once a vendor decides adoption matters more than the lever. [3]
Sources
- Introducing Gemma 3: The most capable model you can run on a single GPU or TPU
Google · Mar 12, 2025
- Google's new Gemma 3 AI model is optimized to run on a single GPU
Ars Technica · Mar 12, 2025
- 'Open' AI model licenses often carry concerning restrictions
TechCrunch · Mar 14, 2025