Eight models, one switch
Alibaba's Qwen team put out eight models on April 29, 2025: two mixture-of-experts models, Qwen3-235B-A22B (235 billion total parameters, 22 billion active) and Qwen3-30B-A3B (30 billion total, 3 billion active), plus six dense models running from 0.6B up to 32B. All eight shipped under Apache 2.0, and all eight could switch between a thinking mode that generates a visible chain of reasoning inside a <think> block and a non-thinking mode that answers directly, toggled with the enable_thinking parameter in code or the /think and /no_think tokens in a chat turn. [1]
The context window scaled with model size: 32,768 tokens for the smaller dense models, extended to 131,072 tokens for the 4B and larger models. Pretraining used roughly 36 trillion tokens across 119 languages, done in three stages, a general 30-trillion-token pass at 4K context, a 5-trillion-token pass weighted toward STEM and code, then a long-context stage that pushed native context to 32K before the YaRN extension took it further. [1]
A dial, not a flag
Atlas interpretation: The part worth separating from the marketing is the thinking budget. Earlier reasoning models mostly treated the chain of thought as on or off. Qwen3 exposed it as a quantity: a developer sets how many tokens the model is allowed to spend reasoning before it has to answer, trading latency and cost against accuracy on a sliding scale rather than a binary choice. Publishing that mechanism, rather than just the toggle, is what made the release a design worth studying instead of a longer benchmark table. [3]
The coordination was the story
Simon Willison's same-day writeup focused less on the benchmarks Alibaba published and more on the release logistics: day-one support across Transformers, llama.cpp, Ollama, LM Studio, mlx-lm, SGLang and vLLM, rather than the usual pattern of weights landing on Hugging Face and every downstream tool catching up over the following week. He noted the size range meant something concrete for ordinary hardware: the 0.6B and 1.7B models fit on a phone, and the 32B model fit on his 64GB Mac with room to spare. [2]
Atlas interpretation: That coordination, not any single benchmark number, is why Qwen3 spread. A model with a permissive license and no runnable path on a laptop is a paper release. Qwen3 had both the license and the path on the day it was announced, which is a logistics achievement most labs releasing open weights still do not bother with. [2]
The switch that got removed
Atlas interpretation: The hybrid switch did not last as Alibaba's house style. Qwen3.8-2.4T-A95B, published sixteen months later, dropped the toggle entirely: its model card states plainly that thinking cannot be disabled and every response begins with a reasoning block by default. The dial Qwen3 introduced survives only as a reasoning_effort parameter that adjusts how much the model thinks, not whether it thinks at all. The industry kept the budget and retired the off switch. [4]
Sources
- Qwen3: Think Deeper, Act Faster
Qwen · Apr 29, 2025
- Qwen 3 offers a case study in how to effectively release a model
Simon Willison · Apr 29, 2025
- Alibaba's Qwen3: Open-weight LLMs with hybrid thinking
TechTalks · Apr 30, 2025
- Qwen/Qwen3.8-2.4T-A95B
Hugging Face · Aug 12, 2026