27 billion parameters, built to fit on a desk
Alibaba published Qwen3.8-27B-FP8 on Hugging Face on August 14, 2026, under the Apache 2.0 license. It is a 27 billion parameter model built on the Qwen3.5 foundation, using 64 layers with a hybrid attention mechanism combining Gated DeltaNet and Gated Attention, a native context window of 262,144 tokens extensible to one million via RoPE scaling, and multi-token prediction training. [1]
The model natively supports image and video understanding alongside text. Simon Willison, testing the quantized release, reported the version he ran occupied 17GB on disk, small enough for a MacBook Pro or an Nvidia DGX Spark, and that it could write code, drive tools and annotate images with bounding boxes. [2]
Atlas interpretation: That size is the point. A 27 billion parameter, vision-capable model that runs on hardware a person can own, rather than rents by the token, is a different kind of release than a frontier flagship, and it is the category the summary above is judging by: not the most capable model overall, but the most capable one most people could actually put on their own machine. [1][2]
Reasoning on by default, and turned up too far
The model card documents flexible thinking control: reasoning mode is on by default but can be disabled per request, with depth adjustable through a reasoning_effort parameter. Willison found the shipped default set to the highest of those levels, which he called "a hilarious default," and demonstrated it by asking the model to draw an SVG of a circle. Instead of returning minimal output, it reasoned at length about the geometric composition of the shape before answering. [1][2]
Atlas interpretation: Willison's recommendation, to run the model on low or no reasoning effort at first rather than accept the default, is a practical note rather than a condemnation of the model's underlying ability. He measured 15 to 30 tokens per second running it locally, well under the 74 to 184 tokens per second he gets from OpenAI's cloud models, and attributed part of the gap to that default reasoning depth rather than to raw model speed. Multi-token prediction, per the model card, is what narrows that gap, and Willison credited it with roughly a 72 percent throughput improvement over running without it. [2][1]
The second Qwen3.8 release inside a week
Qwen3.8-27B-FP8 followed Alibaba's release of Qwen3.8 Max, a 2.4 trillion parameter set of open weights, by two days. [1]
Atlas interpretation: The pairing shows the same lab shipping at both ends of the size scale within the same week, one release aimed at whoever can run a cluster and the other aimed at whoever has a workstation. Willison's review, coming from someone who evaluates open models for personal and small-team use rather than for datacenter deployment, is evidence that the smaller release found its intended audience: people running a capable model themselves, overthinking default and all. [2]
Sources
- Qwen/Qwen3.8-27B-FP8
Hugging Face · Aug 14, 2026
- Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Simon Willison · Aug 16, 2026