A Kimi K3 variant built to use fewer reasoning tokens
Fireworks AI released Ember-1 on September 23 as a research preview on its serverless API. The model starts from Kimi K3 and is post-trained to produce shorter reasoning traces for coding and agent workflows, where generated reasoning can be carried into later turns. Fireworks describes an initial two-week serverless access window, with a decision about making research releases permanent based on demand. [1]
The model is listed as ready on Fireworks’ own model page, under accounts/fireworks/models/ember-1. It accepts text and image input, supports function calling, and has a context length listed at roughly one million tokens. Vercel added it to AI Gateway on September 27, a second route through which developers could call the research preview. [4][2]
On September 24, YFarmX reported sending requests to the Fireworks-served model and measuring its tokenizer signature. That independently confirms a callable artifact and a close tokenizer relationship to the Moonshot model family, but it does not verify Fireworks’ quality or token-saving results. [3]
The efficiency claim rests on Fireworks’ tests
Fireworks reports results on five coding and agent benchmarks against Kimi K3 at its maximum reasoning effort. Ember-1 scored 92.2 percent versus 93.2 percent on SWE-bench Verified while using 15.5 percent fewer tokens, and 82.0 percent versus 80.9 percent on Terminal Bench 2.1 while using 51.9 percent fewer. The table spans different tasks and sample sizes, so the launch’s roughly 40 percent figure is a summary of its own evaluation rather than a fixed saving on every prompt. [1]
Fireworks also reports two customer A/B tests on production coding traffic, with about 35 percent fewer tokens per task at comparable quality. The company names neither customer in the announcement, and the tests and benchmark runs were conducted or reported by Fireworks. Vercel repeats the savings as a Fireworks claim rather than publishing an independent evaluation. [1][2]
Atlas interpretation: Ember-1’s claim is about work done per billed token, not a new base model or a lower token rate. Teams would have to test their own workloads to see whether shorter traces preserve the answers they need; the published customer results do not expose enough detail to generalize that outcome. [1][4]
Sources
- Introducing Ember-1
Fireworks AI · Sep 23, 2026
- Ember-1 from Fireworks now available on AI Gateway
Vercel · Sep 27, 2026
- Fireworks: Ember-1: model identity record
YFarmX · Sep 29, 2026
- Ember-1
Fireworks AI · Sep 29, 2026