Fireworks AI drops Ember-1, a reasoning model that cuts half the tokens

AI reasoning models spend up to 90% of their output tokens talking to themselves, and a lot of that thinking is pure waste.
Fireworks Research released Ember-1, a specialized model built on Kimi K3 that cuts reasoning tokens by 35–50% without sacrificing accuracy. The team ran over 50 training experiments on Fireworks Serverless Training to train the model to cut bloated internal monologues while keeping the self-reflection steps that actually prevent errors.
Why it matters: In multi-turn coding agents, old reasoning steps get re-read and re-billed on every single call, making automation expensive at scale. Ember-1 fixes this context bloat while matching or beating original K3 quality, hitting 82.0% on Terminal Bench 2.1 compared to K3's 80.9%.
Here's the gist: simply turning down reasoning effort on stock K3 breaks accuracy. Fireworks instead retrained Ember-1 across math, software engineering, and tool-use tasks, teaching the model efficient thinking natively.
Turns out AI models, much like software engineers, do better work when they stop overthinking.
Sources
- Fireworks Research — https://fireworks.ai/blog/ember-1
- Hacker News — https://news.ycombinator.com/item?id=49868830

