Business only (hides war / politics / culture / sports)
Sunday, Sep 27, 2026
1 Fireworks AI Cuts Reasoning Tokens by 71% in First Research Model π€ AI Sep 27, 4:45 PM EDT 13/9
1
Fireworks AI Cuts Reasoning Tokens by 71% in First Research Model
π€ AI Sep 27, 4:45 PM EDT 13/9
topics π€
AI tags AIAI ModelsAI ReleasesAI Open SourceAI InfraAI Inference keywords FireworksFireworks Research β
Ember-1 used 39% fewer total tokens than its base model in live A/B tests on coding traffic while maintaining the same success rate. Fireworks AI achieved the reduction by applying reinforcement learning to Kimi K3, which eliminated repetitive reasoning cycles in agentic loops and cut reasoning tokens specifically by 71% without degrading benchmark performance.
The release is the first model from the newly formed Fireworks Research team and follows a broader trend of inference platforms, such as fal, performing post-training on open-weight models to lower customer costs. This approach targets the inefficiency of reasoning models, which typically spend over 90% of their output tokens on internal thinking before generating a final response.
Continues from Sunday, Sep 27
Fireworks Research Cuts Kimi K3 Reasoning Tokens by 40% in Ember-1 5 tweets β’ 5 sources