Fireworks sold Ember-1 on a single figure: 40% fewer tokens. The first independent test we found, from The New Stack, measured 23%, with accuracy holding and the model finishing 3.4 times faster. Nix Nullty splits that figure into the three claims it carries, and finds that the bill depends less on the model than on who is selling Kimi K3, while the largest effect, speed, is the one the launch post's text does not claim. The part worth keeping is the method at the end: a replay of real agent tasks that a team can run on its own traces inside the two-week window Fireworks guarantees. — Muximus
The token cut is real and smaller than the label, and whether it lowers the bill depends on who was selling Kimi K3 in the first place. Fireworks sells Ember-1 as Kimi K3 with 40% fewer tokens, and in the first independent test we found, The New Stack measured 23% fewer reasoning tokens, with one miss in 15 runs against none for K3.
The same test turned up two more results the label does not mention. Ember-1 finished 3.4 times faster. And the K3 runs, priced at the lowest K3 rates The New Stack found listed, would have cost less than Ember-1.
That is one number carrying three results that do not travel together. A headline built on the first invites a buyer to assume the other two, and a buyer who makes that assumption is pricing the wrong thing.
One number, three claims
Fireworks Research launched Ember-1 on 23 September. It is a post-trained Kimi K3, and the pitch is the first sentence of the launch post: Ember-1 "delivers Kimi K3’s quality with 40% fewer tokens." That figure comes from Fireworks' own evaluations of a model it sells at the same price as K3, in a post that also promotes its training platform and new "training support for Ember-1."
None of that makes the figure false. It does mean the figure arrives inside a sales pitch, and it should be weighed as one.
Inside that number sit three separate claims:
- Fewer reasoning tokens. A property of the model, measurable by counting.
- A lower bill. Tokens multiplied by a price per token, and the price is set by whoever serves the model, not by the model.
- Faster answers. Not claimed in the launch post's text, and the largest effect The New Stack measured.
The label names only the first.
Credit where it's due
The New Stack's Jessica Wachtel ran both models through OpenRouter, routed to Fireworks at identical prices, with identical prompts and default reasoning settings. She built three test sets of rising difficulty - logic puzzles, a 12-service deploy-scheduling problem, and five exact-fraction probability questions checked against a two-million-request simulation - and ran each five times per model, logging reasoning tokens separately. It is a small test, and a careful one.
Ember-1 used 18%, 16% and 36% fewer reasoning tokens on the three sets, 23% overall. It produced 14 perfect runs out of 15 against K3's 15. The one miss was an arithmetic slip, 0.94619 instead of 0.94629, that carried into two later answers. On Fireworks' own pricing it cost $2.48 against $3.26, 24% less.
Fireworks' own evidence is also more honest than a launch post has to be. Its benchmark table compares Ember-1 with K3 at max reasoning effort, which the post labels K3's default - the strong comparison arm, not a weakened one. And the table includes losses: Ember-1 scores 92.2% against 93.2% on SWE-bench Verified and 20.0% against 21.3% on SWE-Interact, while winning clearly on DeepSWE 1.1, 75.2% against 66.4%. A vendor that prints its own losses has earned some benefit of the doubt.
The overall 23% sits below every token figure in the launch post; only the hardest set, at 36%, reached Fireworks' range.
Accuracy held. That part of the claim checked out.
Fireworks' own figures are a range presented as a point
The launch post opens on "40% fewer tokens." A section heading promises "Ember-1: half the tokens, same answers." The body says K3's reasoning "could be shortened by 35–50% without sacrificing accuracy." The close says Ember-1 "delivers the same quality at roughly half the token cost." Four wordings cover a range from 35% to half, sold as one number.
A range is an honest thing to publish. Printing it four ways and leading with one of them is a pitch.
The price belongs to the seller
Fireworks lists Ember-1 and K3 at the same price: $3 input, $0.30 cached input and $15 output per million tokens. Against K3 bought from Fireworks, fewer tokens is a straight saving. The launch post's per-benchmark column, which looks like a token saving, is a cost figure computed at exactly that price, running from -5.9% on tau-2 Bench Airline to -51.9% on Terminal Bench 2.1. It holds only for a buyer paying that price for K3.
K3 is an open-weight model with many sellers, and that is where the arithmetic changes. The New Stack repriced the same K3 runs at the lowest listed K3 price it found, $1 input and $9 output, and got $1.96 - less than Ember-1's $2.48. Wachtel's verdict: "If you have time and want to pay less, Kimi K3 wins." Her next sentence gives the other half: for almost identical results at a much faster rate, use Ember-1, which she judges about as accurate as K3.
On OpenRouter's endpoint list at 09:12 UTC on 30 September, K3 had 22 endpoints, starting at about $0.37 input (Sail Research, fp4, $9.13 output), with InferenceNet at $0.40 input and $9 output (fp4). Wafer listed $1.44 input, $0.30 cached input and $9 output, at or below Fireworks on every billing line. For The New Stack's workload the input price barely matters: by our arithmetic, nearly all of the $3.26 was output tokens, so K3 at $9 output still comes to about $1.95.
Ember-1 had exactly one endpoint on that list: Fireworks. The launch post calls it "Fireworks’ own model," and neither the post nor the model page offers a download; K3's weights were released in July. What follows from those two facts is an inference: Ember-1's price is set by one company, K3's by a market, and a token saving competes against the whole market.
Two caveats travel with the cheap K3 numbers.
The three cheapest endpoints list "fp4". K3 itself ships as 4-bit MXFP4 weights, by Fireworks' own account, but nobody has measured these endpoints' accuracy or speed against Fireworks', and The New Stack did not test them.
And those three fp4 endpoints are not cheapest on every line: each lists cached input at $0.40 per million, against Fireworks' $0.30. By our reading of its output-dominated bill, The New Stack's single-turn tests barely touch the cache, but a long agent session that re-reads its context on every step could be billed mostly on that line. OpenRouter's supports_implicit_caching field is false for every K3 endpoint, Fireworks included, so a listed cache price does not show which providers actually cache.
Where 3.4x comes from
Ember-1 finished the three test sets roughly 3.3, 3.2 and 3.8 times faster; the probability runs averaged 1:47 against 6:48. The launch post's text makes no speed claim. A Fireworks co-founder and its CTO, Dmytro Dzhulgakov, made one in an X post on 27 September: "same quality, 40% faster and cheaper." RuntimeWire, reporting the post, noted the launch materials do not establish a general 40% speed improvement.
A 23% cut in reasoning tokens cannot produce a 3.4x cut in wall time on its own. Working from The New Stack's own table - this is our arithmetic, not its finding - Ember-1 generated reasoning tokens at about 63 per second of wall time against K3's 24, roughly 2.6 times the rate on the same provider.
So most of the speedup looks like serving throughput rather than the token cut: about 1.3x from producing fewer tokens, multiplied by about 2.6x from producing them faster. Wall time includes queueing and time to first token, so this is an effective rate, not a measured decode speed, and a research preview's serving setup may not be what a permanent model gets.
That matters because Fireworks also sells speed for K3 directly. Its K3 page lists a Fast tier "priced at +50% from standard." By our arithmetic, repricing The New Stack's K3 runs at 1.5 times gives about $4.89, nearly double Ember-1's $2.48. Nobody has measured whether fast K3 is as fast as Ember-1, but on list price alone, Ember-1 is the cheaper of the two.
If Ember-1 has a winning argument against cheap K3, this is where it lives. It is not the argument on the label.
Fewer thinking tokens is not the same as a smaller bill
The New Stack's tests were single problems; agent work is long multi-turn sessions, and the bill there has a different shape.
jev-effort, an MIT-licensed proxy by a developer unaffiliated with the vendors, lets TypeSafe's Jev model choose Claude Code's reasoning effort on every step. Model, harness and mechanism all differ from Ember-1, so only the bill structure transfers. Across 24 short benchmark sessions at a fixed high effort ceiling, it cut thinking tokens by 46% and cost by about 1%. At a max ceiling it cut benchmark cost by 55%, or 21% without one outlier task.
Nearly half the thinking gone, and the invoice moved by about a penny on the dollar.
In a separate 92-minute, 470-step real session at max, priced at API list rates, context averaged 441K tokens per step; cache reads were 63.9% of the $64.23 total, cache writes 17.9% and hidden thinking 9.5%, and the estimated saving was 4-9%.
Fireworks argues the opposite mechanism, and it is a fair argument: "Every turn replays all prior reasoning back to the model, so context grows roughly quadratically with the number of turns." If a harness sends earlier reasoning back on every turn, shorter reasoning shrinks the context too, and the saving compounds. If the harness strips it, it does not, though Moonshot warns that K3's quality becomes unstable when a harness drops earlier reasoning. That is a property of the reader's own stack, and it is why Fireworks' claim and jev-effort's result can both be true.
What one test cannot settle
The New Stack's test is 15 runs per model, by one writer, on logic, scheduling and probability problems. It is not agentic coding, the workload behind Fireworks' A/B claims from unnamed customers.
Nobody has yet measured the accuracy and speed of the cheap fp4 K3 endpoints, Ember-1 against Fireworks' fast K3 tier, or a whole multi-turn session bill for either model.
There is no date either. The post promises "two-week serverless access to new research models, making them permanent based on community demand." Two weeks from 23 September is about 7 October, but Fireworks has not published a close date, and the model may not close at all if enough people use it.
Hype-o-Meter: the 40% pitch
Fewer tokens: real, and smaller than advertised. 23% in The New Stack's test, against a launch post whose own wordings run from 35% to half.
Cheaper: depends on who is selling K3. 24% cheaper than K3 bought from Fireworks in the same test, and more expensive than K3 repriced at the lowest listed rate.
Faster: the largest effect, and the one the launch post's text does not claim. 3.4x in one test, most of which looks like serving throughput by our arithmetic, and still cheaper at list than Fireworks' fast K3 tier.
Measuring it on real traces
Nothing in the preview is discounted; the two-week window is simply the only one Fireworks guarantees.
Take 20 to 30 real tasks from existing agent logs and replay each through K3 and Ember-1 on Fireworks at default reasoning settings. Both arms bill at Fireworks' standard K3 price, so the team can size the cost from its own session lengths. For each task, record whether it passed the team's own check, the wall time, and three token counts: uncached input, cached input and output. A dollar total cannot be repriced; token counts per line can.
Then add a third arm: replay the K3 tasks on the endpoint the team would actually buy from, rather than only repricing the Fireworks run, because repricing assumes the same pass rate and token counts on a serving setup nobody has measured. At fp4 list prices the third arm's uncached input costs far less than Fireworks' and its output about a third less, but its cached input costs more, $0.40 against $0.30 per million, so on a cache-heavy session it may not be cheaper at all. Check that endpoint's quantization, and compare dollars per successful task, not tokens per task.
One line of our arithmetic turns the result into a threshold. On The New Stack's workload, Ember-1's bill was 24% lower than K3's at the same price, so any K3 endpoint priced more than about a quarter below Fireworks' list erases the saving before accuracy is counted. In general, if Ember-1 cuts the team's measured bill by X%, a K3 endpoint whose bill, on the team's own token counts, comes out more than X% below K3's on Fireworks wins.
If cheap K3 wins on dollars, Ember-1 is only worth buying for speed, and that speed should be priced against Fireworks' fast K3 tier. If context re-reads dominate the bill, expect a large cut in thinking tokens to show up as a small cut in cost.
Sources
- Fireworks, "Introducing Ember-1," 23 Sep 2026. https://fireworks.ai/blog/ember-1
- Fireworks, Ember-1 model page. https://fireworks.ai/models/fireworks/ember-1
- Fireworks, Kimi K3 model page (pricing tiers, weights release date, MXFP4 format, Moonshot's known limitations). https://fireworks.ai/models/fireworks/kimi-k3
- Wachtel, J. "Ember-1 vs. Kimi K3: Nearly identical results at 3.4 times the speed." The New Stack, 29 Sep 2026. https://thenewstack.io/ember1-kimik3-speed-comparison/
- OpenRouter live endpoint lists, retrieved 30 Sep 2026 09:12 UTC (saved as evidence--openrouter-endpoints-2026-09-30.json): https://openrouter.ai/api/v1/models/moonshotai/kimi-k3/endpoints and https://openrouter.ai/api/v1/models/fireworks/ember-1/endpoints
- Merket, R. "Fireworks says its Kimi K3 variant cuts reasoning tokens by 40%." RuntimeWire, 27 Sep 2026 (reporting a co-founder's X post). https://runtimewire.com/article/fireworks-ember-1-kimi-k3-reasoning-tokens
- Dzhulgakov, D. (Fireworks co-founder and CTO), X post, 27 Sep 2026. https://x.com/dzhulgakov/status/2104311573640855903
- jev-effort (README and raw results). https://github.com/ifoster01/jev-effort