Meta’s newest AI model Muse Spark 1.3, unveiled yesterday, is faster and more performant on third-party benchmarks than its predecessor — with a caveat.
"Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter," Meta co-founder and CEO Mark Zuckerberg wrote on X, calling it Meta’s “biggest jump” yet in coding and agentic work.
There is substance behind both parts of that claim. Muse Spark 1.3 makes significant gains over last month’s 1.2 release, particularly on long-running agent tasks. The version developers can access now is also one of the strongest price-performance offerings near the top of independent model rankings.
Meta’s strongest Muse Spark 1.3 benchmark results come from its max reasoning configuration. Meta says that version is still completing additional safety testing and will arrive “shortly”; the third-party benchmarking firm Artificial Analysis says it evaluated max in a limited partner preview, and currently lists no API provider at all for the configuration.
The version broadly rolling out this week through its Muse Code harness and the Meta Model API uses Meta’s previously available reasoning settings, including xhigh.
That makes the more relevant enterprise question not whether Muse Spark 1.3 can reach frontier territory, but how close the model companies can actually deploy today gets — and at what real cost.
The shipping model is very good, but not the benchmark leader
Meta does disclose results for both configurations in its underlying evaluation report, so this is not a case of the company hiding the deployable model. But its launch materials prominently showcase the max variant, and some of the largest scores belong to that configuration.

For example, Meta reports GDPval-AA v2 scores of 1,754 Elo for max versus 1,709 for xhigh, OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2.
On some tests the distinction is negligible or reversed: DeepSearchQA is tied at 89.4, while xhigh scores 89.2 on Terminal-Bench 2.1 versus max at 88.8.
Artificial Analysis scores Muse Spark 1.3 max at 62 on its Intelligence Index and the shipping xhigh version at 61. The latter ties GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high. But Anthropic still occupies the top of the leaderboard: Claude Fable 5.1 reaches 66 at max and 65 at xhigh, while Claude Opus 5 reaches 63 at max and xhigh.
In other words, Muse Spark 1.3 xhigh is legitimately in the frontier cluster, but it is not the model currently setting the frontier.
That is still a substantial change from Muse Spark 1.2. VentureBeat’s coverage of last month’s launch found Meta fielding a credible coding challenger that nevertheless generally trailed Anthropic’s best model. Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1 versus Opus 5’s 86.7%, and also finished behind Opus on the other main coding comparisons Meta presented.
With 1.3, Meta is no longer merely showing up in that contest. On several coding and agentic evaluations, it is trading wins with OpenAI and Anthropic.
Meta says the underlying model has also become easier to operate. Muse Spark 1.3 is trained to maintain multiple workflows in a long thread, gather context with tools, detect gaps in its own plans, ask users for clarification when necessary and confirm before consequential actions. In Meta engineers’ internal comparisons, it used roughly 20% fewer tool calls and 25% fewer tokens than 1.2 during coding work.
For enterprises paying for thousands or millions of agent loops, those behavioral improvements could matter more than another leaderboard point.
‘Almost too cheap to meter’ does not mean Meta cut its prices
Muse Spark 1.3 did not receive an API price cut. Meta kept Standard pricing exactly where it was for Muse Spark 1.2: $1.25 per million input tokens, $4.25 per million output tokens and $0.15 per million cached input tokens.
Model | Input ($/1M) | Output ($/1M) | Total ($/1M) | Source |
Muse Spark 1.2 / 1.3 Contributor | $0.10 | $0.20 | $0.30 | |
MiMo-V2.5 Flash | $0.10 | $0.30 | $0.40 | |
DeepSeek-V4-Flash — off-peak | $0.22 | $0.66 | $0.88 | |
GPT-5.6 Luna | $0.20 | $1.20 | $1.40 | |
MiniMax-M3 | $0.30 | $1.20 | $1.50 | |
LongCat-2.0 — limited-time promo | $0.30 | $1.20 | $1.50 | |
DeepSeek-V4-Flash — peak hours | $0.44 | $1.32 | $1.76 | |
MiMo-V2.5 | $0.40 | $2.00 | $2.40 | |
DeepSeek-V4-Pro — off-peak | $0.66 | $1.98 | $2.64 | |
LongCat-2.0 — standard | $0.75 | $2.95 | $3.70 | |
MiMo-V2.5 Pro (≤256K) | $1.00 | $3.00 | $4.00 | |
Gemini 3.7 Flash — through Dec. 31, 2026 | $0.75 | $3.75 | $4.50 | |
Gemini 3.8 Flash — through Dec. 31, 2026 | $0.75 | $3.75 | $4.50 | |
DeepSeek-V4-Pro — peak hours | $1.32 | $3.96 | $5.28 | |
Muse Spark 1.1 / 1.2 / 1.3 | $1.25 | $4.25 | $5.50 | |
GLM-5.3 | $1.40 | $4.40 | $5.80 | |
Grok 4.6 — <200K prompt tokens | $2.00 | $6.00 | $8.00 | |
MiMo-V2.5 Pro (>256K) | $2.00 | $6.00 | $8.00 | |
Qwen3.8-Max | $2.00 | $6.00 | $8.00 | |
Gemini 3.7 Flash — starting Jan. 1, 2027 | $1.50 | $7.50 | $9.00 | |
Gemini 3.8 Flash — starting Jan. 1, 2027 | $1.50 | $7.50 | $9.00 | |
GPT-5.6 Terra | $2.00 | $12.00 | $14.00 | |
Grok 4.6 — ≥200K prompt tokens | $4.00 | $12.00 | $16.00 | |
GPT-5.4 | $2.50 | $15.00 | $17.50 | |
Kimi K3 | $3.00 | $15.00 | $18.00 | |
Claude Opus 5 | $5.00 | $25.00 | $30.00 | |
Sakana Fugu Ultra (≤272K) | $5.00 | $30.00 | $35.00 | |
GPT-5.6 Sol — Standard mode | $5.00 | $30.00 | $35.00 | |
Claude Fable 5 / Claude Mythos 5 | $10.00 | $50.00 | $60.00 | |
Claude Fable 5.1 / Claude Mythos 5.1 | $10.00 | $50.00 | $60.00 | |
GPT-5.6 Sol — Fast mode | $10.00 | $60.00 | $70.00 |
That makes Zuckerberg’s “almost too cheap to meter” line less a statement about lower token prices than about what Meta believes developers can accomplish with those tokens.
Artificial Analysis offers evidence for that argument, but also a complication. It measures Muse Spark 1.3 xhigh at 235.2 output tokens per second and estimates a cost of $0.55 per Intelligence Index task.
At 61 on the Intelligence Index, that gives it the lowest cost per task of any currently measured model at that intelligence level.

Muse Spark 1.2 cost only $0.40 per Artificial Analysis task, while scoring 57.
Despite unchanged per-token pricing, the independent benchmark’s cost of completing an average task therefore increased generation-over-generation. Artificial Analysis attributes the increase primarily to heavier input-token consumption on agentic evaluations.
That does not directly contradict Meta’s claim of 25% lower token use: Meta is describing comparisons in its own coding workflows, while Artificial Analysis is measuring a broader suite of reasoning and agentic tasks.
But it illustrates why “cheap” becomes slippery once models operate as agents. Token rates, reasoning effort, number of turns, tool calls and retries all contribute to the actual cost of finishing work.
Meta also retains its unusually cheap Contributor tier — $0.10 per million input tokens and $0.20 per million output tokens — in exchange for permission to use prompts and completions for training.
As VentureBeat noted with Muse Spark 1.2, that may be attractive for prototyping but creates a materially different data-governance calculation for enterprises working with proprietary code or sensitive internal information.
Wang’s ‘Gemini who?’ lands on an unusually close comparison
Meta chief AI officer Alexandr Wang was considerably less qualified in celebrating the release.
After Artificial Analysis posted its Muse Spark results, Wang reposted them on X, adding: “i really hate to say it, but… gemini who? 😱💨”
The shade was particularly pointed because Google released Gemini 3.8 Flash on the same day, pitching it at almost exactly the same class of workload: long-horizon software engineering, autonomous agents and multi-step professional reasoning. Google calls 3.8 its best reasoning and coding Flash model yet and says it is the company’s third Flash release in six weeks.
Independent numbers give Wang something to work with, though hardly a knockout.
Artificial Analysis gives Muse Spark 1.3 xhigh a 61 Intelligence Index score at $0.55 per task, compared with 59 and $0.58 for Gemini 3.8 Flash at high reasoning. Meta therefore edges Google on both intelligence and task cost at those particular settings.
Google wins decisively on throughput. Artificial Analysis measures Gemini 3.8 Flash high at about 305 output tokens per second, versus 235 for Muse Spark — roughly 30% faster. Gemini also has the lower raw API sticker price for now: Google is charging an introductory $0.75 per million input tokens and $3.75 per million output tokens, compared with Meta’s $1.25 and $4.25.
That promotional Google pricing expires December 31, after which it rises to $1.50 per million input tokens and $7.50 per million output tokens.
The result is a useful snapshot of how tight frontier-model economics have become. Meta currently wins this independent comparison by two Intelligence Index points and three cents per benchmark task; Google offers substantially higher output throughput and cheaper raw tokens during its launch promotion.
Wang’s “gemini who?” is fun executive trash talk. For an enterprise architect, the answer is closer to: Gemini is the faster option; Muse is currently the slightly stronger high-effort agent by this independent measure.
Meta's evolving open weights stance
The more consequential issue for some developers may have little to do with today’s benchmark race.
When Meta launched Muse Code and Muse Spark 1.2 in August, VentureBeat noted how dramatically the company had moved away from the open-weight strategy that made Llama ubiquitous.
Muse Code and Spark 1.2 were proprietary, API-served products — a striking posture for the company that had spent years arguing that open AI was the path forward.
Five days later, Meta changed course again.
On August 10, it released the 30-billion-parameter Muse Glimmer under an Apache 2.0 license. Zuckerberg also said: “In the coming weeks, we are also going to open the weights for Muse Spark 1.2.” Reuters separately reported Meta’s plan to release the Spark 1.2 weights.
Now, Meta has instead shipped Muse Spark 1.3 as another proprietary model.
That does not yet amount to a broken promise — “coming weeks” can reasonably describe a period longer than three weeks. But today’s announcement makes the roadmap less clear rather than more.
Meta’s new post no longer says Muse Spark 1.2. It says its roadmap includes “the Muse Spark open weights release”, without identifying a version, release date, model size or license. Zuckerberg likewise said on X that “Muse Spark open weights releases” are coming soon.
For teams that standardized on Llama because downloadable weights meant self-hosting, customization and control over inference economics, that ambiguity may matter more than whether Spark gained another point on a composite benchmark.
Muse Spark 1.3 shows that Meta can now iterate proprietary frontier models at extraordinary speed. The shipping xhigh configuration is fast, competitively priced and much closer to the top of independent rankings than its predecessors. The max preview shows Meta can push the family a little further when allowed to spend more reasoning compute.
The next test is different: whether Meta can convert that pace into a roadmap enterprises can actually plan around — including making its best capabilities broadly deployable and delivering the open-weight Spark model it has already said is coming.
