Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

VentureBeat
Published
Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

The short version

  • "Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter," Meta co-founder and CEO Mark Zuckerberg wrote on X…
  • There is substance behind both parts of that claim.
  • Muse Spark 1.3 makes significant gains over last month’s 1.2 release , particularly on long-running agent tasks.
  • The version developers can access now is also one of the strongest price-performance offerings near the top of independent model rankings.
  • Meta’s strongest Muse Spark 1.3 benchmark results come from its max reasoning configuration.

The story

Meta’s newest AI model Muse Spark 1.3, unveiled yesterday, is faster and more performant on third-party benchmarks than its predecessor — with a caveat.

"Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter," Meta co-founder and CEO Mark Zuckerberg wrote on X, calling it Meta’s “biggest jump” yet in coding and agentic work.

There is substance behind both parts of that claim. Muse Spark 1.3 makes significant gains over last month’s 1.2 release, particularly on long-running agent tasks. The version developers can access now is also one of the strongest price-performance offerings near the top of independent model rankings.

Meta’s strongest Muse Spark 1.3 benchmark results come from its max reasoning configuration. Meta says that version is still completing additional safety testing and will arrive “shortly”; the third-party benchmarking firm Artificial Analysis says it evaluated max in a limited partner preview, and currently lists no API provider at all for the configuration.

The version broadly rolling out this week through its Muse Code harness and the Meta Model API uses Meta’s previously available reasoning settings, including xhigh.

That makes the more relevant enterprise question not whether Muse Spark 1.3 can reach frontier territory, but how close the model companies can actually deploy today gets — and at what real cost.

The shipping model is very good, but not the benchmark leader

Meta does disclose results for both configurations in its underlying evaluation report, so this is not a case of the company hiding the deployable model. But its launch materials prominently showcase the max variant, and some of the largest scores belong to that configuration.

For example, Meta reports GDPval-AA v2 scores of 1,754 Elo for max versus 1,709 for xhigh, OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2.

On some tests the distinction is negligible or reversed: DeepSearchQA is tied at 89.4, while xhigh scores 89.2 on Terminal-Bench 2.1 versus max at 88.8.

Artificial Analysis scores Muse Spark 1.3 max at 62 on its Intelligence Index and the shipping xhigh version at 61. The latter ties GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high. But Anthropic still occupies the top of the leaderboard: Claude Fable 5.1 reaches 66 at max and 65 at xhigh, while Claude Opus 5 reaches 63 at max and xhigh.

In other words, Muse Spark 1.3 xhigh is legitimately in the frontier cluster, but it is not the model currently setting the frontier.

That is still a substantial change from Muse Spark 1.2. VentureBeat’s coverage of last month’s launch found Meta fielding a credible coding challenger that nevertheless generally trailed Anthropic’s best model. Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1 versus Opus 5’s 86.7%, and also finished behind Opus on the other main coding comparisons Meta presented.

With 1.3, Meta is no longer merely showing up in that contest. On several coding and agentic evaluations, it is trading wins with OpenAI and Anthropic.

Meta says the underlying model has also become easier to operate. Muse Spark 1.3 is trained to maintain multiple workflows in a long thread, gather context with tools, detect gaps in its own plans, ask users for clarification when necessary and confirm before consequential actions. In Meta engineers’ internal comparisons, it used roughly 20% fewer tool calls and 25% fewer tokens than 1.2 during coding work.

For enterprises paying for thousands or millions of agent loops, those behavioral improvements could matter more than another leaderboard point.

‘Almost too cheap to meter’ does not mean Meta cut its prices

Muse Spark 1.3 did not receive an API price cut. Meta kept Standard pricing exactly where it was for Muse Spark 1.2: $1.25 per million input tokens, $4.25 per million output tokens and $0.15 per million cached input tokens.

Model

Input ($/1M)

Output ($/1M)

Total ($/1M)

Source

Muse Spark 1.2 / 1.3 Contributor

$0.10

$0.20

$0.30

Meta

MiMo-V2.5 Flash

$0.10

$0.30

$0.40

Xiaomi

DeepSeek-V4-Flash — off-peak

$0.22

$0.66

$0.88

DeepSeek

GPT-5.6 Luna

$0.20

$1.20

$1.40

OpenAI

MiniMax-M3

$0.30

$1.20

$1.50

MiniMax

LongCat-2.0 — limited-time promo

$0.30

$1.20

$1.50

LongCat

DeepSeek-V4-Flash — peak hours

$0.44

$1.32

$1.76

DeepSeek

MiMo-V2.5

$0.40

$2.00

$2.40

Xiaomi

DeepSeek-V4-Pro — off-peak

$0.66

$1.98

$2.64

DeepSeek

LongCat-2.0 — standard

$0.75

$2.95

$3.70

LongCat

MiMo-V2.5 Pro (≤256K)

$1.00

$3.00

$4.00

Xiaomi

Gemini 3.7 Flash — through Dec. 31, 2026

$0.75

$3.75

$4.50

Google

Gemini 3.8 Flash — through Dec. 31, 2026

$0.75

$3.75

$4.50

Google

DeepSeek-V4-Pro — peak hours

$1.32

$3.96

$5.28

DeepSeek

Muse Spark 1.1 / 1.2 / 1.3

$1.25

$4.25

$5.50

Meta

GLM-5.3

$1.40

$4.40

$5.80

Z.AI

Grok 4.6 — <200K prompt tokens

$2.00

$6.00

$8.00

xAI

MiMo-V2.5 Pro (>256K)

$2.00

$6.00

$8.00

Xiaomi

Qwen3.8-Max

$2.00

$6.00

$8.00

QwenCloud

Gemini 3.7 Flash — starting Jan. 1, 2027

$1.50

$7.50

$9.00

Google

Gemini 3.8 Flash — starting Jan. 1, 2027

$1.50

$7.50

$9.00

Google

GPT-5.6 Terra

$2.00

$12.00

$14.00

OpenAI

Grok 4.6 — ≥200K prompt tokens

$4.00

$12.00

$16.00

xAI

GPT-5.4

$2.50

$15.00

$17.50

OpenAI

Kimi K3

$3.00

$15.00

$18.00

Moonshot AI

Claude Opus 5

$5.00

$25.00

$30.00

Anthropic

Sakana Fugu Ultra (≤272K)

$5.00

$30.00

$35.00

Sakana AI

GPT-5.6 Sol — Standard mode

$5.00

$30.00

$35.00

OpenAI

Claude Fable 5 / Claude Mythos 5

$10.00

$50.00

$60.00

Anthropic

Claude Fable 5.1 / Claude Mythos 5.1

$10.00

$50.00

$60.00

Anthropic

GPT-5.6 Sol — Fast mode

$10.00

$60.00

$70.00

OpenAI

That makes Zuckerberg’s “almost too cheap to meter” line less a statement about lower token prices than about what Meta believes developers can accomplish with those tokens.

Artificial Analysis offers evidence for that argument, but also a complication. It measures Muse Spark 1.3 xhigh at 235.2 output tokens per second and estimates a cost of $0.55 per Intelligence Index task.

At 61 on the Intelligence Index, that gives it the lowest cost per task of any currently measured model at that intelligence level.

Muse Spark 1.2 cost only $0.40 per Artificial Analysis task, while scoring 57.

Despite unchanged per-token pricing, the independent benchmark’s cost of completing an average task therefore increased generation-over-generation. Artificial Analysis attributes the increase primarily to heavier input-token consumption on agentic evaluations.

That does not directly contradict Meta’s claim of 25% lower token use: Meta is describing comparisons in its own coding workflows, while Artificial Analysis is measuring a broader suite of reasoning and agentic tasks.

But it illustrates why “cheap” becomes slippery once models operate as agents. Token rates, reasoning effort, number of turns, tool calls and retries all contribute to the actual cost of finishing work.

Meta also retains its unusually cheap Contributor tier — $0.10 per million input tokens and $0.20 per million output tokens — in exchange for permission to use prompts and completions for training.

As VentureBeat noted with Muse Spark 1.2, that may be attractive for prototyping but creates a materially different data-governance calculation for enterprises working with proprietary code or sensitive internal information.

Wang’s ‘Gemini who?’ lands on an unusually close comparison

Meta chief AI officer Alexandr Wang was considerably less qualified in celebrating the release.

After Artificial Analysis posted its Muse Spark results, Wang reposted them on X, adding: “i really hate to say it, but… gemini who? 😱💨”

The shade was particularly pointed because Google released Gemini 3.8 Flash on the same day, pitching it at almost exactly the same class of workload: long-horizon software engineering, autonomous agents and multi-step professional reasoning. Google calls 3.8 its best reasoning and coding Flash model yet and says it is the company’s third Flash release in six weeks.

Independent numbers give Wang something to work with, though hardly a knockout.

Artificial Analysis gives Muse Spark 1.3 xhigh a 61 Intelligence Index score at $0.55 per task, compared with 59 and $0.58 for Gemini 3.8 Flash at high reasoning. Meta therefore edges Google on both intelligence and task cost at those particular settings.

Google wins decisively on throughput. Artificial Analysis measures Gemini 3.8 Flash high at about 305 output tokens per second, versus 235 for Muse Spark — roughly 30% faster. Gemini also has the lower raw API sticker price for now: Google is charging an introductory $0.75 per million input tokens and $3.75 per million output tokens, compared with Meta’s $1.25 and $4.25.

That promotional Google pricing expires December 31, after which it rises to $1.50 per million input tokens and $7.50 per million output tokens.

The result is a useful snapshot of how tight frontier-model economics have become. Meta currently wins this independent comparison by two Intelligence Index points and three cents per benchmark task; Google offers substantially higher output throughput and cheaper raw tokens during its launch promotion.

Wang’s “gemini who?” is fun executive trash talk. For an enterprise architect, the answer is closer to: Gemini is the faster option; Muse is currently the slightly stronger high-effort agent by this independent measure.

Meta's evolving open weights stance

The more consequential issue for some developers may have little to do with today’s benchmark race.

When Meta launched Muse Code and Muse Spark 1.2 in August, VentureBeat noted how dramatically the company had moved away from the open-weight strategy that made Llama ubiquitous.

Muse Code and Spark 1.2 were proprietary, API-served products — a striking posture for the company that had spent years arguing that open AI was the path forward.

Five days later, Meta changed course again.

On August 10, it released the 30-billion-parameter Muse Glimmer under an Apache 2.0 license. Zuckerberg also said: “In the coming weeks, we are also going to open the weights for Muse Spark 1.2.” Reuters separately reported Meta’s plan to release the Spark 1.2 weights.

Now, Meta has instead shipped Muse Spark 1.3 as another proprietary model.

That does not yet amount to a broken promise — “coming weeks” can reasonably describe a period longer than three weeks. But today’s announcement makes the roadmap less clear rather than more.

Meta’s new post no longer says Muse Spark 1.2. It says its roadmap includes “the Muse Spark open weights release”, without identifying a version, release date, model size or license. Zuckerberg likewise said on X that “Muse Spark open weights releases” are coming soon.

For teams that standardized on Llama because downloadable weights meant self-hosting, customization and control over inference economics, that ambiguity may matter more than whether Spark gained another point on a composite benchmark.

Muse Spark 1.3 shows that Meta can now iterate proprietary frontier models at extraordinary speed. The shipping xhigh configuration is fast, competitively priced and much closer to the top of independent rankings than its predecessors. The max preview shows Meta can push the family a little further when allowed to spend more reasoning compute.

The next test is different: whether Meta can convert that pace into a roadmap enterprises can actually plan around — including making its best capabilities broadly deployable and delivering the open-weight Spark model it has already said is coming.

Read the full story at VentureBeatOriginal

Related Markets

All Markets
View full chart →
View Full Chart
View full chart →
View Full Chart
View full chart →
View Full Chart

Market data may be delayed. Not financial advice.

How other outlets covered this

Compare all

Alto found this story at 3 outlets. Same event, different framing — compare the headlines.

How this story developed

Full timeline

Alto has tracked this across 9 days of coverage from 3 outlets.

    • Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yetYou are here

Powered by Gab AI

The Story At A Glance

Reading this article now — analysis appears below

Reading the article

💡 AI analysis provides alternative perspectives on current events

Up next

Related coverage from across the outlets Alto indexes.

Questions Alto can answer

From this story — each link opens a live data page or a tool already filled in.

  1. What is $100 from 1990 worth today?CPI-adjusted dollars — result on the next page
  2. Where does a $75,000 household income rank nationally?Census percentile — national and state
  3. What's Alto covering on the Tech & AI desk?Latest headlines on this beat

All toolsAll topicsSource directoryStory timelinesHeadline comparisonSearchMost read

From Gab Shop

Official merchandise. Every order funds free speech infrastructure.

Shop all products

Install Alto on your phone

Add Alto to your home screen for breaking news — no app store, no account.

  1. Step 1Open alto.gab.com in SafariMust be Safari — not Chrome or in-app browsers
  2. Step 2Tap the Share buttonSquare with an arrow, at the bottom of Safari
  3. Step 3Tap "More"If you don’t see Add to Home Screen yet
  4. Step 4Tap "Add to Home Screen"Scroll the share sheet if you need to
  5. Step 5Tap "Add"Alto appears on your home screen like any other app.
gab

Talk Big Tech Where Big Tech Can't Reach

AI, surveillance, and censorship, covered by the people the platforms removed first.

What Makes Gab Different

We're not just another social network. We're a platform built on principles that matter.

Freedom of Speech & Reach

All First Amendment protected speech is welcome. No algorithmic throttling or shadow banning.

Family-Friendly Platform

We maintain a clean environment. Explicit adult content is strictly prohibited.

Western Nations Only

Third-world IPs are blocked. No scammers, no spam farms. Built for Western civilization.

Funded By Users

Our users are our investors and customers. You're not the product being sold.

Battle Tested

A decade of standing strong. Banned from app stores, banks—and still here.

American Owned & Operated

We reject foreign censorship demands. Built by Americans, for free people.

Support Alto & Gab

Alto is funded entirely by readers like you. Your donation helps us continue delivering curated news from a right-wing Christian Nationalist perspective, powered by Gab AI.