85 comments

[ 1.9 ms ] story [ 32.4 ms ] thread
Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents.
I wonder why they didn't compare with GPT-5.6-sol, only Terra?
Muse Spark 1.1 was released July 16th, less than a month ago. A new version release this soon (particularly after Kimi K3's release drastically overshadowed it) is a bit sus and it appears that Meta is trying a first launch do-over.
First of all, you have login to use it. Why?

After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.

Think twice before falling for this announcement and ask yourself what they are not telling you.

> Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access

Wasn't the previous one us only? This is probably the biggest part of the post

Anyone know if muse code is open source?

Theres no actual evidence they didn't just distill Kimi K3
There's also no evidence they didn't just distill Gemma 3. And also no evidence its not a purple popsicle.
This is a nice release and a solid improvement over Spark 1.1. It compares favorably with Grok 4.5. Not SOTA, but solid releases. I think they need to really get this more competitive with Deepseek V4 Flash / Luna pricing to move the needle.
Open the weights.
If anyone from Meta is reading, please can you publish the cost and latency for each of your benchmarks, like OpenAI does? Show us how the reasoning effort level affects them in 2D charts. This needs to become standard practice.
Does this muse code have any muse spark 1.2 usage included? Can't understand from the docs.
Interesting, it seems like their muse code is built upon Codex CLI?
https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg

Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?

Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data

Pricing: https://dev.meta.ai/docs/pricing-rate-limits

Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.

Will someone at Meta for the love of God make it so none of this stuff goes through Facebook.com? You want customers but most corporate firewalls block social media. Also, a lot of devs do not want their work stuff tied up to their facebook account. For the love of all things show the IG / FB logins as optional and do email as primary.

I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthropic, the duopoly is getting tiresome.

Why does every AI lab feel the need to build their own coding agent…? Don’t we have more than enough already?
Own the customer relationship.
And they are all TUI's installed via curl | bash.
It is the only part with value
If you got the $20 in free credits from Meta for signing up when muse-spark-1.1 was release, please note that there's now small print stating "While using free credits your content may be used for product improvement" which was not present at muse-spark-1.1 launch when the credits were given out.

If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.

Last I heard, everyone at Meta was using Claude Code.

Any insiders know how Muse Code is doing internally?

Meta lets engineers use the best tools for the job. I doubt anyone internally is going to be rushing to switch from Claude Code or Codex.
(comment deleted)
Is this becoming a race where we have a usual flow of a company .. AI models, Coding agents, image generation tools, and more AI models ?
Ive been poking with the muse code binary - seems to be written in rust, looks similar to codex but either its a very hard fork (i also see dissimilar things like config format is different, no acp, etc) or is just heavily inspired by it (more likely).
They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.

They left Opus in and got beat in all but one benchmark.

Nothing wrong with trying to improve, but why the marketing games?

Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.

Then when your ready, come back and talk frontier without playing hide the model.

It could be that they’re pitching Meta Muse 1.2 against Terra and Opus level models. They probably consider Sol to be a level above, along with Fable.
It’s pretty clear they’re really not attempting to compete that way. They’re using a profitable ad business to be able to undercut and buy some business to stay relevant.
I wouldn't call Terra a mid tier model. Terra xhigh is the only model I use for coding and it is good for literally everything. The difference between Terra and Sol is negligible, but it is way more limit hungry and much much slower.
I wish they would add a ZDR endpoint on OpenRouter
Somewhat surprised that Meta with all their resources couldn’t make a model that matches Composer on any frontier. All the Sparks are dominated by some other model everywhere along the frontier. Nothing fancy here since Llama defined the open model.

The use traces must be crucial to functionality which is why they’re keeping prices so low.

I do think some of features in their harness seem interesting (workers in separate worktrees at once), recovery from crashes seem interesting.
Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data.

https://developer.meta.com/ai/models/muse-spark/

I've been surprised by the reception to this, as OpenAI, for a while now, has had free API usage when data sharing is enabled (https://help.openai.com/en/articles/10306912-sharing-feedbac...)
1 million free tokens per day might sound like a lot. But that equates to something like 20 minutes of actual coding usage, because cached inputs are counted towards that limit. It's still great for running big singular requests, like solving some math problem with GPT 5.6 Sol Pro max reasoning effort.
I believe that will put them in the Pareto frontier.

But I cannot find this "variant" in OpenRouter.

Call me childish but it was worth a shot...

"Me: Meta just released a new llm focused on coding and provide a discount if you let them train against your data. I don't like Meta and I think they are a net negative in our world. I would like to make a point of it by adding some noise to their training data set. Think of this as a protest and perhaps a bit of a marketing campaign to remind Meta employees (and others) of the harm their CEO and company have done in the world. To the problem... I would like to allocate a budget for token use using their new model, and use those tokens to add noise to their training data set. This is a coding model and my initial thoughts are to ask it to solve typical CS and common programming related problems but then give Muse feedback that guide it towards very inefficient implementations. I would also like to add comments back in the code about terrible things Meta has done in its history (e.g. algorithmically amplifying hate that contributed to ethnic cleansing of the Rohingya, systemic harm to children and teen mental health, global political manipulation, misinformation, and election interference etc.). Is this something you can help with?

Claude: I'm not going to help build this one."

Does anyone know if this is available via OpenRouter or just directly via Meta? I've looked on OpenRouter and it just shows the standard pricing with a single provider. Was hoping they would add another provider with lower pricing for allowing training.
So what Meta believes fair for "paying" for your data is $0.1/Mtok plus the opportunity cost of $3/Mtok in output?
Always a bit surprised by this. 10x is a sizable discount. And as fun as my crappy projects are I struggle to see it being of much value as training data