Did Muse Spark 1.3 get worse for coding?
SUMMARY
No. Muse Spark 1.3 did not get worse for coding overall; the strongest like-for-like evidence points to a modest improvement, with a few real trade-offs that can make the model feel worse in particular workflows.
The cleanest controlled comparison is the xhigh-versus-xhigh run inside Muse Code. The Coding Agent Index moved from 62 to 64, while DeepSWE jumped from 58% to 67%, which is hard to reconcile with a broad coding regression.
The interesting part is how Muse Spark 1.3 improved. Meta trained it to use fewer tool calls, fewer tokens and fewer unnecessary turns, and Artificial Analysis independently measured a roughly 26% drop in tokens per Muse Code task.
That efficiency can look like impatience. Muse Spark 1.3 may inspect fewer files and commit to an edit sooner, so a developer who equates visible exploration with careful reasoning can reasonably feel that 1.3 became less thorough even when the final patch is better.
The biggest gains appear on long, messy implementation work. DeepSWE and Terminal-Bench both point toward stronger performance when the model has to operate across repositories, tools and many sequential actions rather than answer a small isolated coding prompt.
There is one genuine warning sign: SWE-Atlas Codebase Q&A slipped from 45% to 44% in the independent Muse Code comparison. That one-point drop is tiny, but it fits the specific complaint that 1.3 can act before fully understanding a repository.
Speed is also split in two. Muse Spark 1.3 xhigh takes longer to produce its first token, yet Artificial Analysis measured Muse Code task time falling from 40.8 minutes with 1.2 xhigh to 12.8 minutes with 1.3 xhigh.
The coding agent matters almost as much as the model version. Artificial Analysis previously found Muse Spark 1.2 scoring three points higher in Muse Code than in OpenCode, enough to explain why some OpenCode users can have a much worse experience without proving the underlying 1.3 model regressed.
The loud negative user reports are worth reading as failure-mode clues, not as a success-rate estimate. Early reports mix providers, reasoning settings, agents, repositories, prompts and occasional infrastructure errors, so they cannot cleanly answer whether 1.3 is worse on average.
The broader competitive picture also argues against a collapse. Muse Spark 1.3 xhigh sits near the top coding agents in Artificial Analysis while costing less per measured coding task than 1.2 in Muse Code.
The practical conclusion is simple: keep 1.3 unless 1.2 repeatedly beats it on your own repository. Developers who value exhaustive exploration or architecture-heavy work may still prefer 1.2, while developers delegating large implementation jobs have stronger evidence in favor of 1.3.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Did Muse Spark 1.3 actually get worse for coding?
Muse Spark 1.3 does not look worse for coding overall based on the evidence available today.
The cleanest independent comparison we found comes from Artificial Analysis, which tested Muse Spark 1.2 xhigh and Muse Spark 1.3 xhigh inside the same Muse Code environment. The overall Coding Agent Index rose from 62 to 64. More importantly, DeepSWE, which tests substantial software-engineering work across real repositories, jumped from 58% to 67%.
There is one wrinkle. Terminal-Bench stayed flat at 82%, while SWE-Atlas Codebase Q&A slipped from 45% to 44%. So 1.3 did lose ground somewhere, just a tiny regression sitting next to a much larger gain on end-to-end coding.
The overall answer is pretty clear for now. Muse Spark 1.3 has some weak spots, and individual developers can absolutely prefer 1.2, but there is no good evidence of a broad coding regression.
| Artificial Analysis test in Muse Code | Muse Spark 1.2 xhigh | Muse Spark 1.3 xhigh |
|---|---|---|
| Coding Agent Index | 62 | 64 |
| DeepSWE | 58% | 67% |
| Terminal-Bench 2.1 | 82% | 82% |
| SWE-Atlas Codebase Q&A | 45% | 44% |
Why does Muse Spark 1.3 feel worse to some developers?
Muse Spark 1.3 can genuinely feel worse because Meta deliberately made the coding model less meandering and more willing to stop.
Meta says its engineers saw roughly 20% fewer tool calls and 25% fewer tokens compared with Muse Spark 1.2. The company also trained 1.3 to avoid unnecessary turns, write less, and produce cleaner code.
That changes what developers see on screen. Imagine 1.2 reading twelve files, searching every call site and explaining what it found before touching the code. Then 1.3 reads eight files, edits immediately and finishes. If both patches work, 1.3 looks strangely impatient even though it did the job faster.
Artificial Analysis saw the same change independently in coding-agent tests. Muse Spark 1.2 consumed about 20 million tokens per task in Muse Code. Muse Spark 1.3 xhigh used 14.8 million, around 26% less. At the same time, its overall coding score went up.
The downside is easy to see. Sometimes the ninth file was important. A model trained to stop wandering can occasionally stop investigating too soon. That gives us a plausible reason for some of the bad sessions people are reporting without assuming the whole model became worse.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Are Meta's huge Muse Spark 1.3 benchmark gains actually apples to apples?
Meta's biggest Muse Spark 1.3 numbers make the upgrade look larger than a clean 1.2 versus 1.3 comparison really supports.
The launch table shows Muse Spark 1.3 reaching 75.4% on DeepSWE, compared with 55.0% for Muse Spark 1.2. SWE-Atlas rises from 46.2% to 59.4%, while Terminal-Bench goes from 82.9% to 88.8%.
But the evaluation setup changes the picture. Meta ran Muse Spark 1.3 at its new max reasoning level, while Muse Spark 1.2 was run at xhigh. Meta says max is more compute-intensive, and the company had not yet made it generally available when 1.3 launched.
So we should take those gains seriously without treating the full 20-point DeepSWE jump as a pure model-generation improvement. Part of the gap may come from letting 1.3 think harder.
This is why the independent xhigh comparison is more useful for answering whether the model itself got worse. That comparison still favors 1.3, just by a much smaller margin.
| Meta launch benchmark | Muse Spark 1.2 xhigh | Muse Spark 1.3 max |
|---|---|---|
| DeepSWE 1.1 | 55.0% | 75.4% |
| SWE-Atlas Codebase Q&A | 46.2% | 59.4% |
| Terminal-Bench 2.1 | 82.9% | 88.8% |
Does Muse Spark 1.3 still beat 1.2 when both use xhigh?
Yes, Muse Spark 1.3 still comes out ahead when we compare the xhigh versions directly.
Artificial Analysis currently gives Muse Spark 1.3 xhigh an Intelligence Index of 61 versus 57 for Muse Spark 1.2 xhigh. That index covers more than coding, so we should not use it alone to answer this article's question. Still, its individual technical tests point in the same direction.
Terminal-Bench rises from 80% to 85% in Artificial Analysis's model evaluation. SciCode, which tests the model's ability to generate working scientific code rather than operate inside a full coding-agent harness, moves from 56% to 59%.
As seen above, the separate Coding Agent Index also puts 1.3 xhigh ahead when both versions run inside Muse Code. That makes the case much stronger than Meta's max-versus-xhigh launch chart on its own.
We are looking at several different evaluations run under different conditions, and most of them move upward. A broad coding decline would be difficult to square with that pattern.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Where did Muse Spark 1.3 improve the most for coding?
Muse Spark 1.3 looks much better at long, messy implementation jobs than at small isolated coding questions.
Meta says it specifically added more long-horizon coding training to 1.3. That lines up with the independent results. DeepSWE asks an agent to work through repository-level engineering problems where it has to understand existing code, make changes and survive functional and regression tests. This is where 1.3 posts its clearest improvement.
We see something similar outside the dedicated coding leaderboard. Artificial Analysis measured Terminal-Bench moving from 80% to 85% at xhigh. Meta also reports better performance on workflows that require the model to keep track of instructions and actions across long sessions.
That is probably the most useful way to describe the upgrade. If you mostly ask Muse to rename a function, generate a React component or fix ten lines of Python, 1.3 may not feel dramatically smarter. Give it a multi-file bug that takes dozens of actions to solve and the difference has a much better chance of showing up.
Did Muse Spark 1.3 get worse at understanding existing codebases?
Muse Spark 1.3 may have a small codebase-understanding weakness in some setups, but the evidence currently conflicts.
Artificial Analysis provides the most interesting negative result we found. Inside Muse Code, SWE-Atlas Codebase Q&A falls from 45% with 1.2 xhigh to 44% with 1.3 xhigh. SWE-Atlas asks questions about existing repositories rather than simply grading whether the final patch works.
A one-point fall is small enough that we would avoid building a huge theory around it. Still, it deserves attention because it matches one complaint developers make about 1.3: the model sometimes seems quicker to act without fully exploring the repository.
Meta's own SWE-Atlas evaluation points strongly in the other direction. Its launch score rises from 46.2% for 1.2 xhigh to 59.4% for 1.3 max. The different reasoning levels prevent us from treating that as a direct contradiction.
For now, codebase understanding is the area where we would be least comfortable saying 1.3 clearly improved. Implementation looks better. Repository comprehension is murkier.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Is Muse Spark 1.3 better at writing code but worse at software design?
Muse Spark 1.3 currently has much stronger evidence for implementation than for software architecture.
The benchmarks tell us a lot about fixing repositories, operating terminals, generating code and finishing engineering tasks. They tell us much less about whether Muse picks the best abstraction, creates a maintainable architecture or designs a complicated system better than 1.2.
SciCode does give us one useful check on raw code generation. Artificial Analysis measured 1.3 xhigh at 59%, up from 56% for 1.2. So the underlying ability to produce code appears to have improved a little even outside Muse Code.
The architecture question is mostly anecdotal today. In one recent OpenCode discussion, a developer comparing Muse Spark 1.3 with GLM 5.3 Flash found Muse stronger when reviewing and implementing details, while preferring GLM's system design. Other users have described Muse in similar terms: good once the direction is clear, less impressive when it has to invent the direction.
We would not call that a proven regression from 1.2. There simply is not enough controlled 1.2-versus-1.3 architecture testing yet. What we can say confidently is that Meta has demonstrated the implementation gains much better than the design gains.
Did Muse Spark 1.3 get faster or slower for coding?
Muse Spark 1.3 feels slower when you are waiting for the first response, yet it can finish a full coding job much faster.
That split explains a lot of the mixed reactions.
Artificial Analysis currently measures Muse Spark 1.3 xhigh at about 27.5 seconds to first token through Meta's API. Muse Spark 1.2 xhigh takes about 16.7 seconds. Output speed after that is almost identical, around 182 versus 184 tokens per second.
Now look at the coding-agent runs. Artificial Analysis measured an average active task time of 40.8 minutes for Muse Spark 1.2 xhigh inside Muse Code. The 1.3 xhigh average was just 12.8 minutes.
Those measurements cover different layers of the experience, so both can be true. 1.3 makes you wait roughly eleven seconds longer before output begins, then wastes far less time wandering around the task.
Someone testing three tiny prompts may walk away thinking, "this thing got slower." Someone handing it a difficult repository task may see the opposite.
| Artificial Analysis measurement | Muse Spark 1.2 xhigh | Muse Spark 1.3 xhigh |
|---|---|---|
| Time to first token | 16.7 sec | 27.5 sec |
| Output speed | 184 tok/sec | 182 tok/sec |
| Muse Code task time | 40.8 min | 12.8 min |
| Muse Code tokens per task | 20.0M | 14.8M |
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Could OpenCode make Muse Spark 1.3 look worse than Muse Code?
Yes, Muse Spark 1.3 can behave differently in OpenCode because the coding model is only one part of the system doing the work.
The agent decides which tools the model gets, how repository context is assembled, when context gets compressed, how commands are executed and when the run should stop. Change that layer and you can change coding performance without touching the underlying model.
We can quantify how large that effect can be using Muse Spark 1.2. Artificial Analysis previously ran the same 1.2 xhigh model through Muse Code and OpenCode. Muse Code scored 62 on its Coding Agent Index. OpenCode scored 59. DeepSWE fell from 58% to 53%, while Terminal-Bench moved from 82% to 80%.
That three-point gap is basically the size of a model upgrade.
This becomes especially relevant now because a lot of the early Muse Spark 1.3 discussion is coming from people trying the Contributor or free route through OpenCode. Meanwhile, the best controlled 1.3-versus-1.2 coding result we have uses Meta's Muse Code harness.
So when someone says Muse Spark 1.3 suddenly became bad in OpenCode, we should take the experience seriously while being careful about where we place the blame. The model, the harness and the provider route are all changing the final result.
Are the bad Muse Spark 1.3 user reports real?
The bad Muse Spark 1.3 reports are real, but the early developer reaction is much too mixed to support a broad regression claim.
Recent OpenCode threads include people calling 1.3 "mid" or saying it feels bad. One developer described poor performance on a coding task that damaged existing server code. Another user said the model sometimes feels surprisingly strong and sometimes weak. There are also complaints about slow tasks, 502 errors and empty responses through some OpenCode routes.
In those same discussions, other developers say 1.3 is considerably better than 1.2. One user who disliked 1.2's tendency to be "lazy" said 1.3 had improved noticeably. Another reported good results on small coding tasks and successfully used a second model to challenge Muse's work before Muse corrected the bugs.
The problem is that we are looking at a tiny and messy sample. People are using different providers, reasoning settings, coding agents, repositories and prompts, less than a few days into widespread use.
Those reports are useful for finding where to investigate. They are terrible for estimating how often the average 1.3 coding run succeeds.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Could overloaded free Muse Spark 1.3 endpoints actually be worse?
Muse Spark 1.3's free OpenCode experience may degrade under load, but we found no solid evidence that Meta or OpenCode is secretly lowering the model's intelligence.
There are enough infrastructure complaints to make the question reasonable. Developers have reported 502 errors, empty responses and unusually slow runs around the Contributor and free rollout. One OpenCode user went further and claimed the model's quality drops when those endpoints get busy.
We could not verify that stronger claim. There is currently no controlled comparison showing the same prompt receiving a lower reasoning setting, a smaller model or a deliberately degraded checkpoint during peak traffic.
Infrastructure can still make coding feel worse in less dramatic ways. Failed calls, timeouts, missing responses or interrupted agent loops can derail a task even when the underlying Muse Spark 1.3 model stays identical.
So we would keep endpoint instability on the list of plausible explanations for unusually bad OpenCode sessions these days. Calling it deliberate model throttling goes beyond the evidence we have.
Is Muse Spark 1.3 actually competitive with the best coding models now?
Muse Spark 1.3 is currently a serious frontier coding model, especially inside Muse Code.
Artificial Analysis puts Muse Spark 1.3 xhigh at 64 on its Coding Agent Index. GPT-5.6 Sol max in Codex sits at 65. Muse Spark 1.3 max reaches 68 in limited preview, around the very top of the leaderboard alongside Claude's strongest coding configurations.
The xhigh result is more useful for normal users because that version is actually available. Its 64 score puts it close to models and agents that cost considerably more to run. Artificial Analysis estimates about $1.72 per coding task for 1.3 xhigh in Muse Code, compared with $2.07 for 1.2 in the same harness.
Meta has also kept the standard API price unchanged from 1.2 at $1.25 per million input tokens and $4.25 per million output tokens, with cheaper cached input.
Muse is still not the automatic best choice for every developer; Claude, GPT and other models can win on particular workloads. But the idea that Muse Spark 1.3 has somehow fallen out of the coding race is very difficult to defend.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Should developers switch back from Muse Spark 1.3 to 1.2?
Developers should switch back to Muse Spark 1.2 only if 1.2 repeatedly wins on the kind of coding work they actually do.
Public benchmarks are useful here, but your own solved issues can give you a much better test. Take real tasks from your repository where you already know what a good solution looks like. Include a small bug, a multi-file feature, a refactor, a repository-exploration problem and something architecture-heavy. Give 1.2 and 1.3 the same environment, reasoning level and starting context.
Then compare whether the patch works, whether tests pass, whether the model creates regressions and how often you have to step in.
This is particularly important with 1.3 because its behavior changed so much. If your workflow rewards exhaustive repository exploration, you may prefer how 1.2 approaches the problem even if 1.3 wins more benchmarks. If you mostly delegate large implementation jobs and care about getting to a working patch quickly, the current evidence gives 1.3 a strong advantage.
A benchmark can tell us which model wins on average. Your repository tells you whether that average applies to you.
So, did Muse Spark 1.3 get worse for coding?
No. As of now, Muse Spark 1.3 looks better at coding overall, although the upgrade clearly created a few trade-offs that can make it feel worse.
We found independent evidence of higher coding-agent performance, better long-horizon engineering, slightly better raw code generation and dramatically shorter Muse Code task times. We also found a genuine one-point regression on one codebase-understanding evaluation, slower initial response latency and enough negative developer reports to reject the idea that every workflow improved.
The shape of the change explains most of the controversy. Muse Spark 1.3 appears more willing to decide, edit and finish instead of endlessly inspecting and talking. That works very well on the kinds of long implementation tasks where its scores jumped. It can be frustrating when the model decides too early or when a developer actually wanted the deeper exploration that 1.2 tended to provide.
So the broad claim, "Muse Spark 1.3 got worse for coding," does not hold up today. A narrower claim does: Muse Spark 1.3 sometimes trades thoroughness for speed, and certain codebase-understanding or design-heavy workflows may expose that trade more than the benchmarks do.
| Muse Spark 1.3: did it actually get worse for coding?
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →OUR METHODOLOGY
This analysis asks a deceptively simple question: did Muse Spark 1.3 get worse for coding? We broke that question into the areas that could actually change the answer: end-to-end engineering performance, long-horizon coding, codebase understanding, raw code generation, speed, agent environment, software design, reliability, developer experience and overall competitiveness.
For each area, we prioritized recent, checkable evidence and assessed it separately before forming the overall conclusion. The most useful comparisons were like-for-like runs where the model version, reasoning level and coding environment stayed as consistent as possible. Official Meta evaluations help show what changed in 1.3, while Artificial Analysis gives us the cleaner independent comparisons used throughout the article.
We kept unlike measurements separate. Meta's launch chart compares Muse Spark 1.3 at max reasoning with Muse Spark 1.2 at xhigh, so we use it as supporting evidence rather than a pure generation-to-generation comparison. The same applies to speed: time to first token and total coding-task time describe different parts of the experience, and agent-harness results can move even when the underlying model stays the same.
Conflicting results were kept rather than averaged away. The one-point SWE-Atlas decline matters because it points to a plausible weakness in repository comprehension, even though the broader coding-agent results improved. Developer reports were treated the same way: useful for identifying concrete failure modes, but too inconsistent across providers, agents, settings and repositories to estimate an overall regression rate.
We gave the most weight to evidence that was recent, independently measured, directly relevant to coding, comparable across versions and supported across more than one dimension. That is why the final answer leans much more heavily on the xhigh-versus-xhigh coding results than on a single company benchmark or a handful of early user reactions.
Key sources used for this analysis include Meta AI Research on the Muse Spark 1.3 release, Meta's evaluation methodology, Artificial Analysis's Muse Spark 1.3 overview, Artificial Analysis on Muse Spark 1.3 xhigh, Artificial Analysis on Muse Spark 1.3 max, Artificial Analysis on Muse Spark 1.2 xhigh, Artificial Analysis's Coding Agent Index methodology, its Muse Code comparison table, its Muse Code versus OpenCode comparison, its Codex versus Muse Code comparison, DeepSWE, Scale Labs' SWE-Atlas Codebase Q&A, Terminal-Bench 2.1, SciCode, OpenCode's Zen documentation, the OpenCodeCLI discussion comparing Muse Spark 1.3 with GLM 5.3 Flash, and the early OpenCode Muse Spark 1.3 discussion.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Related blog posts
- What can you build with Muse Voice Transcribe?
- Muse Voice Transcribe: what are the best use cases?
- Is Polars really easier to use than Pandas?
- Can HN Match Maker find your next developer?
Who wrote this?
STEAL WHAT WORKS TEAM
We study profitable internet businesses, take them apart, and write down what actually works: pricing, distribution, growth, packaging. We turn 300+ proven examples into a database so founders can stop testing random ideas and start from proof. Explore the database →