Muse Spark 1.3: Meta's new model is now on ChatPlus
For years Meta gave its Llama models away for free. Now it builds a closed model family, Muse, and competes head-on with OpenAI and Anthropic. Here is where Muse Spark 1.3 came from, what it can do, what its numbers mean and which tasks it suits. The model is already available on ChatPlus.

A year and a half ago, Meta was "the company that gives its models away." Millions of people downloaded Llama, startups built products on it, teams fine-tuned it for their own work. But in the race for the smartest model, Meta was clearly behind OpenAI, Anthropic and Google.
Today it has a different strategy and a different model family: Muse. The models are closed, a dedicated division builds them, and new versions ship almost every month. The latest, Muse Spark 1.3, came out on September 2 and is available on ChatPlus starting today. Below: where it came from, what it can do, what is behind its numbers and when to pick it.
How Meta got to Muse
In April 2025 Meta released Llama 4, and the reception was lukewarm. The models were decent but fell short of the promises, and the release of the largest version was postponed. For a company spending tens of billions of dollars a year on AI, that was an unpleasant signal.
That summer Zuckerberg rebuilt the whole effort. Meta created a new division, Meta Superintelligence Labs, hired strong researchers away from OpenAI, Google and other labs, and started building new data centers for training large models. The stated goal was a bold one: "personal superintelligence," an AI that helps each person with their own goals.
The first result was Muse Spark. It launched on April 8, 2026 and went straight into the Meta AI assistant, on the web and in the app. It was a multimodal reasoning model: it could call tools, reason over images and run several agents in parallel. Developers barely got access at first, only through a private preview.
After that, things moved fast:
- July: Muse Spark 1.1 and a public API preview for developers;
- August: Muse Spark 1.2 and Muse Code, Meta's own coding agent that runs right in the terminal;
- September 2: Muse Spark 1.3, the subject of this article.
Three versions in three months is a very fast pace for a large lab.
What Muse Spark 1.3 can do
Meta describes the new version this way: the model is trained for long, multi-step tasks. Behind that phrase are three concrete skills.
First, it keeps track of the work. On a long task the model remembers what it has already done and what came of it, instead of starting each step from scratch. If you ask it to go through a project, fix the tests and update the docs, it will not forget by step three what it changed in step one.
Second, it copes with mess. Real tasks rarely arrive tidy: the spec says one thing, the email thread another, the code a third. Muse Spark 1.3 was trained to notice contradictions like these and work through them rather than pick one version at random.
Third, it asks. When information is missing, the model checks with you instead of filling the gap with guesses. It sounds minor, but guessing is exactly why agents most often go off track.
Meta also highlights coding. Version 1.3 was refined on real work in Muse Code, and according to the company it takes fewer unnecessary steps and writes more compact code. If you use AI as a programming partner, you feel it: less to clean up after the model.
The model understands more than text. You can show it a screenshot, a photo, a diagram or a PDF. Meta mentions an interesting detail: Muse Spark's visual reasoning runs through a real execution environment. Put simply, the model can process an image with code rather than just "look" at it and estimate by eye.
The model's context window is 1,048,576 tokens, more than a million. That is hundreds of pages of text or a large repository in one go. More importantly, the model actually uses that window well, as the numbers below show.
What the numbers say
Meta measured all of the results below itself, in its maximum reasoning mode.
| Test | What it measures | Muse Spark 1.3 |
|---|---|---|
| Terminal-Bench 2.1 | Agent work in a terminal | 88.8% |
| DeepSWE 1.1 | Long engineering tasks in real repositories | 75.4% |
| MRCR v2, 512K–1M | Finding the right detail in a very long context | 98.1% |
| JobBench | Professional work with tools | 64.9% |
| GDPval-AA | Office tasks: reports, spreadsheets, documents | 1754 Elo |
The long-context result is the most impressive. MRCR checks whether a model can find the right fragment among many similar ones hidden in a huge text. Plenty of models formally accept a million tokens but start "losing" the middle at that scale. Muse Spark 1.3 scores 98.1% in the range from half a million to a million tokens, with almost nothing lost.
The gain over the previous version is noticeable too. On JobBench the model rose from 61.6% to 64.9%, and on the office benchmark GDPval-AA from 1615 to 1754 Elo points.
Now the comparison with competitors, which needs a caveat. On GDPval-AA, Muse Spark 1.3 beats GPT-5.6 Sol (1710) but trails Claude Opus 5 (1824). Meta, however, compares itself with the previous generation. Claude Opus 5.5 and GPT-6 Sol have come out since, and Meta's tables do not include them. So the fair summary is this: Muse Spark 1.3 plays in the same league as the flagships but does not beat the newest of them.
What it costs
The price is attractive: $1.25 per million input tokens and $4.25 per million output tokens. For comparison, Claude Opus 5.5 costs $4 and $20, GPT-6 Sol $2 and $10, Grok 4.7 $2 and $6. For performance close to flagship level, that is one of the best deals on the market.
Meta also offers a second, almost free option: 10 cents for input and 20 for output. At that price, though, Meta gets the right to train its models on your prompts and responses. On ChatPlus we connected only the standard tier, so your conversations with Muse Spark are not used for Meta's training.
Where Muse Spark shines, and where to pick something else
Muse Spark 1.3 is strongest where a task is long and tangled. Going through a big project, finding the cause of a bug in someone else's code, reconciling several contradictory documents into one conclusion, pulling what you need out of a thick report: that is its sweet spot. It also works well with screenshots; showing it an interface with an error and asking it to figure things out is a perfectly normal use.
It has weak spots too. It is a reasoning model, so it answers simple questions more slowly than lightweight models: it thinks first, then writes. For translations, summaries and quick lookups, pick something lighter, such as GPT-6 Luna. Meta says almost nothing about creative writing, so if style matters to you, compare it with Claude on your own task.
How to try it on ChatPlus
Muse Spark 1.3 is available on the PRO plan and, like other powerful models, it uses your allowance — the 5-hour limit from which requests to expensive models are deducted. Open a chat with Muse Spark 1.3.
The most useful thing is to compare it with other models on your own task. On ChatPlus you can switch models in the middle of a conversation: ask Muse Spark, switch to Claude or GPT, send the same request and see right away whose answer is better. The model's specs and benchmark results are collected on a dedicated Muse Spark 1.3 page.
In short
Muse Spark 1.3 is the newest model from Meta Superintelligence Labs, released on September 2, 2026. It is strong in long agentic tasks, code and very large contexts: 88.8% on Terminal-Bench 2.1 and almost no loss at a million tokens. It still falls short of the newest Anthropic and OpenAI flagships, but it costs several times less. On ChatPlus the model is already available on the PRO plan.