TL;DR
Meta has shipped Muse Spark 1.3, its strongest model to date, and says it has drawn level with the leading labs. Alexandr Wang, Meta’s chief AI officer, rates it a match for Claude Fable 5.1 and stronger than GPT-5.6 Sol at generating code. Independent scoring broadly supports the claim, with one caveat.
What the numbers show
Vendor benchmark claims deserve the usual scepticism, since labs choose which tests to publish. Here there is outside corroboration: Artificial Analysis scored the model 62 on the Intelligence Index it publishes, placing it ahead of everything OpenAI has shipped and behind only two Anthropic models, Opus 5 and Fable 5.1. Third among frontier models is a genuine change in Meta’s standing — though “caught up” is doing some work when two rivals still sit above it.
Efficiency is the more practical claim. Wang says 1.3 completes comparable work using around a quarter fewer tokens than the August release, holds context across several tasks at once, and copes better with long instructions. It also seeks confirmation before doing anything irreversible — the same containment reflex now visible across the industry.
Pricing stays flat against 1.2, consistent with Zuckerberg’s stated aim of making this family among the cheapest available. Developers get at it via Meta’s own model API, and Wang claims some are already consuming trillions of tokens weekly.
The open-weights question
This is the fourth release in five months, and the cadence is being funded by extraordinary spending — including the $14bn (£11bn) paid last year for a ScaleAI stake and Wang himself, who now runs Superintelligence Labs. Investors want to see a return.
That pressure explains the retreat from the openness Llama was built on. Meta has not decided whether to publish weights for 1.3, and the weights promised for 1.2 have still not appeared, even as Zuckerberg publishes arguments for accessible AI development. A larger model, Watermelon, remains undated.
Looking forward
For UK buyers the interesting variable is price, not the leaderboard. A credible third-place model priced deliberately below Anthropic and OpenAI puts pressure on inference budgets across the market, which matters more to a mid-sized British firm than two index points. The open-weights retreat cuts the other way, though — it narrows the pool of frontier-class models that a company can self-host, which is precisely the gap Cambridge’s Flower Labs is targeting with a British model.