Muse Spark 1.1: A Million-Token Agent That Wants Your Workflow
Muse Spark 1.1 AI is Meta’s new multimodal AI model designed to act as an agent that plans, delegates, and executes complex workflows across coding, computer use, and knowledge tasks while using a one-million-token context window to keep long-running projects and large codebases in working memory and reduce the need for humans to manually restitch information across sessions. The important takeaway is that Meta is no longer chasing the frontier with yet another chat model; it is targeting enterprise AI capabilities where speed, context, and tool control matter more than clever conversation. By orchestrating multi-agent systems to complete complex projects faster, Muse Spark 1.1 focuses on practical automation over flashy demos. That shift, backed by the dedicated Meta Superintelligence Labs, makes this release less about catching up and more about redefining how agentic AI reasoning should work inside real organizations.

Benchmarks: Beating Google, Pressuring Opus and GPT on Real Work
On numbers, Muse Spark 1.1 does not dominate every leaderboard, but it lands where it matters: applied reasoning and tools. The model scores 51 on the Artificial Analysis Intelligence Index, ahead of Google’s Gemini 3.5 Flash at 50 and Gemini 3.1 Pro Preview at 46, effectively tying GPT-5.6 Luna (max) and GLM-5.2 (max) at the top of the mid-tier pack. On Humanity’s Last Exam with tools, it hits 45%, edging past GPT-5.5 at 44% and nearly matching Claude Opus 4.8 at 46%. "Muse Spark 1.1’s score on AA-Omniscience more than quadrupled, from 4 to 18," a jump that signals fewer hallucinations alongside higher accuracy. Agentic tool use is where the model shines: 88.1 on MCP Atlas and 54.7 on JobBench, ahead of both Opus 4.8 and GPT 5.5, while remaining competitive—but not leading—on several coding and computer-use tests. This pattern says Muse Spark 1.1 is tuned for doing work with tools, not winning abstract math contests.

Million-Token Context and Agentic Computer Use: Built for Long-Haul Automation
The most meaningful architectural decision is the jump to a one-million-token context window, up from 262,000 in the original Muse Spark. Muse Spark 1.1 decides on its own what to remember, retrieve, or compress as sessions grow, preserving key details from earlier work instead of forcing humans to recap them. That design is central to agentic AI reasoning: the model can act as a main agent that plans and delegates to subagents, or as a subagent that knows when to escalate tasks back. On computer use, it is trained to decide whether a script or direct UI interaction is faster, to generate batches of actions per step, and to handle workflows that span several applications with minimal human intervention. A demo shows it using a smartphone video to pull product photos and details, then operating a browser to list the item on a marketplace automatically—exactly the kind of end-to-end flow enterprises want from automation rather than isolated prompts.
Coding and Enterprise AI Capabilities: Competitive Where It Counts
Muse Spark 1.1’s AI coding performance is not always the highest score on public benchmarks, but its behavior aligns better with enterprise needs. The model can diagnose complex bugs, implement new features in enterprise systems, execute large code migrations, and support agentic coding workflows such as planning mode, subagent delegation, and context compaction across large repositories. On Terminal-Bench 2.1 and SWE-Bench Pro, it trails Claude Opus 4.8 and GPT 5.5, yet Meta’s internal coding benchmark places it effectively neck-and-neck: 68.3 versus Opus 4.8’s 69.0 and ahead of GPT 5.5 at 67.1, nearly ten points above the original Muse Spark. That jump indicates rapid iteration rather than incremental tuning. More importantly, the model is built to plan and orchestrate actions across external apps and services, zero-shot generalize to new tools and MCP servers, and generate automation scripts that fit into existing enterprise workflows instead of forcing teams to re-architect their systems around the model.

Meta Model API and the Superintelligence Bet
Muse Spark 1.1 matters not only for what it can do, but for how Meta is shipping it. The model is already live in “Thinking” mode inside the Meta AI app and on meta.ai, and it is accessible to developers through the new Meta Model API in public preview. That API is OpenAI-compatible, giving enterprises a hosted path to Muse Spark 1.1 without managing their own infrastructure and signaling a direct grab for developer relationships and usage revenue. According to Artificial Analysis, Meta is pricing it at USD 1.25 (approx. RM5.86) per million input tokens and USD 4.25 (approx. RM19.92) per million output tokens, with cache hits at USD 0.15 (approx. RM0.70) per million, making Muse Spark 1.1 one of the most cost-effective models at its intelligence tier. This push follows a rough Llama 4 release, an internal reshuffle, and the creation of Meta Superintelligence Labs—now headed by Alexandr Wang—to pursue a personal superintelligence vision centered on models that help people achieve goals and take action. The conclusion is clear: Meta is betting that agentic automation, not chat alone, will define the next era of AI competitiveness.






