From Bigger Models To Smarter Teams
Sakana AI’s Fugu Ultra model is a frontier-level multi-agent AI orchestration system that coordinates a pool of language models through one OpenAI compatible API to solve complex, multi-step tasks at performance levels that rival leading single frontier models, signaling a strategic shift from raw model scale toward intelligent AI model coordination as a competitive edge. Sakana AI has announced the launch of Fugu Ultra, a frontier-level orchestration model designed to rival top-tier AI models such as Anthropic’s Fable 5 and Mythos Preview. At the same time, the company has launched Sakana Fugu, a related system aimed at matching frontier performance by coordinating several existing models instead of training one giant network. This is not a side project: the lab is valued at USD 2.65 billion (approx. RM12.2 billion) after its November 2025 Series B, so Fugu is a flagship bet rather than a speculative experiment. In a landscape obsessed with ever-larger models, Sakana is arguing that orchestration is now as important as raw scale.

How Fugu Ultra’s Multi-Agent Coordination Works
Fugu Ultra distinguishes itself by using multi-agent AI orchestration rather than brute-force scaling. It is a language model that can autonomously delegate subtasks to a pool of expert LLMs, including instances of itself, then choose and combine their outputs for the best overall result. Instead of assigning fixed roles or workflows to specific models, Fugu learns to dynamically assemble agents from a pool and route work between them in patterns that aren’t obvious but, according to Sakana, highly efficient. The tagline captures the intent: “One Model to Command Them All,” where a well-orchestrated team of models can outperform any individual one. This coordination logic is grounded in two ICLR-accepted papers: TRINITY, which evolves a lightweight coordinator that assigns Thinker, Worker, and Verifier roles across multiple turns, and the Conductor, which uses reinforcement learning to discover natural-language coordination strategies instead of relying on hand-designed workflows. In other words, the secret sauce is meta-intelligence about how models should collaborate.

Benchmarks: Collective Intelligence Meets Frontier-Level Performance
If multi-agent orchestration sounds like a clever theory, the numbers suggest it works. Across coding, reasoning, scientific, and agentic evaluations, Fugu and Fugu Ultra land near the top of the field. On SWE-Bench Pro — a demanding software engineering benchmark — Fugu Ultra scores 73.7, ahead of Claude Opus 4.8’s 69.2 and GPT-5.5’s 58.6. On LiveCodeBench, Fugu scores 92.9 and Fugu Ultra 93.2, both ahead of Gemini 3.1 Pro’s 88.5. On Humanity’s Last Exam, one of the hardest general-knowledge benchmarks available, Fugu Ultra reaches 50.0, essentially matching Opus 4.8’s 49.8. Importantly, Sakana notes that neither Claude Fable 5 nor Claude Mythos Preview are part of Fugu’s agent pool because they are not publicly accessible, yet Fugu sits shoulder-to-shoulder with those models on several benchmarks. That is the point: orchestrating many strong models allows a startup to rival labs that control the largest single models. It is a quiet challenge to the assumption that only frontier labs can produce frontier performance.
The qualitative results push the argument further. In an AutoResearch experiment where an AI agent autonomously ran 123 training experiments over 14 hours on a single H100 GPU to improve a small language model’s training recipe, Fugu Ultra achieved the best mean validation score across all seeds, ahead of three frontier model baselines. Early users also report that Fugu Ultra sustains persona consistency and thoroughness in long workflows, surfacing more issues and maintaining focus better than other models in tasks like code review and security assessment. These are not shiny demo tricks; they are signs that multi-agent coordination can keep complex procedures on track over many steps — a weak spot for many large, single models. The message: if you care about long-horizon reasoning and detailed technical work, how models cooperate may matter more than how large any one of them is.

One OpenAI-Compatible API, Many Agents, Less Lock-In
Sakana’s architectural gamble would matter far less if the system were painful to use. Instead, Fugu Ultra arrives through a single, OpenAI compatible API, accessible to the general public worldwide for both individual and enterprise customers. The result is delivered through that same API so users are not forced to manage multiple model integrations. For ordinary developers, this makes multi-agent AI orchestration feel like a normal model call, even when Fugu is quietly coordinating several agents behind the scenes. Pricing is also designed to keep orchestration from turning into a billing nightmare: when only one agent is active, users pay the standard rate for that underlying model, but when multiple agents coordinate, Sakana charges a single rate based on the top-tier model involved instead of stacking fees. Fugu Ultra has fixed pricing at USD 5 (approx. RM23) per million input tokens and USD 30 (approx. RM138) per million output tokens, doubling for contexts above 272K tokens. Subscription tiers at USD 20 (approx. RM92), USD 100 (approx. RM460), and USD 200 (approx. RM920) per month give predictable access, with a free second month for anyone subscribing before the end of July 2026. This is how you turn a research idea into something developers can depend on day-to-day.
Vendor flexibility is a feature, not a footnote. Enterprise buyers can control which models participate in Fugu’s pool, excluding providers that conflict with data residency or compliance needs. Sakana AI positions this release as a step toward AI sovereignty, promising ongoing improvements as new models are added to the agent pool. In a market where frontier model competition among Anthropic, OpenAI, and Google is compressing performance differences to a few benchmark points, the ability to access frontier-level performance without locking into a single vendor is strategically attractive. Early real-world use backs this up: an industry researcher used Fugu Ultra for patent landscape analysis across roughly 20 papers and several patents and completed in a few hours a task that would normally take three to four days. According to Sakana, Fugu Ultra is the heavier-duty option designed for long-horizon tasks such as Kaggle competitions, paper reproduction, cybersecurity assessments, and patent and literature investigations. The system is currently unavailable in the EU and EEA while Sakana works toward GDPR compliance, highlighting that orchestration might solve technical lock-in faster than regulatory friction.
A Strategic Pivot: Architecture As The New Battleground
Why is Sakana pushing multi-agent coordination now? Frontier labs are locked in a race to build the best single AI models, each iteration swallowing more compute. Sakana’s broader identity, co-founded by David Ha and Llion Jones, has consistently pursued alternatives to brute-force scaling. Earlier this year, its AI Scientist system became the first AI to have a fully generated paper pass peer review at a machine learning conference, a milestone that later appeared in Nature. Fugu extends the same philosophy to deployment: extract more value from what already exists instead of only chasing larger models. Sakana’s pitch is that collective intelligence, properly orchestrated, can outperform any single model — and that a startup working with different constraints and assumptions than compute-heavy labs might be well-placed to build it. In that sense, Fugu Ultra is not merely a product release; it is a public bet that AI’s future will be shaped by how models collaborate, not only how big they become.
The implications are hard to ignore. If Fugu Ultra’s approach scales, smaller startups will have a credible path to frontier-level performance without owning the largest models. That changes the economics of innovation and could reduce concentration of power among a few labs. It also invites a more modular ecosystem in which organizations mix and match models under a coordinating layer, rather than tying their fate to one vendor’s roadmap. Sakana AI positions Fugu Ultra as part of a journey toward AI sovereignty, where control over architecture and vendor mix becomes a lever for operational and geopolitical security. The frontier race is not over, but Fugu Ultra has added a new lane: instead of asking who has the biggest model, it asks who can orchestrate many good models into a smarter whole. For now, that question makes Sakana one of the most interesting players in the field.






