Multi-agent orchestration: one system to command many models
Multi-agent orchestration is an AI architecture where a coordinating system breaks a complex task into subtasks and routes them to different specialized AI agents, then verifies and fuses their outputs so the user experiences a single unified model. This is the core bet behind Sakana AI’s new Fugu system and the clearest challenge yet to the dominance of giant, monolithic frontier AI models.
Instead of pouring more compute into one huge model, Sakana AI—a lab valued at $2.65 billion (approx. RM12.2 billion) after its November 2025 Series B—released Fugu as a multi-agent orchestration system designed to deliver frontier-model performance while reducing reliance on any single provider. The key claim is bold and quotable: “across a suite of coding, reasoning, scientific, and agentic evaluations, Fugu and Fugu Ultra land near the top of the field.” In other words, coordination, not sheer size, is the differentiator. That architectural bet has consequences for performance, economics, and the politics of AI sovereignty—and it is already forcing a new kind of AI model routing debate.

Fugu Ultra’s collective intelligence: routing over raw scale
Fugu Ultra is Sakana’s clearest argument that careful AI model routing can rival the best frontier AI models without training another mega-model from scratch. Instead of a single neural behemoth, Fugu coordinates a pool of existing models, breaking user prompts into subtasks and sending each piece to a swappable expert agent that is best suited for the job. It is multi-agent orchestration as a product: one OpenAI-compatible API on the surface, but a learned conductor directing a hidden ensemble underneath.
The benchmark story is impressive. On SWE-Bench Pro, a heavy-duty software engineering benchmark, Fugu Ultra scores 73.7, beating Claude Opus 4.8’s 69.2 and GPT-5.5’s 58.6. On LiveCodeBench, Fugu hits 92.9 while Fugu Ultra reaches 93.2—both ahead of Gemini 3.1 Pro’s 88.5. On Humanity’s Last Exam, Fugu Ultra reaches 50.0, essentially matching Opus 4.8’s 49.8. These are not toy metrics; they say that a coordinated team of specialized AI agents can stand shoulder-to-shoulder with the most advanced single models in key domains. If you care about results rather than glamour, that should shake your faith in “bigger is always better.”

Why orchestration is a different AI architecture, not “just a router”
The easy critique is that Fugu is “just another router.” That is wrong, and misunderstanding it means missing where AI architecture is heading. Traditional routers blast the same prompt to multiple models, then compare or blend the outputs. Fugu instead decomposes the task itself: it identifies subtasks, assigns them to appropriate models, and then synthesizes the results—often across multiple turns—using learned coordination strategies rather than hand-written workflows.
Under the hood, Sakana’s TRINITY coordinator assigns models dynamic Thinker, Worker, or Verifier roles across turns, while the Conductor uses reinforcement learning to discover natural-language strategies for telling agents what to do. This means the system is not just routing; it is learning how to orchestrate. On X, co-founder David Ha argues that large-scale monolithic models have had their moment and that solving harder real-world problems will need “collective intelligence” instead. That stance matters: multi-agent routing is not merely a cost trick; it is a fundamentally different architectural approach to scaling capabilities, betting that specialization plus orchestration beats further brute-force scaling.

The sovereignty pitch and its cracks: dependence, price, and speed
Sakana is positioning Fugu as a blueprint for AI sovereignty: a way for governments and companies to avoid dependence on any one frontier provider. The logic is straightforward. Export controls recently forced a major lab to pull its top models Fable 5 and Mythos 5 just three days after launch, proving that access to frontier AI can vanish with one directive. Fugu’s answer is architectural: because it uses a pool of “entirely swappable agents,” if one provider restricts access, the orchestrator can route work to others instead.
Yet early users are already poking holes in this sovereignty story. Some describe Fugu as “a highly advanced router/wrapper” whose reliance on external models means it inherits their political risks. If more than one provider tightens access at the same time, Fugu’s abilities also degrade. Others complain about the economics and user experience: fast burn rates, unnecessarily high prices, and an “extremely slow” API. One developer outside the US says, “it’s nowhere remotely near usable as a day-to-day workhorse.” Those critiques matter because sovereignty is not just about architecture; it is about whether ordinary teams can afford and tolerate the system in daily use.
What it means for startups, users, and the next AI architectures
Fugu’s launch signals that the AI race is splitting into two clear strategies: frontier labs escalating the size and training budgets of single models, and smaller players pushing multi-agent orchestration to squeeze more value from the models the world already has. Sakana Fugu makes that second camp credible. In an AutoResearch setup, Fugu Ultra autonomously ran 123 training experiments over 14 hours on a single H100 GPU and achieved the best mean validation score across all seeds against three frontier baselines. An industry researcher reported finishing a patent landscape analysis covering about 20 papers and several patents in a few hours instead of three to four days. Those are tangible gains for ordinary users: less integration work due to a single OpenAI-compatible API, and the ability to offload long, multi-step research workflows to an orchestration engine.
But there are caveats. Fugu Ultra’s fixed pricing at $5 (approx. RM23) per million input tokens and $30 (approx. RM138) per million output tokens, doubling above 272K tokens, makes it an explicitly premium tool. A launch promotion offers a free second month for anyone subscribing before the end of July 2026, signalling that Sakana knows it must win over skeptical developers. My view: multi-agent routing is not a sideshow; it is the new architecture frontier. Yet Fugu shows that orchestration alone will not save users from tradeoffs in cost, latency, or geopolitical risk. The next decade of AI will likely belong not to one paradigm but to the teams that learn when to use monoliths, when to use specialized AI agents, and when to orchestrate both.






