MilikMilik

How Sakana AI’s Fugu Model Uses Multi-Agent Orchestration to Match Frontier Labs

How Sakana AI’s Fugu Model Uses Multi-Agent Orchestration to Match Frontier Labs
Interest|High-Quality Software

What the Sakana Fugu Model Is—and Why It Matters

The Sakana Fugu model is a multi-agent orchestration system that coordinates a pool of specialist language models, including copies of itself, to solve complex tasks more effectively than any single model by dynamically delegating, checking, and merging their work through an OpenAI-compatible API. Instead of betting everything on a single giant model, Sakana AI built Fugu and Fugu Ultra as conductors for AI model coordination. They are language models that can split work into subtasks, assign those subtasks to different expert models, and then synthesize their outputs into a final answer. This approach targets frontier AI performance on demanding workflows in engineering, science, research, cybersecurity, and data analysis. By focusing on orchestration rather than raw scale, Sakana gives developers and enterprises a way to reach near-frontier capability without locking into a single provider or infrastructure.

How Sakana AI’s Fugu Model Uses Multi-Agent Orchestration to Match Frontier Labs

How Multi-Agent Orchestration Gives Fugu Frontier AI Performance

Fugu Ultra’s core advantage is multi-agent orchestration: it acts as a central language model that can autonomously call other models as Thinkers, Workers, or Verifiers. Guided by the TRINITY architecture, this coordinator assigns roles per turn and adapts as the task progresses, while the Conductor component uses reinforcement learning to discover natural-language strategies for prompting and routing agents. According to Sakana’s benchmark reports, Fugu Ultra scores 73.7 on SWE-Bench Pro, ahead of Claude Opus 4.8’s 69.2 and GPT-5.5’s 58.6. On LiveCodeBench, Fugu and Fugu Ultra reach 92.9 and 93.2, outscoring Gemini 3.1 Pro’s 88.5. On Humanity’s Last Exam, Fugu Ultra reaches 50.0, effectively matching Claude Opus 4.8’s 49.8. These results indicate that coordination and orchestration can bring Sakana Fugu model performance into the same band as frontier systems without relying on a single massive network.

How Sakana AI’s Fugu Model Uses Multi-Agent Orchestration to Match Frontier Labs

Developer Access, Pricing, and Enterprise Control

Sakana delivers both Fugu and Fugu Ultra through a single OpenAI-compatible API, so developers interact with what feels like one model while it quietly manages multiple agents behind the scenes. The standard Fugu tier is tuned for everyday coding, code review, and responsive chat, while Fugu Ultra focuses on longer workflows such as Kaggle projects, cybersecurity assessments, and dense literature reviews. Pricing reflects the orchestration model: when only one agent is active, users pay the rate of that underlying model, and when several models coordinate, Sakana charges a single rate based on the strongest model involved. Fugu Ultra has fixed pricing at USD 5 (approx. RM23) per million input tokens and USD 30 (approx. RM138) per million output tokens, with higher rates for contexts above 272K tokens. Enterprises can also restrict which models participate to meet compliance, residency, or vendor-preference constraints.

How Sakana AI’s Fugu Model Uses Multi-Agent Orchestration to Match Frontier Labs

A New Competitive Playbook for Smaller AI Labs

Fugu signals a different competitive strategy in the frontier AI race: smaller labs competing not on raw model size, but on architecture and multi-agent coordination. Sakana’s founders, David Ha and Llion Jones, have a history of exploring alternatives to brute-force compute scaling, and Fugu continues that line of thinking. Their earlier AI Scientist system produced a fully generated paper that passed peer review, and with Fugu they apply similar ideas to deployment—extracting more value from existing models through orchestration. Early qualitative results show strengths in tasks like patent landscape analysis and autonomous training experiments, where Fugu Ultra has outperformed three frontier model baselines in mean validation score. As top-tier models from large labs cluster within a few benchmark points of each other, Sakana’s bet is that collective intelligence—well-orchestrated and vendor-flexible—can redefine what counts as frontier AI performance.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!