The Pulse
Sakana AI splits Fugu into a cheaper Max and stronger Ultra v2
Sakana AI has released Fugu Max and Fugu Ultra v2, two orchestration systems aimed at lowering inference costs and raising performance on complex tasks. Fugu Max costs $2 per million input tokens, while Ultra v2 posts Sakana’s leading score

AI.info Team ·
Sakana AI is putting two different answers to the same industry argument on sale. The Tokyo-based company says larger, more expensive foundation models are not always the best way to handle real workloads; providers building those models still control much of the market’s strongest software and reasoning capability. Its response is Fugu Max, which routes tasks across a wider pool of open and specialized models, and Fugu Ultra v2, which aims for maximum quality on difficult, multi-step work.
Sakana announced both systems on September 11, 2026, describing them as two versions of the same orchestration architecture rather than unrelated products. Fugu Max is built around cost efficiency, while Fugu Ultra v2 is designed to push benchmark performance higher. Both are available through Sakana’s OpenAI-compatible API, and existing Fugu customers can switch by changing a single model parameter, the company says.
The release advances a product strategy that challenges the idea that a single model should perform every task. Sakana’s Fugu system presents one API endpoint to users but internally selects, delegates to, coordinates and combines multiple agents. The company’s technical report describes Fugu models as language models trained to construct adaptive agentic scaffolds for individual queries, with one variant favoring latency and another favoring answer quality.
Fugu Max targets the cost of calling frontier models
Fugu Max is Sakana’s attempt to make model orchestration economically useful at scale. The system dynamically routes each request to what the company describes as the leanest model capable of solving it, instead of sending every task to the most powerful available model. Sakana says the expanded pool includes open-weight and specialized systems, including NVIDIA’s Nemotron family through its partnership with the chipmaker.
The published price is $2 per million input tokens and $6 per million output tokens. Cached input costs $0.25 per million tokens, and Sakana says the model’s output price is 40% to 60% below the prices of Sonnet 5, GPT 5.6 Terra and Kimi K3. The company also says Fugu Max delivers performance within striking distance of elite models at two to six times lower cost, though those comparisons come from Sakana’s own evaluation and pricing analysis.
Sakana reports that Fugu Max achieves the best overall score on six benchmarks: Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. SWEFish is described as an internal benchmark based on Sakana’s coding challenges and use cases. The company also says Max extends the cost-performance frontier on seven of ten benchmarks it evaluated.
Those claims matter because routing systems can incur costs that a conventional model hides. Each request may trigger several internal model calls, and the final answer can include work performed by the orchestration layer as well as the visible response. Sakana’s pricing documentation says those orchestration tokens count toward the final bill at the same input and output rates, giving developers a clearer accounting of the system’s internal work but also making headline token prices an incomplete measure of total spend.
Ultra v2 trades cost for harder reasoning
Fugu Ultra v2 takes the opposite position. Sakana says it coordinates a deeper pool of agents for complex reasoning, autonomous research and full-stack software development, where latency and token consumption matter less than the quality of the final result. The company reports that Ultra v2 achieves the best or joint-best score on five of eight benchmarks: GDP.pdf, Chartography, SWEFish, DeepSWE and Toolathon.
Its most striking published result comes from Chartography, which tests visual reasoning and data interpretation. Sakana reports a score of 48.3 for Fugu Ultra v2, compared with 27.3 for Opus 5 and 29.5 for Fable 5. On DeepSWE, a software-engineering evaluation, the company reports a score of 74.3. Sakana says Ultra v2 ranks in the top two on seven of the eight benchmarks in its comparison.
The company is also making a narrower claim about what sits behind those results. Sakana says Fugu Ultra v2’s model pool does not include Fable 5, Fable 5.1 or GPT-6-Astra, and that its reported scores do not depend on those proprietary systems. That does not mean the system operates without access to any external models; it means Sakana says the three named models are absent from the pool used for Ultra v2.
Fugu Ultra v2’s training cutoff is August 28, 2026, according to Sakana’s release. Pricing is $5 per million input tokens and $30 per million output tokens, with cached input priced at $0.50 per million. For contexts above 272,000 tokens, the rates rise to $10 per million input tokens, $45 per million output tokens and $1 per million cached input tokens.
The product depends on a changing pool of models
Sakana’s main technical bet is that model selection can become a capability in its own right. A request involving software engineering might be routed toward agents that perform well at code generation and tool use; a scientific question could trigger a different combination. The Fugu technical report says the system can route, delegate and coordinate across specialized agents while presenting the user with the experience of calling one model.
The approach builds on Sakana’s earlier Fugu release in June. That version described two variants: standard Fugu for stronger performance at lower latency, and Fugu Ultra for deeper orchestration over a larger worker pool. Sakana’s technical report says the systems use large-scale fine-tuning, evolutionary algorithms and reinforcement learning to train the orchestration layer and adapt it to different tasks.
Sakana argues that a swappable model pool reduces dependence on any single supplier. If a provider changes an API, withdraws access or limits a model in a particular market, the company says Fugu can substitute another agent without forcing customers to redesign their application interface. The claim is architectural rather than a guarantee: the system still depends on the availability, pricing and terms of the models it can call.
NVIDIA’s Nemotron models are part of that strategy. In a July announcement, Sakana said it would integrate Nemotron as specialized agents inside Fugu, with strengths in coding, tool calling and instruction following. David Ha, co-founder and chief executive of Sakana AI, said: “We’re excited to collaborate with NVIDIA to build the next generation of Fugu orchestration models together, by incorporating leading open-weights models like Nemotron.”
One API, two operating points
For developers, the distinction between Max and Ultra v2 is straightforward even if the systems behind them are not. Max is the lower-cost option for workloads that need strong results across a large volume of requests. Ultra v2 is the higher-priced choice for work where deeper reasoning, longer execution and higher answer quality justify a larger bill.
Sakana says both models are immediately available through its standard OpenAI-compatible API. Existing Fugu applications do not require a migration project; the company says customers can move to either release by changing one parameter. Vercel’s AI Gateway lists both model identifiers and supports access through OpenAI Chat Completions, OpenAI Responses and Anthropic Messages-compatible interfaces, giving developers another route to test the systems.
The API model also keeps the underlying agent pool largely out of view. Users receive a response through a single endpoint rather than choosing each worker model themselves. That reduces application complexity, but it also means customers must trust Sakana’s routing choices and accept that the pool can change as models are added, removed or repriced.
Sakana’s benchmark claims now face production tests
The release gives Sakana a sharper commercial story than its original Fugu launch. Max offers a concrete price position, while Ultra v2 supplies a set of high scores on coding, visual reasoning and agent evaluations. Together, they let the company sell orchestration as a choice between a lower-cost operating point and a higher-capability one instead of presenting a single system with an ambiguous performance-cost tradeoff.
Independent validation will determine how well that distinction holds outside Sakana’s test suite. The company’s own release identifies the benchmarks and comparison models, but it does not establish that every workload will reproduce the same ranking. Some baseline figures come from model providers, and the pool composition, routing behavior and internal token usage can all affect results in ways that a static benchmark score does not show.
For now, the concrete offer is clear: Fugu Max is available at $2 per million input tokens and $6 per million output tokens, while Fugu Ultra v2 starts at $5 and $30. Both systems expose the same OpenAI-compatible interface, but they place Sakana’s orchestration thesis on opposite sides of the buying decision—one asking how cheaply a task can be completed, the other asking how much performance a customer is willing to purchase.