Menu

Categories

Tags

Sakana AI’s Fugu multi-agent system outperforms Anthropic’s top models

June 23, 2026 | Source: sakana | Anthropic, OpenAI | 160 views 0 comments

Sakana AI launches multi-agent system Fugu, beats Fable 5 on reasoning and coding tests
Japanese AI startup Sakana AI has unveiled Sakana Fugu, a multi-agent collaboration system that orchestrates multiple models dynamically through a single API. In several authoritative benchmarks spanning academia, reasoning, and coding, the top-tier Fugu Ultra scored higher than Anthropic’s flagship models Fable 5 and Mythos Preview.

Notably, Fable 5 and other controlled models aren’t part of Fugu’s underlying model pool. Instead, Fugu stitches together publicly available models like GPT-5.5, Gemini-3.1-Pro, and Claude-Opus-4.8, leveraging collective intelligence to outperform any single top-tier model.

Benchmark results show Fugu Ultra scoring 93.2 on LiveCodeBench (vs. Fable 5’s 89.8), 95.5 on GPQA Diamond scientific reasoning (vs. Mythos Preview’s 94.6), and 82.1 on Terminal Bench 2.1 agent tasks (vs. Fable 5’s 80.4). Even the lower-latency standard Fugu approaches or beats traditional monolithic closed-source models on multiple tests.

The performance leap comes from dynamic role assignment under the hood. In common-sense QA, for instance, Fugu Ultra assigns Gemini-3.1-Pro as the aggregator, with GPT-5.5 and a second Gemini as leaf nodes solving problems independently. For math tasks, it switches GPT-5.5 to aggregator to resolve disagreements between Gemini and Opus. In multi-turn coding tasks, Fugu Ultra alternately lets GPT-5.5 write code while Claude-Opus-4.8 handles security audits and debugging. This fine-grained dynamic complementarity lets the whole system far surpass any single agent.

Leave a Reply

Your email address will not be published. Required fields are marked *