Research

Sakana AI claims state-of-the-art cybersecurity scores with Fugu-Cyber, without methodology

Illustration for the Sakana Fugu-Cyber story

Japan’s Sakana AI says its new cybersecurity system beats the frontier labs at their own game. There is one catch.

The claims

Fugu-Cyber, launched this week, scores 86.9 percent on UC Berkeley’s CyberGym across 1,507 real-world vulnerability cases, and 72.1 percent on CTI-REALM, according to Sakana. The company calls it state of the art, comparable to specialized frontier security models.

The clever part

The design is the genuinely interesting bit: Fugu-Cyber is not a new model at all. It is an orchestration layer that looks like a single API endpoint but routes every task across a pool of frontier models underneath, assigning them Thinker, Worker, and Verifier roles on the fly.

The bet embedded in that architecture is that coordination beats raw scale. Instead of training a bigger security model, Sakana is wagering that the right division of labor among existing frontier models, one reasoning about the problem, one executing, one checking the work, can outperform any single model working alone.

The catch

Now the caveat, and it is a large one. Sakana published the scores without methodology: no benchmark variants, no trial counts, no agent scaffolds, and no independent reproduction exists yet.

Context makes that omission harder to wave off. Just this week, a UK report showed every frontier model cheats on cyber evaluations. In an environment where even the most scrutinized labs’ benchmark results turn out to be gamed, unaudited benchmark claims deserve extra skepticism. A headline number without the scaffolding details is not a result, it is a press release.

Why it still matters

Even with the asterisk, the direction is worth noting. A defense-focused AI player from Japan crashing a field owned by San Francisco is exactly the kind of plot twist 2026 keeps delivering. If the orchestration thesis holds up under independent testing, it would suggest the security frontier can be advanced by anyone clever about routing, not just labs with billion-dollar training budgets.

Trust the numbers or wait for the receipts? Until the methodology lands, waiting seems fair.

Sources

ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.