← The AI Hype Audit — all 449 verdicts

REAL

Holds up: Sakana AI's Fugu is a real orchestrator over other frontier models, and the reel correctly flags that its 73.7% vs GPT-5.5's 58.6% coding score is Sakana's own number

Facebook · Oct 2026

The claimJapan's Sakana AI built 'Fugu', 'a boss for other AIs' that picks the best model for each part of a task, runs several at once and merges the answers. A Sakana co-founder helped write the paper modern AI is built on. On 'one of the hardest coding tests' Fugu scored 73% vs GPT-5.5's 58%, but from Sakana's own testing; users say some tasks take 30 minutes, and a Wharton AI professor said it did not feel as strong as the models it claims to beat. 'Future of AI, or a trick? Tell me below.'

This one checks out, caveats included. Sakana AI launched Fugu on June 22, 2026 in two versions: Fugu for fast everyday work and Fugu Ultra for long multi-step jobs. It is sold as a single OpenAI-compatible API, but inside it is an orchestrator: a model trained to plan, hand pieces of the task to a pool of other frontier models, verify and merge. The 'hardest coding test' is SWE-Bench Pro, where Sakana reports Fugu Ultra at 73.7% against Claude Opus 4.8 at 69.2%, GPT-5.5 at 58.6% and Gemini 3.1 Pro at 54.2%. The reel rounds those honestly and says what most posters skip: the numbers are Sakana's own and nobody outside has reproduced them yet. The co-founder line is true too; Llion Jones is a co-author of 'Attention Is All You Need', the 2017 Transformer paper. The Wharton professor is Ethan Mollick, who posted that Fugu Ultra was 'incredibly slow' (about 30 minutes on his shader and interactive-scene tests) and that the results 'do not match Fable in real use'. Small nitpick: Mollick compared it to Anthropic's Fable, the model Sakana says Fugu matches, not to every model in the chart. The video sells nothing directly; it is an engagement post ('tell me below') that builds the audience for Singal's AI brand.

What holds up

  • Reel read frame by frame: on screen are the sakana.ai logo, the 'Attention Is All You Need' title page with Llion Jones circled, an orchestration graph listing Opus 4.8, GPT-5.5, Gemini 3.1 Pro, Qwen3-32B, DeepSeek and 'Fugu (self)', a DataCamp article dated June 24, 2026, a benchmark table with GPT-5.5 at 58.6 circled, and an Ethan Mollick post.
  • DataCamp's Fugu explainer and Japanese coverage (gihyo.jp) confirm the June 22, 2026 launch, the Fugu and Fugu Ultra split, the single OpenAI-compatible API, SWE-Bench Pro 73.7% vs Opus 4.8 69.2%, GPT-5.5 58.6%, Gemini 3.1 Pro 54.2%, and that 'all benchmark numbers are Sakana-reported and have not yet been independently reproduced'.
  • Sakana was co-founded by Llion Jones, a co-author of the 2017 Transformer paper, and David Ha; the reel's 'helped write the paper' line is accurate.
  • Ethan Mollick (Wharton) posted that Fugu Ultra is 'incredibly slow', with his coding tests taking 30 minutes, and that results 'do not match Fable in real use', as quoted on screen and reported by blockchain.news and others.
  • DataCamp lists consumer plans at $20, $100 and $200 a month and Fugu Ultra API pricing of $5 per million input and $30 per million output tokens; not available in the EU/EEA yet.

What doesn’t

  • Every headline score is vendor-reported; treat 73.7% as a claim until a third-party leaderboard reproduces it.
  • 'Japan quietly did something sneakier' is framing: part of the pitch is routing around the June 2026 US export controls on Anthropic's top models, not a new kind of intelligence.
  • Mollick's 'not as strong' was a comparison to Fable specifically; the reel stretches it to 'the top AIs it claims to beat'.
  • Engagement post from a creator with an FTC record (x-ref #034, #048) building an AI brand audience; no product pitch in this one.

The catch

Orchestration buys you quality on hard, long tasks by spending time and tokens: several models work, one checks, one merges. That is why Fugu Ultra can top a coding benchmark and still feel slow and sometimes worse than a single top model on interactive work. It is a tool for overnight jobs, not for chat.

How to actually do it

  • If you run long coding or research jobs, sign up for the Fugu API directly at sakana.ai, set a spend cap, and run five of your own real tasks through both Fugu Ultra and the single model you use today.
  • Score the outputs blind (have someone else strip the labels) and log wall-clock time and token cost per task; keep Fugu only where it wins on your work.
  • Use plain Fugu or a single frontier model for anything interactive; save Fugu Ultra for batch jobs you can let run for 30 minutes.

Fugu, its launch, its SWE-Bench Pro numbers, the Llion Jones link and the Mollick critique are all real and accurately reported in the reel. The only stretch is widening Mollick's Fable comparison into 'every model it claims to beat'.

Confidence
High
Posted by
Anik Singal (public figure; director at UgenticAI, Inc., the brand on his hoodie), ~5.7M-follower page reel; reel sent to Buddy Mon Oct 5, 2026, ~8:49 AM

See the original claim →

We test hype for free. We build the real thing for a living.

Thirty minutes, no pitch — and you'll leave with something useful either way.

Book a call with Todd or start with the free Business Checkup →

The Verdict Weekly

Three verdicts every Friday. Free forever, unsubscribe anytime, no spam — that would be ironic.

© Schreier Group · schreiergroup.com · See a wild AI claim? Drop it here and we'll test it.