← The AI Hype Audit — all 181 verdicts

PARTLY “5 hacks to cut Claude's token usage by 95%” — real repos, marketing math TikTok · Aug 17, 2026

The claim5 hacks to cut Claude's token usage by 95%: the caveman skill (claims 65% fewer output tokens), context-mode on MCP calls (98% less tool output), headroom (95% fewer tokens on JSON), agentmemory (95.2% recall), and model tiering — Fable 5 $10/$50, Opus 5 $5/$25, Sonnet 5 $3/$15, Haiku 4.5 $1/$5, “10x cheaper on Haiku than Fable.”

All four repos exist and are popular — verified Aug 17 via the GitHub API: caveman 98.7k stars, context-mode 19.9k, headroom 66.6k, agentmemory 27.1k. And the price table matches real Anthropic API pricing exactly: Fable 5 $10/$50 per million tokens, Opus 5 $5/$25, Sonnet 5 $3/$15 (currently $2/$10 intro through Aug 31), Haiku 4.5 $1/$5 — Haiku genuinely is 10x cheaper than Fable on input, and hack 5 (pick the smallest model for the job) is the one genuinely load-bearing tip in the post. The problem is every percentage is the repo's own marketing, not an independent benchmark. Caveman's “65% fewer tokens” is literally the repo's tagline — and this audit already ruled on Aug 12 that the maintainer's own real-world benchmarks failed to reproduce it. Headroom's own README says “20% fewer tokens for coding agents, 60–95% fewer for JSON” — the slide cherry-picks the top of the JSON range and presents it as the norm. Context-mode's 98% and agentmemory's “95.2% recall” are self-reported. And stacking the hacks doesn't multiply into “95% total savings” — the headline number is vibes.

What holds up

  • All four repos real and popular: caveman 98.7k / context-mode 19.9k / headroom 66.6k / agentmemory 27.1k stars (GitHub API, Aug 17)
  • The price table is exactly right: Fable 5 $10/$50, Opus 5 $5/$25, Sonnet 5 $3/$15 ($2/$10 intro through Aug 31), Haiku 4.5 $1/$5 per MTok
  • Hack 5 is genuinely sound — routing simple, high-volume work to Haiku 4.5 really is ~10x cheaper on input than Fable 5

What doesn't

  • Every percentage is the repo's own tagline, not an independent benchmark — caveman's 65%, context-mode's 98%, agentmemory's 95.2% are all self-reported
  • This audit already ruled (Aug 12) that the caveman 65% figure was repo-marketing the maintainer failed to reproduce under real benchmarks
  • Headroom's own README says 20% for coding agents, 60–95% for JSON — the slide quotes only the 95% top end as if it were typical
  • The headline “95%” isn't a measurement of anything — the hacks don't stack multiplicatively

The catch

The slide's numbers are all real — they're just quotes from the repos' own ads, dressed as results. The one hack with honest math (use a cheaper model) is the one that needs no repo at all.

How to actually do it

  • Do model tiering first: route simple, high-volume work to Haiku 4.5 ($1/$5) and keep Fable for the genuinely hard calls — a real 10x on the routed traffic
  • Add prompt caching before any token-diet skill — ~90% off repeated context, from Anthropic directly, no third-party repo
  • Treat every “-X%” in a repo tagline as marketing until you've measured it on your own workload — log a week of usage before and after
  • If you install one compressor, read its own README's full range (headroom: 20% for code, 60–95% for JSON) and plan on the bottom of it

Real repos, correct price table, self-reported percentages — the only guaranteed 10x on the slide is the one Anthropic publishes.

Confidence
High
Posted by
an AI-engineer tips account — the tools are real, the percentages are the repos' own taglines

See the original claim →

We test hype for free. We build the real thing for a living.

Thirty minutes, no pitch — and you'll leave with something useful either way.

Book a call with Todd or start with the free Business Checkup →

The Verdict Weekly

Three verdicts every Friday. Free forever, unsubscribe anytime, no spam — that would be ironic.

© Schreier Group · schreiergroup.com · See a wild AI claim? Drop it here and we'll test it.