← The AI Hype Audit — all 181 verdicts
PARTLY “5 hacks to cut Claude's token usage by 95%” — real repos, marketing math
The claim5 hacks to cut Claude's token usage by 95%: the caveman skill (claims 65% fewer output tokens), context-mode on MCP calls (98% less tool output), headroom (95% fewer tokens on JSON), agentmemory (95.2% recall), and model tiering — Fable 5 $10/$50, Opus 5 $5/$25, Sonnet 5 $3/$15, Haiku 4.5 $1/$5, “10x cheaper on Haiku than Fable.”
All four repos exist and are popular — verified Aug 17 via the GitHub API: caveman 98.7k stars, context-mode 19.9k, headroom 66.6k, agentmemory 27.1k. And the price table matches real Anthropic API pricing exactly: Fable 5 $10/$50 per million tokens, Opus 5 $5/$25, Sonnet 5 $3/$15 (currently $2/$10 intro through Aug 31), Haiku 4.5 $1/$5 — Haiku genuinely is 10x cheaper than Fable on input, and hack 5 (pick the smallest model for the job) is the one genuinely load-bearing tip in the post. The problem is every percentage is the repo's own marketing, not an independent benchmark. Caveman's “65% fewer tokens” is literally the repo's tagline — and this audit already ruled on Aug 12 that the maintainer's own real-world benchmarks failed to reproduce it. Headroom's own README says “20% fewer tokens for coding agents, 60–95% fewer for JSON” — the slide cherry-picks the top of the JSON range and presents it as the norm. Context-mode's 98% and agentmemory's “95.2% recall” are self-reported. And stacking the hacks doesn't multiply into “95% total savings” — the headline number is vibes.
What holds up
- All four repos real and popular: caveman 98.7k / context-mode 19.9k / headroom 66.6k / agentmemory 27.1k stars (GitHub API, Aug 17)
- The price table is exactly right: Fable 5 $10/$50, Opus 5 $5/$25, Sonnet 5 $3/$15 ($2/$10 intro through Aug 31), Haiku 4.5 $1/$5 per MTok
- Hack 5 is genuinely sound — routing simple, high-volume work to Haiku 4.5 really is ~10x cheaper on input than Fable 5
What doesn't
- Every percentage is the repo's own tagline, not an independent benchmark — caveman's 65%, context-mode's 98%, agentmemory's 95.2% are all self-reported
- This audit already ruled (Aug 12) that the caveman 65% figure was repo-marketing the maintainer failed to reproduce under real benchmarks
- Headroom's own README says 20% for coding agents, 60–95% for JSON — the slide quotes only the 95% top end as if it were typical
- The headline “95%” isn't a measurement of anything — the hacks don't stack multiplicatively
The catch
The slide's numbers are all real — they're just quotes from the repos' own ads, dressed as results. The one hack with honest math (use a cheaper model) is the one that needs no repo at all.
How to actually do it
- Do model tiering first: route simple, high-volume work to Haiku 4.5 ($1/$5) and keep Fable for the genuinely hard calls — a real 10x on the routed traffic
- Add prompt caching before any token-diet skill — ~90% off repeated context, from Anthropic directly, no third-party repo
- Treat every “-X%” in a repo tagline as marketing until you've measured it on your own workload — log a week of usage before and after
- If you install one compressor, read its own README's full range (headroom: 20% for code, 60–95% for JSON) and plan on the bottom of it
Real repos, correct price table, self-reported percentages — the only guaranteed 10x on the slide is the one Anthropic publishes.
- Confidence
- High
- Posted by
- an AI-engineer tips account — the tools are real, the percentages are the repos' own taglines
We test hype for free. We build the real thing for a living.
Thirty minutes, no pitch — and you'll leave with something useful either way.
Book a call with Todd or start with the free Business Checkup →The Verdict Weekly
Three verdicts every Friday. Free forever, unsubscribe anytime, no spam — that would be ironic.
© Schreier Group · schreiergroup.com · See a wild AI claim? Drop it here and we'll test it.