← The AI Hype Audit — all 427 verdicts

PARTLY

Partly true: Claude does not take video files and the yt-dlp plus FFmpeg frames trick is real, but '1,847 frames in one pass, zero API cost' ignores the image limits and the tokens every frame burns, and the skill sits behind a comment-EYES DM

Facebook · Oct 2026

The claim'Claude can't watch video. It only reads the transcript and misses half.' 'So I gave it eyes. One skill that makes Claude see every single frame': yt-dlp grabs the video, FFmpeg rips the frames, 'and Claude flips through them while it reads.' 'A 30-minute video, decoded frame by frame, in 90 seconds. Running local, zero API cost.' On screen: '1,847 frames seen, one pass' and 'local, $0 API cost.' Caption: 'Every analyze-this-video tool just pulls the words.' 'Comment EYES and I'll DM you the skill, free.'

The premise is right and the recipe is real. Anthropic's vision docs list four image formats (JPEG, PNG, GIF, WebP), say animations are read as their first frame only, and offer no video input, so Claude on its own gets a transcript at best. Pulling the file with yt-dlp, cutting stills with FFmpeg and handing Claude the stills next to the transcript is a sound fix. This audit is graded that way every night, so we know it works. The numbers are where the reel runs ahead of itself. The same docs cap a request at 600 images on the API (100 on 200k-context models) and 20 per message on claude.ai, so 1,847 frames cannot go through in 'one pass'; an agent has to batch them or tile them into contact sheets. Every frame also costs visual tokens: by the published formula a small 640 by 360 still is 299 tokens, which puts 1,847 of them near 552,000 tokens before the transcript. Downloading and slicing are local and free. The reading is not; it runs on Anthropic's servers and comes out of your plan's usage limits or your API bill. 'Every tool just pulls the words' is also too broad, since Gemini's API takes video natively at one frame per second. The 90 seconds is unverified, and the skill itself is behind a comment-for-DM hook, so we could not read it.

What holds up

  • Reel read frame by frame: animated text slides only (no screen recording of the skill running), showing '$ yt-dlp https://youtu.be/...', 'FFMPEG', '1,847 FRAMES SEEN · ONE PASS', 'local · $0 API cost' and 'COMMENT EYES'; the voiceover matches the caption.
  • platform.claude.com vision docs read Oct 4: supported formats are JPEG, PNG, GIF and WebP; 'Animations are unsupported, and only the first frame is used'; limits are 20 images per message on claude.ai, 100 per API request on 200k-context models and 600 on all others; cost is ceil(width/28) x ceil(height/28) visual tokens per image.
  • Arithmetic from that formula: a 640x360 frame is 23 x 13 = 299 tokens, so 1,847 frames is about 552,000 tokens. A 30-minute video sampled once a second is 1,800 frames, which matches the on-screen count.
  • ai.google.dev video-understanding docs read: Gemini accepts video natively, sampled at 1 frame per second, so 'every analyze-this-video tool just pulls the words' does not hold across tools.
  • The skill could not be obtained without commenting for a DM, so its contents, its sampling rate and the '90 seconds' figure are unverified.

What doesn’t

  • 'One pass' and 'every single frame' collide with the documented per-request image caps; real implementations sample and batch.
  • 'Zero API cost' covers the download and the FFmpeg step only; Claude reading frames spends tokens against a subscription limit or an API bill.
  • No demo of the skill running, only motion graphics; the numbers on screen are asserted, not shown.
  • A free skill gated behind comment-for-DM is a follower and lead funnel, not a download link.

The catch

Frames plus transcript is the right way to make Claude 'watch' a video, and you can build it yourself in an evening. Just do not expect every frame of a half-hour video in one go for nothing; sample smartly, tile frames into sheets and budget the tokens.

How to actually do it

  • Download with yt-dlp, then cut stills with FFmpeg at a sane rate: one frame every 5 to 10 seconds, or scene-change detection (select='gt(scene,0.3)'), scaled to about 640 px wide.
  • Tile the stills into 4x4 contact sheets with FFmpeg's tile filter so one image carries 16 moments; a 30-minute video becomes roughly one to two dozen images instead of 1,800.
  • Transcribe the audio locally with Whisper or faster-whisper and give Claude the transcript and the sheets together, asking it to cite timestamps for anything it says it saw.
  • If you need true per-second video understanding, use a model with native video input and compare its notes with Claude's.

The limitation is real and so is the fix: Claude takes images, not video, and yt-dlp plus FFmpeg stills is a working bridge. The headline numbers are not supported by Anthropic's own image limits and token costs, and the skill is gated behind a comment-for-DM hook.

Confidence
High
Posted by
a Claude-workflow creator (small page, ~22K views on the reel) with a comment-EYES hook for a free skill; reel sent to Buddy Sat Oct 3, 2026, ~8:04 AM

See the original claim →

We test hype for free. We build the real thing for a living.

Thirty minutes, no pitch — and you'll leave with something useful either way.

Book a call with Todd or start with the free Business Checkup →

The Verdict Weekly

Three verdicts every Friday. Free forever, unsubscribe anytime, no spam — that would be ironic.

© Schreier Group · schreiergroup.com · See a wild AI claim? Drop it here and we'll test it.