The outcome: in about 10 minutes, your Claude can watch any video on the internet (YouTube, TikTok, IG reels, Looms, local screen recordings) and answer questions about what's actually on screen. Free, open source, runs on your machine.
This is the exact setup I run. I audited every line of this skill's code before installing it, and the same afternoon it was breaking down competitor reels for me frame by frame.
1 — The skill (free, open source)
Repo: https://github.com/bradautomates/claude-video
Built by Brad Bonanno (@bradbonanno on YouTube). MIT licensed. It works by splitting any video into the two things Claude already understands: frames (images) and a timestamped transcript. No expensive video-model API in the middle.
2 — Install in Claude Code (2 commands)
/plugin marketplace add bradautomates/claude-video/plugin install watch@claude-video/watch <any video URL> <your question> — the skill auto-installs its two dependencies (ffmpeg + yt-dlp) on first run and walks you through anything missing.Other surfaces:
- Cursor / Codex / Copilot / Gemini CLI:
npx skills add bradautomates/claude-video -g - claude.ai (web): download
watch.skillfrom the repo's latest GitHub release → Settings → Capabilities → Skills → the + button (turn on "Code execution" first)
3 — The 2-minute upgrade: Groq key (for videos without captions)
YouTube videos usually have free captions, so they cost nothing. TikToks, reels, and screen recordings usually don't — the skill falls back to Whisper transcription for those.
~/.config/watch/.env on the line GROQ_API_KEY=your_key_hereThat's the whole setup. Paste a URL, ask a question, Claude watches it.
4 — My wire-in playbook (the part nobody else gives you)
Installing the skill is step one. The real unlock is wiring it into a system so it runs without you. Here's the pattern I used:
- Build one adapter script that calls the skill's engine as a library and always outputs the same files:
source.mp4,source.srt,transcript.json, hero frames, and amedia.jsonwith cut timestamps. One media layer, every pipeline shares it. - Keep your metrics scraper separate. The watch skill sees the video; it can't see views, likes, saves, or follower counts. I still pull those via Apify. Video understanding + metrics = a complete teardown.
- Point your existing workflows at the adapter. My competitor-teardown, swipe-file canvas, and video-deconstruct flows all call the same adapter now. One afternoon of rewiring replaced three separate scraping mechanisms.
- Let it compound. Every video your AI watches becomes searchable notes in your knowledge base. The system gets smarter every day without you watching anything manually.
💡 Pro tip: for a long video, run /watch <url> --start 2:15 --end 2:45 to zoom into one section instead of burning tokens on a sparse full scan.Want this running in YOUR business?
If you run a clinic or practice and want a system like this doing your content research automatically — watching what works in your niche and turning it into content ideas while you see patients — that's literally what I build. DM me BUILD on Instagram (@dylan_j_watkins) and I'll show you what it looks like.
— Dylan