ffmpeg-skill Review 2026 — 21 Structured Video-Editing Tools That Teach Claude Code, Cursor and Codex to Stop Guessing About FFmpeg
✅ Pros
- • Takes the guessing out of agent video editing: every job starts with probe.py measuring actual duration, fps (with variable-frame-rate detection), resolution, rotation, bit depth, HDR format including Dolby Vision, colour tags and every audio stream — the agent decides from measured facts, not from the file name, and cut.py tries a lossless stream copy before ever re-encoding
- • 21 structured tools with typed argparse arguments instead of shell strings: no filter graph is accepted from the caller, nothing runs through a shell, every tool supports --dry-run (print the commands, write nothing), --json (structured result with a probe of the output), --fast and --progress, and a test runs every tool under --dry-run behind a fake ffmpeg asserting no ffmpeg call happened and no file appeared
- • A machine-readable contract generated from the code that runs: contract --json states input schema, output schema, role, required and conditional FFmpeg capabilities, verification steps, visual-check requirements and idempotency hints for all 21 tools, and the MCP server builds its tools/list from that same contract at start-up — a test keeps them byte-identical, so names and schemas cannot drift from the scripts
- • Verification is built into the workflow, not bolted on: outputs are probed and checked against the destination spec, and when the picture changed the agent runs look.py and inspects a contact sheet — the report is not finished until its Look: line names that image, and the README is explicit that 'inspects' means the calling agent's own vision, not a hidden feature of the skill
- • Real measurements published instead of claims: 92/92 verification steps on a 10-file real-device corpus (GoPro, DJI, iPhone with Dolby Vision, Android screen recordings, HDR10, 24p, Tears of Steel), sync.py offset detection 40/40 within 10 ms on real dialogue with a 1.1 ms max error, silence.py with zero missed gaps across 20 known-gap cases, scenes.py at F1 0.97, and 72/72 graded agent runs across 24 prompts in English and Japanese including honest refusals
- • Audio is a first-class input, not an afterthought: WAV/FLAC/MP3/M4A/OGG/Opus go through probe/cut/join/silence/loudness/audio/sync/check with the same commands as video, cut.py --accurate trims at the sample and reports precision, typed dynamics flags are range-checked before ffmpeg runs, and loudness.py does two-pass EBU R128 to any target
⚠️ Cons
- • Requires a capable FFmpeg build and the discipline to run doctor: plain Homebrew ffmpeg lacks the subtitles, drawtext and zscale filters, so macOS users are told to install ffmpeg-full; caption, graphics and colour tools report usable: no until the right build is present, and the project's own 0.9.0 release broke capability detection on FFmpeg 8 (fixed in 0.9.1 by parsing io-spec tokens instead of fixed-width flag columns)
- • The skill edits, but it does not decide: ffmpeg-skill is explicitly the 'hands' of the kajisho5 video ecosystem — it cuts, measures and exports, while deciding cut points, approving deliverables and judging what makes a highlight interesting belongs to a 'brain' agent in front of it, which means you still need a strong multimodal model to get judgement-grade results
- • Visual decisions lean on the calling agent's vision: look.py only renders a PNG, fit.py's default is a plain centre crop, and scenes.py --highlights ranks by measured proxies (audio energy or duration), never by content — a non-visual caller has to supply the anchoring decision itself
- • Young project with a fast-moving surface: created September 3, 2026 and reviewed at v0.10.0 with 30 commits and 2 contributors in four days; the README advises downstream repos to pin a version by tag or npm rather than tracking main, and sync.py explicitly does audio-to-audio only — no lip-sync or face detection
- • No GUI and no preview player: everything is verified through probes, contact sheets and structured JSON, which is exactly right for agent pipelines but will frustrate anyone who wants to scrub a timeline visually before exporting
Developers and content teams who want Claude Code, Cursor or Codex to cut, clean, caption and export video locally — podcasts, interview clips, Reels and YouTube edits, HDR-to-SDR conversions, multi-cam sync — without uploading footage to a cloud editor, and who are comfortable with the agent reading a contract and inspecting contact sheets instead of dragging clips on a timeline
Free and open source (MIT). Install with npx ffmpeg-skill (Claude Code, Cursor, Codex, or --all); requires FFmpeg 5.0+ on PATH and Python 3.9+ (standard library only — no pip dependencies, no API keys, no cloud). Node 16+ is needed only for the npx installer. macOS: brew install ffmpeg-full; Ubuntu/Debian: sudo apt install ffmpeg; Windows: winget install Gyan.FFmpeg
The Problem: An Agent That ‘Knows FFmpeg’ Still Guesses
Ask a coding agent to edit a video and it will happily pretend to know FFmpeg — then assume a frame rate, pick a codec the container cannot hold, re-encode a file that only needed a stream copy, and report “done” without ever opening the result. The gap is not knowledge; it is that an agent has no reliable way to measure the media it is editing or verify what it produced. kajisho5’s ffmpeg-skill, created September 3, 2026, is an Agent Skill built to close exactly that gap: it teaches Claude Code, Cursor, Codex and any agent that reads SKILL.md a fixed workflow — probe → edit losslessly where possible → check → verify — and ships 21 tools that do the actual work with local ffmpeg/ffprobe. No cloud, no API keys, no Python dependencies beyond the standard library.
In four days it passed 320 stars on an MIT license at v0.10.0, with 30 commits from two contributors and a CI matrix running FFmpeg 6.1 on Ubuntu, 8.x on macOS and 9.x on Windows. The pitch that explains the traction: an agent that “knows FFmpeg” still guesses, and this skill exists to take the guessing out.
The 21 Tools: Structured Operations, Not Shell Strings
The tool surface is organised into five groups, all Python 3.9 standard library, all with --help, --dry-run, --json, a non-zero exit and a reason on stderr on failure:
- Analysis:
probe.pymeasures duration, fps (with variable-frame-rate detection), resolution, codecs, bit depth, HDR format including Dolby Vision, colour space, rotation and every audio stream;scenes.pyfinds scene changes and audio peaks with highlight proposals and emits a cut list;look.pyrenders contact sheets, single frames and side-by-side comparisons as PNG so the agent can see what it made. - Editing:
cut.pydoes in/out or multi-segment cuts — lossless-c copyfirst, re-encode fallback,--accuratefor frame-exact video and sample-exact audio, reportingprecision;join.pyconcatenates with xfade transitions, normalising size, fps and audio;silence.pydetects and removes dead air with a margin around speech;fit.pyfits to a duration (pitch-preserving speed change or trim) and/or aspect ratio with an off-centre subject anchor. - Audio:
audio.pyruns a voice clean-up chain, FFT denoise, typed compressor/limiter/gate, music bed with sidechain ducking, 5.1→stereo and extraction;sync.pyoffsets two recordings by audio cross-correlation at 1 ms resolution with clock-drift correction;loudness.pydoes two-pass EBU R128loudnormto −14 LUFS/−1 dBTP or any target while stream-copying video. - Picture:
caption.pyburns SRT/ASS with full styling plus animated word-by-word karaoke;overlay.pyhandles logos and watermarks;graphics.pydraws lower-thirds, title cards and countdowns from a brand kit;color.pymaps HDR10/HLG/Dolby Vision to SDR BT.709, strips DV layers, applies 3D LUTs and does typed primary correction. - Delivery and orchestration:
export.pyhas presets for YouTube (incl. 4K), Reels, X, ProRes, H.265 and GIF, all tagged BT.709;check.pyreturns PASS/WARN/FAIL against platform specs with the fix for each failure;report.pybuilds a single-file HTML delivery report;render.pyrenders a whole edit from a declarativeproject.json;batch.pyapplies recipes to folders with a content-hash cache;multicam.pyaligns any number of cameras by audio and cuts from a switch list;verify.pyruns the toolchain on real device files.
The Design Principles: What Separates This From a List of One-Liners
The README’s eight design principles are where the project’s seriousness shows. Probe first: no tool decides from the file name. Lossless when possible: cut.py, join.py and loudness.py stream-copy what they do not need to touch; re-encoding happens only when it must. Plan before render: every tool takes --dry-run and a test runs every tool under it behind a fake ffmpeg, asserting no ffmpeg call happened and no file appeared. A contract the agent can read: contract --json states, for every tool, what it takes, what it writes, which FFmpeg components it needs and how the result is verified — generated from the argparse parsers themselves, not maintained beside them. Contract-derived MCP: mcp/server.py is a stdio JSON-RPC transport with no tool table of its own; tools/list and inputSchema are translated from the contract at start-up, so a new flag appears in MCP with no edit to the transport. Capability detection: doctor reads ffmpeg -encoders/-filters/-bsfs and reports which components are available, missing or unknown before a job fails inside ffmpeg. Unknown is not missing: when a listing cannot be read, capabilities are unknown — never missing, never silently available. Verify the result: outputs are probed, checked against the destination spec, and inspected visually when the picture changed.
That last point has an honest boundary the README states bluntly: nothing in this repository detects faces, products, subjects or “the interesting part” of a frame. look.py only renders a PNG; “inspects” means the calling agent’s own vision. When a crop needs to keep an off-centre subject, it is the multimodal agent looking at the contact sheet and choosing the anchor — a non-visual caller gets a plain centre crop by default. scenes.py --highlights ranks by a measured proxy (--rank-by audio or --rank-by duration), never by content.
The Measurements: Verified on Real Devices, Not Synthetic Clips
The testing story is the strongest signal that this is not another FFmpeg wrapper. The project publishes: 92/92 verification steps on a 10-file real-device corpus (GoPro, DJI, iPhone including Dolby Vision, Android screen recordings, HDR10, 24p, Tears of Steel) against local FFmpeg 6.1; 40/40 within 10 ms for sync.py offset detection on real dialogue and music with gain, noise and EQ changes (±30 s offsets, max error 1.1 ms); 0 missed gaps for silence.py across 20 cases with known gaps (≤1 ms leftover); F1 0.97 for scenes.py across 53 hard cuts (precision 0.95, recall 1.00 at the default threshold); sample-exact cut.py --accurate on WAV and FLAC; and 72/72 graded agent runs of 24 prompts — 12 English edits, 8 Japanese, 4 that must be declined — three repeats each, graded by an independent model: routing, honest refusals and user’s language 72/72, report format 71/72, visual check whenever the picture changed 24/24.
The project even documents its own regression honestly: v0.9.0 broke capability detection on FFmpeg 8 because that release shortened the flag column of ffmpeg -filters, and a parser anchored on the old width matched nothing and reported every filter missing. Since 0.9.1, rows are recognised by their io-spec token (A->A, AA->A, |->V, N->N), so flag width, legend and separator no longer matter, and a listing that still cannot be read yields unknown rather than missing. Captured listings from each CI runner are committed as fixtures with provenance, so a new layout is visible before it bites.
Honest Limits, the Ecosystem, and Who It’s For
ffmpeg-skill is deliberately the hands of kajisho5’s wider video-production ecosystem: it cuts, measures and exports files and reports in structured JSON, while deciding cut points, approving deliverables and planning whole edits belongs to sibling repos (video-production-agent, AI-video-production-OS) that read this repo’s contract — the dependency runs one way, and the skill never calls into them. Standalone, it needs no other repo: npx ffmpeg-skill, doctor, contract --json, and you are running. The main operational costs are a capable FFmpeg build (Homebrew’s plain ffmpeg lacks libass/drawtext/zscale, which doctor reports as missing rather than letting caption jobs fail inside ffmpeg) and a multimodal agent to make the visual calls the tools deliberately leave open.
For teams that want agents to cut podcasts, clean interview audio, burn captions, convert HDR footage to SDR or export platform-ready clips without uploading anything to a cloud editor, ffmpeg-skill is the most rigorously engineered answer yet — contract-driven, measured against real footage, and honest about the line between what a tool can verify and what only the agent’s eyes can judge.