Video

21 tools reviewed

7.8
anything2explainer Review 2026 — Turn Any Topic Into a Code-Drawn Explainer Video With an AI Coding Agent

anything2explainer (created September 8, 2026, PolyForm Noncommercial, 1,171 stars) is a Claude Code / Codex skill that turns a topic into a narrated, black-canvas motion-graphics explainer video — every frame drawn in code with Remotion, with TTS voiceover, word-aligned subtitles and a chapter progress bar. This review covers the nine-stage pipeline, the four checkpoints that stop the agent, the 9:16-less single visual style, the CPU-only rendering path, the Raspberry Pi 5 support, the PolyForm Noncommercial licence, and the honest limits: narration frozen after voiceover, heavy parallel-build resource needs, and a reference film that is also the quality bar you have to match.

7.7
ffmpeg-skill Review 2026 — 21 Structured Video-Editing Tools That Teach Claude Code, Cursor and Codex to Stop Guessing About FFmpeg

ffmpeg-skill is an open-source Agent Skill (created 2026-09-03, 320+ stars in four days, MIT license, v0.10.0) that gives Claude Code, Cursor, Codex and any agent that reads SKILL.md a real video editor: 21 structured Python tools that wrap local FFmpeg 5.0+ with a probe-first workflow, typed arguments instead of shell strings, lossless stream-copy cuts where possible, verification after execution, a machine-readable contract that is generated from the code rather than maintained beside it, and an MCP transport whose tool list is derived from that same contract so names and schemas cannot drift. Every job starts with probe.py measuring real duration, fps, resolution, colour and audio layout; cut/join/silence/fit handle editing, audio.py and loudness.py do voice clean-up and EBU R128 loudness, caption/overlay/graphics/color handle captions, lower-thirds, tone mapping and LUTs, export/check/report cover delivery presets (YouTube, Reels, podcast), and render/batch/multicam orchestrate whole projects. The project publishes real measurements: 92/92 verification steps on a 10-file real-device corpus (GoPro, DJI, iPhone Dolby Vision, HDR10, screen recordings), sync.py offset detection 40/40 within 10 ms, silence detection with zero missed gaps, scene detection at F1 0.97, and 72/72 graded agent runs. This review covers the 21 tools, the design principles that separate it from a list of FFmpeg one-liners, the FFmpeg 8 parser fix, the honest boundary where the agent's own vision must make the call, and how it compares with editing in Descript or CapCut.

7.4
whiteboard-animator Review 2026 — The CPU-Only Render Engine That Turns a Static Whiteboard Image Into a Hand-Drawn Video

whiteboard-animator (created September 8, 2026, MIT, ~100 stars) by masihsultani is a Python package that turns a finished whiteboard-style image into a hand-drawn reveal animation with one command and no GPU: pip install whiteboard-animator, then whiteboard-animate sketch.png --duration 8 -o sketch.mp4. It is the render engine behind the Whiteboard format at Kinoslide, released open source so anyone can animate their own images. The engine finds the ink (every connected blob of non-white pixels becomes a component, with a bundled 83 MB CRAFT text-detection ONNX model marking which components are text so words are written rather than traced), orders components the way a hand would (containers before contents, shapes before labels, text in reading order), assigns time slots that scale with the square root of area, gives every pixel a reveal time (strokes follow their skeleton from a real endpoint, closed outlines get one travelling front, fills get an outline pass then an angled sweep or bristled brush strokes, line art with junctions decomposes into sequential pen paths), and streams frames to ffmpeg. Optional narration: the drawing paces itself to an audio file, with a JSON region plan for exact drawing order — or --detect-regions lets Gemini propose the plan from the image and narration text. This review covers the render pipeline, the region-plan JSON format, quality presets, the honest limitations (white-background images only, text-based pacing that does not align to spoken-word timestamps, no audio alignment), and how the engine compares with the full Kinoslide product it powers.