stitch-videos-ffmpeg
SkillMediaStitch video segments with ffmpeg concat, xfade, overlay, audio mux, and export settings. Ships montage.py, a free montage assembler (python3 + ffmpeg, no keys). It takes a JSON EDL of clips and stills, normalizes them and hard-cuts them in order, burns captions from an SRT, a cue list or word timings, and lays a VO over a music bed that ducks under it, mastered to -14 LUFS.
Use stitch-videos-ffmpeg in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add stitch-videos-ffmpeg and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the stitch-videos-ffmpeg skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by gooseworks-ai/goose-skills in skills/ads/packs/video-ad-formats/stitch-videos-ffmpeg/SKILL.md and read by ahel’s review.
Purpose
Stitch video segments with ffmpeg concat, xfade, overlay, audio mux, and export settings.
Implementation status: refactored from existing repository skills. The workflow consolidates behavior that previously lived across larger skills.
Sources: ad-studio, voiceover-product-ad, ugc-product-video, voiceless-music-transformation-reel, product-stopmotion-ad.
Extraction notes: assembly and composite scripts.
Montage helper: scripts/montage.py
A free, deterministic way to cut a montage ad from finished clips. It needs python3 and
ffmpeg/ffprobe only (Pillow is optional, see Captions). It makes no network or provider
calls and needs no keys, so re-cuts cost nothing.
| Step | What it does |
|---|---|
edl | Checks a JSON spec and writes edl.json. Every file must exist and is probed with ffprobe; every cut must fit inside its clip. All problems are listed at once. |
assemble | Scales each cut to one size (default 1080×1920, cover crops to fill, contain pads), one frame rate (default 30), square pixels and yuv420p. Then it hard-cuts them in order with the filter_complex concat filter, never the -f concat demuxer. Each cut gets an exact frame count, so the output length equals the EDL total. |
captions | Burns an SRT, a JSON cue list, or word timings (one cue per --per words) into the video. The cue text is kept exactly as given (wrapping only turns spaces into line breaks). |
mix | Lays the VO over the video and adds an optional music bed that ducks under the VO (sidechain compression, 20:1). Then it masters to -14 LUFS / -1 dBTP with two-pass loudnorm. |
run | Runs all four steps from one spec and writes manifest.json. |
Each step prints a one-line JSON summary. Exit codes: 0 ok, 1 ffmpeg failed, 2 bad
spec or arguments, 3 a tool is missing (ffmpeg, ffprobe, or a caption renderer).
Run it
After gooseworks fetch stitch-videos-ffmpeg (or as a dependency of a format package), the
scripts are in /tmp/gooseworks-scripts/stitch-videos-ffmpeg/scripts/:
S=/tmp/gooseworks-scripts/stitch-videos-ffmpeg/scripts # or this folder's scripts/ in a checkout
python3 $S/montage.py run --spec montage.json --out edits/master.mp4 --workdir edits/work
# or one step at a time
python3 $S/montage.py edl --spec montage.json --out edits/edl.json
python3 $S/montage.py assemble --edl edits/edl.json --out edits/body.mp4
python3 $S/montage.py captions --video edits/body.mp4 --words audio/vo.words.json \
--respell '{"symbiotic": "synbiotic"}' --color "#FFE800" --out edits/captioned.mp4
python3 $S/montage.py mix --video edits/captioned.mp4 --vo audio/vo.mp3 \
--music audio/bed.mp3 --out edits/master.mp4
Spec format
Paths are relative to the spec file. Times are seconds or a timecode string ("SS",
"MM:SS", "HH:MM:SS", optional .ms).
{
"output": {"width": 1080, "height": 1920, "fps": 30, "fit": "cover"},
"words": "audio/vo.words.json",
"clips": [
{"file": "clips/scene-01.mp4", "label": "hook", "word_range": [0, 4]},
{"file": "clips/scene-02.mp4", "label": "feature", "in": 0.2, "out": 1.6},
{"file": "clips/scene-03.mp4", "label": "reaction", "in": "00:00.5", "duration": 0.9},
{"file": "clips/scene-04.mp4", "label": "b-roll", "t_in": 4.1, "t_out": 5.0},
{"file": "clips/scene-05.mp4"},
{"file": "overlays/landing-page.png", "label": "landing-page", "duration": 1.5,
"pan": {"from": {"x": 0.5, "y": 0.2, "zoom": 1.0}, "to": {"x": 0.5, "y": 0.7, "zoom": 1.6}}},
{"file": "brand/end-card.png", "label": "end-card", "duration": 2.0}
],
"clip_audio": "drop",
"captions": {"words": "audio/vo.words.json", "per": 1, "respell": {"symbiotic": "synbiotic"},
"style": {"color": "#FFE800", "y": 0.5}},
"audio": {"vo": "audio/vo.mp3", "music": "audio/bed.mp3", "music_start": 14.2}
}
How long each cut is (first match wins):
word_range: [i, j]cuts on the VO's word boundaries. It needs a top-levelwordsfile, used as the transcriber wrote it: a flat[{text|word, start, end}]list,{words: [...]}(OpenAI, ElevenLabs;spacingentries are skipped), Whisper's{segments: [{words: [...]}]}(goose-studio'stranscribe-audio-fal), or fal's{chunks: [{text, timestamp: [s, e]}]}. Words with no time are skipped, and leading spaces are stripped. The cut runs from the start of wordito the start of wordj + 1, or to the last word's end. The first cut starts at 0 so it covers the lead-in. Consecutive ranges tile the VO with no gaps.t_in/t_outgive a window on the timeline. The cut ist_out - t_inlong. A window that does not start where the previous cut ended is reported as a gap or overlap warning;--strictturns warnings into errors.in+out, orin+duration, trim the source clip.- With nothing set, the whole clip plays (from
in, default 0).
Every cut starts at in in its source (default 0). A still image (.png, .jpg, .webp)
needs a length; pan zooms and pans across it. The still is first cover-cropped to the output
aspect (9:16 by default), and x/y are the window centre as a fraction of that cropped image
(zoom ≥ 1). A video clip may be up to 0.1 s short of its cut, and then its
last frame is held. Anything shorter is an error.
clip_audio (--clip-audio on assemble): drop (default) or keep. keep keeps each
clip's own sound and fills silence under stills and silent clips.
Captions
- Sources:
srt,cues(a list or a JSON file of{start, end, text}), orwords(any of the word-file shapes above) plusperand an optionalrespellmap.respellswaps a misheard word for the locked spelling and keeps the punctuation around it. Overlapping cues: the later one wins. - Look: each cue is drawn whole, in one colour with an outline. With
per: 1that is one word at a time, each held until the next. There is no active-word (karaoke) highlight. - Timing: a cue shows from the first output frame at or after its start until the first frame at or after its end (a 1 ms caption grid; ffmpeg older than 5 falls back to 1/25 s and says so in the step's warnings).
- Style:
font,font_size(px, default 4.5% of the height),color,outline_color,outline(px),y(centre of the caption block as a fraction of the height, default 0.72) andmax_width(default 0.86 of the width). - Renderers:
autouses Pillow when it is installed, so captions look the same on every machine. Otherwise it uses libass, which needs an ffmpeg with theassfilter. Homebrew's defaultffmpeghas no libass and nodrawtext, so on a Mac runpython3 -m pip install pillow. With neither, the step exits 3 and says so.
Mix defaults (all overridable)
| Setting | Default | Meaning |
|---|---|---|
vo_lufs | -16 | VO level before mixing (static gain from a measured pass) |
music_lufs | -24 | Bed level between VO lines, before ducking |
duck_threshold / duck_ratio | 0.02 / 20 | The bed ducks while the VO is above the threshold. ffmpeg caps the ratio at 20. |
duck_attack / duck_release | 20 / 400 ms | How fast the bed dips and comes back |
vo_start, music_start | 0 | Where each track enters on the timeline. vo_start moves only the VO audio, not word_range cuts or word captions, so keep it 0 when those come from the same VO. |
music_fade_in / music_fade_out | 0 / 1 s | A bed shorter than the video loops (with a warning) |
target_lufs / target_tp | -14 / -1 | Master loudness. off skips it. |
keep_video_audio | false | Mix the video's own audio in too (not ducked) |
A VO that runs past the end of the video is cut and reported as a warning, so lengthen the EDL (for example, the end card) instead.
Tests
tests/test_stitch_montage.py makes tiny synthetic clips, a still and two tones with ffmpeg,
then checks: the output length equals the EDL total to the frame, size/fps/pixel shape are
normalized, captions appear only on their cue frames with the exact text, the bed ducks
under the VO (about 14 dB on the fixture), and the master is at -14 ±1 LUFS. It runs in CI
(media-tests) and locally:
python3 -m pytest -q skills/ads/packs/video-ad-formats/stitch-videos-ffmpeg/tests
Related
- The older scripts here (
composite.py,composite_final.py/.sh,normalize_clip.sh,voiceless/composite.py) are unchanged. mix-masterremains the multi-clip VO + SFX mix. Amontage.pymaster is already at -14 LUFS, so a later -14 LUFS finishing pass barely changes it.caption-burnremains the place for plate, seam and hook-card caption styles.
Inputs
- A clear user brief or source asset path.
- Brand, product, audience, platform, and approval constraints when relevant.
- Required credentials or provider access for any external service used by this skill.
- Output directory or test-run directory where artifacts should be saved.
Workflow
- Read the brief and confirm all required inputs are present.
- Load any referenced files in this skill folder only when they are needed.
- Run the provider, script, or planning workflow described by this skill.
- Save outputs under the requested output folder or
skills/test-runs/<timestamp>/<skill-name>/during tests. - Write or update a
manifest.jsonfor executable runs with status, provider, outputs, warnings, and errors.
Output
- Primary artifact or written plan requested by the skill.
manifest.jsonfor executable runs.verification.mdor a short verification summary that names the checks performed.- Any generated source assets, intermediate files, or final exports in the run folder.
Quality Checks
- Required files exist and paths in the manifest are valid.
- Output matches the requested format, platform, duration, dimensions, or text structure.
- Brand claims, captions, on-screen text, and CTAs follow the provided brand rules.
- Provider failures, skipped integrations, and human-review needs are explicit.
Failure Modes
- Missing credentials, provider access, or source files.
- Output does not match requested dimensions, duration, structure, or brand constraints.
- Generated media contains artifacts, unreadable text, unsafe claims, or caption collisions.
- Scaffolded skills cannot run production workflows until implementation details are added.
Signals
- GitHub stars
- 1k
- Forks
- 211
- Last commit
- Oct 2026
ahel review
K1binfo
installs-packagesK1binfo
installs-packages (in scripts/montage.py)K5info
obfuscation (in tests/test_stitch_montage.py)
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
stitch-videos-ffmpeg- Source
- github.com/gooseworks-ai/goose-skills
github.com/gooseworks-ai/goose-skills
Related picks
Skill · autonnel
The pick for Landing Pagelanding-page-copy
Skill · leadmagic
The pick for Landing Pagelanding-page-audit
Skill · hashgraph-online
The pick for Landing Pageimagegen-frontend-web
Skill · leonxlnx
The pick for Landing Pagepython-performance-optimization
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Python