Case study · Content Factory · Local video pipeline · Python / ffmpeg / diffusion

Cut to the voice, never the other way round

A narrated-documentary pipeline that runs entirely on one desk — script in, finished film out, no cloud APIs and no per-minute billing. Every number below is a measurement, and six of them overturned the obvious implementation. Two of those were only caught by looking at a picture; none of them showed up in a passing test.

Hardware
RTX 5080 (16 GB) · i9-14900K · 64 GB · Windows 11
Stack
Python · ffmpeg · ComfyUI · Qwen-Image & FLUX.1-schnell · Kokoro-82M
Delivered
12 films + 2 vertical shorts · 857 frames reviewed by eye
Licence
Apache-2.0 models throughout — commercially clean by construction
Cost
$0.00 per minute of finished video
FIG 01

The shape of the thing

Five steps and one gate, talking to each other through a single file. The ordering is the whole design: voice runs before pictures, because a shot lasts exactly as long as its narration takes to speak — read from the rendered wav's sample count, never estimated from a word count.

Pipeline — storyboard.json is the only contract between steps

storyboard.json — the single source of truth STEP 1 script cuts placed by hand STEP 2 voice Kokoro-82M STEP 3 frames Qwen / FLUX / Commons STEP 3b review a person, looking STEP 4 assemble ffmpeg concat STEP 5 publish delivery folder audio/*.wav measured duration → frames/*.png saved after every frame review.json fingerprints the frames <slug>.mp4 + burned subtitles GATE step 4 refuses to run on frames nobody has looked at
Each step is independently runnable and resumable — step 3 writes the storyboard back after every frame, so a 90-shot render killed at shot 60 costs one shot on resume, not sixty. Nothing a project produces lives in the repository; script, narration, frames and clips all travel with the finished video.
01 — THE PICTURE

Asking for more pixels returns a worse picture, not a bigger one

The obvious way to get a 2304×1296 frame is to ask the model for 2304×1296. That shipped a whole video before anyone measured it. FLUX is trained around a megapixel; asked for three, it stops composing a scene and starts fusing and duplicating local structure.

Composition size vs. whether the picture holds — same prompts, same seeds, five sizes

≈1.5 MP ceiling
1344 × 7681.03 MP
correct
1536 × 8641.33 MP — current
correct
1600 × 8961.43 MP
correct
1792 × 10241.84 MP
drifting
2048 × 11522.36 MP
wrong
2304 × 12962.99 MP — old default
broken
At 2.99 MP a dozen keys around a keyhole came back as one melted mass, a valve wheel as unreadable pulp, a seahorse fused into the calipers holding it. Every one of them was correct at 1.03. This is not the model being weak at hard prompts — it is the model being asked for a canvas it cannot hold.

Time per frame, by route — steady state, checkpoint resident

compose 1536×864→ Lanczos → 2304×1296
7.1 s
direct 2304×1296old default — broke the frames
13.1 s
compose → 4× ESRGANsecond model on a 16 GB card
24–31 s
The correct route is also the fastest — which is exactly why the old measurement went unquestioned for so long. It was a real measurement of the wrong quantity: it timed the routes and never asked whether the picture was still right. At 1.5× enlargement, plain Lanczos and a 4× ESRGAN are indistinguishable in a 1:1 crop.
02 — THE CAMERA

Oversampling the source made the shake worse

ffmpeg's zoompan rounds its crop window to whole source pixels. A slow move needs a fractional advance per frame, so it advances 0 px on some frames and 1 px on others, and the irregular rhythm reads as camera shake. More pixels is the intuitive fix. It does not work — the rounding is in the window position, not the sampling.

Judder — high-frequency residual of the frame-difference series; lower is smoother

zoompan, 2× oversampled
0.186
zoompan, 4× oversampledtwice the pixels, worse result
0.264
Pillow float crop boxresize(box=…) takes a float rect
0.035
The metric itself has a trap worth stating: plain std/mean of the frame-difference series triples once easing is on, because eased motion is deliberately slow-fast-slow. Subtract a moving average first and measure only the residual. Getting this wrong once already produced a false regression report — the one time a measurement looked alarming, the measurement was broken, not the code.
03 — THE CLOCK

A second of silence was being measured as narration

Kokoro leaves roughly 0.40 s of silence before the first word and 0.59 s after the last one, on every utterance regardless of length. Step 2 measured the whole file, so that dead air was the shot's duration: it set the pace of the video, it was invisible to the tests, and no amount of tuning the tail could reach it — the tail was being added on top of it.

Same words, same voice — project “clothes”

as shippedsilence measured as speech
11.3 min
re-cut, trimmed50 dB floor, 30 ms margin
10.1 min
1.2 minutes of the running time was padding. On eight-second shots it read as breathing room and nobody noticed. Cut-driven pacing runs two-second shots, where the same second is a third of the shot.

Where a cut may fall — pitch at the seam

at clause boundarieswhere a comma could sit
+0.0 st
every six wordsignoring grammar
+9.6 st
Break inside a noun phrase and the pitch jumps nearly ten semitones — the voice asking a question the writing never wrote. This is why cuts are placed by hand in the script; only their duration is measured.

The invariant everything hangs off

shot.duration == audio_sec + tail. The narration track concatenates shot audio plus exactly that tail; step 4 places every boundary at the running sum of the same value. Change one and you must change both — and the smoke test, which drives steps 2 and 4 for real and asserts the finished file's duration against the storyboard's prediction, is what says you did not. Current baseline: 1 ms drift.

04 — THE PACING

Dropping the camera move paid for everything else

Ken Burns across flat vector art has no parallax to reveal — the drawing just slides, which reads as a slideshow rather than a camera. Switching to hard cuts removed the motion, and three savings fell out of it that were not the point.

What changed when the camera stopped moving

PropertyPannedCutWhy
Frame size2304 × 12961920 × 1080headroom existed only for the zoom; a still frame is shown 1:1
Assemblyxfade chainconcatthe chain opens every clip at once and adds a filter per boundary
Audio/video drift8 ms0 msa cut consumes nothing, so there is no surplus arithmetic
Tail pause0.35 s0.12 / 0.34 sper shot now — a comma and a full stop need different gaps
Shot count77≈235≈7 words a shot instead of a paragraph
GPU time / video≈1.5 h≈4.5 hthe bill — and short shots are what make the lower resolution affordable
The two decisions pay for each other: three times as many frames, each a third cheaper, assembled by a linear operation instead of a filter graph that was already uncomfortable at 90 inputs. Cuts are placed by hand — automatic chunking targets a word count, which cannot see where a joke lands or where an idea turns.
05 — THE PROMPT

Naming a thing is a request for it, repeated once per shot

The negative prompt does nothing on FLUX.1-schnell — it is guidance-distilled and runs at cfg 1.0, so the negative conditioning is never applied. There is no way to ask for an absence. The only lever is the positive prompt, and a style prompt reaches every shot in the video.

Four batches, one mistake — each phrase was in the style prompt, so it ran once per shot

screen print texture, hand printed, subtle paper grain
→
signatures and edition numbers — “Nzainful”, “S0/20IG 1918” — and deckled paper borders
a print has a signature, so the model drew one · adding “unsigned” put them back
white blob faces, dot eyes, mitten hands
→
a person in all 268 shots — somebody sitting in a vat of snail shells
fixed by splitting cast_prompt out of base_prompt — it now reaches only shots marked as having people
one clear subject centred in a tall vertical frame
→
a mounted print: a white-bordered panel standing inside the scene
14 of 20 frames · the fix was to delete the instruction — 864×1536 already is the instruction
one clear subject standing full height
→
a standing human figure, in three of three probe frames that asked for none
3 of 3 · “standing full height” is a description of a person
Two more failure modes share the shape. Words with two senses: “bloom” returned a flower, “field” returned farmland, “shell” returned a seashell with a pearl. The vacuum: strip a prompt of every physical object and the model fills it from the only concrete nouns left in the text — which live in the style prompt, so “mid-century modern illustration” arrived as literal mid-century furniture and tower blocks. Write the scene; let the narration carry the abstraction.
06 — THE GATE

A scene lands; a diagram does not

857 frames across ten videos, every one reviewed by eye. The re-prompt rate was not evenly spread, and it sorted cleanly by what the script was about.

Frames re-prompted, by subject class

ideas with no physical form things and places
money, time, conspiracy
22–29%
gold, dragons, cities, clothes
12–17%
salt, dark, fire
10–11%
Almost every rejected frame in the first group had the same shape of prompt — “a stylised composition of exchange, accounting and storage” — and came back as a pleasant abstract landscape with no relation to the idea. The style prompt already supplies the flatness; when the shot prompt also describes an abstraction, nothing in the whole prompt names a thing that exists.

Why the gate is a gate and not a good intention

Step 4 refuses to run until a review file exists and matches the frames on disk; regenerate one frame and the review goes stale again, naming the shot. Every visual defect this project has shipped or nearly shipped was invisible to the test suite and obvious in a picture — oversized subtitles, mismatched asset backdrops, mid-phrase subtitle breaks, fake body copy, invented signatures, and a whole video's worth of frames broken by generating above the resolution ceiling.

07 — THE SECOND SHAPE

Turning it ninety degrees without touching the clock

A vertical profile, switched on by one environment variable. The split that makes it safe: geometry is a per-profile concern and the timeline arithmetic is not, and that boundary is enforced rather than trusted — a profile that names a timing constant refuses to load.

Geometry, transposed — the megapixel ceiling is respected rather than re-tested

HORIZONTAL · 16:9 1536 × 864 1.33 MP composed multiple of 16 · exact 16:9 ×1.25 1920 × 1080 delivered VERTICAL · 9:16 864 ×1536 1.33 MP ×1.25 1080 ×1920 delivered LOCKED — a profile that names one of these refuses to load FPS · SHOT_TAIL_SEC · SENTENCE_TAIL_SEC · CROSSFADE_SEC · MIN_SHOT_SEC · TRIM_* · STILL_FRAMES the audio/video lock is a shared invariant, not a per-shape decision
1.33 MP transposed, both sides still a multiple of 16, and a clean 1.25× enlargement with no reframing. The guard matters because the smoke test runs under the default profile — a profile that broke the lock in one aspect ratio would go on passing.

One take — silence at joins inside a single utterance

as the model returned it rebuilt
join 1
1066 ms
join 1, rebuilt
388 ms
join 2
989 ms
join 2, rebuilt
376 ms
Sending the whole script as one call does not make it one utterance: the TTS splits its own input to stay under a token limit, and pads every piece it speaks. One run in the log, three takes glued end to end — invisible to every check in the repo, audible immediately. Each chunk is now trimmed and the join rebuilt to a sentence's own beat, with the word timestamps moved to match.

Subtitle line width — Arial 58 px, 960 px between margins

28 characters
690 px
30 characters
722 px
32 characterschosen
783 px
34 characters
847 px
36 characters
911 px
Derived by proportion from the horizontal profile, the limit came out at 26 — and too narrow is the expensive direction. It does not compress the text, it adds a third line and strands single words on a cue of their own, both of which the style guide calls a fault. Measured instead of reasoned.

First-pass frames accepted, vertical — 20 shots each, same profile, same model

mid-century printcosmology script
13 / 20
cel cartoonhuman-perception script
18 / 20
The gap is partly the writing — the second script was written after the prompt traps above were understood — and partly the style. The poster margins that survived in 3 of 20 mid-century frames appeared in 0 of 20 cartoon frames, which settles what caused them: a tall canvas plus the words “mid-century modern” is a travel poster, and a travel poster has margins. Not the aspect ratio.
BY THE NUMBERS

What it costs to run

~20 minper 12-minute film, mostly unattended
146words / minute, measured narration pace
857frames reviewed by eye
1 msaudio / video drift, smoke baseline
2virtualenvs, deliberately
0cloud API calls

Licensing is a design constraint, not an afterthought. FLUX.1-schnell, Qwen-Image and Kokoro-82M are all Apache-2.0. FLUX.1-dev is better and is forbidden here — its licence bars commercial use. Real images are fetched from Wikimedia Commons and the Met and sorted into tiers by an allowlist: public domain and CC0 pass freely, CC-BY passes with credit, CC-BY-SA is excluded because the share-alike term can arguably reach the finished video, and unknown is rejected — never assumed permissive. Being wrong that way costs a generated image; being wrong the other way costs a takedown on a monetised video.