Everyone with a podcast has been told to cut it into short vertical clips. The advice is
sound and the work is miserable: scrubbing a sixty-minute timeline looking for thirty good
seconds is an afternoon nobody has.
Automation cannot do the whole job, but it can do the tedious half.
What a machine can actually detect
Without understanding the words, the one signal a machine can read reliably is
energy — how loud each moment is relative to the rest.
That turns out to correlate with interesting moments more often than you would expect.
People get louder when they laugh, when they interrupt, when they emphasise, when they get
annoyed, and when they finally say the thing they have been circling for ten minutes. Energy
finds all of those.
It also finds coughs, microphone bumps and someone knocking a mug over. Which is the point
of the caveat further down.
How the moments get picked
The recording is measured one second at a time, so every moment has its own level rather
than one average for the whole file. That distinction matters more than it sounds: a summary
figure for a sixty-minute episode tells you nothing about where anything happened.
From there:
- The loudest moments are taken in order.
- A minimum gap is enforced so two clips never cover the same passage. Without it, one
animated thirty-second exchange produces five nearly identical clips. - Each clip is built around its peak with lead-in and follow-through — a few seconds
before, the rest after — so the listener hears what led up to the moment rather than
landing in the middle of it. - Short fades at both ends stop a clip starting with a click or cutting off mid-word.
Why fifteen seconds
Long enough to carry one complete thought, short enough for every vertical video platform,
and short enough that a viewer will finish it. Every clip being the same length also makes them
easy to batch: same template, same captions, same publishing routine.
The honest limitation
Energy tells you where something happened. It cannot tell you whether it was
interesting. A cough, a laugh at a terrible joke and a genuinely brilliant one-liner look
identical to a level meter.
So treat the output as a shortlist, not a publishing queue. Ten candidate clips take seconds
to produce and about three minutes to listen through. Two or three will be usable. That is a
good trade against an afternoon of scrubbing — but it is a shortlist, and the listening
step is not optional.
A related honesty point about sensitivity controls: on a lively episode with plenty of loud
moments, a strict setting and a loose one pick the same top ten, because the loudest moments are
the loudest either way. The control earns its place on quiet or uneven recordings, where a
strict setting comes back with two clips and a looser one finds eight.
What to do with the clips
You get audio, not video — the visual layer is still yours to add. The usual routes:
- A waveform or static image with burned-in captions, assembled in whatever editor you
already use. - A talking-head crop, if you filmed the session.
- The timestamps alone, used as a map back into your video project.
That last one is underrated. Even if you never publish a single audio clip, knowing that the
five most animated moments of the episode were at 12:04, 19:30, 27:15, 41:52 and 48:07 saves
the entire search.
Where it fits in the routine
Publish the episode first, then mine it. Clean up any steady background noise, master the
audio to a consistent loudness, publish, then run the finished file through clip cutting and
work from the shortlist. The clips inherit whatever quality the master has, so doing it in that
order means you are not cutting up a noisy file.
Back catalogues are where this pays off most. Fifty old episodes nobody has time to mine
become fifty shortlists, and the good moments in them are just as good as they were the week
they went out.
Try it yourself. The tool is free, needs no account and deletes your file afterwards.




