The Social Media Clip Cutter takes a long recording and returns short clips ready for social platforms. It’s worth understanding how it decides where to cut — both to use it well and to know what it can’t do.
How it works
The tool measures the loudness of your recording second by second, then looks for high-energy moments: sections that are markedly more intense than their surroundings. Around each detected peak it cuts a clip of exactly 15 seconds — five seconds before the peak and ten after — with a short fade in and out, keeps at least ten seconds between clips so they don’t overlap or repeat the same moment, ranks candidates by energy, and keeps the strongest ones. You choose whether you want 5, 10 or 15 clips.
You get the clips individually as MP3s, plus a single mashup of all of them, packaged with a metadata file listing the timestamp each clip came from — so you can always find the moment in the original.
A sensitivity setting (Soft, Normal, Aggressive) controls how wide the candidate pool is. One honest note: on a recording with plenty of lively material, all three settings often return the same clips, because the highest-energy moments are the highest-energy moments regardless of where you set the bar. The setting earns its keep on quiet or unevenly recorded material.
Why a fixed loudness threshold wouldn’t work
This is the part most “auto clip” descriptions skip. Speech has a much narrower dynamic range than people assume — in one real episode we measured, the median level was −18.5 dB and the loudest peak −14.9 dB, a total spread of 3.6 dB. A rule like “anything more than 4 dB above the median” sits above even the loudest moment in that file and returns nothing at all. That’s why detection works on relative ranking within your specific file rather than a fixed decibel value.
The honest limitation: loud is not the same as good
Energy detection finds intensity — laughter, raised voices, animated discussion. Often that correlates with your best moments. But your sharpest insight may have been delivered calmly, and a loud moment can be two people talking over each other. A volume-based algorithm cannot tell the difference, and we’d rather say that plainly than pretend otherwise.
The realistic workflow: candidates, not final picks
Treat the output as a shortlist. The tool has done the tedious part — scanning the full recording and pulling out moments worth reviewing, with timestamps. Your part is a few minutes of listening: keep the clips that stand on their own, discard the ones that need context, and if you know a great calm moment the ranking missed, its neighbours’ timestamps make it easy to locate.
A clip works on social platforms when someone with zero context understands it. That judgment stays with you; the scanning no longer has to.
One practical tip
Run your episode through noise removal and loudness normalization before cutting clips — the clips inherit whatever audio quality the source has, and a clip is often someone’s very first contact with your show. Our post on the full post-production workflow covers the order in detail.
