Kategori: Podcast Production

  • A Realistic Podcast Post-Production Workflow: Clean, Master, Clip

    A Realistic Podcast Post-Production Workflow: Clean, Master, Clip

    You’ve recorded an episode. What happens between that raw file and something you’d actually publish? This post walks through the three-step order we recommend, and — just as important — why the order matters.

    Step 1: Remove background noise first

    Start with the Background Noise Remover. Home recordings almost always carry a constant noise floor: air conditioning, computer fans, room hum. Removing it first matters because every later step — especially loudness normalization — amplifies whatever is already in the file. If you normalize a noisy recording, you normalize the noise too.

    We measured this while tuning the tool. On our test file, running the cleanup with loudness normalization switched on inside the same pass cut the signal-to-noise improvement from +11.1 dB to +2.2 dB, and left the residual noise floor at −37.4 dBFS instead of −52.0 dBFS. Normalization raises quiet content, and the quietest content in a raw recording is the noise. That is exactly why it belongs in a separate, later step.

    Pick the gentlest aggressiveness level that solves your problem. Higher levels remove more noise but can leave the voice sounding processed. We measured the difference between all four levels — see our post on noise removal levels for the numbers.

    Step 2: Normalize loudness

    Next, run the cleaned file through the Podcast Audio Studio. It applies two-pass loudness normalization to a target you choose — Podcast (−16 LUFS with a −1.5 dBTP true-peak ceiling), YouTube or Spotify (both −14 LUFS, −1.0 dBTP) — trims silence from the start and end of the file, and adds a fade in and out.

    Loudness consistency is the single most audible difference between a produced-sounding episode and a raw one — more than any effect or EQ. If the terms in that paragraph are unfamiliar, we explain LUFS and true peak here.

    One honest note on what it does not do: it trims silence at the two ends of the file, not long pauses in the middle. Removing a rambling gap halfway through the episode is editing, not mastering.

    Step 3: Cut clips last

    Finally, run the finished episode through the Podcast Clip Cutter to generate short clips for social platforms. Doing this last means the clips inherit the cleaned, normalized audio — you never want to promote your show with a clip that still has fan hum in it.

    The clip cutter finds high-energy moments automatically. Treat its output as candidates, not final picks: the loudest moment is often a good one, but not always the best one. Listen before you post. More on how energy detection works, and where it falls short.

    Why this order and not another?

    Noise removal before normalization: normalization raises quiet content, including noise. Clean first, then raise — the measurement above is what that costs when you get it backwards.

    Clips last: any change you make to the master after cutting clips would force you to re-cut them.

    What this workflow doesn’t do

    It doesn’t edit content. Removing filler words, cutting a rambling section, rearranging segments — that’s editorial work these tools don’t attempt. What they do automate is the mechanical layer: noise, loudness, end-trimming, fades, clip extraction. For many episodes, that’s most of the time spent in post.

    All three tools are free, need nothing installed, and accept files up to 200 MB and three hours long, with a limit of 10 jobs per day. Your files are deleted from the server automatically: finished output after two hours, the uploaded source after one.

  • How Much Noise Should You Remove? We Measured All Four Levels

    How Much Noise Should You Remove? We Measured All Four Levels

    Most noise-removal tools give you a strength slider and no guidance. We think you should know what each position actually does before you commit, so we ran the same noisy recording through all four levels of our Background Noise Remover and measured the results.

    The four levels, measured

    On our test recording, measured noise-floor reduction was:

    • Soft: −8.6 dB — noticeable cleanup, voice character fully preserved
    • Normal: −11.8 dB — the recommended default for typical home recordings
    • Aggressive: −22.1 dB — for recordings with prominent constant noise
    • Brutal: −43.5 dB — maximum removal, for recordings that would otherwise be unusable

    A quick reference point: a 10 dB reduction is perceived as roughly halving the loudness of the noise, so even the Soft setting makes a clearly audible difference. Speech level stayed essentially untouched across all four levels.

    Why the levels differ at all

    One thing worth being concrete about, because it explains the tool’s behaviour: the file’s actual noise floor is measured first, and each level then sets the filter’s threshold relative to that measurement rather than to a fixed number. A fixed threshold is what most simple denoisers use, and it’s why they barely touch some files — in our testing a fixed setting moved the noise floor by only 1–2 dB on a recording whose noise sat higher than the value assumed. Measuring first is what makes the same “Normal” meaningful across different recordings.

    What the levels cost you

    Noise removal is always a trade. The filters estimate what’s noise and subtract it — push them harder and they start subtracting parts of the voice too. In practice:

    Soft and Normal keep the voice natural on almost all material. Aggressive can slightly dull room ambience and the very top end of the voice; on most spoken-word content this is acceptable. Brutal will audibly process the voice — use it when the alternative is throwing the recording away, not as a default.

    How to choose in 30 seconds

    Process a short section at Normal and listen. Still hearing the noise? Step up one level. Voice sounding underwater or metallic? Step down. The goal is the lowest level that makes the noise stop drawing attention to itself — not silence at any cost. A quiet, steady noise floor under a clear voice is fine; listeners don’t notice it. They notice artifacts.

    What this tool won’t fix

    Constant, steady noise — hum, fans, hiss — is what it’s built for. Intermittent sounds are different: keyboard clacks, a door slamming, a dog barking mid-sentence. Those overlap with speech in both time and frequency, and removing them cleanly is editorial work, not filtering. If your recording’s problem is interruptions rather than a noise floor, no strength setting here will solve it.

    Try it on your own recording

    Upload a file and compare the levels yourself — the before/after preview plays ten seconds of each, so you can decide by ear rather than by our numbers. The tool is free, with a limit of 10 jobs per day.

    Once the noise is gone, loudness is the next step, in that order and not the other way round — see our full post-production workflow for why.

  • Podcast Transcripts: Why Text Still Does the Heavy Lifting

    Podcast Transcripts: Why Text Still Does the Heavy Lifting

    Search engines don’t listen to audio. However good an episode is, its content is invisible to search until it exists as text somewhere. That’s the entire case for transcripts, and it’s a strong one — no growth-hack statistics required.

    What a transcript actually does for an episode

    It makes the episode findable. A transcript page can be indexed, so the things you actually said — names, tools, techniques, opinions — become searchable text. Without it, the only indexable content is your episode title and description.

    It makes the episode quotable. Listeners who want to share a point can copy the words. Writers who want to reference your episode can cite it. Neither happens with audio alone.

    It makes the episode accessible. For deaf and hard-of-hearing audiences, and for anyone in a situation where they can read but not listen, the transcript is the episode.

    Producing one without typing it yourself

    The AI Speech to Text tool takes an audio or video file and returns a timed transcript. It accepts MP3, WAV, M4A, AAC, OGG and FLAC audio as well as MP4, MOV, WebM and MKV video, and gives you the result three ways: plain text for reading, and SRT or VTT if you want the timings attached. It labels speakers as it goes, and can optionally translate the transcript into another language.

    Automatic transcription gets you most of the way; plan a quick read-through pass for the things automatic systems reliably miss: proper names, brand names, technical terms. Fix those, leave the rest.

    Publishing it

    Put the transcript on the same page as the episode’s player and show notes, as normal page text — not inside an image, not as a downloadable file only. One episode, one URL, with the audio and its text together. If the transcript is long, a heading per segment helps readers skim and gives the page useful structure.

    Timestamps: keep them or strip them?

    For a published transcript page, light timestamps at section boundaries are useful; a timestamp on every line is noise for readers. Keep the fully timed version for yourself — the SRT or VTT export is exactly what you need later to find quotable moments or cut clips from the right places.

  • Auto-Clipping a Podcast: What Energy Detection Can and Can’t Find

    Auto-Clipping a Podcast: What Energy Detection Can and Can’t Find

    The Social Media Clip Cutter takes a long recording and returns short clips ready for social platforms. It’s worth understanding how it decides where to cut — both to use it well and to know what it can’t do.

    How it works

    The tool measures the loudness of your recording second by second, then looks for high-energy moments: sections that are markedly more intense than their surroundings. Around each detected peak it cuts a clip of exactly 15 seconds — five seconds before the peak and ten after — with a short fade in and out, keeps at least ten seconds between clips so they don’t overlap or repeat the same moment, ranks candidates by energy, and keeps the strongest ones. You choose whether you want 5, 10 or 15 clips.

    You get the clips individually as MP3s, plus a single mashup of all of them, packaged with a metadata file listing the timestamp each clip came from — so you can always find the moment in the original.

    A sensitivity setting (Soft, Normal, Aggressive) controls how wide the candidate pool is. One honest note: on a recording with plenty of lively material, all three settings often return the same clips, because the highest-energy moments are the highest-energy moments regardless of where you set the bar. The setting earns its keep on quiet or unevenly recorded material.

    Why a fixed loudness threshold wouldn’t work

    This is the part most “auto clip” descriptions skip. Speech has a much narrower dynamic range than people assume — in one real episode we measured, the median level was −18.5 dB and the loudest peak −14.9 dB, a total spread of 3.6 dB. A rule like “anything more than 4 dB above the median” sits above even the loudest moment in that file and returns nothing at all. That’s why detection works on relative ranking within your specific file rather than a fixed decibel value.

    The honest limitation: loud is not the same as good

    Energy detection finds intensity — laughter, raised voices, animated discussion. Often that correlates with your best moments. But your sharpest insight may have been delivered calmly, and a loud moment can be two people talking over each other. A volume-based algorithm cannot tell the difference, and we’d rather say that plainly than pretend otherwise.

    The realistic workflow: candidates, not final picks

    Treat the output as a shortlist. The tool has done the tedious part — scanning the full recording and pulling out moments worth reviewing, with timestamps. Your part is a few minutes of listening: keep the clips that stand on their own, discard the ones that need context, and if you know a great calm moment the ranking missed, its neighbours’ timestamps make it easy to locate.

    A clip works on social platforms when someone with zero context understands it. That judgment stays with you; the scanning no longer has to.

    One practical tip

    Run your episode through noise removal and loudness normalization before cutting clips — the clips inherit whatever audio quality the source has, and a clip is often someone’s very first contact with your show. Our post on the full post-production workflow covers the order in detail.