Blog

  • A Complete Subtitle Workflow: Generate, Translate, Convert

    A Complete Subtitle Workflow: Generate, Translate, Convert

    Subtitles involve three separate jobs that often get lumped together: creating them, translating them, and delivering them in the format a platform accepts. Each is one step here.

    Step 1: Generate timed subtitles from the video

    The AI Video Subtitle Generator listens to your video and produces subtitles with timings attached, which you can download as an SRT file — or, if you’d rather not deal with subtitle files at all, as a copy of your video with the subtitles burned in.

    Auto-generated subtitles get the timing and most words right; plan one review pass for names, technical terms, and any place where two people talk at once — the predictable weak spots of every automatic system. Review is fastest now, in the original language, before any translation multiplies the text.

    Step 2: Translate, keeping the timing

    Translating subtitles is not just translating text: every line has a start time, an end time, and a length that has to stay readable. The AI Subtitle Translator takes SRT and VTT files and translates them while preserving their timing, so the translated line appears exactly when the original did. One source file can become as many language versions as you need.

    Step 3: Deliver the format the platform wants

    Two file formats dominate: SRT (the older, universally accepted one) and VTT (the web-native one used by HTML5 players and many web platforms). They contain the same essential information, so conversion is lossless and instant — SRT to VTT and VTT to SRT both exist as one-click tools. Keep one master format for editing and convert copies on the way out.

    Already have the audio transcribed?

    If you’ve run the episode or video through AI Speech to Text, you already have SRT and VTT exports of the same content — the same file serves as a transcript page and as a subtitle track.

    Why bother with subtitle files at all when platforms auto-caption?

    Control and portability. A subtitle file you generated and reviewed is correct once, everywhere you publish — auto-captions are regenerated by each platform, mistakes included, and can’t be fixed centrally. Subtitles also make the video watchable with the sound off and readable by search engines where the platform indexes them.

  • The PDF Toolkit: Which Tool for Which Job

    The PDF Toolkit: Which Tool for Which Job

    PDF problems are small, specific and annoying precisely because you hit them five minutes before a deadline. Here’s the map of Bubixo’s PDF tools, organized by the situation you’re in — every tool listed is live on the site.

    “Several PDFs need to be one document.”

    Merge PDF. Contracts plus annexes, chapters into a manuscript, scans into one file.

    “One PDF needs to be several.”

    Split PDF to break a document apart, Extract PDF Pages to pull out just the pages you need, Delete PDF Pages to remove the ones you don’t.

    “It’s too big to send.”

    Compress PDF. Scanned documents compress especially well — scans are images, and images carry most of a PDF’s weight.

    “I need to edit the text.”

    PDF to Word, edit in your word processor, then Word to PDF to lock it back down for sending.

    “I need pages as images, or images as a PDF.”

    PDF to JPG for slides-as-images and previews; JPG to PDF to turn photographed pages into a shareable document.

    “A page is sideways.”

    Rotate PDF. Common with scans; fixes orientation permanently instead of relying on your viewer to remember.

    “It asks for a password I legitimately have.”

    Unlock PDF removes the password from a document you’re authorized to open, so you stop retyping it. It is not a tool for opening documents that aren’t yours.

    Two habits that prevent PDF pain

    Keep the editable original — the Word file, the export source. Converting back from PDF is always the lossier direction.

    And when preparing a document that will be scanned, scan at a sensible resolution: extreme scan quality is the main reason PDFs end up needing compression at all.

    If your document problem is actually a “I’d rather listen to this” problem, see turning a PDF or EPUB into an audiobook.

  • Video File Too Big? Compress, Convert or Cut – Which You Need

    Video File Too Big? Compress, Convert or Cut – Which You Need

    “The file is too big” has three different causes, and each has a different fix. Picking the right one saves both quality and time.

    Cause 1: The video is simply heavy → compress it

    Same video, same length, same format — you just need fewer megabytes. The Video Compressor re-encodes the video more efficiently, trading a controlled amount of visual quality for a much smaller file. This is the answer for “it won’t fit in the upload limit” and “it’s eating my storage”.

    You choose a quality level (High, Medium or Low) and optionally scale the output down to 1080p, 720p or 480p, which is often the bigger win for screen recordings and talking-head footage. Uploads go up to 500 MB and 30 minutes.

    Cause 2: The format is wrong → convert it

    A platform or editor rejects the file outright: it wants MP4 and you have MKV, or you have a MOV from an iPhone that a Windows tool won’t open. The Video Format Converter converts between MP4, WebM, MOV, AVI and MKV, with uploads up to 500 MB.

    Converting doesn’t primarily shrink the file — it changes the container so the file is accepted. Sometimes converting an old format to a modern one also reduces size as a side effect, but if size is your actual problem, compression is the direct answer.

    Cause 3: You only need part of it → cut it

    The recording is an hour, the useful part is four minutes. The Video Cutter trims the file to a time range without touching the rest. This is the highest-quality “size reduction” there is: the part you keep isn’t re-encoded harder, there’s just less of it.

    Combining them

    The common real-world chain: cut first (drop the dead footage), then compress (shrink what remains), and convert only if the destination demands a specific format. Cutting first also makes every later step faster, because the file entering it is smaller.

    The same logic applies on the audio side, where the order matters for quality rather than speed — see our guide to audio formats for the convert-down-not-up rule.

  • Turning a PDF or EPUB into an Audiobook: What to Check

    Turning a PDF or EPUB into an Audiobook: What to Check

    Generating audio for a whole book is a different job from converting a paragraph. The AI Audiobook Creator takes a PDF or EPUB — up to 1,500 pages — and produces a narrated MP3 audiobook with AI voices. Because the input is long, preparation and review matter more than with any other audio tool. Here’s what to check at each end.

    Before: the source file decides the result

    EPUB is usually the better source. EPUB files carry real text and chapter structure. PDFs vary: a PDF made from a text document works well, but a PDF that is actually scanned page images contains no readable text — check by trying to select and copy a sentence. If you can’t, the file needs OCR before it can be narrated.

    Strip what shouldn’t be read aloud. Page headers and footers, page numbers, footnotes, long tables and URLs all become noise when narrated. The cleaner the source, the fewer awkward moments in the audio.

    Mind the front matter. A narrated copyright page and table of contents is a rough first minute for a listener. Consider starting the conversion from the introduction.

    Choosing voices

    A voice you like for one paragraph needs to hold up for hours. Listen to a longer sample before committing to a full book, and prefer clarity over character — a slightly plain voice ages better over ten chapters than a heavily stylized one.

    The tool does more than read everything in one voice. It analyses the book to identify the narrator and the individual characters who speak, then assigns each of them a distinct voice, deliberately avoiding giving similar voices to two characters who appear in the same scene. For fiction with a lot of dialogue this is the difference between a readable audiobook and a confusing one — but it also means the character detection is worth reviewing before you commit to a long render.

    After: spot-check like a proofreader

    Nobody re-listens to a whole audiobook before publishing, and you don’t need to. Check the first minutes of several chapters, plus any section with unusual content — names, numbers, quoted dialogue, non-English words. These are where text-to-speech is most likely to stumble, and they cluster predictably.

    A note on rights

    Only convert documents you have the rights to: your own writing, public-domain works, or material whose licence permits it. Converting someone else’s book for personal listening and publishing the result are very different situations legally — when in doubt, don’t publish.

    If your source is an article rather than a book, the workflow is shorter and cheaper — see turning blog posts into audio.

  • Which Audio Tool Do You Actually Need? A Decision Guide

    Which Audio Tool Do You Actually Need? A Decision Guide

    Tool lists describe features. Problems don’t arrive as features — they arrive as “my recording sounds bad” or “I need this in a different format”. This page maps problems to tools, in plain language.

    “My recording has constant background noise — hum, fan, hiss.”

    Background Noise Remover. Four strength levels; start at Normal. It removes steady noise floors, not one-off interruptions like a door slam. We measured what each level removes.

    “My episodes come out quieter than other shows.”

    Podcast Audio Studio. Normalizes loudness to a consistent target (Podcast, YouTube or Spotify), trims silence at the start and end, adds fades. Run it on every episode as the final mastering step. What LUFS and true peak mean.

    “I need short clips from a long recording for social media.”

    Social Media Clip Cutter. Finds high-energy moments and cuts them into ready 15-second clips. Review before posting — it finds loud moments, which are often but not always your best ones.

    “This audio file is in the wrong format.”

    Audio Format Converter for MP3, WAV, M4A, OGG and FLAC. Which format to pick, and why. For video files, see the video tools below.

    “I want to turn written text into speech.”

    → Two steps: TTS Text Prep Studio first — free, runs in your browser, cleans the text and counts the characters — then AI Text to Speech to generate the audio, which shows you the price before you confirm. The full article-to-audio workflow.

    “I want to turn speech into text.”

    AI Speech to Text. Timed transcripts from audio or video, as plain text, SRT or VTT — for transcript pages, subtitles, or finding quotes in long recordings.

    “I want an audiobook from a whole document.”

    AI Audiobook Creator. Built for book-length PDFs and EPUBs rather than short passages, and it can give different characters different voices. What to check before and after.

    “My video file is too big, or the wrong format, or too long.”

    → Three different problems with three different answers: Video Compressor, Video Format Converter, Video Cutter. How to tell which one you need.

    “I need subtitles.”

    AI Video Subtitle Generator to create them, AI Subtitle Translator to translate them, and SRT to VTT or VTT to SRT to deliver the right format. The three-step subtitle workflow.

    “I have a PDF problem.”

    → Merging, splitting, compressing, converting, rotating or unlocking — the PDF toolkit map covers all eleven tools.

    A note on order

    If your situation involves several of these, the order that works: fix noise first, normalize loudness second, cut clips or convert formats last. Cleaning and loudness affect the whole file, so everything downstream inherits their result. The reasoning, with measurements.

  • MP3, WAV, M4A, OGG or FLAC? A Practical Guide to Audio Formats

    MP3, WAV, M4A, OGG or FLAC? A Practical Guide to Audio Formats

    Audio formats exist because there are two competing goals: perfect quality and small files. Every format is a different answer to that trade-off. Here’s the short version of when each one makes sense.

    The five formats, briefly

    MP3 — the universal one. Lossy compression: it discards audio detail the ear is least likely to miss, in exchange for small files. Every device and platform plays it. For spoken word at a reasonable bitrate, the quality loss is not something listeners will notice. Default choice for publishing podcasts and sharing recordings.

    WAV — the working one. Uncompressed audio, large files, zero quality loss. The right format while you’re still working on a recording — every edit and export stays pristine. The wrong format for distribution, purely because of size.

    M4A (AAC) — the modern lossy one. Same idea as MP3, more efficient compression: comparable quality in somewhat smaller files. Excellent support on phones and modern platforms. The typical output of many recording apps.

    OGG — the open one. Lossy like MP3 and AAC, an open standard, used by some platforms and games. You’ll more often need to convert from it than to it.

    FLAC — the archival one. Compressed but lossless: files shrink meaningfully (unlike WAV) yet decode back to the exact original. The right format for archiving masters you may need again. Playback support is narrower than MP3 or M4A.

    The two rules that prevent most mistakes

    Convert down, not up. Converting a lossy file to WAV or FLAC does not restore lost quality — it just makes a big file with the same lossy audio inside. Keep originals in the highest quality you have; export downward from there.

    Avoid repeated lossy-to-lossy hops. Each lossy re-encode discards a little more. One conversion is fine; five generations of MP3 to M4A to MP3 audibly degrade.

    When you actually need a converter

    A platform demands a format you don’t have; a collaborator sends OGG or FLAC your editor won’t open; you want a small MP3 of a huge WAV; you’re archiving to FLAC.

    The Audio Format Converter converts between all five of these formats. For MP3 you choose the bitrate directly (128, 192, 256 or 320 kbps); for M4A and OGG you pick a quality level on a 0–9 scale, with the equivalent bitrate shown next to it so the choice means something. WAV and FLAC are lossless and have no quality setting to make. You can also convert channels to mono or stereo, and keep or strip the file’s existing metadata tags.

    Uploads go up to 200 MB and three hours, with a limit of 10 conversions per day, and it’s free. Note it’s for audio files — video files are a different job, covered in our guide to compressing, converting and cutting video.

  • Auto-Clipping a Podcast: What Energy Detection Can and Can’t Find

    Auto-Clipping a Podcast: What Energy Detection Can and Can’t Find

    The Social Media Clip Cutter takes a long recording and returns short clips ready for social platforms. It’s worth understanding how it decides where to cut — both to use it well and to know what it can’t do.

    How it works

    The tool measures the loudness of your recording second by second, then looks for high-energy moments: sections that are markedly more intense than their surroundings. Around each detected peak it cuts a clip of exactly 15 seconds — five seconds before the peak and ten after — with a short fade in and out, keeps at least ten seconds between clips so they don’t overlap or repeat the same moment, ranks candidates by energy, and keeps the strongest ones. You choose whether you want 5, 10 or 15 clips.

    You get the clips individually as MP3s, plus a single mashup of all of them, packaged with a metadata file listing the timestamp each clip came from — so you can always find the moment in the original.

    A sensitivity setting (Soft, Normal, Aggressive) controls how wide the candidate pool is. One honest note: on a recording with plenty of lively material, all three settings often return the same clips, because the highest-energy moments are the highest-energy moments regardless of where you set the bar. The setting earns its keep on quiet or unevenly recorded material.

    Why a fixed loudness threshold wouldn’t work

    This is the part most “auto clip” descriptions skip. Speech has a much narrower dynamic range than people assume — in one real episode we measured, the median level was −18.5 dB and the loudest peak −14.9 dB, a total spread of 3.6 dB. A rule like “anything more than 4 dB above the median” sits above even the loudest moment in that file and returns nothing at all. That’s why detection works on relative ranking within your specific file rather than a fixed decibel value.

    The honest limitation: loud is not the same as good

    Energy detection finds intensity — laughter, raised voices, animated discussion. Often that correlates with your best moments. But your sharpest insight may have been delivered calmly, and a loud moment can be two people talking over each other. A volume-based algorithm cannot tell the difference, and we’d rather say that plainly than pretend otherwise.

    The realistic workflow: candidates, not final picks

    Treat the output as a shortlist. The tool has done the tedious part — scanning the full recording and pulling out moments worth reviewing, with timestamps. Your part is a few minutes of listening: keep the clips that stand on their own, discard the ones that need context, and if you know a great calm moment the ranking missed, its neighbours’ timestamps make it easy to locate.

    A clip works on social platforms when someone with zero context understands it. That judgment stays with you; the scanning no longer has to.

    One practical tip

    Run your episode through noise removal and loudness normalization before cutting clips — the clips inherit whatever audio quality the source has, and a clip is often someone’s very first contact with your show. Our post on the full post-production workflow covers the order in detail.

  • Podcast Transcripts: Why Text Still Does the Heavy Lifting

    Podcast Transcripts: Why Text Still Does the Heavy Lifting

    Search engines don’t listen to audio. However good an episode is, its content is invisible to search until it exists as text somewhere. That’s the entire case for transcripts, and it’s a strong one — no growth-hack statistics required.

    What a transcript actually does for an episode

    It makes the episode findable. A transcript page can be indexed, so the things you actually said — names, tools, techniques, opinions — become searchable text. Without it, the only indexable content is your episode title and description.

    It makes the episode quotable. Listeners who want to share a point can copy the words. Writers who want to reference your episode can cite it. Neither happens with audio alone.

    It makes the episode accessible. For deaf and hard-of-hearing audiences, and for anyone in a situation where they can read but not listen, the transcript is the episode.

    Producing one without typing it yourself

    The AI Speech to Text tool takes an audio or video file and returns a timed transcript. It accepts MP3, WAV, M4A, AAC, OGG and FLAC audio as well as MP4, MOV, WebM and MKV video, and gives you the result three ways: plain text for reading, and SRT or VTT if you want the timings attached. It labels speakers as it goes, and can optionally translate the transcript into another language.

    Automatic transcription gets you most of the way; plan a quick read-through pass for the things automatic systems reliably miss: proper names, brand names, technical terms. Fix those, leave the rest.

    Publishing it

    Put the transcript on the same page as the episode’s player and show notes, as normal page text — not inside an image, not as a downloadable file only. One episode, one URL, with the audio and its text together. If the transcript is long, a heading per segment helps readers skim and gives the page useful structure.

    Timestamps: keep them or strip them?

    For a published transcript page, light timestamps at section boundaries are useful; a timestamp on every line is noise for readers. Keep the fully timed version for yourself — the SRT or VTT export is exactly what you need later to find quotable moments or cut clips from the right places.

  • Turning Blog Posts into Audio: Preparation, Cost and Workflow

    Turning Blog Posts into Audio: Preparation, Cost and Workflow

    A blog post and an audio version of it are not the same text. Written articles are full of things that read fine and sound wrong: URLs, digits, parenthetical asides, bullet lists. If you feed an article straight into a text-to-speech engine, you pay to generate audio you’ll end up regenerating. Here’s the workflow that avoids that.

    Step 1: Prepare the text (free, in your browser)

    The TTS Text Prep Studio cleans a text for speech before any audio is generated. It converts numbers into words (English or Turkish), replaces URLs and e-mail addresses with markers, collapses repeated whitespace, turns line breaks into spaces, straightens curly quotes, normalises dashes, and optionally strips special characters while keeping ordinary punctuation. It also counts the characters, so you know the size of the job before you start.

    It runs entirely in your browser — the text is never sent to a server at this stage.

    One deliberate omission worth knowing about: it shows you character and request counts, not a price in currency. Text-to-speech is billed per character and every plan converts that differently, so a made-up money figure here would be worse than no figure at all.

    Step 2: See the cost before you generate

    Generation is priced by text length. The AI Text to Speech tool shows the character count and the resulting price for your exact text before you confirm anything — no surprises after the fact. This is also why step 1 matters: cleaning the text first means you only pay for words that belong in the audio.

    Step 3: Pick the voice before committing to a long text

    Voices that sound similar in a one-line demo diverge over ten minutes of listening. Generate a short passage with a few different voices and choose by ear, using one paragraph of your article rather than the whole thing.

    Our AI Text Translator & Voice Compare tool is built for exactly this side-by-side comparison. Be aware of one limitation before you click: its interface is currently in Turkish only, though the voices it previews are not.

    Step 4: Generate, then spot-check

    Generate the full article, then listen to the beginning, one section from the middle, and the end. The most common issues are pronunciation of names and technical terms — if one matters to your content, consider rewording it or spelling it phonetically in the source text and regenerating just that section.

    What to expect honestly

    An audio version does not automatically create an audience — it gives the audience you already have another way to consume the same content, and it makes your articles accessible to people who can’t or don’t read on screen. Publish a few, see whether your readers use them, and decide from your own data whether to continue. That’s a cheap experiment precisely because the preparation step is free and the generation cost is visible up front.

    If your source is a whole book rather than an article, that’s a different job with different preparation — see turning a PDF or EPUB into an audiobook.

  • Podcast Loudness Explained: LUFS, True Peak and Processing Order

    Podcast Loudness Explained: LUFS, True Peak and Processing Order

    If you’ve ever exported an episode that sounded fine on your headphones but came out much quieter than other shows, loudness normalization is the missing step. Here’s the shortest useful explanation of the three numbers involved.

    LUFS: perceived loudness

    LUFS measures loudness the way human hearing perceives it — averaged over the whole program, weighted toward the frequencies we’re sensitive to. It’s a better yardstick than peak level, because two files can have identical peaks and wildly different perceived loudness.

    For spoken-word content, a target around −16 LUFS is the widely used convention — it’s the level podcast apps are generally mixed around, so an episode normalized there sits comfortably next to other shows in a listener’s queue. Broadcast television uses a different standard, EBU R128, which targets −23 LUFS; you’ll sometimes see the two conflated online. For podcasts, around −16 is the practical target.

    True peak: the safety ceiling

    Normalizing raises your audio toward the target, and the loudest instants need somewhere to go. True peak measures the actual maximum the waveform reaches — including the “inter-sample” peaks that appear when digital audio is converted back to an analog signal. Keeping true peak safely below 0 prevents clipping and distortion on playback devices and after lossy encoding to MP3 or AAC.

    How far below 0 depends on where the audio is going. Our Podcast Audio Studio pairs each loudness target with the ceiling normally used alongside it: the Podcast preset is −16 LUFS with a −1.5 dBTP ceiling, while the YouTube and Spotify presets are −14 LUFS with −1.0 dBTP.

    One honest limitation of any loudness target

    Loudness and peak ceiling can pull against each other. Making a file louder means applying gain, and gain raises the peaks by the same amount. On a very quiet or very peaky recording, a −14 LUFS target and a −1 dBTP ceiling cannot both be satisfied — something has to give. Our tool keeps the ceiling, gets as close to the loudness target as the ceiling allows, and tells you when it had to stop short, rather than quietly crushing the audio to hit a number.

    Order matters: clean first, normalize second

    Loudness normalization raises quiet content. If your recording still has background noise, normalization raises the noise right along with the voice.

    We have a measurement for this. On our test file, performing the loudness step inside the noise-removal pass instead of separately afterwards cut the signal-to-noise improvement from +11.1 dB to +2.2 dB, and the residual noise floor ended up at −37.4 dBFS rather than −52.0 dBFS. Run noise removal first, then normalize the cleaned file as a separate step.

    What the Podcast Audio Studio does with all this

    It applies two-pass loudness normalization to the target you pick with true-peak protection, trims silence from the start and end of the file, and adds a fade at each end. You upload a file, it returns a normalized MP3 or WAV — the numbers above are the whole story of what happens in between.

    Do I need to think about this every episode?

    No — that’s the point of automating it. What’s worth internalizing is just the order (clean, then normalize) and the reason your episodes should all pass through the same normalization step: consistency across episodes is what listeners experience as “professional”. The full three-step workflow is here.