Promptrift

AI tools

Free AI Transcription Tools With Speaker Labels, Tested on Real Interviews

We ran real two- and three-person interviews through the free ai transcription tool with speaker labels options that actually exist, and logged where each one breaks.

Mockup of an interview transcript with color-coded Speaker 1 and Speaker 2 labels next to a waveform.

For real interviews with two or more people talking, three free setups actually produce usable speaker labels: Otter.ai’s free plan, OpenAI’s Whisper paired with pyannote.audio for diarization, and Fireflies.ai’s free plan. None of them get speaker labels perfectly right out of the box — crosstalk and fast interruptions trip up all three — but each one is good enough to hand you a transcript where “Speaker 1” and “Speaker 2” are correctly split most of the way through.

If you want something that works in a browser with zero setup, Otter.ai is the more reliable pick for interviews under 30 minutes. If you’d rather have no minute cap and don’t mind a bit of command-line work, Whisper plus pyannote is the only option in this list that won’t eventually ask for a card number.

How Speaker Labels Actually Work

“Speaker labels” don’t come from the tool recognizing who’s talking. They come from a diarization model listening to the audio, splitting it into short segments, and clustering those segments by acoustic voice signature — pitch, timbre, speaking rhythm. When two segments cluster together, the tool calls them the same speaker. It has no idea who that speaker actually is; “Speaker 1” is just cluster number one, and you rename it to a real name yourself.

That’s why diarization quality depends heavily on things that have nothing to do with the AI model itself:

  • How distinct the voices are. Two people with similar pitch and cadence get confused more often than a mismatched pair.
  • Mic setup. A single mic picking up two people at different distances produces volume differences the model can mistake for a speaker change.
  • Overlap. When people talk over each other, most diarization models pick one voice and drop the other, or invent a third speaker that doesn’t exist.
  • Speaker count. Accuracy drops noticeably once you go past three people in the same room.

Every tool below runs on some version of this mechanism. The differences are in how much of that error correction the tool does for you, and how big the free allowance is before it asks you to pay.

What “Free” Usually Means in Practice

Free transcription tools tend to limit you in one of three ways: a monthly minute cap, a storage cap on how long recordings stay accessible, or a feature wall where diarized export (SRT/VTT with speaker tags, not just plain text) sits behind the paid tier. A few tools are genuinely free with no catch, but that usually means self-hosted and technical rather than a polished web app. Worth keeping in mind before you commit an hour-long interview to any one platform’s free plan.

The Tools, Tested on Real Interviews

OpenAI Whisper + pyannote.audio (self-hosted)

Whisper itself is a free, open-weight transcription model — but it doesn’t produce speaker labels on its own. You pair it with pyannote.audio, a separate open-source diarization model, to get the “who said what” layer. Running both requires a Python environment, a free Hugging Face account, and accepting pyannote’s model license before it will download.

There’s no minute cap and no account wall once it’s set up, which makes it the only truly unlimited option here. In testing, transcription accuracy on accented English was the best of anything on this list. Diarization held up well for two clear voices recorded on separate tracks, but got shakier with three speakers of similar pitch, occasionally merging two people into one label mid-conversation. On CPU it’s slow — a 40-minute interview took well over the length of the recording itself to process; a GPU cuts that down dramatically.

Otter.ai (free plan)

Otter runs in the browser or app and transcribes close to real time. Its free tier has been capped at 300 minutes of transcription per month with a 30-minute limit per single recording for a long stretch now, though it’s worth checking their current pricing page since limits like this shift without much notice.

Speaker labels appear automatically and can be renamed to real names, and re-editing a mislabeled segment is a couple of clicks. On a clean two-person interview recorded with a decent mic, Otter’s diarization was the most consistent of anything tested. It started mixing up speakers during fast back-and-forth exchanges — the kind where people finish each other’s sentences — more than the other tools.

Fireflies.ai (free plan)

Fireflies works two ways: as a bot that joins a live Zoom or Google Meet call, or as a file upload for recordings you already have. Its free tier is storage-capped rather than minute-capped, meaning older transcripts eventually age out rather than you hitting a hard monthly wall.

Speaker detection was noticeably better in live-call mode, where Fireflies can use separate audio channels per participant, than in upload mode with a single mixed-down file. Uploaded interviews with fast speakers sometimes had two people merged under one label for stretches of the conversation.

tl;dv (free plan)

tl;dv is built around recording meetings directly through the tool rather than accepting arbitrary audio uploads, and its free tier gives unlimited transcription for calls captured that way. It’s a good fit if your “interview” is happening over Zoom or Meet already, less useful if you’re transcribing a recorded phone call or an in-person conversation — worth double-checking their current upload support before assuming it fits your workflow.

What doesn’t work: YouTube auto-captions and generic “free unlimited” transcribers

YouTube’s auto-captions don’t do real diarization — they break audio into lines based on pauses, not speaker changes, so a two-person interview reads as one continuous voice. A lot of the “100% free, no signup” transcription sites floating around are Whisper under the hood, with diarization quietly locked behind a paid upgrade once you try to export.

Where Free Tools Break Down on Real Interviews

Across all of them, the same failure patterns showed up:

  • Interruptions and overlapping speech get flattened into one speaker or split into a false third one.
  • Accuracy drops once you’re past three speakers with similar vocal range.
  • A single mic capturing people at different distances confuses volume-based speaker detection.
  • Mid-conversation language or accent switching (common in bilingual interviews) increases mislabeling.

None of this is unique to the free tiers — paid diarization has the same weak points, just with better error-recovery tooling around it.

Picking Between Them

For a one-off interview under 30 minutes, Otter’s zero-setup browser flow is the fastest path to a usable transcript. For recurring interviews where you don’t want to think about a monthly cap, Whisper plus pyannote is the only option that stays free indefinitely, at the cost of a technical setup. For interviews that happen as live video calls, Fireflies or tl;dv fold transcription into the recording step itself.

After the Transcript: What to Do With It

A correctly labeled transcript is still just raw text. If the interview runs long, feeding the whole thing into a chat model for a summary can hit a context limit before you’re done — we cover how to work around that separately. If the goal is turning the interview into published content, the same AI-assisted approach used to convert a long video into a blog post applies just as well to a transcript with speaker labels already attached. And if you’re doing this regularly enough that copy-pasting transcripts into a summarizer gets old, it’s the same automation pattern behind sending AI meeting summaries straight to Slack — just swap the meeting bot for whichever transcription tool you land on here.