Skip to content

No account · No credits · Free exports

Transcribe audio to text

Upload an MP3, WAV, M4A or FLAC and get the words back with timestamps — nothing to install, nothing to sign up for.

Audio to text
Whisper large-v3
  1. File
  2. Transcribe
  3. Transcript

Set it yourself for short clips, heavy accents, or mixed languages.

Output

English translates as it transcribes, in one pass. It's the only target.

Free, no account. Up to 20 min per file, 2 per hour.

The catch

What "free" usually means

Most free transcription tools are a trial wearing a different word. Here's the shape it usually takes, and what happens here instead.

ComparisonTypicallyHere
AccountEmail required after the first fileNever
Free allowanceCredits or minutes that run outNo credits, no total cap
Downloading itExport is the paid featureTXT, SRT and VTT, free
Accuracy claimA precise-sounding percentageThe model name, so you can check
The real limitDiscovered when you hit it20 minutes per file, 2 per hour

TXT, SRT and VTT all export free, with no account at any point. There's no paid tier here, so there's nothing for a limit to push you toward. The two that exist are there because transcription runs on GPU time that costs real money per minute, and capping it per person keeps it working for everyone.

More on how that pattern works, and five questions worth asking any transcription tool before you upload — free transcription without signing up.

How it works

Three steps, no settings to learn

  1. 01

    Upload

    AAC, AIFF, FLAC, M4A, MP3, OGG and WAV, up to 80MB and 20 minutes.

  2. 02

    Set the language

    Or leave it on auto-detect, which handles clear single-language audio fine.

  3. 03

    Read or export

    Click any line to hear it. Copy the text, or download TXT, SRT or VTT.

Languages

Auto-detect, or tell it yourself

Detection reads the opening seconds of the audio, which is exactly why it fails in predictable ways. A thirty-second clip gives it little to work with. A recording that opens in English before switching languages gets labelled English. Two languages alternating throughout get whichever came first.

Setting the language yourself removes the guess entirely, and costs nothing — same model either way. It's worth doing for short clips, strong accents, and mixed-language audio. Whisper's larger models are also noticeably stronger on lower-resource languages than the smaller ones most free tools run, which matters if you're working in Nepali, Hindi, Bengali or Urdu.

Non-English audio can also come back as English. Choosing English output translates as it transcribes, in one pass — you don't transcribe first and translate after.

Detect automatically
Clear speech, one language, over a minute.
Set it manually
Short clips, heavy accents, two languages mixed.
Translate to English
Any source language. English is the only target.

Exports

TXT, SRT or VTT

Same words in all three. What differs is whether timing travels with them, and how it's written.

FormatUse it forTiming
TXTReading, searching, pasting into notesNone
SRTVideo editors, YouTube, most caption uploadsNumbered blocks, comma before ms
VTTHTML5 video, via a <track> elementWEBVTT header, period before ms

That comma-versus-period difference is the most common reason a caption file silently fails to load. If an editor accepts the file but shows nothing, check that first.

Honest limits

What this won't do

Worth knowing before you upload rather than after.

No speaker labels
An interview comes back as continuous text, not "Speaker 1 / Speaker 2". Two people talking over each other is also the hardest case for any transcription engine.
No editor
You get the transcript and the export. Corrections happen in whatever you paste it into.
20 minutes per file
Longer recordings need splitting first with the Silence Splitter, which cuts at natural pauses.
Noisy audio stays noisy
Nothing cleans the recording before transcribing it. Run it through the Voice Cleaner first — that improves a transcript more than any setting here.

For the fuller version — what degrades accuracy, when to set the language, and how to handle long recordings — read the transcription accuracy guide.

Questions

Frequently asked questions

Limits and model last verified . Whisper large-v3 · 20 minutes per file · 2 per hour.