No account · No credits · Free exports
Transcribe audio to text
Upload an MP3, WAV, M4A or FLAC and get the words back with timestamps — nothing to install, nothing to sign up for.
- File
- Transcribe
- Transcript
Set it yourself for short clips, heavy accents, or mixed languages.
English translates as it transcribes, in one pass. It's the only target.
Free, no account. Up to 20 min per file, 2 per hour.
The catch
What "free" usually means
Most free transcription tools are a trial wearing a different word. Here's the shape it usually takes, and what happens here instead.
| Comparison | Typically | Here |
|---|---|---|
| Account | Email required after the first file | Never |
| Free allowance | Credits or minutes that run out | No credits, no total cap |
| Downloading it | Export is the paid feature | TXT, SRT and VTT, free |
| Accuracy claim | A precise-sounding percentage | The model name, so you can check |
| The real limit | Discovered when you hit it | 20 minutes per file, 2 per hour |
TXT, SRT and VTT all export free, with no account at any point. There's no paid tier here, so there's nothing for a limit to push you toward. The two that exist are there because transcription runs on GPU time that costs real money per minute, and capping it per person keeps it working for everyone.
More on how that pattern works, and five questions worth asking any transcription tool before you upload — free transcription without signing up.
How it works
Three steps, no settings to learn
01
Upload
AAC, AIFF, FLAC, M4A, MP3, OGG and WAV, up to 80MB and 20 minutes.
02
Set the language
Or leave it on auto-detect, which handles clear single-language audio fine.
03
Read or export
Click any line to hear it. Copy the text, or download TXT, SRT or VTT.
Languages
Auto-detect, or tell it yourself
Detection reads the opening seconds of the audio, which is exactly why it fails in predictable ways. A thirty-second clip gives it little to work with. A recording that opens in English before switching languages gets labelled English. Two languages alternating throughout get whichever came first.
Setting the language yourself removes the guess entirely, and costs nothing — same model either way. It's worth doing for short clips, strong accents, and mixed-language audio. Whisper's larger models are also noticeably stronger on lower-resource languages than the smaller ones most free tools run, which matters if you're working in Nepali, Hindi, Bengali or Urdu.
Non-English audio can also come back as English. Choosing English output translates as it transcribes, in one pass — you don't transcribe first and translate after.
- Detect automatically
- Clear speech, one language, over a minute.
- Set it manually
- Short clips, heavy accents, two languages mixed.
- Translate to English
- Any source language. English is the only target.
Exports
TXT, SRT or VTT
Same words in all three. What differs is whether timing travels with them, and how it's written.
| Format | Use it for | Timing |
|---|---|---|
| TXT | Reading, searching, pasting into notes | None |
| SRT | Video editors, YouTube, most caption uploads | Numbered blocks, comma before ms |
| VTT | HTML5 video, via a <track> element | WEBVTT header, period before ms |
That comma-versus-period difference is the most common reason a caption file silently fails to load. If an editor accepts the file but shows nothing, check that first.
Honest limits
What this won't do
Worth knowing before you upload rather than after.
- No speaker labels
- An interview comes back as continuous text, not "Speaker 1 / Speaker 2". Two people talking over each other is also the hardest case for any transcription engine.
- No editor
- You get the transcript and the export. Corrections happen in whatever you paste it into.
- 20 minutes per file
- Longer recordings need splitting first with the Silence Splitter, which cuts at natural pauses.
- Noisy audio stays noisy
- Nothing cleans the recording before transcribing it. Run it through the Voice Cleaner first — that improves a transcript more than any setting here.
For the fuller version — what degrades accuracy, when to set the language, and how to handle long recordings — read the transcription accuracy guide.
Next
More free tools
Questions
Frequently asked questions
Limits and model last verified . Whisper large-v3 · 20 minutes per file · 2 per hour.