Free AI Vocal Remover
Upload a song and remove vocals with AI to create an instrumental or acapella. No sign-up, nothing to install.
- File
- Separate
- Result
GPU-accelerated AI
Real source separation, not a basic center-channel filter.
No install
Nothing to download. Upload, process, download in your browser.
Free
No sign-up, no watermark. Up to 80MB per upload.
What you get
- Vocals
- Lead and backing vocals, isolated from the instrumentation around them — usable as an acapella on its own.
- Instrumental
- The full mix with vocals removed, ready as a karaoke backing track or a base to build a remix around.
- Preview and download
- Both tracks play directly in the browser once separation finishes, and each downloads independently — grab one, the other, or both.
What the output actually is
- Format
- WAV (RIFF), 16-bit signed PCM
- Bitrate
- Lossless — about 1,411 kbps at 44.1 kHz stereo
- Sample rate
- 44,100 Hz — fixed
- Channels
- 2 (stereo) — fixed
- Inherited from your file
- None of it
That last row is worth reading twice if you work at 48 kHz. Demucs operates at 44.1 kHz in stereo internally, so the output rate and channel count are fixed no matter what you upload — a 48 kHz file comes back at 44.1 kHz, a mono file comes back as two channels, and a 24-bit file comes back at 16-bit. That is how the model pipeline works rather than a choice we made, and it is true of every tool built on Demucs, including the ones that don't mention it. Drop a stem into any editor and check.
You can also verify the models: standard runs htdemucs at 0.25 overlap, Studio Quality runs htdemucs_ft — four fine-tuned instances, ensembled — at 0.5. Higher overlap means more redundant computation across chunk boundaries, which is where the artifacts on longer tracks tend to show up. Between the ensemble and the overlap, Studio Quality is roughly five times the compute of standard, which is where the extra minute goes.
Who is this for?
Producers pulling an instrumental to sample or build on, DJs extracting an acapella for a mashup, singers practicing over a clean backing track, music teachers preparing karaoke material for students, and content creators needing an instrumental bed all use this tool for the same underlying job — splitting a mix into vocal and instrumental stems.
How to remove vocals from a song
- Upload an AAC, AIFF, FLAC, M4A, MP3, OGG, or WAV file.
- AI source separation splits the track into vocal and instrumental components, usually 20 seconds to 1 minute, depending on length and server load.
- Download the result directly in your browser, no install needed.
Need a track from YouTube first? Grab it with our YouTube to WAV converter and upload the result here, or skip the step entirely with the YouTube Vocal Remover.
Remove vocals from a song online
AudioForges lets you remove vocals from a song online without installing audio software. Upload an AAC, AIFF, FLAC, M4A, MP3, OGG, or WAV file and the AI separation model creates two tracks: an isolated vocal stem and an instrumental with the vocals removed.
You can use the instrumental for karaoke or practice, or use the isolated vocal as an acapella for remixing, sampling, and mashups. Each result can be previewed and downloaded separately after processing.
How AI vocal removal works
This tool uses real AI audio-source-separation processing to split a track into vocals and instrumental — not a simple center-channel filter, which only partially removes vocals and often damages the mix.
A center-channel filter works by cutting whatever's panned dead-center in the stereo mix — that catches lead vocals in many commercial mixes, but it also strips out anything else placed centrally (kick, bass, snare) and leaves behind any vocal element that isn't perfectly centered. AI source separation instead analyzes the audio's learned characteristics of what a voice sounds like versus an instrument, which is why it can isolate vocals regardless of where they sit in the stereo field, and why it produces a cleaner instrumental as a result. The model processes and outputs full stereo audio, and splits a track into exactly two stems — vocals and instrumental — rather than separating individual instruments like drums or bass on their own.
AudioForges processes the AI separation workload on GPU-accelerated infrastructure. A single track usually takes 20 seconds to 1 minute, and usage is rate-limited per IP address so it stays free and available for everyone. There is no app or plugin to install — you upload through the browser and the separation runs on the server.
The model is htdemucs — the published Hybrid Transformer Demucs, not something wrapped and renamed. Studio Quality runs htdemucs_ft, a "bag of four": four instances of the same architecture, each fine-tuned toward one stem, then ensembled. That is where the extra minute goes and why the stems come out cleaner.
Both stems come back as WAV. One thing worth saying because it's counter-intuitive: asking for two stems instead of four doesn't make this cheaper to run. The model separates all four sources internally either way and sums three of them into the instrumental — vocal removal is the same amount of work as a full stem split, just with different files kept.
Want the fuller breakdown of how this compares to older methods and where separation still struggles? Read How AI Vocal Removal Actually Works.
Standard vs. Studio Quality
| Comparison | Standard | Studio Quality |
|---|---|---|
| Processing time | 20 sec–1 min | 1–2 min |
| Model | htdemucs | htdemucs_ft |
| Separation quality | Good for most tracks | Noticeably cleaner on both stems |
| Usage limit | 6 per hour | 2 per houron the free tier |
| Cost | Free, always | A free allowance each month, shared with the other Studio Quality tools, then 1 credit per run |
| Best for | Quick previews, casual use | Sampling, remixing, anything going into a final mix |
Studio Quality uses a larger, ensembled model rather than a single pass, which is why it takes longer — the trade-off is worth it when the stems are headed into an actual production, not just a quick check.
AI vocal removal isn't perfect
Separation quality depends heavily on the source track. Choir or group vocals confuse the model since it has multiple overlapping vocal-like sources to untangle instead of one. Heavy distortion can share enough spectral character with a distorted or screamed vocal that the two get separated less cleanly. Live recordings with crowd noise or stage bleed give the model a messier signal to work from than a controlled studio mix. None of these make separation fail outright — they just tend to leave more audible traces behind than a clean studio recording would.
GPU acceleration changes the infrastructure the separation runs on, not the difficulty of the underlying problem — source quality and arrangement still determine the final result.
Instrumental vs. acapella
An instrumental is the track with vocals removed — everything except the voice. An acapella is the reverse: just the isolated vocal, with the instrumentation removed. Both come from the same underlying separation process, just keeping the opposite stem. Karaoke and remixing usually call for the instrumental; sampling a vocal hook or building a mashup usually calls for the acapella.
Common uses
- Karaoke & practice
- Get an instrumental to sing or play along with.
- Remixing & sampling
- Isolate an acapella or a clean instrumental bed to build on. Check the key first with our Key & BPM Finder if you're building something new around the sample.
- DJ mashups
- Pull an acapella from one track to lay over the instrumental of another.
- Cover reference
- Hear the instrumentation clearly without the original vocal in the way.
Next
More free tools
Stem Splitter
Split a track into vocals, drums, bass, and other stems.
Key & BPM Finder
Detect musical key, tempo, and Camelot notation.
YouTube Vocal Remover
Paste a YouTube link and get vocal and instrumental stems directly.
Audio to MIDI Converter
Transcribe a melody or vocal line into a downloadable MIDI file.
YouTube to WAV
Convert YouTube videos to lossless WAV audio for sampling and DJ sets.
Copyright & fair use
You are responsible for ensuring you have the right to process any track you upload — for personal practice, content you own, or material you have permission to use. AudioForges does not host or distribute the tracks processed through this tool.