How to Remove Background Noise from Audio
Background noise removal isn't magic — it's a tradeoff between how much noise gets pulled out and how much of the wanted audio gets damaged in the process. Understanding that tradeoff is what separates a clean result from one that sounds worse than the noise you started with.
How FFT-based noise reduction works
Most denoisers, including FFT-based ones, work by analyzing the audio's frequency content over time and identifying a noise profile — the frequencies where hiss, hum, or static consistently sit. It then reduces energy at those frequencies throughout the file, on the assumption that steady background noise occupies roughly the same frequency range the whole way through, while the wanted audio (voice, instruments) moves around more.
That assumption is usually solid for genuinely steady noise — tape hiss, fan hum, electrical buzz. It breaks down when the noise overlaps heavily with frequencies your wanted audio also uses, which is exactly what happens when you push reduction strength too far.
Why aggressive settings cause warbling
Warbling — a fluttery, underwater-sounding artifact — shows up when the denoiser starts removing energy from frequencies the wanted audio actually needs, not just the noise. At low-to-moderate strength, the algorithm can be conservative about where the line sits between noise and signal. Push it higher and it gets more aggressive about cutting anything resembling the noise profile, which starts carving into the wanted audio's own frequency content too — especially on music, where instruments and vocals legitimately occupy a wide frequency range that can overlap with the noise being targeted.
This is why a default, moderate strength setting works well for most recordings, and why raising it should be a response to audibly remaining noise, not a default move toward "more is better."
General denoising vs. a speech-specific preset
A general-purpose denoiser has to work across music, field recordings, and speech, which means it can't make assumptions specific to any one of them — you control the strength directly and accept the resulting tradeoff yourself. A speech-tuned preset can afford to be more targeted, since it only needs to preserve one type of signal: the human voice's frequency range. That lets it combine noise reduction with other speech-specific steps — like cutting low-frequency rumble and normalizing loudness — in a fixed chain that's already tuned for exactly that content.
In practice: if your source is specifically a podcast, phone recording, or interview, a speech-tuned chain will usually outperform manually dialing in a general denoiser. If your source is music, a field recording, or anything else where the frequency content is less predictable, a general-purpose denoiser with direct strength control is the better fit.
A practical approach
Start at the default strength and only increase it if noise is still clearly audible — don't jump straight to an aggressive setting expecting a cleaner result. If you hear warbling or a hollowed-out quality after processing, that's a sign the strength was pushed past what the source material could tolerate; back it off rather than trying to fix it with more processing.
Our Noise Remover gives you direct control over reduction strength for music and general audio. For speech-only recordings, the Voice Cleaner runs a fixed, speech-optimized chain that usually needs no manual tuning at all.