How do I edit a voice recording so it sounds professional?
A fixed order — clean, cut, correct tone, control dynamics, then set loudness. Applied in that sequence, five modest adjustments produce a recording most people would call broadcast quality.
- Difficulty
- intermediate
- Time
- 25 min
- Read
- 4 min
Short answer
Remove noise and rumble first, then cut out the mistakes and long pauses, then use a high-pass filter and gentle equalisation to fix the tone, then compress moderately to even out the level, then normalise the whole thing to the loudness target for wherever it is going. Do it in that order — processing before cutting wastes effort and compressing before de-noising amplifies the noise.
Voice processing sounds like a dark art and is a short, fixed recipe. The reason results vary is almost entirely the order things are done in and how heavily each step is applied.
Step by step
- Work on a copy and keep the original.Every step here is destructive in some editors. Keep the untouched recording so you can start again when a heavy hand becomes obvious later.
- Remove steady noise first.Take a noise profile from a passage where nobody is speaking, and apply reduction gently — enough to take the edge off the hiss, not enough to hear the voice swim. Doing this before compression matters, because compression will otherwise amplify the noise between words.
- Apply a high-pass filter.Roll off everything below about 80Hz for a male voice or 100Hz for a female one. It removes rumble, traffic and plosive energy that contributes nothing but muddiness, and it makes everything afterwards cleaner.
- Cut the content.Remove false starts, stumbles, long pauses and anything repeated. Shorten pauses rather than removing them entirely — speech with no breath sounds unnatural and rushed. Cut on the silence between words, not through a word's tail.
- Correct the tone with equalisation.Small moves only. A gentle cut of a couple of decibels around 200 to 400Hz reduces boxiness. A small lift above 5kHz adds clarity if the recording is dull. Anything more than three or four decibels usually means the microphone or placement was wrong.
- De-ess before compressing.Compression raises quiet detail including sibilance, so a de-esser placed before it has less work to do and sounds more natural.
- Compress moderately.A ratio around three to one with the threshold set so the loudest phrases are reduced by three to six decibels evens out the difference between a quiet aside and an emphatic sentence. Heavy compression makes a voice sound flat and makes the room noise pump audibly.
- Set the loudness last.Normalise the finished file to a loudness target rather than a peak level: around minus sixteen LUFS for a mono spoken podcast, around minus fourteen for music streaming platforms. Check the specific platform, because this is the number that decides whether your recording sounds as loud as everyone else's.
- Check the peaks after normalising.Leave at least one decibel of peak headroom below full scale so that the encoding used for streaming does not clip. A limiter set at minus one is the standard safety net.
- Listen on two different systems.Headphones and a phone speaker, at least. A voice that sounds warm on headphones and disappears on a phone speaker usually needs less bass and a little more upper-mid presence.
Tips
- Almost every problem is better solved by re-recording one sentence than by processing the whole file. If a take is bad, record it again.
- Use the same processing chain for every episode or every video. Consistency between recordings matters more to listeners than any individual setting.
- Automatic voice enhancement tools built into editing software are genuinely good starting points now. Use one, then compare against the untreated original and decide honestly.
Common mistakes
- Compressing before removing noise — Compression raises the quiet parts, which is exactly where the noise lives. Denoise first and the compressor has clean material to work with.
- Making large equalisation moves — Cuts and boosts of six or more decibels are almost always compensating for a bad recording, and they introduce their own artificiality. Fix placement at the source instead.
- Normalising to peak instead of loudness — Peak normalisation makes the single loudest instant reach the ceiling, which says nothing about perceived volume. Two files normalised to the same peak can differ enormously in how loud they sound.
If it doesn't work
Voice sounds boxy or muddy
Cause: Excess energy in the low mids from room reflections or proximity — Fix: High-pass filter, and a gentle two to three decibel cut somewhere between 200 and 400Hz.
Voice sounds dull and distant
Cause: Missing upper-mid presence, or too far from the microphone — Fix: A small lift between 3 and 6kHz, and record closer next time.
Background noise pumps up and down
Cause: Heavy compression, or automatic gain, raising the noise floor between phrases — Fix: Reduce the compression ratio, denoise before compressing, and turn off any automatic gain at the source.
Sounds fine on headphones, inaudible on a phone
Cause: Too much low frequency and not enough presence — Fix: High-pass more aggressively and add a little upper-mid, then check again on the small speaker.
Podcast quieter than every other one
Cause: Not normalised to a loudness standard — Fix: Measure and normalise to the platform's LUFS target with a limiter at minus one decibel.
Questions people ask
Which software should I use?
The free open-source audio editors do everything described here, as do the audio tools built into video editors. The specific application matters far less than applying the steps in order and with restraint.
Should I use an automatic mastering service?
For spoken word they are reasonable and consistent, and they will not rescue a recording with echo or clipping. Use one to save time on a clean recording, not to fix a bad one.