Mix Guide
How to Mix Vocals: A Step-by-Step Guide for Rap and Pop
Seimix Team · 12 min read
The first thing a listener locks onto in any song is the vocal. It doesn't matter how good the beat is: if the vocal doesn't cut through, if the highs are harsh, or if the voice disappears underneath the instrumental, the whole song sounds amateur. Here's the good news: mixing vocals isn't some magic talent. It's a process with a clear order and a clear logic. Follow the right steps in the right order and you can push even a vocal you recorded at home up to a professional level.
In this guide I'll walk you through mixing a vocal from start to finish: starting with tuning, then cleaning up the low end, cutting the mud in the mids, using compression to bring the vocal forward, taming harsh 's' sounds, adding brightness and air, creating depth with reverb and delay, and finally balancing every line with automation. I'll give you concrete frequency and ratio targets for both rap and pop. The goal isn't to memorize numbers, it's to understand what you're actually doing.
Before You Mix: A Clean Source and the Right Order
Fifty percent of a vocal mix is done before you open a single plugin. Even the best plugins can't rescue a bad recording, so start by fixing the source. Make sure the take isn't clipping, and keep your recording level peaking around -6 dB so you leave yourself headroom to process later. If you captured several takes, comp them together by stitching the best lines into one clean performance.
Next, do the rough cleanup: cut the silence between phrases, don't delete breaths entirely but soften them by pulling them down 4-6 dB, and clean up any mouth clicks and plosives. This editing stage feels tedious, but it makes the rest of the mix dramatically easier.
Order matters a lot. Follow this general chain: tuning, then editing and cleanup, then high-pass and subtractive EQ (cutting the mud), then compression, then de-esser, then additive EQ (presence and air), then light saturation, then reverb and delay, and finally automation. This order isn't random: you clean up and control the sound first, then add color and depth on top.
- Keep your recording peaks around -6 dB and leave headroom.
- Comp your takes: stitch the best lines from multiple passes.
- Don't delete breaths, tame them by pulling them down 4-6 dB so it still sounds natural.
- Clean up clips, plosives, and mouth clicks before you open any plugin.
Pitch and Tuning: Lock In the Notes
Put pitch correction at the very front of the chain, because everything that comes after (especially compression) amplifies the sound exactly as it is. If you don't fix a flat note first, you'll only push it further forward. You can tune two ways: real-time automatic correction, or note-by-note manual correction. The manual method sounds more natural but takes time; the automatic method is fast and adds character depending on the style.
On rap and trap vocals, a hard, fast-snapping autotune (a short retune speed) is usually what creates that signature 'gliding' effect. In melodic rap and modern pop it's a stylistic choice too. On classic pop and R&B, though, the goal should be to fix the off notes without ever hearing the effect. Slow the retune speed down a bit so the vibrato and natural slides between notes are preserved.
Don't forget to set the right key and scale. Pick the wrong key and the tuning engine will pull notes to the wrong pitches, leaving the vocal more out of tune than before. If you don't know the song's key, listen to the instrumental or grab it from a reference.
- Always tune before compression, at the front of the chain.
- Use a fast retune for rap/trap and a slow retune for natural pop.
- Set the correct key; the wrong key wrecks the tuning completely.
- Over-tuning sounds robotic, so keep it light if you don't want the effect.
High-Pass and Low-End Cleanup
The useful information in a human voice generally starts above 80-100 Hz. Everything below that is mostly room rumble, mic stand vibration, AC hum, and plosive energy. Add a high-pass (low-cut) filter to clean the vocal up: 80-100 Hz is a good starting point for male vocals, and 100-120 Hz for female vocals.
Set the filter by ear, not by eye. Sweep it up slowly, and when you hear the body of the vocal start to thin out and lose that 'chest' feel, back it off just a touch. The goal is to remove the pointless low end, not to kill the vocal's fullness. In rap, where you want the vocal sitting close and clear right up front, cutting a little higher (100-120 Hz) is common.
This step barely registers on small speakers and phones, but it makes a huge difference in the full mix: it clears space for the kick and bass, reduces overall muddiness, and stops the compressor from wrestling with pointless low-end energy.
- Low-cut from 80-100 Hz on male vocals, 100-120 Hz on female.
- Set the filter by ear; back off once the body starts to thin.
- In rap you can cut a little higher for extra clarity.
- Plosives drop off noticeably along with the high-pass.
Cut the Mud: 200-500 Hz and the Midrange
If a vocal sounds 'muddy,' 'boxy,' or 'nasal,' the culprit is usually the low mids. The 200-400 Hz range gives a vocal its body, but too much of it turns into mud. Sweep through this area with a narrow EQ band (medium Q), find the spot that bothers you most, and cut it 2-4 dB. Don't overdo the cut, or the vocal will end up thin and weak.
Around 500-800 Hz is where 'boxy' tone lives. If you recorded in a room with bad acoustics, a small cut here (1-3 dB) will open the sound up. The 1-4 kHz range is where harshness lives. Some vocals have an aggressive ringing around 2-3 kHz, and you can pull that down a few dB with a narrow band too.
The golden rule of subtractive EQ: cut first, boost later. Once you've cleaned out the frequencies that bother you, the vocal already sounds far clearer and more professional. Adding brightness comes in the next steps.
- Cut the mud 2-4 dB at 200-400 Hz, but don't kill the body.
- For a boxy sound, try a 1-3 dB cut at 500-800 Hz.
- If it's harsh, pull 2-3 kHz down a few dB with a narrow band.
- Cut first, boost later: brightness sits better on a clean foundation.
Compression: Bring the Vocal Forward
Compression is the key to keeping the vocal out in front of the beat. In a raw performance, there's a big level gap between the loudest and quietest lines, so some words vanish and others shout. The compressor evens out that gap and locks the vocal to a steady level that sits up front. For a starting point, try a 3:1 ratio, a medium attack (10-30 ms), and a release that breathes with the track (60-120 ms). Aim for the gain reduction to move around 3-6 dB on the loudest lines.
Attack time defines the character. Set the attack too fast and you'll crush the vocal's initial transient and make it dull; leave it a little more open and you keep the punch at the front of each word, so the vocal sounds more alive. Rap and trap vocals want tighter control: here, pushing to 4:1 or even 6:1 and shortening the attack a bit glues the vocal to your chest.
Instead of hammering the vocal with a single compressor, use serial compression (two compressors in a row, each set gently): let each one take 2-3 dB. The total control is the same, but the result breathes and stays natural. If you want the vocal even more forward and dense, use parallel (New York) compression: blend a heavily crushed copy quietly underneath the clean signal.
- Starting point: 3:1 ratio, 10-30 ms attack, 60-120 ms release.
- Keep gain reduction moving 3-6 dB on the loudest lines.
- For rap/trap, 4:1 to 6:1 gives tighter control.
- Instead of one heavy compressor, try serial or parallel compression.
De-Esser: Tame Harsh 'S' and 'Sh' Sounds
As you add compression and brightness, sibilant sounds like 's' and 'sh' start jumping out uncomfortably. They stab at the ear, especially on headphones and phones. That's exactly what a de-esser is for: it turns down only the frequency band where those harsh sounds live, and only at the moment they happen, leaving the rest of the vocal untouched.
Focus the de-esser in the 6-9 kHz range. On female and thinner-sounding vocals, sibilance usually sits higher (7-9 kHz); on deeper male vocals, it sits a little lower (5-7 kHz). To find the exact spot, use the de-esser's 'listen/solo' mode: monitor only the band being reduced and target the frequency where you hear the most 'ssss.'
Don't overdo the amount. The goal isn't to erase the 's,' just to take the edge off it. Too much de-essing makes the vocal lispy, like the singer is slurring. Usually 3-6 dB of reduction is plenty. Place the de-esser after compression and after your brightness EQ, because those are the moves that expose the sibilance in the first place.
- Focus the de-esser at 6-9 kHz; higher for thin voices, lower for deep ones.
- Use solo/listen mode to find the frequency with the most 'ssss.'
- 3-6 dB of reduction is enough; overdo it and the vocal gets lispy.
- Order: de-ess after compression and air EQ.
Presence and Air: Clarity and Brightness
Once you've cleaned up and controlled the vocal, it's time to brighten it. Two ranges matter most: presence and air. The 3-5 kHz range is the presence band; a wide 2-4 dB boost here pushes the vocal in front of the beat, sharpens the words, and improves intelligibility. In rap, where getting the lyrics across clearly is critical, this band is especially valuable.
The 10-16 kHz range is the 'air' band. A gentle boost with a wide shelf here (1-3 dB) gives the vocal that expensive, open, 'breathing' quality. In pop and R&B, it makes the vocal sparkle like glass. But watch out: adding air also increases sibilance, so you'll need to balance it with your de-esser.
In additive EQ, use wide bands (low Q). Narrow, aggressive boosts make the sound artificial and harsh. Always compare against a professional reference track: next to it, is your vocal dull, or is it too shrill? Make that call by comparison, not by ear alone.
- For clarity, boost the 3-5 kHz presence band 2-4 dB with a wide band.
- For air, add 1-3 dB with a shelf from 10-16 kHz.
- Adding air raises sibilance; balance it with the de-esser.
- In additive EQ, wide (low-Q) bands sound more natural.
Create Depth with Reverb and Delay
A dry vocal with no effects sits glued to the dead center of your head and feels flat. Reverb and delay place the vocal in a space and give it depth and dimension. But the most common amateur mistake is overdoing it and drowning the vocal in fog. The rule: you should feel the effect, but not be able to pick it out on its own.
For reverb, start with short and medium spaces (a plate or a room). In rap, a short, tight reverb and/or a slap delay (a single echo, 80-160 ms) keeps the vocal dry but dimensional. In pop, longer, wider reverbs are common. Two tricks to keep reverb from getting muddy: put a high-pass on the reverb return (cut below 300 Hz) so mud doesn't build up, and push the reverb back with a pre-delay so the lead vocal stays clear.
Delay adds depth and rhythm. A tempo-synced delay (say 1/4 or dotted 1/8) leaves nice echoes at the ends of phrases. EQ the delay return too, thinning out the lows and the extreme highs, to separate it from the lead vocal and keep it cleaner. Want the vocal more 'up front'? Pull the effects back. Want it more 'atmospheric'? Push them up. It's entirely a stylistic balance.
- In rap, a short reverb + slap delay (80-160 ms) keeps it dry.
- High-pass the reverb return at 300 Hz so mud doesn't build up.
- Tempo-sync the delay (1/4 or dotted 1/8) and EQ it thinner.
- You should feel the effect but not hear it as a separate source.
Automation, Phone Checks, and Final Touches
Compression balances the level roughly, but the final polish comes from automation. Listen through the vocal line by line: if a word is still getting lost, ride that spot up 1-2 dB by hand; if a line shouts, pull it down. Push the vocal forward a little in the choruses and pull it back in the verses. You can automate effects too: opening up the delay on the last word of a line to leave a 'throw' echo is a popular and effective touch.
Always check your mix on different systems. Most people will hear the song on a phone, a laptop speaker, and cheap earbuds. Small speakers barely reproduce the low end at all, so the body and clarity of the vocal are what carry it there. Play your mix through a phone speaker: is the vocal disappearing, are the highs stabbing? Then collapse it to mono and listen, because plenty of systems play in mono and some of your stereo effects can vanish there.
Finally, compare against a reference and rest your ears. Take a professional song you love, match it to the same loudness, and compare your vocal against it: is yours duller, more shrill, or more buried? Your ears fatigue in about 20 minutes, so take a break and listen again the next day with fresh ears. All of these steps (tuning, EQ, compression, de-essing, presence, reverb, automation) take hours to do by hand and require real experience. If you want to speed this up, or hit a professional reference level in one click, you can upload your vocal and let Seimix mix it automatically. It applies these steps with measured settings for you, and you can fine-tune from there.
- Ride the level by hand, 1-2 dB, line by line with automation.
- Push the vocal forward in the chorus, pull it back slightly in the verse.
- Always check on a phone speaker and in mono.
- Compare against a reference at matched loudness, and rest your ears.
Don’t want to do it by hand?
Upload your track to Seimix, pick a preset or just say what you want — get a mixed file back in minutes. 7-day free trial.
Frequently asked questions
What order should I mix vocals in?
The general chain is: pitch/tuning first, then editing and cleanup, then high-pass and a subtractive EQ to cut the mud, then compression, de-esser, an EQ to add presence/air, light saturation, reverb and delay, and finally automation. This order lets you clean up and control the sound first, then layer color and depth on top. The most critical rule: tuning and subtractive EQ must come before compression.
What's the difference between mixing rap and pop vocals?
In rap, the vocal should be dry, clear, and way out in front of the beat, so the high-pass is set a little higher (100-120 Hz), the compression is more aggressive (4:1-6:1), and short reverbs with slap delay are common. The 3-5 kHz presence band is invaluable for lyric clarity. In pop, longer reverbs, more 'air' (10-16 kHz), doubling, and atmosphere take center stage. The core steps are the same; the doses and character differ.
Do I have to use autotune?
No, it's not required. Autotune is used two ways: as a stylistic effect (that gliding sound in trap and modern pop) or simply to fix off notes. If the performance is already clean, you might not use it at all, or keep it very light. If you don't want the effect, slow the retune speed down so vibrato and natural slides are preserved. Be careful not to pick the wrong key, or the tuning will make things worse.
Why does my vocal get buried under the beat?
The three most common causes: not enough compression (the level swings too much and words get lost), the 200-500 Hz mud not being cut, and a weak 3-5 kHz presence band. The fix: stabilize the level with roughly 3:1 compression, cut the low-mid mud a few dB, and boost the presence 2-4 dB with a wide band. Then ride the lost words up by hand, line by line, with automation.
What exactly does a de-esser do and where does it go?
A de-esser turns down sibilant sounds like 's' and 'sh' only at the moment they occur, leaving the rest of the vocal untouched. That way it removes the ear-stabbing harshness without making the vocal lispy. It focuses on the 6-9 kHz range, and usually 3-6 dB of reduction is enough. In the chain it goes after compression and after your brightness/air EQ, because those are the moves that expose the sibilance in the first place.
Reverb is making my vocal muddy, what should I do?
It's probably too long, too loud, and unfiltered. Try three things: put a high-pass at 300 Hz on the reverb return so mud doesn't build up in the lows, shorten the reverb time (plate/room is better for rap), and push the reverb back a bit with pre-delay so the lead vocal stays clear. If you can pick the effect out as a separate source, you've got too much of it, so pull the amount back.
Should I mix the vocal myself or use an automatic tool?
Both make sense. If you want to learn and have full control, working through the steps in this guide by hand is the best path, and your ears will develop. But the process takes time and experience. If you want a fast, consistent, reference-level result, you can upload your vocal to Seimix and let it mix automatically, then fine-tune from there. Most producers combine the two: a fast automatic base with a personal touch on top.
More: All guides