Mix Guide
How to Use Autotune: The Logic and Settings Behind Pitch Correction
Seimix Team · 11 min read
Autotune is one of those tools almost everyone making music at home has heard of, but few actually use well. Some treat it like a magic button that rescues any take; others turn up their noses and insist real singers don't touch it. Both are wrong. Dialed in correctly, autotune becomes an invisible helper the listener never notices; pushed on purpose, it becomes the signature effect that defines trap and pop. The difference comes down to setting a handful of parameters while actually understanding what each one does.
In this guide we'll walk through the logic behind pitch correction, why choosing the right key and scale matters more than anything else, how to move between a natural sound and a full effect with retune speed, what formants are, and the mistakes people make most often. We'll lean into rap and pop vocals in particular, since those are the two styles that generate the most questions.
Let's clear one thing up from the start: autotune is a pitch correction tool, not a mix. It puts the note in the right place, but making a vocal sound tight, clear and professional still takes a de-esser, EQ, compression and proper level balancing. At the end we'll cover how to bring those two worlds together.
What Is Autotune and How Does Pitch Correction Work?
Autotune constantly measures the frequency (the pitch) of the recorded vocal, figures out which "correct" note the sung note is closest to at that moment, and pulls the voice toward it. So if the note you sang landed a little flat or a little sharp, it slides the voice toward the target note the tool has decided on. That's the whole idea at its core; everything else is just controlling how hard and how fast that pull happens.
Two settings decide everything. The first is the scale, meaning which notes count as "correct." The second is retune speed, meaning how quickly the voice snaps to the target note. Get those two right and you're halfway home. Beyond that, most tools give you two working modes: real-time automatic mode (it corrects on the fly as the singer performs, ideal for recording and live use) and graphical/manual mode (you edit each note by hand, one at a time; it gives the cleanest, most natural result but takes time).
Keep this in mind: autotune works like a guessing engine. The cleaner and clearer the pitch it hears, the more accurately it corrects. Feed it a muddy, muffled take or one that drifts between two notes and it gets confused, producing weird sounds (artifacts).
- Automatic mode is fast and practical, perfect for live use and demos; when you want the smoothest result, correct note by note in graphical/manual mode.
- Pitch correction changes only the note, not the level (loudness); autotune won't turn up a vocal that was recorded too quiet.
- Put the tool at the very front of the vocal chain, before EQ and compression, so it reads a clean signal.
Choosing the Right Key and Scale: The Most Critical Step
This is the step people skip most often, yet it's the one that decides everything. Pick the wrong key and the tool will drag even the notes you sang correctly onto a wrong note, leaving the vocal out of tune and unpleasant to hear. Most of the recordings people complain "autotune ruined" actually have one problem: the wrong scale.
The easiest way to find the song's key is to start from the instrumental. Whoever made the beat usually knows the key (something like "F# minor" or "A minor"). If you don't know it, use the root note of the bassline or the chords as a reference; most modern trap and pop tracks sit in a minor key. Once you've settled the key, you also have to pick the scale: major is brighter, minor is darker and by far the most common in rap and trap.
Critical warning: never leave the scale on chromatic (all 12 notes open) unless you truly need to. In a chromatic scale the tool treats every note as "correct," so it won't fix out-of-tune notes at all; it's almost as if protection is switched off. The safest approach is to build a custom scale that only unlocks the notes the song actually uses. That way the tool never lands on a wrong note.
- If you're unsure, find the instrumental's key first; many free tuner/analysis tools will name the beat's key automatically.
- When in doubt, skip chromatic and build a custom scale with only that song's notes; that's the most accurate correction.
- If the chorus and verse move to different chords, defining a separate scale per section can save the vocal.
- If a wrong note keeps getting dragged to something wrong, the fix isn't more autotune, it's choosing the right key.
Retune Speed: The Dial Between Natural and the T-Pain Effect
Retune speed sets how quickly the voice snaps to the target note, and on its own it's the single setting that most shapes autotune's "character." Think of the value in milliseconds: a low value (0-10 ms) glues the voice to the note instantly, erasing the natural slides and vibrato in between; that famous robotic, glassy T-Pain / trap sound comes from right here. A high value (40-100 ms) gives the voice time to "walk" toward the note, preserving vibrato and natural transitions so the listener never notices the correction.
Here are practical starting points: for a transparent, natural pop vocal, 20-50 ms is a good range. If you only want to clean up minor imperfections, 40-80 ms is safer. For a defined but restrained modern feel, 10-20 ms. For a full-on, audibly robotic effect, 0-10 ms, usually paired with "full quantize / 100% correction."
Remember this: retune speed doesn't mean anything in isolation; it only makes sense alongside the song's tempo and delivery speed. In a fast rap flow with frequent note changes, a low value stays cleaner; in a ballad with long, held notes, a high value sounds far more natural and musical.
- For a natural, unnoticeable result, start at 20-50 ms and fine-tune by ear.
- For a defined trap/T-Pain effect, use 0-10 ms and open the correction amount all the way.
- On a singer with beautiful vibrato, too low a speed kills the vibrato; keep it high so it can breathe.
- One value won't fit every section; if needed, set the verse and chorus separately.
How Much Autotune for Which Genre?
The amount and character of autotune changes a lot by genre; "the same setting for everything" is the biggest rookie mistake. In trap and modern rap, obvious autotune is now an aesthetic choice, not a flaw. Low retune speed (0-15 ms) and full correction are common here, and it's pushed even harder on adlibs. The goal isn't to correct the voice secretly, it's to deliberately create that glassy, automated texture.
In pop, the goal is usually transparency: there's correction, but you can't hear it. A speed of 20-40 ms, moderate correction and a good source recording get the job done. In R&B and soul, naturalness and emotion come first; to preserve the vibrato and slides, high speed (50-100 ms) and, where possible, working by hand in graphical mode to touch only the problem notes is best.
In acoustic, folk and traditional styles, keep autotune to a minimum. Here, fixing a few wrong notes by hand in graphical mode sounds far cleaner than running automatic correction across the whole track. Heavy automatic correction kills the soul of the performance in these genres.
- Trap/rap: 0-15 ms, obvious and full correction; the effect is a choice, no need to hide it.
- Pop: 20-40 ms, transparent; nobody should notice the autotune.
- R&B/soul: 50-100 ms or by hand in graphical mode; preserve vibrato and emotion.
- Acoustic/folk: barely touch it, just fix the obvious wrong notes one by one.
Natural or Effect? Two Different Philosophies
You can use autotune with two completely different intentions, and you should decide which one you want from the start. The first is the transparent (natural) approach: the goal is to correct the voice while the listener never realizes it. The second is the effect approach: the correction is meant to be heard, that robotic texture becoming part of the song. Both are legitimate; the settings are just different.
For a transparent result: high retune speed (30-80 ms), a light "humanize" setting if needed, touching only the imperfect notes, and a good source recording. The golden rule here is to leave it alone unless it needs fixing. The less correction you apply, the more alive the voice stays. The best transparent autotune is the one where you can't tell whether it's even on.
For the effect: low retune speed (0-10 ms), full correction, a correct and tight scale lock, and usually a bit of reverb/delay on the vocal to bring that character out. What you're after here isn't "a little exaggeration" but a controlled, deliberate one; even the effect has a limit, though, and we'll get to that in the mistakes section below.
- Decide up front: do you want to hide it or show it off? The settings change accordingly.
- If you want it transparent, apply the least-intervention rule; don't touch every note.
- If you want the effect, keep the scale tight so note jumps give that characteristic "leap" sound.
- Don't mix the two approaches in the same song; unless it's a deliberate design (natural verse, effect chorus), it just sounds inconsistent.
What Is a Formant and Why Does It Matter?
A formant is the set of resonances that gives a voice its identity; it carries the timbre, the "whose voice is this" quality. When you push the pitch up or down by a large amount, the correction tool can unintentionally shift the formants too. The result is a voice that turns cartoonishly high (the chipmunk / Mickey Mouse effect) or strangely thick.
That's why most tools have a formant correction option. If you're making big pitch moves, especially shifting the voice by more than a couple of semitones, keep it on; that way the natural character of the voice survives even as the note changes. On small corrections the difference is minor, but leaving it on is generally the safer bet.
Formant is also a creative tool. Shifting it down on purpose gives the voice a thicker, more mature character; shifting it up gives a thinner, younger or more artificial one. On adlibs and backing vocals, playing with formants adds nice texture.
- If you're shifting the voice by more than a couple of semitones, always turn formant correction on.
- If you hear a chipmunk/weird high-pitched quality, the formant setting is the first place to look.
- Use formant shift creatively: nudge it on adlibs and backing vocals to change character.
- Don't over-shift formants on the lead vocal; the voice easily turns artificial and synthetic.
The Downsides of Overdoing It and the Most Common Mistakes
Autotune's biggest trap is the "more is better" fallacy. Correct every note fully and erase every slide, and you get a technically in-tune but lifeless, plastic, emotionless vocal. Part of the charm of the human voice hides in exactly those small, controlled imperfections. If you want a transparent result, keeping correction light is almost always better.
The second big mistake is trying to use autotune to rescue a bad take. On a recording that was badly out of tune to begin with, drifting between two notes, the tool keeps hesitating and produces strange glitches and warbles. The rule is clear: garbage in, garbage out. Autotune polishes a good performance; it doesn't rescue a bad one. The fix is to go back and re-record.
Other common mistakes: an unnecessarily wide/chromatic scale (it lands on wrong notes), using an extremely low retune speed for every genre (unintended robotic sound), and letting autotune work on breaths, consonants and sounds like "s/f" (the pitch is undefined in those spots and the tool goes haywire). Leaving breaths and consonants uncorrected, or turning correction off in those regions, sounds far cleaner.
- Don't correct every note fully; a few natural slides keep the vocal alive and human.
- Don't force an out-of-tune take with autotune; re-record it and the result is far better.
- Correcting breaths and consonants usually creates glitches; leave those regions uncorrected.
- When in doubt, intervene less; undoing overcooked autotune is harder than adding naturalness back later.
Before Autotune: Getting the Source Recording Right
The best autotune is the one you barely need. The better the recording, the less correction it takes and the more transparent it stays. So before you reach for the knobs, invest in the source. Stay at the right distance from the mic (usually 15-20 cm with a pop filter), hear the instrumental clearly in your headphones so you can lock the pitch, and warm up your voice before you sing.
If your pitch is shaky, the answer isn't stronger autotune, it's singing the same part a few times and picking the best take. Recording a tough chorus line by line (punch-in) reduces the flaws from the start. Give the tool a clean, steady pitch and it'll correct accurately and naturally.
Watch the recording environment too: on a vocal cut in a reverberant room, the tool can't read the pitch clearly. If you can, record in a corner with soft surfaces and little reflection. That alone noticeably improves the autotune result.
- Warm up your voice before recording and hear the instrumental clearly in your headphones; that's how you nail the pitch.
- Don't force tough sections in a single pass; do a few takes and pick the best, or record line by line.
- Recording in a reverberant room fools autotune; choose a soft, low-reflection space.
- Keep a steady distance from the mic; constantly moving in and out throws off both the level and the pitch reading.
After Autotune: Sitting the Vocal in the Mix
You've fixed the pitch, but the job isn't done. Autotune only puts the note in place; making the vocal clear, tight and professional is a separate stage. Think in order: first a de-esser around 6-9 kHz to soften harsh "s/sh" sibilance, then an EQ that cuts a few dB of mud at 200-400 Hz and opens a bit of brightness at 8-12 kHz, then compression at around a 3:1 ratio with a medium attack to keep the level consistent. This chain sits the vocal clearly in front of the beat.
Never forget the small-speaker reality. Most people will hear your song through a phone speaker, a laptop or cheap earbuds. Always check your mix on a phone speaker too; a vocal that sounds great on studio monitors but disappears on a phone means you need to bring it further forward. Finally, keep the loudness target of streaming platforms in mind; most normalize to around -14 LUFS, so there's no point in mastering excessively loud.
As you can see, the autotune part is a small slice of the whole. Key choice, retune settings, de-esser, EQ, compression, level and loudness balance... doing every one of them by hand takes time and experience. That's exactly where Seimix comes in: you upload your vocal, and from pitch correction to mix and mastering balance it applies these steps for you, measured and consistent. If you want a release-ready, balanced result without the manual work, try what you've learned once and let automation handle the rest.
- Order matters: de-esser (6-9 kHz) → EQ (cut the 200-400 Hz mud) → compression (3:1, medium attack).
- Always check your mix on a phone speaker; that's where the real listener is.
- Don't master excessively loud; platforms already normalize to ~-14 LUFS, so keep it dynamic.
- If you'd rather not build this whole chain by hand, upload your vocal to Seimix and let it handle pitch, mix and mastering balance automatically.
Don’t want to do it by hand?
Upload your track to Seimix, pick a preset or just say what you want — get a mixed file back in minutes. 7-day free trial.
Frequently asked questions
Will autotune correct a vocal automatically with no effort?
Partly. In automatic mode, once you pick the right key and scale, the tool corrects the notes instantly. But a genuinely good result still needs the right retune speed, formant setting, and a de-esser/EQ/compression afterward. If you'd rather not build that whole chain by hand, Seimix lets you upload your vocal and applies these steps automatically, from pitch correction to mix balance.
What retune speed should I use for a natural sound?
For a transparent, unnoticeable result, 20-50 ms is a good starting point. If you're only cleaning up minor imperfections, 40-80 ms is safer. If you want a defined trap/T-Pain effect, do the opposite: use 0-10 ms and open the correction all the way. Always make the final call by ear.
Can I use autotune without knowing the song's key?
You can, but the result will be bad. Pick the wrong key and the tool drags even the correct notes to something wrong. The best move is to find the instrumental's key; if you don't know it, rather than leaving it chromatic, build a custom scale with only the notes the song uses for the most accurate correction.
Is autotune a must on rap and trap vocals?
It's not a must, but in modern trap/rap obvious autotune is now an aesthetic choice. Low retune speed (0-15 ms) and full correction are common here, especially on adlibs. You don't need to hide the effect; that glassy, automated texture has become the genre's signature. Even so, a clean source recording noticeably improves the result.
Can I rescue a bad, out-of-tune recording with autotune?
No. Autotune polishes a good performance; it doesn't rescue a broken one. On a badly out-of-tune take, the tool hesitates and produces strange glitches. Garbage in, garbage out. The fix isn't more autotune, it's re-recording the same part more cleanly.
What does formant correction do, and when should I turn it on?
A formant is the set of resonances that gives your voice its character. When you shift the pitch by a large amount, the formant shifts too and the voice turns cartoonishly high (the chipmunk effect). If you're pulling the voice by more than a couple of semitones, turn formant correction on so the natural character survives even as the note changes.
What does it take to make a vocal professional after autotune?
Pitch correction alone isn't enough. In order, a de-esser at 6-9 kHz, an EQ that cuts the 200-400 Hz mud, and compression around 3:1 sit the vocal up front. Check the mix on a phone speaker too, and don't try to exceed the ~-14 LUFS platforms target. To get this entire chain done automatically, you can use Seimix.
More: All guides