You hit Start and hear a four-second chord — several notes sounding together, all in the same piano-like tone color. Your job is to count how many separate notes are actually sounding, then tap a number from 1 to 6. You don't have to wait for the clip to end before answering, and if you're not sure, you get one free replay of the exact same chord before you have to commit. Then the next one plays, and you do it again — five clips total.
It sounds simple until you try it. A two-note chord is usually obvious. By three or four notes, especially in a close voicing, the individual pitches start to blur into a single, thicker-sounding tone rather than reading as separate notes — which is exactly the perceptual effect this section is measuring.
Your ear doesn't naturally hear a chord as a stack of separate pitches — it hears one fused sound and has to do active work to pull it back apart into voices. Psychologists call this general area auditory scene analysis, a term built largely on the work of researcher Albert Bregman, and it's exactly the mechanism this section exercises: pulling a mixed acoustic signal apart into the separate sound sources that produced it.
For a working musician, this is a genuinely practical skill, not an academic one. It's what lets an arranger hear all four lines of a vocal harmony independently instead of just "the chord." It's what lets a mix engineer notice there's a doubled part buried under a pad instead of one part. It's what a conductor is doing when they catch that the second violins are a beat late inside a dense string texture. Composers and arrangers who are strong at this can often hear a full voicing internally before writing a single note down. It's a different skill from pitch discrimination — you can have excellent relative pitch and still find dense voicings genuinely hard to pull apart, because this section isn't asking "which way did it move," it's asking "how many separate things am I hearing at once."
Again, I'll describe exactly what the code does, because that's the whole reason to write a page like this.
There are five clips per run. The number of voices in each one is fixed at two, three, five, three, and four — so across the five questions you're tested on a 2, two separate 3-voice chords, a 4, and a 5. The order those five come in is shuffled every time you take the test, so you can't count on the chords getting steadily denser — you might get the five-voice chord first.
The chords themselves aren't random note clusters. Each voice count draws from a set of real chord shapes — a two-voice clip might be a fifth, a third, or a sixth; a three-voice clip might be a major or minor triad, or a sus2/sus4 shape; four voices might be a seventh chord or a six chord; five voices might be a 9th chord or an open voicing. The root note is chosen at random each time within a roughly three-octave range, so you're never memorizing a specific pitch — but the shapes themselves are always real harmonic structures, not atonal noise. Every voice in every chord uses the exact same underlying sustained tone — there's no difference in timbre between the notes to lean on. The only thing separating one voice from another is pitch. That's deliberate: it keeps the section testing pure harmonic parsing rather than "can you tell a piano note from a different-sounding note."
Scoring gives you real partial credit. Get the exact count right and you earn four points. Miss by exactly one voice — say you counted three when it was actually four — and you still get two points. Miss by two or more and the question scores zero. Five questions at four points each caps the section at 20 points, the same ceiling as every other section. A miss doesn't end the run; you get feedback showing what the actual count was and move straight to the next clip.
The "Listen again" button is free — using it doesn't cost you any points — but it's capped at exactly one replay per question, and it plays back the identical chord you already heard rather than generating a new one. There's no adaptive difficulty here either: every player gets the same five target counts, just in a different order, so there's no staircase quietly backing off after a miss or ramping up after a streak.
Like every other section, Count the Voices contributes up to 20 of the 100 points in your final MusIQ score, weighted evenly alongside pitch sensitivity and the other three sections.
If you only plug in headphones for one section, make it this one. Counting voices depends on hearing harmonic detail — the individual overtones and beating between close pitches that let your brain register "these are two things" instead of "this is one slightly thick thing." That detail lives mostly in the upper part of the frequency range, and it's precisely what gets damaged first by compressed audio and cheap speakers.
Laptop and phone speakers are small, and small drivers can't reproduce a wide, even frequency range — they roll off or resonate unevenly, which smears exactly the fine harmonic separation you need to pull one voice apart from another. Bluetooth headphones add a second layer of damage: most Bluetooth audio codecs are lossy, meaning they discard some of the original signal to fit it through a limited wireless connection, and the detail they cut is disproportionately the quiet, high-frequency, densely-packed information — a good description of what a five-note chord's harmonic content actually looks like. Stack that compression on top of an already-compressed music source and a lot of the information this section asks you to use is simply gone before it reaches your ears. Wired headphones deliver the full, uncompressed signal, and in a genuinely quiet room, you'll consistently hear more separate voices than you will on a laptop across the room.
A low score here doesn't mean you're a bad listener generally, and it definitely doesn't mean you can't play or write music well — plenty of skilled instrumentalists, and even composers, find dense voicings hard to parse by ear alone, especially if most of their listening has been to a single melodic line at a time. It means that, specifically, pulling a stack of simultaneous pitches apart into separate voices is a skill you haven't built yet — which is good news, because it's genuinely trainable.
The most effective practice isn't listening to more music in general — it's deliberately isolated listening. Play a chord on a piano or guitar, then play each note of it separately right after, and go back and forth until you can predict what you're about to hear before the full chord sounds. Try singing or humming just the top note of a chord while it's playing, then just the bottom note, then an inner note — actively picking one voice out of the texture is a completely different skill from just recognizing a chord's overall color. Choral and string-quartet recordings are good practice material precisely because the individual lines are voiced clearly and don't hide behind drums or distortion. I've built a fuller version of this kind of practice into the ear training guide, and if you want the cognitive-science side of why your brain fuses sounds together in the first place, that's covered on the music and the brain page.
The answer buttons run from 1 to 6, but the five clips in a given run only ever target two, three (twice), four, and five simultaneous voices — you won't get a one-voice or six-voice clip in this version, even though you can technically answer either.
No — every voice in every chord is generated from the same sustained tone sample, just pitched differently. There's no timbre difference to lean on; the only cue separating one voice from another is pitch.
No. Miss the exact count by exactly one voice and you still bank half credit — two of the four possible points for that question. Miss by two or more and it scores zero.
Nothing in points — it's free — but you only get one replay per clip, and it's the same chord again, not a new one.
Because a real recording gives you extra cues this section deliberately removes — different instruments, different stereo positions, different entrances. Here every voice shares the same tone color and starts at the same instant, so pitch separation is the only tool you've got, which is exactly what makes it a purer test of the skill.