"FSIQ 95 (90–100)" appears on nearly every report, and the number in parentheses is one of the most misread things in assessment. It is not the range the student scored in. It's a statement about how precise the test is — and understanding it changes how confidently you can say anything about a single score.
No test is perfectly reliable. If a student took the same test again — different day, different mood, a lucky or unlucky guess or two — the score would shift a little. The score you got is one good estimate of the student's true ability, not a perfectly exact reading of it. The confidence interval is the band around that estimate where the true score most likely sits.
So "Standard Score 95, 90% confidence interval 90–100" means, roughly: the student's obtained score is 95, and we can be about 90% confident their true ability falls somewhere between 90 and 100. The band reflects the imprecision of the measurement — how much wiggle there is in the tool — not variation in how the student performed.
The most common mistake is to read "95 (90–100)" as "the student scored between 90 and 100." They didn't. They scored 95. The 90–100 isn't a range of performance — it's the margin of error around that single obtained score.
The distinction matters because the wrong reading quietly changes the story. "She scored between 90 and 100" sounds like her performance bounced around; it didn't. She produced one score, 95, and the band is the test telling you how much to trust that 95. Treating the interval as a performance range invites vague conclusions ("somewhere in the 90s") when the actual finding is specific ("95, give or take a few points of measurement error").
Confidence interval width isn't arbitrary — it's driven by reliability. The more reliable a score, the tighter the band; the less reliable, the wider. This is why band width is genuinely informative: it tells you, at a glance, how much to trust each score on the page.
A composite or index score is built from several subtests, and combining measurements averages out some of the noise — so composites tend to be more reliable and carry narrower intervals. Individual subtests are single, shorter measurements; they're less reliable and get wider bands. This is a big part of why subtest scores deserve more caution than composites: not only are they narrower slices of ability, the test itself is telling you it measured them less precisely.
When you see a wide confidence interval, the test is signaling "read this one carefully." A subtest scaled score with a wide band shouldn't be interpreted with the same confidence as a composite with a tight one, even if both look like clean single numbers on the report. The band is the test's own estimate of how much rope to give the score.
Confidence intervals matter most when you're about to compare two scores and call one a strength and the other a weakness. If the two scores' confidence intervals overlap substantially, the difference between them may not be as solid as the point values suggest — the true scores could actually be closer together, or even reversed.
Say Working Memory comes in at 88 (82–96) and Processing Speed at 95 (89–102). The point scores differ by 7, which looks like a real gap. But the bands overlap across most of the 89–96 range — meaning the true difference could be much smaller. That doesn't prove there's no difference; it means the gap isn't as clear-cut as "88 versus 95" makes it sound, and a confident "this is a relative weakness" claim deserves a second look.
This is also a good argument for showing the intervals, not just listing them. Two dots at 88 and 95 look decisively different. The same two dots with their error bars drawn make the overlap visible — and the appropriate caution obvious to everyone looking at the page, not just the person who did the math.
Reports use different confidence levels — most commonly 90% or 95%, occasionally 68%. A higher level means you can be more confident the true score falls inside the band, but the band is wider to earn that confidence; a lower level gives a tighter band at the cost of being wrong more often about where the true score sits. Neither is "correct" — they're different trade-offs between precision and certainty, and which one appears is usually set by the test or the examiner's convention. The important thing when reading a report is simply to note which level is being used, since a 90% and a 95% interval for the same score won't be the same width.
A confidence interval is the report being honest about its own precision. It says: here is the best single estimate, and here is how much room to leave around it. Read it as measurement error, not performance range; treat wide bands as caution flags; and when two bands overlap, let the language reflect that the gap is less certain than the numbers imply.
This pairs closely with two other guides: score types (the metrics the intervals are expressed in) and descriptive ranges (why a band that straddles two labels is worth extra care).
Plot scores with their confidence intervals drawn in — so precision and overlap are visible at a glance, not buried in parentheses. Free, nothing stored.
Open the Normal Curve Plotter More score guides