Assessment reports mix half a dozen score metrics, each with its own mean and scale. They all describe the same thing — where a score falls relative to a norm group — just in different units. Once you can convert between them, the whole report reads as one language.
Almost every standardized score is a standard score in the general sense: it expresses how far a raw score sits from the average of a norm group, measured in standard deviations. The different metrics — standard scores, scaled scores, T-scores, z-scores, stanines, NCEs — are just different choices of where to put the mean and how wide to make each step. Learn the mean and standard deviation of each, and you can move between them freely.
These rows all describe the same points on the normal curve, expressed in each metric. The percentile is the common translator — it's what every scale ultimately points to.
| Description | SD from mean | Standard (100/15) | Scaled (10/3) | T (50/10) | z (0/1) | Percentile |
|---|---|---|---|---|---|---|
| Well above average | +2 | 130 | 16 | 70 | +2.0 | 98th |
| Above average | +1 | 115 | 13 | 60 | +1.0 | 84th |
| High end of average | +⅔ | 110 | 12 | 57 | +0.7 | 75th |
| Average | 0 | 100 | 10 | 50 | 0.0 | 50th |
| Low end of average | −⅔ | 90 | 8 | 43 | −0.7 | 25th |
| Below average | −1 | 85 | 7 | 40 | −1.0 | 16th |
| Well below average | −2 | 70 | 4 | 30 | −2.0 | 2nd |
Percentiles are rounded to the values most commonly cited; different sources may round slightly differently.
The most common metric on cognitive and academic reports. Composite and index scores are almost always standard scores: a WISC-V Full Scale IQ, a WJ-V cluster, a BASC-3 composite. Because the mean is 100 and each standard deviation is 15, roughly 68% of people score between 85 and 115 — that is, within one standard deviation of the mean. That part is fixed by the math.
What isn't fixed is the label. Where "Average" begins and ends is a publisher's choice, not a statistical fact. Some descriptive systems call the whole 85–115 band "Average"; others reserve "Average" for a narrower 90–110 and split off "Low Average" and "High Average" at the edges. So the same standard score of 87 might be labeled "Average" by one test and "Low Average" by another — the number is comparable across tests, but the word attached to it is not. Keep the general rule of thumb in mind (around 85–115 is the broad average range), but always check the specific test's manual for how it defines the bands.
Used for subtest scores on the Wechsler batteries and many language and processing measures. The mean is 10 and each standard deviation is 3, so the average range runs from 7 to 13. A scaled score of 10 is exactly average; a 7 is one SD below, a 13 one SD above.
A common point of confusion: subtest scaled scores and composite standard scores appear side by side on the same report but live on completely different scales. A scaled score of 10 and a standard score of 100 mean the same thing — dead average — even though the numbers look nothing alike. Reading them as if they were on one scale (thinking a "10" is alarmingly low next to a "100") is a frequent mistake.
The standard metric for behavior, social-emotional, and personality measures — BASC-3, Conners, the MMPI, and most rating scales. Mean 50, standard deviation 10, so the average range is 40 to 60.
One thing to watch with T-scores: on many clinical scales, higher is a greater concern, not a better result. A T-score of 70 on an anxiety or hyperactivity scale means the concern is well above typical — the opposite direction from a cognitive score, where higher is stronger. Always check what the scale is measuring before reading a high T-score as good or bad. (More on behavior & rating-scale scores →)
The z-score is the rawest expression of the same idea: it states, directly, how many standard deviations a score is from the mean. A z of 0 is average; +1 is one SD above; −1.5 is one and a half below. Every other metric is really a z-score dressed up with a friendlier mean and scale. z-scores are used mostly in research and whenever you need to compare across measures on different metrics, because they strip everything down to the common unit.
A nine-point scale (the name comes from "standard nine"), with a mean of 5 and standard deviation of 2. Stanines compress the whole distribution into nine broad bands — 1 is lowest, 9 is highest, 5 is average. They're an older format still seen in some group-administered educational testing. The trade-off is deliberate: stanines are simple, but coarse. A single stanine covers a wide span of percentiles, so two students in the same stanine can differ meaningfully.
Normal Curve Equivalents look almost like percentiles — they run 1 to 99 and meet percentiles at 1, 50, and 99 — but they're built on an equal-interval scale (mean 50, SD 21.06). That equal spacing is the whole point: unlike percentiles, NCEs can be averaged and compared arithmetically, which is why they're used in federal education program reporting. They are easy to mistake for percentiles at a glance; the values in between diverge, so it's worth confirming which one a report is actually using.
A percentile rank states the percentage of the norm group a score met or exceeded: a score at the 37th percentile did as well as or better than 37% of same-age peers. Percentiles are the most intuitive metric for a lay audience, and they're the natural bridge between all the others — every scale above ultimately points to a percentile.
This is also why percentiles and standard scores are worth reporting together: the standard score preserves the equal-interval spacing, and the percentile makes the standing intuitive. Each covers the other's weakness.
Once the metrics click into place, a line like "Scaled 7, SS 85, T 40, 16th percentile" stops being four numbers and becomes one fact stated four ways: one standard deviation below average. That fluency is what lets you read a mixed report smoothly, compare scores that arrive in different metrics, and translate any of them into the plain language a parent or teacher will actually follow.
Two companion guides build directly on this one: descriptive ranges (the "Average / Low Average" labels attached to these bands) and confidence intervals (the precision around any single score).
Mix standard scores, scaled scores, T-scores, and percentiles on a single normal curve — the plotter converts and places them for you. Free, nothing stored.
Open the Normal Curve Plotter More score guides