The age at which an obtained raw score is the average. Like the grade equivalent, it's an estimate that's easily misread and generally not recommended for reporting. See the guide →
A descriptive band on many behavior rating scales (for example, BASC-3), typically starting around a T-score of 60, indicating scores somewhat elevated above the typical range — a signal to look further, not a diagnosis.
The broad middle of a distribution — on a 100/15 standard-score scale, roughly 85 to 115. Where "Average" begins and ends as a label is a publisher's choice and varies by test. See the guide →
The highest (ceiling) and lowest (floor) scores a test or subtest can produce. A score at the ceiling or floor may mean the test ran out of room to measure the student's actual ability in that direction.
The highest concern band on many behavior scales, typically around a T-score of 70 or above, meaning the ratings strongly resemble those of the concern group. It is a screening flag, not a diagnosis.
A summary score built from several subtests (for example, a Full Scale IQ or a Working Memory Index). Because it combines measurements, it's usually more reliable — and carries a narrower confidence interval — than any single subtest. See the guide →
The band around an obtained score where the true score most likely falls. It reflects the test's measurement error, not the range the student scored in. See the guide →
The plain-language label a test attaches to a band of scores ("Average," "Low Average," "Extremely Low"). A description of a score band chosen by the publisher — not a fixed category the child belongs to. See the guide →
Concern bands used on some behavior scales (for example, Conners-4), serving the same role as "At-Risk" and "Clinically Significant" — higher means more concern on a problem scale.
The grade level at which an obtained raw score is the average. It looks like a grade level but isn't a statement of what grade-level work a student can do — and it exaggerates in both directions. See the guide →
The average of a group of scores — the center of the normal distribution. Each score metric sets its mean at a chosen value: 100 for standard scores, 10 for scaled, 50 for T-scores, 0 for z-scores.
An equal-interval score (mean 50, SD 21.06) that runs 1–99 and meets the percentile at 1, 50, and 99. Used in federal program reporting because, unlike percentiles, NCEs can be averaged. Easily mistaken for a percentile. See the guide →
The representative sample a test was standardized on. Every standardized score compares a student to this group — so a score always means "relative to these peers," usually same-age or same-grade.
The symmetrical, bell-shaped distribution that most standardized scores are built on. About 68% of scores fall within one standard deviation of the mean, 95% within two, and 99.7% within three.
The percentage of the norm group a score met or exceeded — a score at the 37th percentile did as well as or better than 37% of peers. Intuitive, but not evenly spaced: percentiles bunch up in the middle and stretch at the extremes. See the guide →
The unconverted count of items correct (or points earned) before it's compared to the norm group. On its own it means little; it becomes interpretable once converted to a standard score, percentile, and so on.
How consistently a test measures — how similar scores would be on retesting. Higher reliability produces narrower confidence intervals; it's the reason composites carry tighter bands than subtests. See the guide →
A standardized score with a mean of 10 and standard deviation of 3, used for subtests on the Wechsler batteries and many other measures. A scaled 10 means the same thing as a standard score of 100 — dead average. See the guide →
A measure of spread — how far scores typically sit from the mean. It's the "step size" of each metric: 15 points per SD for standard scores, 3 for scaled, 10 for T-scores.
In the specific sense, a score with mean 100 and SD 15 — the most common metric for composite and index scores. In the general sense, any score expressed in standard-deviation units. See the guide →
A nine-point scale (mean 5, SD 2) that compresses the distribution into nine broad bands. Simple but coarse — a single stanine spans a wide range of percentiles. See the guide →
An individual task measuring a narrow skill, several of which combine into a composite. Subtests are shorter, less reliable measurements, so they carry wider confidence intervals and deserve more cautious interpretation than composites.
A standardized score with a mean of 50 and standard deviation of 10, standard on behavior, social-emotional, and personality measures. On many clinical scales, a higher T-score means more concern, not a better result. See the guide →
The score a student would get if a test measured perfectly, with no error. We never observe it directly — the obtained score estimates it, and the confidence interval is the band where it most likely sits.
The non-elevated band on a behavior rating scale — the reassuring result. On a problem scale, "Typical" means the ratings look like most children's, i.e. not suggestive of the concern.
The most basic standardized score (mean 0, SD 1): it states directly how many standard deviations a score is from the mean. Every other metric is a z-score rescaled to a friendlier mean and step size. See the guide →
Plot real scores on a normal curve and see how the metrics, bands, and intervals line up. Free, nothing stored.
Open the Normal Curve Plotter Back to all guides