4.4 Standard Normal Distribution and z-Scores
The purpose of this section is to build a standard reference scale for continuous probability calculations.
We will focus on the standard normal random variable \(Z\) and on using \(z\)-scores. The general normal distribution will be introduced in the next section.
Key Terms
standard normal random variable
z-score
standardization
cumulative probability
left-tail probability
right-tail probability
standard normal table
symmetry
4.4.1 Learning Objectives
After completing this section, you should be able to:
explain what a \(z\)-score represents;
compute a \(z\)-score from a value, mean, and standard deviation;
interpret positive, negative, and zero \(z\)-scores;
recognize the standard normal reference variable \(Z\);
use a cumulative standard normal table;
calculate left-tail, right-tail, and interval probabilities for \(Z\);
use symmetry and complements to simplify calculations;
find approximate \(z\)-values from cumulative probabilities.
4.4.2 The z-Score: A Standard Distance
Suppose a quantity has mean \(\mu\) and standard deviation \(\sigma\). For an observed value \(x\), define its z-score as
A \(z\)-score measures the distance from the mean in standard deviation units.
\(z>0\): the value is above the mean;
\(z<0\): the value is below the mean;
\(z=0\): the value equals the mean;
\(|z|\) tells how many standard deviations the value is from the mean.
For example, suppose a measurement has
If \(x=13\), then
Thus, 13 is 1.5 standard deviations above the mean.
If \(x=8\), then
Thus, 8 is 1 standard deviation below the mean.
Note
A \(z\)-score has no physical unit. The original unit cancels when dividing by the standard deviation.
The relationship can also be reversed. From
we obtain
This form will become useful when we later convert a standardized value back to an original measurement scale.
4.4.3 Standardization Does Not Automatically Mean Normality
The calculation
can be used to describe the relative position of a value on many different measurement scales.
However, computing a \(z\)-score does not by itself imply that the underlying variable has a normal distribution.
In this section, we separately study a specific reference random variable, called the standard normal random variable.
The next section will explain when an original random variable can be connected to this standard normal reference curve for probability calculations.
4.4.4 The Standard Normal Random Variable
A standard normal random variable is denoted by
It has
and
Equivalently,
The usual notation is
For now, treat \(N(0,1)\) as the name of this standard reference distribution. We will develop the general normal family in the next section.
Its density is symmetric about zero. For reference, the density is
We will not calculate probabilities by integrating this function by hand. Instead, we will use its cumulative probability function.
The horizontal scale is already measured in standard deviations. Therefore, \(z=1\) means one standard deviation above the mean and \(z=-2\) means two standard deviations below the mean.
Useful reference areas are
and
These are useful landmarks for checking whether a probability calculation is reasonable.
4.4.5 The Standard Normal CDF
Define the cumulative distribution function
Because \(Z\) is continuous,
Thus,
Geometrically, \(\Phi(z)\) is the area under the standard normal curve to the left of \(z\).
The cumulative probability can also be written as
This integral does not have a simple elementary antiderivative. Therefore, probabilities are normally obtained from a standard normal table or software.
4.4.6 Reading a Cumulative z-Table
This handout assumes a standard normal table whose entries are
Always check the heading of a table before using it because some tables use a different convention.
Suppose we want
To locate \(z=1.48\):
use \(1.4\) to locate the row;
use \(0.08\) to locate the column;
read the entry at their intersection.
The table gives
Therefore,
A negative value is read in the same way if the table includes negative \(z\)-scores. For example,
so
4.4.7 Three Basic Probability Patterns
Most standard-normal calculations reduce to three patterns.
Pattern 1: Left of a value
This is read directly from a cumulative table.
Pattern 2: Right of a value
The total area is 1, so
Pattern 3: Between two values
For \(a<b\),
The same rules work whether \(a\) and \(b\) are positive, negative, or on opposite sides of zero.
4.4.8 Example: Left-Tail Probability
Find
From the table,
Therefore,
Interpretation: about 93.06% of the standard normal area lies to the left of \(z=1.48\).
4.4.9 Example: Right-Tail Probability
Find
The table gives
Use the complement:
A right-tail probability for a positive \(z\) should be relatively small, which is consistent with this answer.
Now consider
From the table,
Therefore,
4.4.10 Example: Probability Between Two z-Scores
Find
From the cumulative table,
and
Therefore,
The same procedure works if the interval crosses zero.
For example,
Using
and
we obtain
4.4.11 Symmetry of the Standard Normal Curve
The standard normal curve is symmetric about zero.
Therefore,
Equivalently,
For example,
Symmetry is especially useful if a table contains only positive \(z\)-scores.
Another useful result is
The probability outside a symmetric interval is
For example, since
we obtain
and
These values will appear frequently in later statistical inference topics.
4.4.12 Very Large Positive or Negative z-Scores
Some printed tables stop near \(z=3.5\) or \(z=4.0\).
If \(z\) is far to the right, then
If \(z\) is far to the left, then
For example,
when rounded to four decimal places.
Thus,
Likewise,
These approximations reflect extremely small tail areas.
4.4.13 From a Probability Back to a z-Score
Sometimes the probability is known and the corresponding \(z\)-value is unknown.
Suppose
Search the body of the cumulative table for a value near \(0.9500\) and read the corresponding row and column.
This gives approximately
Similarly,
corresponds to
This reverse lookup is the basis for finding critical values and percentiles later in the course.
4.4.14 Training with z-Scores Before Probability
Before using a probability table, be comfortable interpreting the standardized scale itself.
Suppose a measurement has
Then:
\(x\) |
\(z\) |
Interpretation |
|---|---|---|
50 |
0 |
at the mean |
56 |
1 |
one standard deviation above the mean |
44 |
-1 |
one standard deviation below the mean |
62 |
2 |
two standard deviations above the mean |
38 |
-2 |
two standard deviations below the mean |
For example,
The value 62 is therefore two standard deviations above the mean.
This interpretation does not require a probability calculation.
4.4.15 A Reliable Workflow for Standard Normal Problems
When working with \(Z\), use the following procedure.
Identify the requested region.
Decide whether the probability is left of, right of, between, or outside specified \(z\)-values.
Sketch the region mentally or on paper.
This helps determine whether the answer should be small or large.
Read cumulative probabilities.
Use
\[\Phi(z)=P(Z<z).\]Apply the appropriate probability rule.
left tail: \(\Phi(z)\);
right tail: \(1-\Phi(z)\);
interval: \(\Phi(b)-\Phi(a)\);
symmetric tails: use symmetry when helpful.
Check the answer.
A probability must lie between 0 and 1 and should agree with the shaded region you intended to calculate.
4.4.16 Quick Practice
Use a cumulative standard normal table or software.
1. Find
2. Find
3. Find
4. Find
5. A value is \(1.5\) standard deviations below its mean. What is its \(z\)-score?
6. A measurement has \(\mu=80\) and \(\sigma=5\). Find the \(z\)-score of \(x=92\).
Answers
- \[P(Z<-1.72)=0.0427.\]
- \[P(Z>1.63) = 1-0.9484 =0.0516.\]
- \[P(-0.93<Z<0.55) = 0.7088-0.1762 =0.5326.\]
- \[P(Z<-1.96\text{ or }Z>1.96) \approx0.0500.\]
- \[z=-1.5.\]
- \[z = \frac{92-80}{5} =2.4.\]
4.4.17 Common Mistakes
Mistake 1: Reversing the numerator
Use
not \((\mu-x)/\sigma\).
Mistake 2: Ignoring the sign of z
A negative \(z\) is below the mean. A positive \(z\) is above the mean.
Mistake 3: Treating a cumulative table as a right-tail table
If the table gives \(\Phi(z)=P(Z<z)\), a right-tail probability requires
Mistake 4: Adding instead of subtracting for an interval
For \(a<b\),
Mistake 5: Worrying about < versus <=
Because \(Z\) is continuous,
Mistake 6: Assuming every z-score follows the standard normal distribution
A \(z\)-score is a standardized position. Additional distributional information is required before standard normal probabilities can be attached to an arbitrary standardized variable.
4.4.18 Summary
A \(z\)-score converts distance from the mean into standard deviation units:
The standard normal reference variable has
with mean 0 and standard deviation 1.
Its cumulative distribution function is
The main probability rules are
and
Symmetry gives
The key skill from this section is to become comfortable thinking on the z-scale and translating shaded regions into cumulative probabilities.
In the next section, we will use this standard reference scale to work with the general normal distribution.