Unit 1 · Topic 1.9 · about 25 minutes
Comparisons of the Distributions for One Quantitative Variable
Compare two or more distributions on center, variability, shape and unusual features in context, and use z-scores to compare individual values.
Predict first
Maya scored 88 on a chemistry test where the class mean was 80 and the standard deviation was 4. She scored 92 on an English test where the class mean was 85 and the standard deviation was 10. Compared with her classmates, on which test did she do better?
Comparing distributions
Comparing two distributions uses the same four features as describing one: shape, center, variability and unusual features. The difference is that every statement has to compare. "The median for seniors is 6.5 hours" describes one group. "The median for seniors, 6.5 hours, is lower than the median for freshmen, 8 hours" compares them.
A health class asked random samples of 25 freshmen and 25 seniors how many hours they slept on the last school night. Two dotplots drawn on the same scale make the comparison almost automatic.
Freshmen: hours of sleep on the last school night (25 students)
Seniors: hours of sleep on the last school night (25 students)
In context:
- Center. The seniors slept less. Their median was 6.5 hours, compared with 8 hours for the freshmen.
- Variability. The seniors' sleep was more spread out: an IQR of 1.5 hours against 1.25 for the freshmen, and a range of 4.25 hours against 3.5.
- Shape. Both distributions are roughly symmetric and unimodal.
- Unusual features. Neither group has outliers or gaps. The three lowest values, all under 5.5 hours, belong to seniors.
Dotplots, stem-and-leaf plots and histograms can show all of these features, clusters and gaps included. Boxplots show center, variability, outliers and skewness, but they cannot show clusters or gaps.
Back-to-back stem-and-leaf plots
A back-to-back stemplot puts two groups on one shared column of stems. One group's leaves run to the right of the stems, as usual. The other group's leaves run to the left, so they are read from the stem outward: the row 8 2 | 6 on the left side means 62 and 68. Here are the scores on the same quiz in two class periods.
Quiz scores, 19 students per period. Key: 2 | 8 | 1 means 82 for Period 2 and 81 for Period 6.
Period 2's scores bunch in the 70s and 80s, while Period 6's spread all the way from the 50s to the high 90s. The centers are close, with medians of 82 and 79, but Period 6 is far more variable: its IQR is 23 points against 13 for Period 2.
Comparing with boxplots and summary statistics
Dani can drive to school on the highway (Route A) or on back roads (Route B). She timed the trip on 20 school days for each route. Side-by-side boxplots on one axis, plus the summary statistics, show the whole comparison.
Drive times to school, 20 days on each route
| Route | Mean | SD | Min | Median | Max | |||
|---|---|---|---|---|---|---|---|---|
| A (highway) | 20 | 19.7 | 5.9 | 14 | 16 | 18 | 20.5 | 38 |
| B (back roads) | 20 | 24.1 | 1.8 | 21 | 23 | 24 | 25 | 28 |
Worked exampleWhich route should Dani take?
Use the boxplots and the summary statistics to compare the two routes, then recommend one.
Center. Route A is usually faster. Its median, 18 minutes, is 6 minutes less than Route B's median of 24, and its mean is lower too.
Variability. Route A is far less predictable. Its IQR is 4.5 minutes against 2 for Route B, and its standard deviation is about three times as large.
Shape and outliers. Route A is skewed right. Its upper fence is , so the 31-minute and 38-minute days are outliers, most likely traffic jams. Route B is roughly symmetric with no outliers: its fences are 20 and 28, and every time falls inside them.
Decide in context. On a typical day Route A saves about 6 minutes. But on 2 of its 20 days it took longer than Route B's slowest day. If being on time matters most, as on the morning of an exam, Route B's consistency is worth the extra minutes.
Route A for speed on a typical day, Route B for reliability. A good comparison uses center, variability, shape and outliers together.
z-scores
So far the comparisons have been between whole distributions. Sometimes the question is about one value: how unusual is it, or which of two values from different distributions is more impressive? The tool is a standardized score, or z-score, which measures how many standard deviations a value falls above or below the mean:
Here is the data value, is the population mean and is the population standard deviation. A positive z-score means the value is above the mean, a negative one means it is below, and is exactly at the mean. When the population mean and standard deviation are unknown, use the sample mean and sample standard deviation in their place.
z-scores compare positions within one distribution, such as two students' results on the same test, and between distributions with different centers, spreads or even units.
Worked exampleWhich result is more impressive?
In a school's fitness test, the mile times of all ninth graders have mean minutes and standard deviation minutes. Their push-up counts have mean and standard deviation . Jordan ran a 7.4-minute mile and did 40 push-ups. Compared with the other ninth graders, which result is more impressive?
Mile. . Jordan's time is 1.5 standard deviations below the mean, and for a race time, below the mean means faster.
Push-ups. . Jordan's count is 2 standard deviations above the mean.
Compare. Both results beat the average. The push-ups are 2 standard deviations better than the mean, while the mile is 1.5 standard deviations better, so the push-ups stand out more.
The push-ups. Mind the direction: for a time, a negative z-score is the good side.
Check your understanding
The daily high temperatures in a city in July have mean degrees Fahrenheit and standard deviation degrees. On one day the high was 97 degrees. What is the z-score of that day's high?
Two friends took different college entrance exams. Ava scored 1,310 on an exam with mean 1,050 and standard deviation 200. Ben scored 29 on an exam with mean 21 and standard deviation 5. Who scored better relative to the other test takers?
Side-by-side boxplots of battery life, in hours, for two phone models are built from these five-number summaries. Model X: 8, 11, 13, 14, 16. Model Y: 9, 10, 12, 16, 22. Which statement is supported?
In a random sample of employees at a company, commute times have a mean of 32 minutes and a standard deviation of 8 minutes. One employee's commute has a z-score of . How long is that employee's commute?
Use the two dotplots of hours of sleep near the top of this lesson. Which claim do they support?
Course alignment, for teachers
AP Statistics topic 1.9, Unit 1: Exploring One-Variable Data and Collecting Data.
- Skill 3.B: Calculate summary statistics, relative positions of points within a distribution, and predicted responses.
- Skill 4.A: Describe and compare tabular and graphical representations of data, as well as summary statistics.
- Skill 4.B: Justify a claim based on statistical calculations and results.
- Skill 4.C: Describe distributions and compare relative positions of points within a distribution.