Unit 2 · Topic 2.12 · about 25 minutes
Sampling Distributions and the Central Limit Theorem
Describe a sampling distribution from a simulation, tell it apart from the population and from one sample's data, and use the central limit theorem to predict its shape.
Predict first
A customer help line's call lengths are strongly skewed to the right: most calls take a few minutes, and a handful run past half an hour. You take a random sample of 40 calls, find the mean length, and repeat that a thousand times. What shape will the histogram of those thousand sample means have?
Three distributions to keep apart
Questions about sampling go wrong when distributions get mixed up, so name them carefully.
- The population distribution holds the values of the variable for every individual in the population: the lengths of all 10,000 calls the help line took last month.
- The distribution of one sample's data holds the values in a single random sample, such as the lengths of 40 calls.
- The sampling distribution of a statistic is the distribution of the values of that statistic, such as the sample mean, for all possible samples of a given size from a given population.
The first two describe individual calls. The third describes a statistic, one number per sample.
Every possible sample
A sampling distribution is easiest to see when the population is tiny. Five plants in a classroom are 12, 15, 20, 24 and 34 centimeters tall, a population mean of 21 cm. There are exactly 10 different samples of 2 plants. The table lists each one with its sample mean, and those 10 means are the sampling distribution of for samples of size 2 from this population.
| Sample | Heights | |
|---|---|---|
| 1 | 12, 15 | 13.5 |
| 2 | 12, 20 | 16 |
| 3 | 12, 24 | 18 |
| 4 | 12, 34 | 23 |
| 5 | 15, 20 | 17.5 |
| 6 | 15, 24 | 19.5 |
| 7 | 15, 34 | 24.5 |
| 8 | 20, 24 | 22 |
| 9 | 20, 34 | 27 |
| 10 | 24, 34 | 29 |
Sampling distribution of the sample mean, samples of 2 plants
The 10 sample means average exactly 21, the population mean, and they spread out less than the plants do: no sample mean is as small as 12 or as large as 34.
Simulating a sampling distribution
Real populations are far too big for that. The help line's 10,000 calls can form about different samples of 40, more than the number of atoms in the observable universe. So you simulate. Assume what the population looks like, or assume values for its parameters. Generate a large number of random samples from it, compute the statistic for each sample, and record it. The distribution of the recorded values approximates the sampling distribution.
The three histograms below come from one such simulation: the population of 10,000 call lengths, then the means of 1,000 random samples of size 2, then the means of 1,000 random samples of size 40.
The population: 10,000 call lengths
Strongly skewed to the right. The mean is about 6.1 minutes, but the longest call ran 52.7 minutes.
Means of 1,000 random samples of 2 calls
Averaging just two calls barely helps. The sample means are still clearly skewed to the right.
Means of 1,000 random samples of 40 calls
With 40 calls per sample the means form a single mound that is close to symmetric, with a slight right skew left over from the population.
Worked exampleDescribing the simulated sampling distribution
Describe the simulated sampling distribution of the sample mean for samples of 40 calls, and compare it with the population of call lengths.
Shape. Unimodal and roughly symmetric, close to a normal curve, with a slight skew to the right. The population is strongly skewed to the right.
Center. The sample means center at about 6.0 minutes, close to the population mean of about 6.1 minutes.
Variability. 891 of the 1,000 sample means fall between 4.5 and 7.5 minutes, and none is above 10.4 minutes. Individual calls range from under a minute to 52.7 minutes.
Say what it means in context. The mean length of a random sample of 40 calls is far more predictable than the length of any one call, and its distribution is approximately normal even though call lengths are not.
Approximately normal, centered near the population mean of about 6 minutes, and far less variable than individual call lengths.
The central limit theorem
The pattern in those histograms has a name. The central limit theorem (CLT) says that the sampling distribution of a mean of a random sample has a shape that can be approximated by a normal distribution, and the larger the sample, the better the approximation.
Read that carefully. It is about the distribution of the sample mean across many samples. It does not say the population becomes normal, and it does not say the data inside one sample become normal. How large a sample is large enough depends on how skewed the population is, and 4.1 Sampling Distributions for Sample Means gives the working rule.
Randomization distributions
Experiments assign treatments at random instead of sampling at random, and they have their own version of this simulation. Twenty volunteers were randomly assigned, 10 to learn a list of words with a new memory technique and 10 with their usual method. The technique group recalled a mean of 15.7 words and the usual group 13.0, a difference of 2.7 words.
Could random assignment alone produce a gap that big? To find out, pool the 20 responses, reassign them at random to two groups of 10, and record the difference in means. Repeat many times. The resulting randomization distribution shows how the statistic varies from one random assignment to another, and it approximates the sampling distribution of the difference.
1,000 random reassignments of the 20 memory scores
The reassigned differences center at 0. The observed difference of 2.7 words sits far out in the right tail: only 6 of the 1,000 reassignments produced a difference that large.
Topic 3.6 turns this kind of comparison into a p-value.
Check your understanding
A researcher selects one random sample of 50 students from a large high school and records each student's height. Which of these is the sampling distribution of the sample mean height?
The amounts customers spend at a hardware store are skewed to the right. A manager takes many random samples of 100 receipts and computes the mean of each sample. Which best describes the distribution of those sample means?
A student takes one random sample of 20 values from a population that is skewed to the right, then takes a single much larger random sample of 2,000 values from the same population. How should a histogram of the data in the larger sample look?
In an experiment, 30 volunteers are randomly assigned to two treatments, and the difference in mean responses between the groups is recorded. How is a randomization distribution for that difference built?
A student wants to simulate the sampling distribution of the sample proportion for random samples of 50 people from a population in which 30% have a library card. Which plan does that?
Course alignment, for teachers
AP Statistics topic 2.12, Unit 2: Probability, Random Variables, and Probability Distributions.
- Skill 4.C: Describe distributions and compare relative positions of points within a distribution.