Unit 3 · Topic 3.1 · about 15 minutes
Estimators
Calculate a point estimate from sample data, and justify whether an estimator is biased or unbiased by looking at where its values center.
Predict first
A ride-share company has the wait time for all 12,000 rides in a city last month. An analyst who only gets a random sample of 30 rides uses the longest wait in her sample as her estimate of the longest wait of the month. If she did this with sample after sample, how would her estimates behave?
Estimators and estimates
You almost never know a parameter. What you have is a sample, and you use a statistic from that sample to estimate the parameter. A statistic used this way is a point estimator. The number it produces from your data is the point estimate.
The sample proportion is the point estimator of the population proportion . If 84 of 200 randomly chosen voters favor a new library tax, the point estimate of the proportion of all voters who favor it is . It is called a point estimate because it is a single number. In 3.3 Constructing a Confidence Interval for a Population Proportion you will put a margin of error around it.
Every parameter you meet in this course has a matching statistic, and the notation from 1.2 Variables keeps the pairs straight.
| Population parameter | Point estimator from a sample |
|---|---|
| Proportion, | Sample proportion, |
| Mean, | Sample mean, |
| Standard deviation, | Sample standard deviation, |
Worked exampleTwo point estimates from one sample
A school nurse randomly selects 120 of the 1,450 students at a high school. For each one she records whether the student got a flu shot this fall and how many hours the student slept last night. Of the 120 students, 51 got a flu shot. Their sleep times have a mean of 6.8 hours and a standard deviation of 1.2 hours. Give a point estimate for the proportion of all students at the school who got a flu shot, and one for their mean hours of sleep last night.
Name the parameters. is the proportion of all 1,450 students who got a flu shot this fall. is the mean hours of sleep last night for all 1,450 students.
Match each one to its estimator. The sample proportion estimates , and the sample mean estimates .
Calculate. and hours. The 1.2 hours is also a point estimate (of ), but the question did not ask for it.
Say what the numbers are. Both are statistics. A different random sample of 120 students would give slightly different values, which is why later lessons attach a margin of error.
Estimate that 0.425 of all students at the school got a flu shot, and that they slept 6.8 hours last night on average.
Unbiased means right on average
Nearly every estimate misses the parameter by a little, so missing is not what makes an estimator bad. What matters is whether the misses lean one way. An estimator is unbiased if, on average, its value does not underestimate or overestimate the parameter. An estimator whose values tend to land on one side of the parameter is biased.
To judge this, imagine every possible random sample of the same size, with the estimate computed from each one. If those values center on the parameter, the estimator is unbiased. You rarely see every sample, but a simulation gets close: start with a population whose parameter you know, draw many random samples, and look at where the estimates pile up.
Here is that simulation for the ride-share company. The computer drew 500 random samples of 30 rides and computed two estimators from each sample.
| Estimator | Parameter it estimates | True value | Average of the 500 estimates |
|---|---|---|---|
| Mean wait in the sample, | Mean wait of all 12,000 rides | 5.6 minutes | 5.6 minutes |
| Longest wait in the sample | Longest wait of all 12,000 rides | 31.0 minutes | 13.7 minutes |
Longest wait in each of 500 random samples of 30 rides
The month's longest wait was 31.0 minutes, at the far right of the axis. Two samples happened to include that ride and tied it. Not one sample beat it.
Reading the evidence
The sample mean is unbiased. Its 500 values average 5.6 minutes, the same as the mean of all 12,000 rides. The sample maximum is biased. Its values center around 13.7 minutes when the truth is 31.0, and none of them can land above the truth.
The sample proportion behaves like the sample mean. When the data come from a random sample, the values of center on . That is why is the estimator behind the intervals and tests for proportions in this unit, and 3.2 Sampling Distributions for Sample Proportions shows exactly where its values center and how far they spread.
Bias can come from how the sample was chosen
A good formula is not enough on its own. You can count on the sample proportion being unbiased only when the sample is random. Suppose a radio host asks listeners to call in and uses the proportion of callers who oppose a new stadium to estimate the proportion of all listeners who oppose it. Listeners who are angry about the stadium are the ones most likely to pick up the phone, so sample after sample this estimate would come out too high. That is the systematic error from 1.12 Potential Problems with Sampling, seen from the estimator's side.
Sort it
Each card describes an estimator and what it is used to estimate. Put it in the right bin, then check.
Unbiased
Biased
Check your understanding
A random sample of 250 of the 3,100 seniors at a state university finds that 185 already have a job offer. What is the point estimate of the proportion of all seniors at the university who already have a job offer? Give your answer as a decimal.
The 1,000 mice in a research colony have a mean weight of 24.0 grams. A researcher compares three methods of estimating that mean. For each method she draws 2,000 random samples of 10 mice and computes an estimate from each sample. The averages of the 2,000 estimates are:
- Method A: 22.6 grams
- Method B: 24.0 grams
- Method C: 25.3 grams
Which conclusion do the simulations support?
A student says, "The sample proportion is an unbiased estimator of the population proportion, so the from my random sample of 60 people must equal ." What is wrong with this reasoning?
A town council member wants to estimate the proportion of all town residents who support building a skate park. She uses the proportion of supporters among the 212 residents who emailed her office about the park. Which statement best describes this estimator?
Sixty swimmers race the 50-meter freestyle at a meet. A coach takes a random sample of 8 of the swimmers and uses the fastest time among them to estimate the fastest time of all 60. How would this estimator behave over many random samples?
Course alignment, for teachers
AP Statistics topic 3.1, Unit 3: Inference for Categorical Data: Proportions.
- Skill 3.D: Calculate means, standard deviations, and parameters for probability distributions.
- Skill 4.B: Justify a claim based on statistical calculations and results.