Unit 3 · Topic 3.1 · about 15 minutes

Estimators

Calculate a point estimate from sample data, and justify whether an estimator is biased or unbiased by looking at where its values center.

Predict first

A ride-share company has the wait time for all 12,000 rides in a city last month. An analyst who only gets a random sample of 30 rides uses the longest wait in her sample as her estimate of the longest wait of the month. If she did this with sample after sample, how would her estimates behave?

Estimators and estimates

You almost never know a parameter. What you have is a sample, and you use a statistic from that sample to estimate the parameter. A statistic used this way is a point estimator. The number it produces from your data is the point estimate.

The sample proportion p^ is the point estimator of the population proportion p. If 84 of 200 randomly chosen voters favor a new library tax, the point estimate of the proportion of all voters who favor it is p^=84200=0.42. It is called a point estimate because it is a single number. In 3.3 Constructing a Confidence Interval for a Population Proportion you will put a margin of error around it.

Every parameter you meet in this course has a matching statistic, and the notation from 1.2 Variables keeps the pairs straight.

Parameters and the statistics that estimate them
Population parameterPoint estimator from a sample
Proportion, pSample proportion, p^
Mean, μSample mean, x‾
Standard deviation, σSample standard deviation, s

Worked exampleTwo point estimates from one sample

A school nurse randomly selects 120 of the 1,450 students at a high school. For each one she records whether the student got a flu shot this fall and how many hours the student slept last night. Of the 120 students, 51 got a flu shot. Their sleep times have a mean of 6.8 hours and a standard deviation of 1.2 hours. Give a point estimate for the proportion of all students at the school who got a flu shot, and one for their mean hours of sleep last night.

  1. Name the parameters. p is the proportion of all 1,450 students who got a flu shot this fall. μ is the mean hours of sleep last night for all 1,450 students.

  2. Match each one to its estimator. The sample proportion p^ estimates p, and the sample mean x‾ estimates μ.

  3. Calculate. p^=51120=0.425 and x‾=6.8 hours. The 1.2 hours is also a point estimate (of σ), but the question did not ask for it.

  4. Say what the numbers are. Both are statistics. A different random sample of 120 students would give slightly different values, which is why later lessons attach a margin of error.

Answer.

Estimate that 0.425 of all students at the school got a flu shot, and that they slept 6.8 hours last night on average.

Unbiased means right on average

Nearly every estimate misses the parameter by a little, so missing is not what makes an estimator bad. What matters is whether the misses lean one way. An estimator is unbiased if, on average, its value does not underestimate or overestimate the parameter. An estimator whose values tend to land on one side of the parameter is biased.

To judge this, imagine every possible random sample of the same size, with the estimate computed from each one. If those values center on the parameter, the estimator is unbiased. You rarely see every sample, but a simulation gets close: start with a population whose parameter you know, draw many random samples, and look at where the estimates pile up.

Here is that simulation for the ride-share company. The computer drew 500 random samples of 30 rides and computed two estimators from each sample.

500 simulated random samples of 30 rides
EstimatorParameter it estimatesTrue valueAverage of the 500 estimates
Mean wait in the sample, x‾Mean wait of all 12,000 rides5.6 minutes5.6 minutes
Longest wait in the sampleLongest wait of all 12,000 rides31.0 minutes13.7 minutes

Longest wait in each of 500 random samples of 30 rides

050100150Number of samples8 to 10: 3610 to 12: 13212 to 14: 14014 to 16: 9316 to 18: 5018 to 20: 2720 to 22: 1222 to 24: 424 to 26: 126 to 28: 128 to 30: 230 to 32: 28101214161820222426283032Longest wait in the sample (minutes)

The month's longest wait was 31.0 minutes, at the far right of the axis. Two samples happened to include that ride and tied it. Not one sample beat it.

Reading the evidence

The sample mean is unbiased. Its 500 values average 5.6 minutes, the same as the mean of all 12,000 rides. The sample maximum is biased. Its values center around 13.7 minutes when the truth is 31.0, and none of them can land above the truth.

The sample proportion behaves like the sample mean. When the data come from a random sample, the values of p^ center on p. That is why p^ is the estimator behind the intervals and tests for proportions in this unit, and 3.2 Sampling Distributions for Sample Proportions shows exactly where its values center and how far they spread.

Bias can come from how the sample was chosen

A good formula is not enough on its own. You can count on the sample proportion being unbiased only when the sample is random. Suppose a radio host asks listeners to call in and uses the proportion of callers who oppose a new stadium to estimate the proportion of all listeners who oppose it. Listeners who are angry about the stadium are the ones most likely to pick up the phone, so sample after sample this estimate would come out too high. That is the systematic error from 1.12 Potential Problems with Sampling, seen from the estimator's side.

Sort it

Each card describes an estimator and what it is used to estimate. Put it in the right bin, then check.

Unbiased

Biased

Check your understanding

1

A random sample of 250 of the 3,100 seniors at a state university finds that 185 already have a job offer. What is the point estimate of the proportion of all seniors at the university who already have a job offer? Give your answer as a decimal.

2

The 1,000 mice in a research colony have a mean weight of 24.0 grams. A researcher compares three methods of estimating that mean. For each method she draws 2,000 random samples of 10 mice and computes an estimate from each sample. The averages of the 2,000 estimates are:

  • Method A: 22.6 grams
  • Method B: 24.0 grams
  • Method C: 25.3 grams

Which conclusion do the simulations support?

3

A student says, "The sample proportion is an unbiased estimator of the population proportion, so the p^ from my random sample of 60 people must equal p." What is wrong with this reasoning?

4

A town council member wants to estimate the proportion of all town residents who support building a skate park. She uses the proportion of supporters among the 212 residents who emailed her office about the park. Which statement best describes this estimator?

5

Sixty swimmers race the 50-meter freestyle at a meet. A coach takes a random sample of 8 of the swimmers and uses the fastest time among them to estimate the fastest time of all 60. How would this estimator behave over many random samples?

Lab

Sampling Distribution Machine

Run the ride-share idea yourself. Pick a population, draw a thousand samples, and compare the center of the simulated statistics to the parameter. Then switch to proportions and change the sample size: the spread of the values of p^ changes, but their center stays on p.

Open the full Sampling Distribution Machine lab

Course alignment, for teachers

AP Statistics topic 3.1, Unit 3: Inference for Categorical Data: Proportions.

  • Skill 3.D: Calculate means, standard deviations, and parameters for probability distributions.
  • Skill 4.B: Justify a claim based on statistical calculations and results.