Unit 1 · Topic 1.12 · about 20 minutes
Potential Problems with Sampling
Spot the bias in a sampling method, name it, and explain whether it pushes the estimate too high or too low.
Predict first
A radio station asks listeners to call in and say whether the city should build a new stadium. Of the 1,200 people who call, 80% say no. Does this suggest that most of the city's residents oppose the stadium?
What bias is
Bias in a sampling method is a systematic error. Something about how the sample is chosen, or about how the answers are collected, makes the statistic come out consistently larger, or consistently smaller, than the parameter it is meant to estimate. Random error is the luck of the draw, and it shrinks as samples get bigger. Bias does not shrink, because it is built into the method.
The two dotplots below show the difference. Suppose 40% of a school's students support a later start time. Each dot is the result of one poll that collected 100 responses. In the first plot every poll is an SRS of 100 students. In the second, every poll is an online form posted on the school's social media page, where students who want a later start are the most eager to respond.
60 polls, each an SRS of 100 students
Centered on the true value, 0.40. Individual polls miss by a few points in both directions.
60 polls, each 100 responses to a voluntary online form
Consistently too high. Every one of these polls lands above 0.40, and collecting more responses per poll would not pull the center back down.
Four kinds of bias
Voluntary response bias can occur when a sample consists entirely of volunteers. People who choose to respond usually care more about the topic than people who do not, and often feel more negatively. The radio call-in poll and the online form are both voluntary response samples.
Undercoverage bias can occur when the sampling method leaves out part of the population, or makes part of it less likely to be chosen. A survey about after-school jobs handed out during seventh period misses the seniors who leave early for work, so it likely underestimates the proportion of students with jobs.
Nonresponse bias can occur when some of the individuals chosen for the sample do not respond, and the nonrespondents differ from the respondents in ways that matter for the study. Mail a survey about weekly work hours to 500 randomly chosen adults, and the people working the longest hours are the least likely to find time to send it back. The mean from the returned surveys will likely be too low.
Response bias can occur when answers, or measurements, tend to differ from the true value in one direction. Leading or confusing wording causes it: "Do you agree that the school's slow, outdated wifi needs an upgrade?" pushes students toward yes. So do self-reported answers on topics where people shade the truth, such as hours spent studying (reported too high) or phone use in class (reported too low).
Convenience samples
Any nonrandom sampling method invites bias, because it does not use chance to select the individuals. The most common is a convenience sample, built from whoever is easiest to reach: your friends, or the first 50 shoppers through the door. People who are easy to reach tend to have something in common, and that something can be related to their answers.
Sort it
Name the most likely bias. Tap a description, then tap its bin.
Voluntary response
Undercoverage
Nonresponse
Response
Worked exampleDiagnosing a biased survey
To estimate the mean number of hours per week that students at a large high school spend on extracurricular activities, the activities director hands out a survey at the end of the fall sports banquet. 210 students fill it out. Identify a likely bias and its direction.
Look at who could be chosen. Only students at the sports banquet could take the survey. Students who play no fall sport, or who skipped the banquet, had no chance of being included.
Name it. That is undercoverage bias. It is also a convenience sample: the director surveyed whoever was in the room.
Explain how it happens. Everyone at the banquet plays at least one sport, so these students spend more time on activities than a typical student does.
State the direction. The sample mean will likely be too high as an estimate of the mean for all students at the school.
Undercoverage of students who do not play a fall sport, which likely makes the sample mean overestimate the mean weekly activity hours for the whole school.
Check your understanding
A city randomly selects 800 households and mails each one a survey about how many hours per week its adults spend volunteering. Only 260 surveys come back. What is the most likely problem?
A survey asks, "Since texting while driving causes thousands of crashes, do you support a law banning phone use while driving?" What kind of bias is most likely?
To estimate the proportion of a town's adult residents who use the public library, a researcher surveys a random sample of people leaving the town's grocery stores on weekday mornings. Which statement best describes a likely bias?
An online poll collects 50,000 responses. A news anchor says that a sample this large guarantees an accurate estimate. What is wrong with that reasoning?
A health class has just finished a unit on the dangers of too much sugar. The teacher asks each student to write down how many sugary drinks they had last week, on a signed form handed in to her. How is the class mean likely to compare with the true mean?
Course alignment, for teachers
AP Statistics topic 1.12, Unit 1: Exploring One-Variable Data and Collecting Data.
- Skill 2.A: Identify information to answer a question or solve a problem.