Unit 3 · Topic 3.10 · about 25 minutes
Constructing a Confidence Interval for the Difference Between Two Population Proportions
Estimate how far apart two population proportions are with a confidence interval, from naming the parameter to the finished interval.
Predict first
A dental clinic randomly assigns 400 patients with upcoming appointments to two groups of 200. One group gets a text reminder the day before; the other gets no reminder. Of the reminder group, 82% show up, compared with 71% of the no-reminder group. Which range would you trust to contain the true difference in show-up rates (reminder minus no reminder)?
The procedure and its parameter
The procedure is the two-sample z-interval for a difference between population proportions. It estimates , the difference between two population proportions, using the difference in sample proportions, .
Defining the parameter takes more care than it did for one proportion. A complete definition says that it is a difference in proportions, names the response variable, and names both populations (or both treatments) in context. " is the difference in proportions" tells a reader almost nothing. Compare: ", where is the proportion of all patients like these who would show up for an appointment after a text reminder and is the proportion who would show up with no reminder."
Building the interval
Every confidence interval in this course is a point estimate plus or minus a margin of error. Here the point estimate is , and the interval is
The square root is the standard error of . It is the standard deviation formula from topic 3.9, with the sample proportions standing in for and , which you do not know. The margin of error is the critical value times the standard error, , and depends on the confidence level exactly as it did for one proportion in topic 3.3: 1.645 for 90%, 1.960 for 95% and 2.576 for 99%.
In the reminder study, . At 95% confidence the margin of error is , which is where 0.11 plus or minus 0.082 came from.
| Condition | Two independent random samples | Randomized experiment |
|---|---|---|
| Randomization | A separate random sample from each population | Treatments randomly assigned to the experimental units |
| 10% | and | Not needed |
| Normality | Observed successes and failures, , , and , all at least 10 | The same four observed counts, all at least 10 |
Worked examplePodcast listening by age
A media research group surveyed a random sample of 1,200 U.S. adults ages 18 to 29 and a separate random sample of 1,500 U.S. adults ages 50 to 64. In the younger sample, 552 said they had listened to a podcast in the past week. In the older sample, 435 said so. Construct and interpret a 95% confidence interval for the difference between the proportions of these two age groups who listened to a podcast in the past week.
Name it. Two-sample z-interval for , where is the proportion of all U.S. adults ages 18 to 29 who listened to a podcast in the past week and is the proportion of all U.S. adults ages 50 to 64 who did.
Check conditions. Randomization: two independent random samples, one from each age group. 10%: 1,200 and 1,500 are far less than 10% of U.S. adults in either age group. Normality: the observed counts are 552 listeners and 648 non-listeners in the younger sample, and 435 and 1,065 in the older sample, all at least 10.
Calculate. and , so the point estimate is 0.17. Then and the margin of error is . The interval is , which is about .
Interpret in context. We are 95% confident that the interval from 0.134 to 0.206 contains the true difference between the proportion of U.S. adults ages 18 to 29 and the proportion of U.S. adults ages 50 to 64 who listened to a podcast in the past week (younger minus older).
About 0.134 to 0.206: the younger group's rate is plausibly between about 13 and 21 percentage points higher.
Which group is group 1?
Either order works, as long as you say which one you chose. Subtracting the other way flips the signs and swaps the endpoints: the podcast interval for is . That is the same information read in reverse. The trouble starts when the order is never stated, because then a negative endpoint could mean either group came out ahead.
A habit that prevents it: write the subtraction into the parameter's name, as in , and use subscripts a reader can decode without a key.
Check your understanding
A sports medicine clinic wants to estimate how much the proportion of high school soccer players who have had an ankle sprain differs between girls and boys. It selects a random sample of 300 girls and a separate random sample of 300 boys who play high school soccer in the state. Which procedure fits this goal?
In a random sample of 400 adults in Riverton, 248 said they had visited a public library in the past year. In a separate random sample of 500 adults in Lakeside, 265 said so. Which is a 90% confidence interval for , the proportion of Riverton adults minus the proportion of Lakeside adults who visited a library in the past year?
A clinic randomly assigns 60 patients to a new physical therapy routine and 60 to the standard routine. Of the new-routine patients, 51 regain full range of motion, compared with 42 of the standard-routine patients. Which statement about the conditions for a two-sample z-interval is correct?
A researcher selects a random sample of voters under 30 and a separate random sample of voters 65 and older in one state, to estimate the difference in the proportions who voted by mail in the last election. Which is the best definition of the parameter?
In a randomized experiment, 74 of 120 students who studied with flashcard app A and 57 of 120 students who studied with app B passed a vocabulary quiz. Find the margin of error for a 99% confidence interval for . Round to three decimal places.
Course alignment, for teachers
AP Statistics topic 3.10, Unit 3: Inference for Categorical Data: Proportions.
- Skill 2.C: Identify appropriate statistical inference methods.
- Skill 3.E: Calculate appropriate statistical inference method results.
- Skill 4.E: Justify the use of a chosen statistical inference method by verifying conditions.