Unit 3 · Topic 3.11 · about 20 minutes
Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions
Read an interval for a difference in proportions correctly and use it to say which claims about two populations the data support.
Predict first
A school district randomly assigns 600 students to one of two summer reading programs, 300 to each. A 95% confidence interval for , the difference between the proportions of students like these who would meet their reading goal in Program A and in Program B, is . What does the interval say about whether the programs differ?
What the interval says
A confidence interval is built from samples, so it may or may not contain the true difference . You never find out which. What you can describe is how much to trust the method that built it.
The standard interpretation of a C% interval from to is: we are C% confident that the interval from to contains the true difference between the proportion for population 1 and the proportion for population 2. Then fill in the details. Name both populations, the response variable, and the order of subtraction. ("Captures" works as well as "contains".)
For the reading programs: we are 95% confident that the interval from to 0.14 contains the true difference between the proportions of students like these who would meet their reading goal in Program A and in Program B (A minus B).
What the 95% means
The confidence level describes the method, not this one interval. In repeated random sampling with the same sample sizes from the same populations, about 95% of the intervals built this way would contain the true difference , and about 5% would miss it. For an experiment like the reading study, picture repeating the random assignment instead of the sampling.
So it is wrong to say there is a 95% probability that is between and 0.14. Once the interval is computed, it either contains the true difference or it does not. The 95% is the success rate of the process that produced it.
Using the interval to judge a claim
An interval is a set of plausible values for . To judge a claim, check whether the values the claim depends on are inside it. Zero matters most, because means no difference at all.
- If the interval contains 0, there is insufficient evidence to conclude that the two population proportions differ.
- If the interval does not contain 0, there is convincing evidence that they differ, and the sign of the values tells you which proportion is larger.
| Interval | Contains 0? | What it supports |
|---|---|---|
| No, all positive | Convincing evidence that , plausibly by 4 to 13 percentage points | |
| No, all negative | Convincing evidence that , plausibly by 2 to 11 percentage points | |
| Yes | No convincing evidence of a difference; could plausibly be 3 points lower or 8 points higher |
The same reasoning works for any value a claim names, not only 0. A claim that is at least 5 percentage points higher than needs every plausible value to be at least 0.05. If the interval reaches below 0.05, the data do not back that claim, even when they do show some difference.
Worked exampleOne interval, two claims
A city water department wants more customers on paperless billing. It randomly assigns 3,000 customers to receive one of two emails: 1,500 get the old email (A) and 1,500 get a redesigned email (B). Within a month, 345 of the B customers and 270 of the A customers sign up. The department's communications manager claims that the redesigned email works better and that it raises the sign-up rate by at least 5 percentage points. Use a 95% confidence interval to evaluate both claims.
Name it. Two-sample z-interval for , where and are the proportions of customers like these who would sign up for paperless billing after getting the redesigned email and after getting the old email.
Check conditions. Randomization: the emails were randomly assigned. The 10% condition does not apply to a randomized experiment. Normality: the observed counts are 345 sign-ups and 1,155 non-sign-ups for email B, and 270 and 1,230 for email A, all at least 10.
Calculate. and . Then , and gives about .
Interpret. We are 95% confident that the interval from 0.021 to 0.079 contains the true difference between the proportions of customers like these who would sign up after the redesigned email and after the old email (B minus A).
First claim. Every value in the interval is positive, so 0 is not a plausible difference. The data give convincing evidence that the redesigned email produces a higher sign-up rate than the old one.
Second claim. The interval includes every value from 0.021 up to 0.05, so increases smaller than 5 percentage points are plausible. The data do not give convincing evidence that the increase is at least 5 points, even though the observed difference was exactly 0.05.
The interval supports the first claim but not the second.
Check your understanding
Researchers take a random sample of teens and a separate random sample of adults in a large city. A 90% confidence interval for , the proportion of teens minus the proportion of adults who used a ride-share app in the past month, is . Which interpretation is correct?
Researchers took a random sample of teens and a separate random sample of adults in a large city, and built a 90% confidence interval of for , the difference in the proportions who used a ride-share app in the past month. What does the 90% confidence level mean?
A college takes random samples of students living on its north campus and its south campus. A 95% confidence interval for , the difference in the proportions of students who skip breakfast on weekdays (north minus south), is . Which conclusion is justified?
In a driving simulator experiment, 1,200 volunteers are randomly assigned to approach either a standard stop sign or a flashing LED stop sign, 600 to each. A 99% confidence interval for , the proportion of drivers like these who would come to a full stop at the standard sign minus the proportion at the LED sign, is . Which claim does the interval support?
An online store randomly shows half of its visitors a new checkout page and the other half the old one. A 95% confidence interval for , the difference in the proportions of visitors who use a coupon code (new page minus old page), is . A manager claims the new page raised coupon use by more than 10 percentage points. Do the data support the claim?
Course alignment, for teachers
AP Statistics topic 3.11, Unit 3: Inference for Categorical Data: Proportions.
- Skill 4.F: Interpret results of statistical inference methods.
- Skill 4.G: Justify a claim based on statistical inference method results.