Unit 3 · Topic 3.12 · about 20 minutes

Setting Up a Test for the Difference Between Two Population Proportions

Recognize when a two-sample z-test fits, write its hypotheses in context, and verify its conditions before any calculation.

Predict first

A college surveys a random sample of 100 first-year students and a separate random sample of 300 upper-level students. Of the first-year students, 45 use the campus gym at least once a week; of the upper-level students, 90 do. Suppose the two groups really have the same proportion of weekly gym users. What is the best estimate of that shared proportion?

Comparing two proportions with a test

Topics 3.10 and 3.11 estimated how far apart two proportions are. A significance test asks a yes-or-no question instead: do the data give convincing evidence that the two population proportions differ, or that a particular one is larger? The procedure is the two-sample z-test for the difference between two population proportions. It works with data from two independent random samples or from a randomized experiment with two treatments.

Define the parameters before you write anything else. Each definition needs the response variable and its population in context. "p1 is the proportion of all licensed drivers in the state under 25 who texted while driving in the past month" does the job. "p1 is the proportion for group 1" does not.

Writing the hypotheses

The null hypothesis says there is no difference:

H0:p1=p2or equivalentlyH0:p1−p2=0

The alternative hypothesis comes from the question you are trying to answer, and you choose it before looking at the data. Each hypothesis can be written either way, and both forms mean the same thing, so use whichever reads more naturally.

Matching the research question to the alternative hypothesis
The question asks whetherHaWritten as a difference
p1 is larger than p2Ha:p1>p2Ha:p1−p2>0
p1 is smaller than p2Ha:p1<p2Ha:p1−p2<0
the proportions differ, in either directionHa:p1≠p2Ha:p1−p2≠0

Conditions, with one change

The randomization condition and the 10% condition read exactly as they did for intervals:

  • Randomization condition: the data come from two independent random samples or a randomized experiment.
  • 10% condition: when sampling without replacement, n1≤0.10N1 and n2≤0.10N2. It is unnecessary when the data come from a randomized experiment.

The normality condition is where the test differs. The test is carried out assuming H0 is true, so it uses the combined, or pooled, proportion of successes:

p^c=n1p^1+n2p^2n1+n2=total successestotal sample size

Then n1p^c, n1(1−p^c), n2p^c and n2(1−p^c) must all be at least 10. For the gym survey, p^c=0.3375 gives 33.75, 66.25, 101.25 and 198.75, so the condition is met.

Sort it

Which procedure fits each study? Tap a card, then tap its bin.

One-sample z-test for a proportion

Two-sample z-test for a difference

Two-sample z-interval for a difference

Worked exampleSetting up a test about texting and driving

A state transportation agency selects a random sample of 250 licensed drivers under age 25 and a separate random sample of 400 licensed drivers age 25 or older from its license records. In anonymous surveys, 70 of the younger drivers and 76 of the older drivers admit to texting while driving in the past month. The agency wants to know whether the data give convincing evidence that the proportion who text while driving is higher for drivers under 25. Set up the appropriate test, but do not carry it out.

  1. Identify the procedure and parameters. A two-sample z-test for the difference between two population proportions. Let pY be the proportion of all licensed drivers in the state under 25 who texted while driving in the past month, and pO the proportion of all licensed drivers in the state 25 or older who did.

  2. State the hypotheses. The agency asks whether the younger drivers' proportion is higher, so the test is one-sided: H0:pY=pO and Ha:pY>pO. Written as a difference, H0:pY−pO=0 and Ha:pY−pO>0.

  3. Check randomization and 10%. The data come from two independent random samples. The state has far more than 2,500 licensed drivers under 25 and far more than 4,000 drivers 25 or older, so each sample is no more than 10% of its population.

  4. Check normality with the pooled proportion. p^c=70+76250+400=146650≈0.2246. Then 250(0.2246)≈56.2, 250(0.7754)≈193.8, 400(0.2246)≈89.8 and 400(0.7754)≈310.2, all at least 10.

Answer.

All three conditions are met, so the two-sample z-test with Ha:pY>pO is appropriate. The next lesson carries out a test like this one.

Check your understanding

1

A researcher wants to know whether the proportion of adults who get a flu shot is lower in a state's rural counties than in its urban counties. She takes a random sample of adults from the rural counties and a separate random sample from the urban counties. Let pR and pU be the proportions of all adults in the rural and urban counties who get a flu shot. Which hypotheses fit her question?

2

In an experiment, 24 of 120 seedlings given a new fertilizer and 15 of 100 seedlings given the standard fertilizer wilt within a week. A two-sample z-test will test whether the proportions that wilt differ. Which is the correct check of the normality condition for this test?

3

A city wants to know whether the proportion of residents who support a new bike lane differs between residents who own a car and residents who do not. It takes a random sample from each group. Which significance test should the city use?

4

A researcher randomly assigns 200 volunteers to use either a new sleep app or a standard alarm for two weeks, then records whether each volunteer reports feeling rested. Before running a two-sample z-test, which statement about the conditions is correct?

5

Why does the normality condition for a two-sample z-test use the pooled proportion p^c instead of the two separate sample proportions?

Practice

Practice until it is automatic

Some problems ask for the pooled proportion a test uses. Others ask for the standard error of an interval from topic 3.10, which does not pool.

Comparing two proportions practice page

Course alignment, for teachers

AP Statistics topic 3.12, Unit 3: Inference for Categorical Data: Proportions.

  • Skill 2.C: Identify appropriate statistical inference methods.
  • Skill 2.E: Identify the null and alternative hypotheses.
  • Skill 4.E: Justify the use of a chosen statistical inference method by verifying conditions.