Unit 4 · Topic 4.8 · about 20 minutes

Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means

Interpret a two-sample t-interval and its confidence level in context, and use the interval to justify or reject a claim about two population means.

Predict first

A garden center compares two fertilizers in a randomized experiment. A 95% confidence interval for μA−μB, the difference in mean tomato yield per plant (Fertilizer A minus Fertilizer B), is (-0.18, 0.46) kilograms. Is there convincing evidence that the two fertilizers produce different mean yields?

Reading the sign of the interval

An interval for μ1−μ2 is about a difference, so where it sits relative to 0 carries the message.

  • Both endpoints positive. Every plausible value has μ1 larger. The interval gives convincing evidence that μ1>μ2, by somewhere between the two endpoints.
  • Both endpoints negative. Every plausible value has μ1 smaller. The interval gives convincing evidence that μ2 is larger.
  • One endpoint on each side of 0. A difference of 0 is a plausible value, so there is not convincing evidence of a difference between the population means.

The order of subtraction decides which of the first two you are in. An interval of (−3.1, −0.4) for Brand A minus Brand B says exactly what (0.4, 3.1) for Brand B minus Brand A says. Reading the sign without checking the order is how correct arithmetic turns into a wrong conclusion.

Sort it

Each card is a 95% confidence interval for μA−μB. Tap a card, then tap the conclusion it supports.

Mean A is larger

Mean B is larger

No convincing difference

Two statements, two meanings

The intervals come from samples, so any one of them may or may not contain the true difference in population means. As with one mean in Topic 4.3, there are two separate things to interpret.

The interval. We are C% confident that the interval from a to b contains the true difference in the population means, described with both populations and the order of subtraction.

The confidence level. In repeated random sampling with the same sample sizes from the same populations, about C% of the intervals created would capture the true difference between the two population means.

Worked exampleIs self-checkout faster?

A grocery chain compares checkout times for orders of 10 to 20 items. From its records it takes independent random samples of 50 self-checkout transactions and 45 cashier transactions. A 95% confidence interval for μS−μC, the difference in mean checkout time in seconds (self-checkout minus cashier), is (-4.1, 18.7). The marketing team wants to advertise that self-checkout is faster on average. Interpret the interval and the confidence level, and decide whether the interval supports the team's claim.

  1. Interpret the interval. We are 95% confident that the interval from -4.1 to 18.7 seconds contains the true difference in mean checkout time (self-checkout minus cashier) for all of the chain's orders of 10 to 20 items.

  2. Interpret the level. If the chain took many pairs of random samples of 50 and 45 transactions and built a 95% interval from each pair, about 95% of those intervals would capture the true difference in mean checkout time.

  3. Translate the claim. Self-checkout faster on average means μS<μC, so μS−μC<0.

  4. Judge it. The interval contains negative values, 0 and positive values. Because 0 is plausible, there is not convincing evidence of any difference in mean checkout time, so the interval does not support the claim. The interval also reaches up to 18.7 seconds, so self-checkout being slower on average is plausible too. In fact most of the interval, and its center of 7.3 seconds, lie on the slower side.

Answer.

The interval does not support the advertising claim. A difference of 0 is plausible, and so is self-checkout being slower on average.

When the interval does support a claim

In the note-taking experiment from Topic 4.7, the 95% interval for μH−μL (by hand minus laptop) was about (0.43, 5.82) points. The whole interval is above 0, so there is convincing evidence that students like those in the study have a higher mean recall score when they take notes by hand. Because the methods were randomly assigned, the difference can be credited to the note-taking method.

A claim about the size of the difference needs a closer look. "Hand notes raise mean recall by at least 1 point" is not supported, because plausible values between 0.43 and 1 are inside the interval. "Hand notes raise mean recall by less than 8 points" is supported, because every plausible value is below 8.

Check your understanding

1

Independent random samples of first-year students at two large universities give a 90% confidence interval of (1.2, 4.6) hours for μN−μS, where μN and μS are the mean weekly hours of paid work for all first-year students at North University and South University. Which is a correct interpretation?

2

Independent random samples of two brands of AA batteries give a 95% confidence interval of (-3.1, -0.4) hours for μA−μB, the difference in mean battery life (Brand A minus Brand B). Which conclusion is supported?

3

A randomized experiment compares two study apps. A 95% confidence interval for μ1−μ2, the difference in mean quiz score (App 1 minus App 2) for students like those in the study, is (-2.3, 4.1) points. Which statement is correct?

4

A city compares the mean response times of fire crews in two districts using independent random samples of emergency calls, and reports a 99% confidence interval for μ1−μ2. What does "99% confidence" mean?

5

A school district runs a randomized experiment comparing two reading programs. A 95% confidence interval for μN−μO, the difference in mean reading gain in points (new program minus old program) for students like those in the study, is (2.1, 7.9). The publisher claims the new program raises the mean gain by more than 5 points. What does the interval provide?

Course alignment, for teachers

AP Statistics topic 4.8, Unit 4: Inference for Quantitative Data: Means.

  • Skill 4.F: Interpret results of statistical inference methods.
  • Skill 4.G: Justify a claim based on statistical inference method results.