Unit 4 · Topic 4.7 · about 25 minutes

Constructing a Confidence Interval for the Difference Between Two Population Means

Estimate the difference between two population means from independent samples or a randomized experiment, with conditions checked and the interval interpreted.

Predict first

Three studies compare reaction times. Which one calls for a two-sample t-interval rather than a paired analysis?

Two independent groups

When the data come from two separate groups, either two independent random samples or two groups formed by random assignment in an experiment, there is no pairing to use. The point estimate of μ1−μ2 is the difference in sample means, x‾1−x‾2, and the procedure is the two-sample t-interval for a difference between two population means:

(x‾1−x‾2)±t∗s12n1+s22n2

The square root is the standard error of the difference: the formula from Topic 4.6 with the sample standard deviations s1 and s2 in place of σ1 and σ2. The margin of error is t∗ times that standard error.

The parameter is μ1−μ2, and defining it means naming both populations, the response variable and which group is subtracted from which. For an experiment on volunteers, the populations are people like the ones in the study under each treatment.

Degrees of freedom

The critical value t∗ comes from a t-distribution, but the degrees of freedom are no longer just n−1. Technology computes them from both sample sizes and both standard deviations, and the result is usually not a whole number. Whatever it is, it falls between the smaller of n1−1 and n2−1 at the low end and n1+n2−2 at the high end.

Use the technology value when you can; on a TI-84, 2-SampTInt reports it (answer No to Pooled). If you are working from the printed table, the smaller of n1−1 and n2−1, called the conservative degrees of freedom, gives a slightly larger t∗ and so a slightly wider interval, which errs on the safe side. AP scoring guidelines have accepted either choice when the work shows which one was used.

Three conditions

  • Randomization condition: the data come from two independent random samples or from a randomized experiment.
  • 10% condition: when sampling without replacement, each sample is at most 10% of its population, n1≤0.10N1 and n2≤0.10N2. This condition is unnecessary when the data come from a randomized experiment.
  • Sample data condition: both samples have at least 30 observations, or both population distributions are approximately normal. If either sample has fewer than 30, both sample distributions should be free from strong skewness and outliers.

The last rule catches people: one small sample means you graph both samples, not just the small one.

Recall quiz scores (out of 30) by note-taking method

By handLaptop152025Recall score (points)

Both groups are roughly symmetric with no outliers, so the sample data condition is met even though each group has only 16 students.

Summary statistics for the note-taking experiment
GroupnMeanStandard deviation
By hand1620.8753.649
Laptop1617.7503.804

Worked exampleHand notes versus laptop notes

A psychology teacher recruits 32 student volunteers and randomly assigns 16 to take notes by hand and 16 to take notes on a laptop during the same 15-minute video lecture. A week later, all 32 take a 30-point recall quiz without their notes. The boxplots and summary statistics are above. Construct and interpret a 95% confidence interval for the difference in mean recall score.

  1. Name it. Two-sample t-interval for μH−μL, where μH is the true mean recall score for students like these who take notes by hand and μL is the true mean recall score for students like these who take notes on a laptop.

  2. Check conditions. Random: the volunteers were randomly assigned to the two methods. 10%: not needed, because this is a randomized experiment. Sample data: both groups have fewer than 30 students, so both boxplots must be free of strong skewness and outliers, and they are.

  3. Calculate. x‾H−x‾L=20.875−17.750=3.125 and SE=3.649216+3.804216≈1.318. Technology gives df≈29.95 and t∗≈2.042, so the margin of error is about 2.69 and the interval is about (0.43, 5.82) points. With the conservative df=15, t∗=2.131 and the interval is about (0.32, 5.93).

  4. Interpret in context. We are 95% confident that the interval from 0.43 to 5.82 points contains the true difference in mean recall score (by hand minus laptop) for students like those in this study.

Answer.

About 0.43 to 5.82 points, for the mean recall score by hand minus the mean recall score on a laptop. Topic 4.8 takes up what this interval says about the claim that the method matters.

Check your understanding

1

A nutrition researcher selects independent random samples of 35 adults from City A and 40 adults from City B and records each adult's daily sugar intake in grams. She wants to estimate how much the mean daily sugar intake differs between the two cities. Which procedure is appropriate?

2

Independent random samples give n1=18, s1=4.2 and n2=22, s2=5.6. Find the standard error of x‾1−x‾2. Round to three decimal places.

3

A two-sample t-interval is built from independent random samples of sizes n1=15 and n2=22. Which value could be the degrees of freedom reported by technology?

4

In an experiment, 38 volunteers are randomly assigned to a new allergy medicine (18 people) or a standard one (20 people), and the hours of symptom relief are recorded. A boxplot for the new-medicine group is roughly symmetric, but the standard group's boxplot is strongly skewed to the right with two high outliers. Which statement about the conditions for a two-sample t-interval is correct?

5

A consumer group tests independent random samples of Brand X and Brand Y AA batteries from store shelves and records how many hours each one lasts in a digital camera. It builds a two-sample t-interval for μX−μY. Which is the best definition of the parameter?

Practice

Practice until it is automatic

New numbers every time. Each one is checked the moment you answer, with the full working shown.

Comparing two means practice page

Course alignment, for teachers

AP Statistics topic 4.7, Unit 4: Inference for Quantitative Data: Means.

  • Skill 2.C: Identify appropriate statistical inference methods.
  • Skill 3.E: Calculate appropriate statistical inference method results.
  • Skill 4.E: Justify the use of a chosen statistical inference method by verifying conditions.