Unit 4 · Topic 4.2 · about 30 minutes

Constructing a Confidence Interval for a Population Mean or Population Mean Difference

Estimate a population mean, or the mean difference in paired data, with a t-interval you can justify and interpret in context.

Predict first

In Topic 3.3, a 95% interval for a proportion used the critical value z∗=1.960. For a mean you almost never know the population standard deviation σ, so you estimate it with the sample standard deviation s. With s standing in for σ, what should happen to the critical value for 95% confidence?

t-distributions

When you replace σ with s, the standardized value x‾−μs/n no longer follows the standard normal curve. It follows a t-distribution, also called Student's t-distribution.

The t-distributions are a family of symmetric, bell-shaped, standardized curves centered at 0. Each one is identified by its degrees of freedom (df), which depend on the sample size. For one sample, df=n−1. Compared with the standard normal curve, a t-distribution has a lower, narrower peak and fatter tails, so more of its area sits far from 0. With few degrees of freedom the difference is large. As the degrees of freedom increase, the t-distribution looks more and more like the standard normal curve, and the critical values close in on the familiar z∗.

Critical values t∗ for 95% confidence shrink toward z∗=1.960 as the sample grows
Sample size nDegrees of freedomt∗ for 95% confidence
324.303
652.571
11102.228
21202.086
31302.042
1211201.980
Standard normalnone1.960

The one-sample t-interval

To estimate a population mean μ when σ is unknown, use the one-sample t-interval for a population mean:

x‾±t∗sn

The point estimate is the sample mean x‾. The standard error, SE=s/n, estimates how far sample means typically land from μ. It is the σ/n formula from Topic 4.1, with s in place of σ. The margin of error is t∗×SE, where t∗ is the critical value for the central C% of the t-distribution with df=n−1. On a TI-84, invT(0.975, df) gives t∗ for 95% confidence; the printed t table gives the same values.

State the parameter in context: the mean of the response variable for the population you sampled. "The mean caffeine content of all 16-ounce cold brews sold by the chain" names a parameter. "The mean" does not.

Three conditions

A one-sample t-interval for a population mean requires three conditions:

  • Randomization condition: the data come from a random sample or a randomized experiment.
  • 10% condition: when sampling without replacement, n≤0.10N.
  • Sample data condition: the population distribution is approximately normal, or n≥30, or, if n<30, the sample data are free from strong skewness and outliers.

The third one is new. With proportions you counted successes and failures. With means you look at the data. For a small sample, graph it with a dotplot, boxplot or histogram and look for strong skew or outliers. Mild unevenness is fine, because small samples from normal populations rarely look perfectly symmetric. An outlier or a long tail means the condition is not met, and the interval may not deserve the confidence level printed next to it.

Caffeine in 12 randomly chosen 16-ounce cold brews

190200210220230Caffeine (mg)

No strong skew and no outliers, so the sample data condition is met even though n<30.

Worked exampleA 95% interval for mean caffeine

A food-safety lab buys one 16-ounce cold brew at each of 12 randomly selected stores in a large coffee chain and measures its caffeine. The dotplot above shows the results: x‾=210.92 mg and s=13.33 mg. Construct and interpret a 95% confidence interval for the mean caffeine content of all 16-ounce cold brews sold by the chain.

  1. Name it. One-sample t-interval for μ, the mean caffeine content, in mg, of all 16-ounce cold brews sold by the chain.

  2. Check conditions. Random: the cups came from 12 randomly selected stores. 10%: 12 cups is far less than 10% of all the 16-ounce cold brews the chain sells. Sample data: n=12<30, but the dotplot shows no strong skewness and no outliers.

  3. Calculate. df=12−1=11, so t∗=2.201. SE=13.3312≈3.848 mg, so the margin of error is 2.201(3.848)≈8.47 mg. The interval is 210.92±8.47, which is about (202.45, 219.39) mg.

  4. Interpret in context. We are 95% confident that the interval from 202.45 mg to 219.39 mg contains the true mean caffeine content of all 16-ounce cold brews sold by the chain.

Answer.

About 202.45 mg to 219.39 mg. The interval estimates the mean for all of the chain's 16-ounce cold brews, not the caffeine in any single cup.

Paired data: one sample of differences

Sometimes each value in one set is matched with exactly one value in the other, such as the same student measured before and after a program, or twins split between two treatments. This is a matched pairs design, and the two sets of values are dependent. Do not treat them as two separate samples. Subtract within each pair to get one sample of differences, and use the one-sample t-interval for a population mean difference:

x‾d±t∗sdn

Here x‾d and sd are the mean and standard deviation of the differences, n is the number of pairs, and df=n−1. The parameter is μd, the population mean difference, and you must say which way you subtracted: after minus before is a different parameter from before minus after. The three conditions are checked on the differences, so for the sample data condition you graph the differences, not the two original lists.

Reading speed, in words per minute, for 10 randomly selected students in a district summer reading program
StudentBeforeAfterDifference (after minus before)
118219614
22052149
316718518
42212265
519420713
617619014
72102133
818820113
919921617
101721819

Differences in reading speed, after minus before

51015Difference (words per minute)

The 10 differences show no outliers and no strong skew, so the sample data condition is met for the differences.

Worked exampleA paired interval for the mean gain

The table shows reading speeds for a random sample of 10 of the roughly 600 students in a district's summer reading program, before and after the program. Construct and interpret a 95% confidence interval for the mean change in reading speed.

  1. Name it. One-sample t-interval for μd, the mean difference in reading speed (after minus before), in words per minute, for all students in the district's summer program.

  2. Check conditions. Random: the 10 students were randomly selected. 10%: 10 is less than 10% of the 600 students. Sample data: there are only 10 differences, and their boxplot shows no strong skewness and no outliers.

  3. Calculate. The differences have x‾d=11.5 and sd≈4.905. With df=9, t∗=2.262, so the interval is 11.5±2.262(4.90510)=11.5±3.51, which is about (7.99, 15.01).

  4. Interpret in context. We are 95% confident that the interval from 7.99 to 15.01 words per minute contains the true mean difference in reading speed (after minus before) for all students in the district's summer program.

Answer.

About 7.99 to 15.01 words per minute, for the population mean of after minus before. Topic 4.3 takes up what this interval can and cannot support as a claim.

Check your understanding

1

How does a t-distribution with 4 degrees of freedom compare with the standard normal distribution?

2

A food scientist measures the sodium content of a random sample of 16 cans from a large production run of soup. The sample shows no strong skew or outliers. Which critical value should she use for a 90% confidence interval for the mean sodium content of all cans in the run?

3

A researcher wants a confidence interval for the mean commute distance of employees at a large company. She records the distances for a random sample of 14 employees. A dotplot of the 14 distances shows most values between 3 and 15 miles and one value at 62 miles. Which statement about the conditions is correct?

4

A sports scientist measures the vertical jump of 15 randomly selected players from a large volleyball league, once in their usual shoes and once in a new shoe, in random order. She wants to estimate the mean improvement from the new shoe. Which procedure and parameter fit?

5

A random sample of 20 households in a large city used a mean of 312 gallons of water per day, with standard deviation 64 gallons. The sample shows no strong skew or outliers. Find the margin of error for a 95% confidence interval for the mean daily water use of all households in the city. Round to one decimal place.

Practice

Practice until it is automatic

New numbers every time. Each one is checked the moment you answer, with the full working shown.

Confidence intervals for a mean practice page

Course alignment, for teachers

AP Statistics topic 4.2, Unit 4: Inference for Quantitative Data: Means.

  • Skill 2.C: Identify appropriate statistical inference methods.
  • Skill 3.E: Calculate appropriate statistical inference method results.
  • Skill 4.C: Describe distributions and compare relative positions of points within a distribution.
  • Skill 4.E: Justify the use of a chosen statistical inference method by verifying conditions.