Unit 4 · Topic 4.5 · about 30 minutes

Carrying Out a Test for a Population Mean or Population Mean Difference

Calculate the test statistic and p-value for a one-sample or paired t-test, interpret the p-value, and state a justified conclusion in context.

Predict first

A one-sample t-test on 10 observations gives a test statistic of 2.1, so df=9, with a one-sided alternative. If you found the p-value from the standard normal curve instead of the t-distribution, how would the two p-values compare?

The test statistic and the p-value

The test statistic for a one-sample t-test measures how far the sample mean is from the null value, in standard errors:

t=x‾−μ0s/n,df=n−1

When the null hypothesis is true, this statistic has a t-distribution with n−1 degrees of freedom. For matched pairs, use the mean and standard deviation of the differences, x‾d and sd, with μ0=0 and n equal to the number of pairs.

The p-value is the probability of getting a test statistic as extreme as the one observed, or more extreme, in the direction of the alternative hypothesis, assuming the null hypothesis is true:

  • for Ha:μ>μ0, the area to the right of t;
  • for Ha:μ<μ0, the area to the left of t;
  • for Ha:μ≠μ0, the area beyond t in both tails, which is twice the one-tail area.

Technology gives the area directly (tcdf on a TI-84, or a T-Test that does the whole calculation). The printed t table only brackets the p-value between two columns, which is still enough to compare it with α.

Wait times for 18 randomly selected urgent care visits

15202530Wait time (minutes)

No strong skew and no outliers, so the sample data condition is met for these 18 values.

Worked exampleIs the wait longer than advertised?

An urgent care clinic advertises an average wait of 20 minutes. A local reporter suspects the true mean wait is longer. She randomly selects 18 of the more than 1,000 visits in the clinic's records for the past three months; their wait times are in the dotplot above, with x‾=22.94 minutes and s=5.84 minutes. Do the data give convincing statistical evidence, at α=0.05, that the mean wait is more than 20 minutes?

  1. Procedure and hypotheses. One-sample t-test for a population mean, where μ = the mean wait time, in minutes, of all visits to the clinic in the past three months. H0:μ=20 and Ha:μ>20.

  2. Conditions. Random: the 18 visits were randomly selected from the clinic's records. 10%: 18 is less than 10% of the more than 1,000 visits. Sample data: n=18<30, and the dotplot shows no strong skewness and no outliers.

  3. Calculate. t=22.94−205.84/18=2.941.377≈2.14 with df=17. The p-value is the area to the right of 2.14 under the t-distribution with 17 degrees of freedom, about 0.024.

  4. Conclude. Because the p-value of 0.024 is less than α=0.05, reject H0. There is convincing statistical evidence that the true mean wait time of all visits to the clinic in the past three months is more than 20 minutes.

Answer.

t≈2.14 with a p-value of about 0.024. Reject H0: there is convincing evidence that the mean wait is longer than the advertised 20 minutes.

What the p-value says, and what it does not

A p-value measures how surprising the data would be if the null hypothesis were true. In context: assuming the true mean wait time is 20 minutes, there is about a 0.024 probability of getting a test statistic of 2.14 or larger in a random sample of 18 visits.

It is not the probability that the null hypothesis is true, and 1 minus the p-value is not the probability that the alternative is true. The whole calculation starts by assuming H0, so it cannot turn around and grade H0.

The formal decision compares the p-value with the significance level α, which is chosen before the data are collected. If the p-value is less than or equal to α, reject H0: there is convincing statistical evidence for the alternative. If the p-value is greater than α, fail to reject H0: there is not convincing statistical evidence for the alternative. Failing to reject never shows that H0 is true.

A conclusion stated this way is the statistical reasoning behind the answer to the question that started the study. The reporter asked whether visits to this clinic over the past three months averaged more than 20 minutes, and the test gives convincing evidence that they did, with a p-value of 0.024 as the reason.

Worked exampleA paired test that falls short

A coach randomly selected 14 of the more than 300 members of a running club. Each ran 200 meters once after a sports drink and once after water, in random order. For each runner, the difference is water time minus sports-drink time, so a positive difference means faster after the drink. The 14 differences have x‾d=0.110 seconds and sd=0.287 seconds, and their dotplot shows no strong skew or outliers; the full setup is in Topic 4.4. At α=0.05, is there convincing evidence that 200-meter times are lower after the sports drink, on average?

  1. Procedure and hypotheses. One-sample t-test for a population mean difference, where μd = the mean of (water time minus sports-drink time), in seconds, for all members of the club. H0:μd=0 and Ha:μd>0.

  2. Conditions. Random: the runners were randomly selected, and the order of the drinks was randomly assigned. 10%: 14 is less than 10% of the more than 300 members. Sample data: only 14 differences, but their dotplot shows no strong skewness and no outliers.

  3. Calculate. t=0.110−00.287/14=0.1100.0767≈1.43 with df=13. The p-value, the area to the right of 1.43, is about 0.088.

  4. Conclude. Because the p-value of 0.088 is greater than α=0.05, fail to reject H0. There is not convincing statistical evidence that the true mean of (water time minus sports-drink time) is greater than 0 for members of the club. In plain words, the data do not convincingly show that times are lower after the drink, on average.

Answer.

Fail to reject H0. The sampled runners were a little faster after the drink, but a mean difference of 0.110 seconds would not be unusual if the drink made no difference at all.

Check your understanding

1

A random sample of 25 bags of a brand of trail mix has a mean of 15.6 ounces and a standard deviation of 0.8 ounces. For a test of H0:μ=16 against Ha:μ<16, find the test statistic. Round to two decimal places.

2

A transit agency tests H0:μ=4 against Ha:μ>4, where μ is the mean morning delay, in minutes, of all its buses. A random sample of 30 buses gives a p-value of 0.012. Which is a correct interpretation of the p-value?

3

A granola bar label says 190 calories. A lab measures a random sample of 15 bars from a large batch and tests H0:μ=190 against Ha:μ≠190, where μ is the mean calorie content of all bars in the batch. The p-value is 0.072. Which conclusion is appropriate at α=0.05?

4

A one-sample t-test with 21 degrees of freedom has a two-sided alternative hypothesis and a test statistic of t=1.85. The area to the right of 1.85 under the t-distribution with 21 degrees of freedom is 0.039. What is the p-value?

5

A researcher measures the reaction times of 16 randomly selected drivers twice, once when rested and once after a night of short sleep, and tests Ha:μd>0, where each difference is (short-sleep time minus rested time). The test statistic is t=2.31. How should the p-value be found?

Practice

Practice until it is automatic

New numbers every time. Each one is checked the moment you answer, with the full working shown.

Significance tests for a mean practice page

Course alignment, for teachers

AP Statistics topic 4.5, Unit 4: Inference for Quantitative Data: Means.

  • Skill 3.E: Calculate appropriate statistical inference method results.
  • Skill 4.F: Interpret results of statistical inference methods.
  • Skill 4.G: Justify a claim based on statistical inference method results.