Unit 3 · Topic 3.8 · about 20 minutes
Potential Errors When Performing Tests
Identify and interpret Type I and Type II errors in context, find their probabilities from the significance level and the power, and explain what changes those probabilities.
Predict first
A quality engineer tests whether a machine's defect rate has risen above its usual level. She can use a significance level of 0.01 or 0.10, with the same sample size either way. If the defect rate really has risen, which choice makes her test more likely to miss it?
Two ways to be wrong
A test ends in one of two decisions, and the truth is one of two things, so there are four possible outcomes. Two of them are mistakes.
A Type I error occurs when the test finds convincing statistical evidence that the alternative hypothesis is true, because the p-value was small, but the alternative is not true. A Type II error occurs when the test does not find convincing statistical evidence for the alternative hypothesis, because the p-value was not small, but the alternative is true.
Neither error means someone made an arithmetic slip. Both happen with perfect calculations, because random samples sometimes mislead. A fair coin can land heads 15 times in 20 flips, and a test that then calls the coin unfair has made a Type I error without anyone doing anything wrong.
| Decision | is true | is true |
|---|---|---|
| Reject | Type I error | Correct decision |
| Fail to reject | Correct decision | Type II error |
How likely each error is
The probability of a Type I error is the significance level, . The researcher sets it before collecting any data, usually at a small value such as 0.01, 0.05 or 0.10.
The power of a test is the probability that it correctly rejects a false null hypothesis: the top right cell of the table. The probability of a Type II error is . Unlike , power is not one fixed number. It depends on the true value of the parameter, because a big departure from the null value is easier to detect than a small one. A well-planned study keeps the probability of a Type II error small and the power large, for example and power .
What raises power
Power goes up, and the probability of a Type II error goes down, when any one of these changes while the others stay the same:
- the sample size increases;
- the standard error decreases;
- the true value of the parameter is farther from the null value;
- the significance level increases.
The last one has a cost, since a larger is a larger probability of a Type I error. Of the things a researcher controls, a larger sample is the way to lower the Type II error probability without raising the Type I error probability, and it is also how the standard error comes down. The table shows the pattern for one test.
| Sample size | True | True |
|---|---|---|
| 50 | 0.17 | 0.41 |
| 100 | 0.26 | 0.64 |
| 200 | 0.41 | 0.89 |
| 400 | 0.64 | 0.99 |
Read down a column and the power climbs with the sample size. Read across a row and a true proportion of 0.60 is easier to detect than 0.55, because it is farther from 0.50. The significance level matters too: with 100 observations and a true proportion of 0.60, the power is 0.37 at , 0.64 at and 0.77 at . You do not need to calculate values like these. Read them for the pattern.
Worked exampleErrors in a drug safety test
A drug company will apply to sell a new allergy medicine only if it finds convincing evidence that fewer than 2% of patients who take it have a serious side effect. It tests against , where is the proportion of all patients taking the medicine who have a serious side effect. Describe each type of error and its consequence, decide which is more serious, and give the probability of each error if the company uses and the test has power 0.85 against a true side effect rate of 1%.
Type I error. The company finds convincing evidence that fewer than 2% of patients have a serious side effect, when in fact the rate is not below 2%. Consequence: a medicine that is not as safe as required goes on sale.
Type II error. The company does not find convincing evidence that fewer than 2% of patients have a serious side effect, when in fact fewer than 2% do. Consequence: a medicine that meets the safety standard is shelved, and patients lose a treatment.
Which is worse. A Type I error puts patients at risk, so it is the more serious one here. Because is the probability of a Type I error, that is an argument for a small significance level such as 0.01.
Probabilities. . If the true side effect rate is 1%, .
A Type I error approves a medicine that is not safe enough; a Type II error shelves one that is. The Type I error is more serious, which argues for a small , and a large enough trial keeps the Type II error probability down at the same time.
Let the consequences set the plan
In some studies a Type I error is the more serious mistake, and in others a Type II error is. Weigh the consequences before the study starts, because they decide two things. The consequences of a Type I error should influence the significance level, since is the probability of that error. The consequences of a Type II error should influence the sample size, since the sample size influences the probability of that error.
Now flip the situation around. An aircraft maintenance team testing whether more than 1% of a shipment of bolts are cracked would worry most about a Type II error: missing a real problem and installing cracked bolts. That points toward a larger and a bigger sample.
Check your understanding
A city will launch a curbside composting program only if it finds convincing evidence that more than 30% of its households would use it. It tests against , where is the proportion of all households in the city that would use the program. Which describes a Type I error?
A city tests against , where is the proportion of all households in the city that would use a curbside composting program. Its survey gives a p-value of 0.012, and at the city rejects . Which error could the city have made?
A test of against uses . The power of the test against a true proportion of 0.40 is 0.72. If the true proportion really is 0.40, what is the probability that the test makes a Type II error?
A researcher plans a test of against with a random sample of 200 and . Which change, with nothing else different, would increase the power of the test?
A water utility tests against , where is the proportion of all homes it serves whose tap water is above the safe limit for lead. If it finds convincing evidence for , it will replace old pipes. Which error is more serious here, and what does that suggest about the significance level?
Course alignment, for teachers
AP Statistics topic 3.8, Unit 3: Inference for Categorical Data: Proportions.
- Skill 2.D: Identify types of errors and relationships among components in statistical inference methods.
- Skill 3.C: Calculate and estimate expected counts, percentages, probabilities, and intervals.
- Skill 4.D: Interpret statistical calculations and results to assess meaning or a claim.