Login
📚 Mostly Harmless Statistics
Chapters ▾
⇩ Download ▾

8.3 Hypothesis Test for One Mean

There are three methods used to test hypotheses:

The Traditional Method (Critical Value Method)

There are five steps in hypothesis testing when using the traditional method:

  1. Identify the claim and formulate the hypotheses.
  2. Compute the test statistic.
  3. Compute the critical value(s) and state the rejection rule (the rule by which you will reject the null hypothesis (H0).
  4. Make the decision to reject or not reject the null hypothesis by comparing the test statistic to the critical value(s). Reject H0 when the test statistic is in the critical tail(s).
  5. Summarize the results and address the claim using context and units from the research question.

Steps ii and iii do not have to be in that order so make sure you know the difference between the critical value, which comes from the stated significance level α, and the test statistic, which is calculated from the sample data.

Note: The test statistic and the critical value(s) come from the same distribution and will usually have the same letter such as z, t, or F. The critical value(s) will have a subscript with the lower tail area (zα,z1α,zα/2) or an asterisk next to it (z*) to distinguish it from the test statistic.

You can find the critical value(s) or test statistic in any order, but make sure you know the difference when you compare the two. The critical value is found from α and is the start of the shaded area called the critical region (also called rejection region or area). The test statistic is computed using sample data and may or may not be in the critical region.

The critical value(s) is set before you begin (a priori) by the level of significance you are using for your test. This critical value(s) defines the shaded area known as the rejection area. The test statistic for this example is the z-score we find using the sample data that is then compared to the shaded tail(s). When the test statistic is in the shaded rejection area, you reject the null hypothesis. When your test statistic is not in the shaded rejection area, then you fail to reject the null hypothesis. Depending on if your claim is in the null or the alternative, the sample data may or may not support your claim.

The P-value Method

Most modern statistics and research methods utilize this method with the advent of computers and graphing calculators.

There are five steps in hypothesis testing when using the p-value method:

  1. Identify the claim and formulate the hypotheses.
  2. Compute the test statistic.
  3. Compute the p-value.
  4. Make the decision to reject or not reject the null hypothesis by comparing the p-value with α. Reject H0 when the p-value ≤ α.
  5. Summarize the results and address the claim.

The ideas below review the process of evaluating hypothesis tests with p-values:

The Confidence Interval Method (results are in the same units as the data)

There are four steps in hypothesis testing when using the confidence interval method:

  1. Identify the claim and formulate the hypotheses.
  2. Compute confidence interval.
  3. Make the decision to reject or not reject the null hypothesis by comparing the p-value with α. Reject H0 when the hypothesized value found in H0 is outside the bounds of the confidence interval. We only will be doing a two-tailed version of this.
  4. Summarize the results and address the claim.

For all 3 methods, Step i is the most important step. If you do not correctly set up your hypotheses then the next steps will be incorrect.

The decision and summary would be the same no matter which method you use. Figure 8-12 is a flow chart that may help with starting your summaries, but make sure you finish the sentence with context and units from the question.

Flow chart for writing test summaries. When the claim is in H0, reject H0? yes leads to there is enough evidence to reject the claim that…, and no leads to there is not enough evidence to reject the claim that…. When the claim is in H1, reject H0? yes leads to there is enough evidence to support the claim that…, and no leads to there is not enough evidence to support the claim that….

Figure 8-12

The hypothesis-testing framework is a very general tool, and we often use it without a second thought. If a person makes a somewhat unbelievable claim, we are initially skeptical. However, if there is sufficient evidence that supports the claim, we set aside our skepticism and reject the null hypothesis in favor of the alternative.

8.3.1 Z-Test

When the population standard deviation is known and stated in the problem, we will use the z-test.

Example 8-5 used the traditional critical value method. With the onset of computers, this method is outdated and the p-value and confidence interval methods are becoming more popular.

Most statistical software packages will give a p-value and confidence interval but not the critical value.

TI-84: Press the [STAT] key, go to the [TESTS] menu, arrow down to the [Z-Test] option and press the [ENTER] key. Arrow over to the [Stats] menu and press the [ENTER] key. Then type in value for the hypothesized mean (µ0), standard deviation, sample mean, sample size, arrow over to the , <, > sign that is in the alternative hypothesis statement then press the [ENTER] key, arrow down to [Calculate] and press the [ENTER] key. Alternatively (If you have raw data in a list) Select the [Data] menu and press the [ENTER] key. Then type in the value for the hypothesized mean (µ0), type in your list name (TI-84 L1 is above the 1 key).

TI-84 Z-Test input screen with Stats selected: μ0 = .8535, σ = .06, x̄ = .8472, n = 50, alternative hypothesis μ < μ0 highlighted, with Calculate and Draw options below.

Press the [STAT] key, go to the [TESTS] menu, arrow down to either the [Z-Test] option and press the [ENTER] key. Arrow over to the [Stats] menu and press the [ENTER] key. Then type in value for the hypothesized mean (µ0), standard deviation, sample mean, sample size, arrow over to the , <, > sign that is in the alternative hypothesis statement then press the [ENTER] key, arrow down to [Calculate] and press the [ENTER] key. Alternatively (If you have raw data in a list) Select the [Data] menu and press the [ENTER] key. Then type in the value for the hypothesized mean (µ0), type in your list name (TI-84 L1 is above the 1 key).

TI-84 Z-Test results screen: μ < .8535, z = −.7424621202, p = .2289036235, x̄ = .8472, n = 50.

The calculator returns the alternative hypothesis (check and make sure you selected the correct sign), the test statistic, p-value, sample mean, and sample size.

TI-89: Go in to the Stat/List Editor App. Select [F6] Tests. Select the first option Z-Test. Select Data if you have raw data in a list, select Stats if you have the summarized statistics given to you in the problem. If you have data, press [2nd] Var-Link, the go down to list1 in the main folder to select the list name. If you have statistics then enter the values. Leave Freq:1 alone, arrow over to the , <, > sign that is in the alternative hypothesis statement then press the [ENTER]key, arrow down to [Calculate] and press the [ENTER] key. The calculator returns the test statistic and the p-value.

Four TI-89 screens for a z-test: the Tests menu with 1:Z-Test highlighted, a Choose Input Method dialog with Stats selected, the Z Test input dialog with μ0 = .8535, σ = .06, x̄ = .8472, n = 50 and alternate hypothesis choices μ ≠ μ0, μ < μ0, μ > μ0, and the results screen showing μ < μ0, z = −.742462, P Value = .228904, x̄ = .8472, n = 50, σ = .06.

What is the p-value?

The p-value is the probability of observing an effect as least as extreme as in your sample data, assuming that the null hypothesis is true. The p-value is calculated based on the assumptions that the null hypothesis is true for the population and that the difference in the sample is caused entirely by random chance.

Recall the example at the beginning of the chapter.

8.3.2 T-Test

When the population standard deviation is unknown, we will use the t-test.

Z Versus T

The z and t-tests are easy to mix up. Sometimes a standard deviation will be stated in the problem without specifying if it is a population’s standard deviation σ or the sample standard deviation s. If the standard deviation is in the same sentence that describes the sample or only raw data is given then this would be s. When you only have sample data, use the t-test.

Figure 8-16 is a flow chart to remind you when to use z versus t.

Flow chart asking is σ known? Yes leads to use the z sub α/2 values and σ in the formula; no leads to use the t sub α/2 values and s in the formula. A footnote reads: if n < 30, the variable must be normally distributed.

Figure 8-16

Use Figure 8-17 as a guide in setting up your hypotheses. The two-tailed test will always have a not equal ≠ sign in H1 and both tails shaded. The right-tailed test will always have the greater than > sign in H1 and the right tail shaded. The left-tailed test will always have a less than < sign in H1 and the left tail shaded.

Hypothesis testing table with three columns: two-tailed test (H0: μ = μ0, H1: μ ≠ μ0, normal curve with both tails shaded), right-tailed test (H1: μ > μ0, right tail shaded), and left-tailed test (H1: μ < μ0, left tail shaded). Below, phrases when the claim is in the null hypothesis: = (is equal to, is exactly the same as, has not changed from), ≤ (is less than or equal to, is at most, is not more than, within), ≥ (is greater than or equal to, is at least, is not less than); and when the claim is in the alternative hypothesis: ≠ (is not, is not equal to, is different from, has changed from), > (more than, greater than, above, higher than, longer than, bigger than, increased), < (less than, below, lower than, shorter than, smaller than, decreased, reduced).

Figure 8-17

One-Tailed Versus Two-Tailed Tests

Most software packages do not ask which tailed test you are performing. Make sure you look at the sign in the alternative hypothesis to and determine which p-value to use. The difference is just what part of the picture you are looking at. In Excel, the critical value shown is for a one-tail test and does not specify left or right tail. The critical value in the output will always be positive, it is up to you to know if the critical value should be a negative or positive value. For example, Figures 8-21, 8-22, and 8-23 uses df = 9, α = 0.10 to show all three tests comparing either the test statistic with the critical value or the p-value with α.

Two-Tailed Test

The test statistic can be negative or positive depending on what side of the distribution it falls; however, the p-value is a probability and will always be a positive number between 0 and 1. See Figure 8-21.

t-distribution curve for a two-tailed test labeled total p-value area = 0.28268 and α = 0.10. Blue cross-hatched areas beyond the test statistics TS = −1.143 and TS = 1.143 each represent p-value/2 = 0.14134, and green areas beyond the critical values CV = −1.833 and CV = 1.833 each represent α/2 = 0.05.

Figure 8-21

Right-Tailed Test

If we happened to do a right-tailed test with df = 9 and α = 0.10, the critical value t1-α = 1.383 will be in the right tail and usually the test statistic will be a positive number. See Figure 8-22.

t-distribution curve for a right-tailed test: the blue cross-hatched right-tail area beyond the test statistic TS = 1.143 is the p-value = 0.14134, and the green area beyond the critical value CV = 1.383 is α = 0.10.

Figure 8-22

Left-Tailed Test

If we happened to do a left-tailed test with df = 9 and α = 0.10, the critical value tα = –1.383 will be in the left tail and usually the test statistic will be a negative number. See Figure 8-23.

t-distribution curve for a left-tailed test: the blue cross-hatched left-tail area beyond the test statistic TS = −1.143 is the p-value = 0.14134, and the green area beyond the critical value labeled CV = −1.833 is α = 0.10.

Figure 8-23

Adapted from Mostly Harmless Statistics by Rachel Webb (Portland State University), hosted on LibreTexts (stats.libretexts.org) and licensed under CC BY-SA 4.0. Changes were made. License: CC-BY-SA-4.0.