Login
📚 Mostly Harmless Statistics
Chapters ▾

8 Hypothesis Tests for One Population

8.01: Introduction

A statistic is a characteristic or measure from a sample. A parameter is a characteristic or measure from a population. We use statistics to generalize about parameters, known as estimations. Every time we take a sample statistic, we would expect that estimate to be close to the parameter, but not necessarily exactly equal to the unknown population parameter. How close would depend on how large a sample we took, who was sampled, how they were sampled and other factors. Hypothesis testing is a scientific method used to evaluate claims about population parameters.

A statistical hypothesis is an educated conjecture about a population parameter. This conjecture may or may not be true. We will take sample data and infer from the sample if there is evidence to support our claim about the unknown population parameter.

The null hypothesis (H0, pronounced “H-naught” or “H-zero”), is a statistical hypothesis that states that there is no difference between a parameter and a specific value, or that there is no difference between two parameters. The null hypothesis is assumed true until there is sufficient evidence otherwise.

The alternative hypothesis (H1 or Ha, pronounced “H-one” or “H-ā”), is a statistical hypothesis that states that there is a difference between a parameter and a specific value, or that there is a difference between two parameters. H1 is always the complement of H0.

The researcher decides the probability that the test is true by setting the level of significance, also called the significance level. We use the Greek letter α, pronounced “alpha,” to represent the significance level. The level of significance is the probability that the null hypothesis is rejected when it is actually true. Note: like in the previous chapter, 1 – \(\alpha\) is the confidence level.

When doing your own research, you should set up their hypotheses and choose the significance level before analyzing the sample data.

When reading a word problem, your first step is to identify the parameter(s), for example μ, you are testing and which direction (left, right, or two-tail) test you are being asked to perform. For this course, the homework problems will state the researcher’s claim; usually this is the alternative hypothesis.

The null-hypothesis is always set up as a parameter equal to some value (called the test value) or equal to another parameter. The null hypothesis is assumed true unless there is strong evidence from the sample to suggest otherwise. Similar to our judicial system that a person is innocent until the prosecutor shows enough evidence that they are not innocent.

For example, an investment company wants to build a new food cart. They know from experience that food carts are successful if they have on average more than 100 people a day walk by the location. They have a potential site to build on, but before they begin, they want to see if they have enough foot traffic. They observe how many people walk by the site every day over a month. The investors want to be very careful about setting up in a bad location where the food cart will fail, rather than the missed opportunity build in a prime location. We have two hypotheses. For an average of more than 100 people, we would write this in symbols as μ > 100. This claim needs to go into the alternative hypothesis since there is no equality, just strictly greater than 100. The complement of greater than is μ ≤ 100. This has a form of equality (≤) so needs to go in the null hypothesis.

We then would set up the hypotheses as:

  • H0: μ ≤ 100 (Do not build)
  • H1: μ > 100 (Build).

When performing the hypothesis test, the test statistic assumes that the parameter is the null hypothesis equal to some value. This still implies that the parameter could be any value less than or equal to 100 but our hypothesis test should be written as:

  • H0: μ = 100
  • H1: μ > 100

Either notation is fine, but most textbooks will always have the = sign in the null hypothesis. The null hypothesis is based off historical value, a claim or product specification.

Signs are Important

When there is a greater than sign (>) in the alternative hypothesis, we call this a right-tailed test. If we had a less than sign (<) in the alternative hypothesis, then we would have a left-tailed test. If there were a not equal sign (≠) in the alternative hypothesis, we would have a two-tailed test. The tails will determine which side the critical region will fall on the sampling distribution. Note that you should never have an =, ≤ or ≥ sign appear in the alternative hypothesis.

There are three ways to set up the hypotheses for a population mean μ:

Two-tailed test Right-tailed test Left-tailed test

\(\begin{array}{lll}
\mathrm{H}_{0}: \mu=\mu_{0} & \mathrm{H}_{0}: \mu=\mu_{0} & \mathrm{H}_{0}: \mu=\mu_{0} \\
\mathrm{H}_{1}: \mu \neq \mu_{0} & \mathrm{H}_{1}: \mu>\mu_{0} & \mathrm{H}_{1}: \mu<\mu_{0}
\end{array}\)

or

Two-tailed test Right-tailed test Left-tailed test

\(\begin{array}{lll}
\mathrm{H}_{0}: \mu=\mu_{0} & \mathrm{H}_{0}: \mu \leq \mu_{0} & \mathrm{H}_{0}: \mu \geq \mu_{0} \\
\mathrm{H}_{1}: \mu \neq \mu_{0} & \mathrm{H}_{1}: \mu>\mu_{0} & \mathrm{H}_{1}: \mu<\mu_{0}
\end{array}\)

where μ0 is a placeholder for the numeric test value.

  • The null-hypothesis of a two-tailed test states that the mean μ is equal to some value μ0.
  • The null-hypothesis of a right-tailed test implies that the mean μ is less than or equal to some value μ0.
  • The null-hypothesis of a left-tailed test implies that the mean μ is greater than or equal to some value μ0.

Look for key phrases in the research question to help you set up the hypotheses. Make sure that the =, ≤ and ≥ sign always go in the null hypothesis. The ≠, > and < sign always go in the alternative hypothesis. Look for these phrases in Figure 8-1 to help you decide if you are setting up a two-tailed test (first column), a right-tailed test (second column), or left-tailed test (third column).

When you read a question, it is essential that you identify the parameter of interest. The parameter determines which distribution to use. Make sure that you can recognize and distinguish which parameter you are making a conjecture about: mean = µ, proportion = p, variance = σ2 , standard deviation = σ. There will be more parameters in later chapters.

Do not use the sample statistics, like \(\overline{ x }\) or \(\hat{p}\), in the hypotheses. We are not making any inference about the sample statistics. We know the value of the sample statistic. We use the sample statistics to infer if a change has occurred in the population.

For example, if we were making a conjecture about the percent or proportion in a population we would have the hypotheses:

H0: p = p0

H1: p ≠ p0.

Setting up the hypotheses correctly is the most important step in hypothesis testing. Here are some example research questions and how to set up the null and alternative hypotheses correctly; in a later section, we will perform the entire hypothesis test.

Use Figure 8-1 as a guide in setting up your hypotheses. The first column shows the hypotheses and how to shade in the distribution for a two-tailed test, with common phrases in the claim. The two-tailed test will always have a not equal ≠ sign in the alternative hypothesis and both tails shaded. The second column is for a right-tailed test. Note that the greater than > sign always will be in the alternative hypothesis and the right tail is shaded. The third column is for a left-tailed test. The left-tailed test will always have a less than < sign in the alternative hypothesis and the left tail shaded in.

Hypothesis Testing Common Phrases

clipboard_e05ff69adb91bfe44a87416af84b595b4.png

Figure 8-1

State the hypotheses in both words and symbols for the following claims.

  1. The national mean salary for high school teachers is $61,420. A random sample of 30 teacher’s salaries had a mean of $49,850. A new director for a graduate teacher education program (GTEP) believes that the average salary of a teacher in Oregon is significantly less than national average.
  2. A high school principal is looking into assigning parking spaces at their school if the proportion of students who own their own car is more than 30%. The principal does not have the time to ask all 1,200 students at their school so instead takes a random sample of 70 students and found that 33% owned their own car.
  3. A teacher would like to know if the average age of students taking evening classes is different from the university’s average age of 26. They sample 40 students from a random sample of evening classes and found the average age to be 27.
  4. Solution

    a) The key phrase in the claim is “less than.” The less than sign < is only allowed in the alternative hypothesis and we are testing against the national average.

    H0: The national mean salary is $61,420.

    H1: The GTEP director believes the mean salary in Oregon is less than $61,420.

    H0: μ = 61420

    H1: μ < 61420

    b) The key phrase in the claim is “more than.” The greater than sign > is only allowed in the alternative hypothesis. This is about a proportion, not a mean, so use the parameter p.

    H0: The principal will not assign parking spaces if 30% or less of students own a car.

    H1: The principal will assign parking spaces if more than 30% of students own a car.

    H0: p = 0.3

    H1: p > 0.3

    c) The key word in the claim is “different.” The not equal sign ≠ is only allowed in the alternative hypothesis.

    H0: The population mean age is 26 years old.

    H1: The evening students’ mean age is believed to be different from 26 years old.

    H0: μ = 26

    H1: μ ≠ 26

    Once we collect sample data we need to find out how far away the sample statistic can be from the hypothesized parameter to say that a statistically significant change has occurred.

Suppose a manufacturer of a new laptop battery claims the mean life of the battery is 900 days with a standard deviation of 40 days. You are the buyer of this battery and you think this claim is inflated. You would like to test your belief because without a good reason you cannot get out of your contract. You take a random sample of 35 batteries and find that the mean battery life is 890 days. What are the hypotheses for this question?

Solution

You have a guess that the mean life of a battery is less than 900 days. This is opposed to what the manufacturer claims. There really are two hypotheses, which are just guesses here – the one that the manufacturer claims and the one that you believe. For this problem:

H0: μ = 900, since the manufacturer says the mean life of a battery is 900 days.

H1: μ < 900, since you believe the mean life of the battery is less than 900 days.

Note that we do not put the sample mean of 890 in our hypotheses.

Is the sample mean of 890 days small enough to believe that you are right and the manufacturer is wrong? We would expect variation in our sample data and every time we take a new sample, the sample mean will most likely be different. How far away does the sample mean have to be from the product specification to verify our claim was correct? The sample data and the answer to these questions will be answered once we run the hypothesis test.

If you calculated a sample mean of 435, you would definitely believe the population mean is less than 900. However, even if you had a sample mean of 835 you would probably believe that the true mean was less than 900. What about 875? Or 893? There is some point where you would stop being so sure that the population mean is less than 900. That point separates the values of where you are sure or pretty sure that the mean is less than 900 from the area where you are not so sure.

How do you find that point where the sample mean is close enough to the hypothesized population mean? How close depends on how much error you want to make. Of course, you do not want to make any errors, but unfortunately, that is unavoidable in statistics since we are not measuring the entire population. You need to figure out how much error you made with your sample. Take the sample mean, and find the probability of getting another sample mean less than it, assuming for the moment that the manufacturer is right. The idea behind this is that you want to know what is the chance that you could have come up with your sample mean even if the population mean really is 900 days.

You want to find P(\(\bar{X}\) < 890 | H0 is true) = P(\(\bar{X}\) < 890 | μ = 900). For short, we will call this probability the p-value or simply p.

To compute this p-value, you need to know how the sample mean is distributed. Since the sample size is at least 30 you know the sample mean is approximately normally distributed, by the Central Limit Theorem (CLT). Remember \(\mu_{\bar{x}}=\mu\) and \(\sigma_{\bar{x}}=\frac{\delta}{\sqrt{n}}\). Before calculating the probability, it is useful to see how many standard deviations away from the mean the sample mean is. Using the formula for the z-score for CLT, \(z=\frac{\bar{x}-\mu}{\left(\frac{\sigma}{\sqrt{n}}\right)}\) we can compare this z-score to a zscore based on how sure we want to be of not making a mistake.

Using our sample mean we compute the z-score: \(z=\frac{\bar{x}-\mu_{0}}{\left(\frac{\sigma}{\sqrt{n}}\right)}=\frac{890-900}{\left(\frac{40}{\sqrt{35}}\right)}=-1.479\). This sample mean is more than one standard deviation away from the mean. Is that far enough?

Look at the probability P(\(\bar{X}\) < 890 | H0 is true) = P(\(\bar{X}\) < 890 | μ = 900) = P(Z < –1.479).

Using the TI Calculator normalcdf(-1E99,890,900,40/\(\sqrt{35}\)) \(\approx\) 0.0696.

Alternatively, in Excel use =NORM.DIST(890,900,40/SQRT(35),TRUE) \(\approx\) 0.0696.

Hence the p-value = 0.0696.

A picture is always useful. Figures 8-2 shows the populations distribution. Figure 8-3 shows the sampling distribution of the mean.

clipboard_e891eca4ce1b5f259f5b69d86677f1ab7.png

Figure 8-2

clipboard_eb7f0073afc4c73b75f64df4f7b794491.png

Figure 8-3

There is approximately a 6.96% chance that you could find a sample mean less than 890 when the population mean is 900 days. This is small but not really small. But how do you quantify really small? Is 5% or 10% or 15% really small? How do you decide? That depends on your field of study and the importance of the situation and will be answered in the next section.

To understand the process of a hypothesis test, you need to first understand what a hypothesis is, which is an educated guess about a parameter. Once you have the alternative hypothesis, you collect data and use the data to decide to see if there is enough evidence to show that the alternative hypothesis is true. However, in hypothesis testing you actually assume something else is true, the null hypothesis, and then you look at your data to see how likely it is to get an event that your data demonstrates with that assumption. If the event is very unusual, then you might think that your assumption is actually false. If you are able to say this assumption is false, then your alternative hypothesis could be true. You assume the opposite of your alternative hypothesis is true and show that it cannot be true. If this happens, then your alternative hypothesis is probably true. All hypothesis tests go through the same process. Once you have the process down, then the concept is much easier.

When setting up your hypotheses make sure the parameter, not the statistic, is used in the hypotheses. The equality always goes in the null hypothesis H0 and the alternative hypothesis Ha will be a left-tailed test with a less than sign <, a two-tailed test with a not equal sign \(\neq\), or a right-tailed test with a greater than sign >.

“‘But alright,’ went on the rumblings, ‘so what's the alternative?’

‘Well,’ said Ford, brightly but slowly, ‘stop doing it of course!’”

(Adams, 2002)

8.02: Type I and II Errors

How do you quantify really small? Is 5% or 10% or 15% really small? How do you decide? That depends on your field of study and the importance of the situation. Is this a pilot study? Is someone’s life at risk? Would you lose your job? Most industry standards use 5% as the cutoff point for how small is small enough, but 1%, 5% and 10% are frequently used depending on what the situation calls for.

Now, how small is small enough? To answer that, you really want to know the types of errors you can make in hypothesis testing.

The first error is if you say that H0 is false, when in fact it is true. This means you reject H0 when H0 was true. The second error is if you say that H0 is true, when in fact it is false. This means you fail to reject H0 when H0 is false.

Figure 8-4 shows that if we “Reject H0 ” when H0 is actually true, we are committing a type I error. The probability of committing a type I error is the Greek letter \(\alpha\), pronounced alpha. This can be controlled by the researcher by choosing a specific level of significance \(\alpha\).

clipboard_eec1caec9e13e3465d5a0c53094848700.png

Figure 8-4

Figure 8-4 shows that if we “Do Not Reject H0 ” when H0 is actually false, we are committing a type II error. The probability of committing a type II error is denoted with the Greek letter β, pronounced beta. When we increase the sample size this will reduce β. The power of a test is 1 – β.

A jury trial is about to take place to decide if a person is guilty of committing murder. The hypotheses for this situation would be:

  • \(H_0\): The defendant is innocent
  • \(H_1\): The defendant is not innocent

The jury has two possible decisions to make, either acquit or convict the person on trial, based on the evidence that is presented. There are two possible ways that the jury could make a mistake. They could convict an innocent person or they could let a guilty person go free. Both are bad news, but if the death penalty was sentenced to the convicted person, the justice system could be killing an innocent person. If a murderer is let go without enough evidence to convict them then they could possibly murder again. In statistics we call these two types of mistakes a type I and II error.

Figure 8-5 is a diagram to see the four possible jury decisions and two errors.

clipboard_e3c10ea812a7425f19e1c849bec82e74c.png

Figure 8-5

Type I Error is rejecting H0 when H0 is true, and Type II Error is failing to reject H0 when H0 is false.

Since these are the only two possible errors, one can define the probabilities attached to each error.

\(\alpha\) = P(Type I Error) = P(Rejecting H0 | H0 is true)

β = P(Type II Error) = P(Failing to reject H0 | H0 is false)

An investment company wants to build a new food cart. They know from experience that food carts are successful if they have on average more than 100 people a day walk by the location. They have a potential site to build on, but before they begin, they want to see if they have enough foot traffic. They observe how many people walk by the site every day over a month. They will build if there is more than an average of 100 people who walk by the site each day. In simple terms, explain what the type I & II errors would be using context from the problem.

Solution

The hypotheses are: H0: μ = 100 and H1: μ > 100.

Sometimes it is helpful to use words next to your hypotheses instead of the formal symbols

  • H0: μ ≤ 100 (Do not build)
  • H1: μ > 100 (Build).

A type I error would be to reject the null when in fact it is true. Take your finger and cover up the null hypothesis (our decision is to reject the null), then what is showing? The alternative hypothesis is what action we take.

If we reject H0 then we would build the new food cart. However, H0 was actually true, which means that the mean was less than or equal to 100 people walking by.

In more simple terms, this would mean that our evidence showed that we have enough foot traffic to support the food cart. Once we build, though, there was not on average more than 100 people that walk by and the food cart may fail.

A type II error would be to fail to reject the null when in fact the null is false. Evidence shows that we should not build on the site, but this actually would have been a prime location to build on.

The missed opportunity of a type II error is not as bad as possibly losing thousands of dollars on a bad investment.

What is more severe of an error is dependent on what side of the desk you are sitting on. For instance, if a hypothesis is about miles per gallon for a new car the hypotheses may be set up differently depending on if you are buying the car or selling the car. For this course, the claim will be stated in the problem and always set up the hypotheses to match the stated claim. In general, the research question should be set up as some type of change in the alternative hypothesis.

Controlling for Type I Error

The significance level used by the researcher should be picked prior to collection and analyzing data. This is called “a priori,” versus picking α after you have done your analysis which is called “post hoc.” When deciding on what significance level to pick, one needs to look at the severity of the consequences of the type I and type II errors. For example, if the type I error may cause the loss of life or large amounts of money the researcher would want to set \(\alpha\) low.

Controlling for Type II Error

The power of a test is the complement of a type II error or correctly rejecting a false null hypothesis. You can increase the power of the test and hence decrease the type II error by increasing the sample size. Similar to confidence intervals, where we can reduce our margin of error when we increase the sample size. In general, we would like to have a high confidence level and a high power for our hypothesis tests. When you increase your confidence level, then in turn the power of the test will decrease. Calculating the probability of a type II error is a little more difficult and it is a conditional probability based on the researcher’s hypotheses and is not discussed in this course.

“‘That's right!’ shouted Vroomfondel, ‘we demand rigidly defined areas of doubt and uncertainty!’”

(Adams, 2002)

Visualizing \(\alpha\) and β

If \(\alpha\) increases that means the chances of making a type I error will increase. It is more likely that a type I error will occur. It makes sense that you are less likely to make type II errors, only because you will be rejecting H0 more often. You will be failing to reject H0 less, and therefore, the chance of making a type II error will decrease. Thus, as α increases, β will decrease, and vice versa. That makes them seem like complements, but they are not complements. Consider one more factor – sample size.

Consider if you have a larger sample that is representative of the population, then it makes sense that you have more accuracy than with a smaller sample. Think of it this way, which would you trust more, a sample mean of 890 if you had a sample size of 35 or sample size of 350 (assuming a representative sample)? Of course, the 350 because there are more data points and so more accuracy. If you are more accurate, then there is less chance that you will make any error.

By increasing the sample size of a representative sample, you decrease β.

  • For a constant sample size, n, if \(\alpha\) increases, β decreases.
  • For a constant significance level, \(\alpha\), if n increases, β decreases.

When the sample size becomes large, point estimates become more precise and any real differences in the mean and null value become easier to detect and recognize. Even a very small difference would likely be detected if we took a large enough sample size. Sometimes researchers will take such a large sample size that even the slightest difference is detected. While we still say that difference is statistically significant, it might not be practically significant. Statistically significant differences are sometimes so minor that they are not practically relevant. This is especially important to research: if we conduct a study, we want to focus on finding a meaningful result. We do not want to spend lots of money finding results that hold no practical value.

The role of a statistician in conducting a study often includes planning the size of the study. The statistician might first consult experts or scientific literature to learn what would be the smallest meaningful difference from the null value. They also would obtain some reasonable estimate for the standard deviation. With these important pieces of information, they would choose a sufficiently large sample size so that the power for the meaningful difference is perhaps 80% or 90%. While larger sample sizes may still be used, the statistician might advise against using them in some cases, especially in sensitive areas of research.

If we look at the following two sampling distributions in Figure 8-6, the one on the left represents the sampling distribution for the true unknown mean. The curve on the right represents the sampling distribution based on the hypotheses the researcher is making. Do you remember the difference between a sampling distribution, the distribution of a sample, and the distribution of the population? Revisit the Central Limit Theorem in Chapter 6 if needed.

If we start with \(\alpha\) = 0.05, the critical value is represented by the vertical green line at \(z_{\alpha}\) = 1.645. Then the blue shaded area to the right of this line represents \(\alpha\). The area under the curve to the left of \(z_{\alpha / 2}\) = 1.96 based on the researcher’s claim would represent β.

clipboard_e7c65b0c521321075f8c809c2fab3b9ac.png

Figure 8-6

clipboard_e22a57fa8d6d91b8a21696f00278d2230.png

Figure 8-7

If we were to change \(\alpha\) from 0.05 to 0.01 then we get a critical value of \(z_{\alpha / 2}\) = 2.576. Note that when \(\alpha\) decreases, then β increases which means your power 1 – β decreases. See Figure 8-7.

This text does not go over how to calculate β. You will need to be able to write out a sentence interpreting either the type I or II errors given a set of hypotheses. You also need to know the relationship between \(\alpha\), β, confidence level, and power.

Hypothesis tests are not flawless, since we can make a wrong decision in statistical hypothesis tests based on the data. For example, in the court system, innocent people are sometimes wrongly convicted and the guilty sometimes walk free, or diagnostic tests that have false negatives or false positives. However, the difference is that in statistical hypothesis tests, we have the tools necessary to quantify how often we make such errors. A type I Error is rejecting the null hypothesis when H0 is actually true. A type II Error is failing to reject the null hypothesis when the alternative is actually true (H0 is false).

We use the symbols \(\alpha\) = P(Type I Error) and β = P(Type II Error). The critical value is a cutoff point on the horizontal axis of the sampling distribution that you can compare your test statistic to see if you should reject the null hypothesis. For a left-tailed test the critical value will always be on the left side of the sampling distribution, the right-tailed test will always be on the right side, and a two-tailed test will be on both tails. Use technology to find the critical values. Most of the time in this course the shortcut menus that we use will give you the critical values as part of the output.

8.2.1 Finding Critical Values

A researcher decides they want to have a 5% chance of making a type I error so they set α = 0.05. What z-score would represent that 5% area? It would depend on if the hypotheses were a left-tailed, two-tailed or right-tailed test. This zscore is called a critical value. Figure 8-8 shows examples of critical values for the three possible sets of hypotheses.

clipboard_eb9ca3f2fa72ae8e0e0186541560d1157.png

Figure 8-8

Two-tailed Test

If we are doing a two-tailed test then the \(\alpha\) = 5% area gets divided into both tails. We denote these critical values \(z_{\alpha / 2}\) and \(z_{1-\alpha / 2}\). When the sample data finds a z-score (test statistic) that is either less than or equal to \(z_{\alpha / 2}\) or greater than or equal to \(z_{1-\alpha / 2}\) then we would reject H0. The area to the left of the critical value \(z_{\alpha / 2}\) and to the right of the critical value \(z_{1-\alpha / 2}\) is called the critical or rejection region. See Figure 8-9.

clipboard_e7a6daefb1bf296ee0ee1389fd3cfdeb5.png

Figure 8-9

When \(\alpha\) = 0.05 then the critical values \(z_{\alpha / 2}\) and \(z_{1-\alpha / 2}\) are found using the following technology.

Excel: \(z_{\alpha / 2}\) =NORM.S.INV(0.025) = –1.96 and \(z_{1-\alpha / 2}\) =NORM.S.INV(0.975) = 1.96

TI-Calculator: \(z_{\alpha / 2}\) = invNorm(0.025,0,1) = –1.96 and \(z_{1-\alpha / 2}\) = invNorm(0.975,0,1) = 1.96

Since the normal distribution is symmetric, you only need to find one side’s z-score and we usually represent the critical values as ± \(z_{\alpha / 2}\).

Most of the time we will be finding a probability (p-value) instead of the critical values. The p-value and critical values are related and tell the same information so it is important to know what a critical value represents.

Right-tailed Test

If we are doing a right-tailed test then the \(\alpha\) = 5% area goes into the right tail. We denote this critical value \(z_{1-\alpha}\). When the sample data finds a z-score more than \(z_{1-\alpha}\) then we would reject H0, reject H0 if the test statistic is ≥ \(z_{1-\alpha}\). The area to the right of the critical value \(z_{1-\alpha}\) is called the critical region. See Figure 8-10.

clipboard_e8a4056c54332f7e0695328df084a0342.png

Figure 8-10

When \(\alpha\) = 0.05 then the critical value \(z_{1-\alpha}\) is found using the following technology.

Excel: \(z_{1-\alpha}\) =NORM.S.INV(0.95) = 1.645 Figure 8-10

TI-Calculator: \(z_{1-\alpha}\) = invNorm(0.95,0,1) = 1.645

Left-tailed Test

If we are doing a left-tailed test then the \(\alpha\) = 5% area goes into the left tail. If the sampling distribution is a normal distribution then we can use the inverse normal function in Excel or calculator to find the corresponding z-score. We denote this critical value \(z_{\alpha}\).

When the sample data finds a z-score less than \(z_{\alpha}\) then we would reject H0, reject Ho if the test statistic is ≤ \(z_{\alpha}\). The area to the left of the critical value \(z_{\alpha}\) is called the critical region. See Figure 8-11.

clipboard_ec4666de6d263a6bb55405555c4b54b6a.png

Figure 8-11

When \(\alpha\) = 0.05 then the critical value \(z_{\alpha}\) is found using the following technology.

Excel: \(z_{\alpha}\) =NORM.S.INV(0.05) = –1.645

TI-Calculator: \(z_{\alpha}\) = invNorm(0.05,0,1) = –1.645

The Claim and Summary

The wording on the summary statement changes depending on which hypothesis the researcher claims to be true. We really should always be setting up the claim in the alternative hypothesis since most of the time we are collecting evidence to show that a change has occurred, but occasionally a textbook will have the claim in the null hypothesis. Do not use the phrase “accept H0” since this implies that H0 is true. The lack of evidence is not evidence of nothing.

There were only two possible correct answers for the decision step.

i. Reject H0

ii. Fail to reject H0

Caution! If we fail to reject the null this does not mean that there was no change, we just do not have any evidence that change has occurred. The absence of evidence is not evidence of absence. On the other hand, we need to be careful when we reject the null hypothesis we have not proved that there is change.

When we reject the null hypothesis, there is only evidence that a change has occurred. Our evidence could have been false and lead to an incorrect decision. If we use the phrase, “accept H0” this implies that H0 was true, but we just do not have evidence that it is false. Hence you will be marked incorrect for your decision if you use accept H0, use instead “fail to reject H0” or “do not reject H0.”

8.03: Hypothesis Test for One Mean

There are three methods used to test hypotheses:

The Traditional Method (Critical Value Method)

There are five steps in hypothesis testing when using the traditional method:

  1. Identify the claim and formulate the hypotheses.
  2. Compute the test statistic.
  3. Compute the critical value(s) and state the rejection rule (the rule by which you will reject the null hypothesis (H0).
  4. Make the decision to reject or not reject the null hypothesis by comparing the test statistic to the critical value(s). Reject H0 when the test statistic is in the critical tail(s).
  5. Summarize the results and address the claim using context and units from the research question.

Steps ii and iii do not have to be in that order so make sure you know the difference between the critical value, which comes from the stated significance level \(\alpha\), and the test statistic, which is calculated from the sample data.

Note: The test statistic and the critical value(s) come from the same distribution and will usually have the same letter such as z, t, or F. The critical value(s) will have a subscript with the lower tail area \((z_{\alpha}, z_{1–\alpha}, z_{\alpha / 2})\) or an asterisk next to it (z*) to distinguish it from the test statistic.

You can find the critical value(s) or test statistic in any order, but make sure you know the difference when you compare the two. The critical value is found from α and is the start of the shaded area called the critical region (also called rejection region or area). The test statistic is computed using sample data and may or may not be in the critical region.

The critical value(s) is set before you begin (a priori) by the level of significance you are using for your test. This critical value(s) defines the shaded area known as the rejection area. The test statistic for this example is the z-score we find using the sample data that is then compared to the shaded tail(s). When the test statistic is in the shaded rejection area, you reject the null hypothesis. When your test statistic is not in the shaded rejection area, then you fail to reject the null hypothesis. Depending on if your claim is in the null or the alternative, the sample data may or may not support your claim.

The P-value Method

Most modern statistics and research methods utilize this method with the advent of computers and graphing calculators.

There are five steps in hypothesis testing when using the p-value method:

  1. Identify the claim and formulate the hypotheses.
  2. Compute the test statistic.
  3. Compute the p-value.
  4. Make the decision to reject or not reject the null hypothesis by comparing the p-value with \(\alpha\). Reject H0 when the p-value ≤ \(\alpha\).
  5. Summarize the results and address the claim.

The ideas below review the process of evaluating hypothesis tests with p-values:

  • The null hypothesis represents a skeptic’s position or a position of no difference. We reject this position only if the evidence strongly favors the alternative hypothesis.
  • A small p-value means that if the null hypothesis is true, there is a low probability of seeing a point estimate at least as extreme as the one we saw. We interpret this as strong evidence in favor of the alternative hypothesis.
  • The p-value is constructed in such a way that we can directly compare it to the significance level (\(\alpha\)) to determine whether to reject H0. We reject the null hypothesis if the p-value is smaller than the significance level, \(\alpha\), which is usually 0.05. Otherwise, we fail to reject H0.
  • We should always state the conclusion of the hypothesis test in plain language use context and units so non-statisticians can also understand the results.

The Confidence Interval Method (results are in the same units as the data)

There are four steps in hypothesis testing when using the confidence interval method:

  1. Identify the claim and formulate the hypotheses.
  2. Compute confidence interval.
  3. Make the decision to reject or not reject the null hypothesis by comparing the p-value with \(\alpha\). Reject H0 when the hypothesized value found in H0 is outside the bounds of the confidence interval. We only will be doing a two-tailed version of this.
  4. Summarize the results and address the claim.

For all 3 methods, Step i is the most important step. If you do not correctly set up your hypotheses then the next steps will be incorrect.

The decision and summary would be the same no matter which method you use. Figure 8-12 is a flow chart that may help with starting your summaries, but make sure you finish the sentence with context and units from the question.

clipboard_e1140413bcf9562500c92f289898037f4.png

Figure 8-12

The hypothesis-testing framework is a very general tool, and we often use it without a second thought. If a person makes a somewhat unbelievable claim, we are initially skeptical. However, if there is sufficient evidence that supports the claim, we set aside our skepticism and reject the null hypothesis in favor of the alternative.

8.3.1 Z-Test

When the population standard deviation is known and stated in the problem, we will use the z-test.

The z-test is a statistical test for the mean of a population. It can be used when σ is known. The population should be approximately normally distributed when n < 30.

When using this model, the test statistic is \(Z=\frac{\bar{x}-\mu_{0}}{\left(\frac{\sigma}{\sqrt{n}}\right)}\) where µ0 is the test value from the H0.

M&Ms candies advertise a mean weight of 0.8535 grams. A sample of 50 M&M candies are randomly selected from a bag of M&Ms and the mean is found to be \(\overline{ x }\) = 0.8472 grams. The standard deviation of the weights of all M&Ms is (somehow) known to be σ = 0.06 grams. A skeptic M&M consumer claims that the mean weight is less than what is advertised. Test this claim using the traditional method of hypothesis testing. Use a 5% level of significance.

Solution

By letting \(\alpha\) = 0.05, we are allowing a 5% chance that the null hypothesis (average weight that is at least 0.8535 grams) is rejected when in actuality it is true.

1. Identify the Claim: The claim is “M&Ms candies have a mean weight that is less than 0.8535 grams.” This translates mathematically to µ < 0.8535 grams. Therefore, the null and alternative hypotheses are:

H0: µ = 0.8535

H1: µ < 0.8535 (claim)

This is a left-tailed test since the alternative hypothesis has a “less than” sign.

We are performing a test about a population mean. We can use the z-test because we were given a population standard deviation σ (not a sample standard deviation s). In practice, σ is rarely known and usually comes from a similar study or previous year’s data.

2. Find the Critical Value: The critical value for a left-tailed test with a level of significance \(\alpha\) = 0.05 is found in a way similar to finding the critical values from confidence intervals. Because we are using the z-test, we must find the critical value \(z_{\alpha}\) from the z (standard normal) distribution.

This is a left-tailed test since the sign in the alternative hypothesis is < (most of the time a left-tailed test will have a negative z-score test statistic).

clipboard_e30dde96c4893d599ab1beb1114682332.png

Figure 8-13

First draw your curve and shade the appropriate tale with the area \(\alpha\) = 0.05. Usually the technology you are using only asks for the area in the left tail, which in this case is \(\alpha\) = 0.05. For the TI calculators, under the DISTR menu use invNorm(0.05,0,1) = –1.645. See Figure 8-13.

For Excel use =NORM.S.INV(0.05).

3. Find the Test Statistic: The formula for the test statistic is the z-score that we used back in the Central Limit Theorem section \(z=\frac{\bar{x}-\mu_{0}}{\left(\frac{\sigma}{\sqrt{n}}\right)}=\frac{0.8472-0.8535}{\left(\frac{0.06}{\sqrt{50}}\right)}=-0.7425\).

4. Make the Decision: Figure 8-14 shows both the critical value and the test statistic. There are only two possible correct answers for the decision step.

i. Reject H0

ii. Fail to reject H0

clipboard_e7d1b91a777de3955030ec329dee1860c.png

Figure 8-14

To make the decision whether to “Do not reject H0” or “Reject H0” using the traditional method, we must compare the test statistic z = –0.7425 with the critical value zα = –1.645.

When the test statistic is in the shaded tail, called the rejection area, then we would reject H0, if not then we fail to reject H0. Since the test statistic z ≈ –0.7425 is in the unshaded region, the decision is: Do not reject H0.

5. Summarize the Results: At 5% level of significance, there is not enough evidence to support the claim that the mean weight is less than 0.8535 grams.

Example 8-5 used the traditional critical value method. With the onset of computers, this method is outdated and the p-value and confidence interval methods are becoming more popular.

Most statistical software packages will give a p-value and confidence interval but not the critical value.

TI-84: Press the [STAT] key, go to the [TESTS] menu, arrow down to the [Z-Test] option and press the [ENTER] key. Arrow over to the [Stats] menu and press the [ENTER] key. Then type in value for the hypothesized mean (µ0), standard deviation, sample mean, sample size, arrow over to the \(\neq\), <, > sign that is in the alternative hypothesis statement then press the [ENTER] key, arrow down to [Calculate] and press the [ENTER] key. Alternatively (If you have raw data in a list) Select the [Data] menu and press the [ENTER] key. Then type in the value for the hypothesized mean (µ0), type in your list name (TI-84 L1 is above the 1 key).

clipboard_e84bcfd52053cf54117018344a8499a9d.png

Press the [STAT] key, go to the [TESTS] menu, arrow down to either the [Z-Test] option and press the [ENTER] key. Arrow over to the [Stats] menu and press the [ENTER] key. Then type in value for the hypothesized mean (µ0), standard deviation, sample mean, sample size, arrow over to the \(\neq\), <, > sign that is in the alternative hypothesis statement then press the [ENTER] key, arrow down to [Calculate] and press the [ENTER] key. Alternatively (If you have raw data in a list) Select the [Data] menu and press the [ENTER] key. Then type in the value for the hypothesized mean (µ0), type in your list name (TI-84 L1 is above the 1 key).

clipboard_e40992627be2973329e831eef7c18a74d.png

The calculator returns the alternative hypothesis (check and make sure you selected the correct sign), the test statistic, p-value, sample mean, and sample size.

TI-89: Go in to the Stat/List Editor App. Select [F6] Tests. Select the first option Z-Test. Select Data if you have raw data in a list, select Stats if you have the summarized statistics given to you in the problem. If you have data, press [2nd] Var-Link, the go down to list1 in the main folder to select the list name. If you have statistics then enter the values. Leave Freq:1 alone, arrow over to the \(\neq\), <, > sign that is in the alternative hypothesis statement then press the [ENTER]key, arrow down to [Calculate] and press the [ENTER] key. The calculator returns the test statistic and the p-value.

clipboard_e2e00e29f95abe7de3f5f9594417245e7.png


What is the p-value?

The p-value is the probability of observing an effect as least as extreme as in your sample data, assuming that the null hypothesis is true. The p-value is calculated based on the assumptions that the null hypothesis is true for the population and that the difference in the sample is caused entirely by random chance.

Recall the example at the beginning of the chapter.

Suppose a manufacturer of a new laptop battery claims the mean life of the battery is 900 days with a standard deviation of 40 days. You are the buyer of this battery and you think this claim is inflated. You would like to test your belief because without a good reason you cannot get out of your contract. You take a random sample of 35 batteries and find that the mean battery life is 890 days. Test the claim using the p-value method. Let \(\alpha\) = 0.05.

Solution

We had the following hypotheses:

H0: μ = 900, since the manufacturer says the mean life of a battery is 900 days.

H1: μ < 900, since you believe the mean life of the battery is less than 900 days.

The test statistic was found to be: \(Z=\frac{\bar{x}-\mu_{0}}{\left(\frac{\sigma}{\sqrt{n}}\right)}=\frac{890-900}{\left(\frac{40}{\sqrt{35}}\right)}=-1.479\).

The p-value is P(\(\overline{ x }\) < 890 | H0 is true) = P(\(\overline{ x }\)< 890 | μ = 900) = P(Z < –1.479).

On the TI Calculator use normalcdf(-1E99,890,900,40/\(\sqrt{35}\)) \(\approx\) 0.0696. See Figure 8-15.

clipboard_e4da98f35311c37ecda88f5e7fe48fe2e.png

Figure 8-15

Alternatively, in Excel use =NORM.DIST(890,900,40/SQRT(35),TRUE) \(\approx\) 0.0696.

clipboard_e8cf70bc4aa8ffbbe69077034f968135b.png

The TI calculators will easily find the p-value for you.

clipboard_e53381c2d9a67d1ae7089a5ec2247d34b.png

Now compare the p-value = 0.0696 to \(\alpha\) = 0.05. Make the decision to reject or not reject the null hypothesis by comparing the p-value with \(\alpha\). Reject H0 when the p-value ≤ α, and do not reject H0 when the p-value > \(\alpha\). The p-value for this example is larger than alpha 0.0696 > 0.05, therefore the decision is to not reject H0.

Since we fail to reject the null, there is not enough evidence to indicate that the mean life of the battery is less than 900 days.


8.3.2 T-Test

When the population standard deviation is unknown, we will use the t-test.

The t-test is a statistical test for the mean of a population. It will be used when σ is unknown. The population should be approximately normally distributed when n < 30.

When using this model, the test statistic is \(t=\frac{\bar{x}-\mu_{0}}{\left(\frac{s}{\sqrt{n}}\right)}\) where µ0 is the test value from the H0. The degrees of freedom are df = n – 1.

Z Versus T

The z and t-tests are easy to mix up. Sometimes a standard deviation will be stated in the problem without specifying if it is a population’s standard deviation σ or the sample standard deviation s. If the standard deviation is in the same sentence that describes the sample or only raw data is given then this would be s. When you only have sample data, use the t-test.

Figure 8-16 is a flow chart to remind you when to use z versus t.

clipboard_e81d4ef166161821bd95dc4f6bbe41a1b.png

Figure 8-16

Use Figure 8-17 as a guide in setting up your hypotheses. The two-tailed test will always have a not equal ≠ sign in H1 and both tails shaded. The right-tailed test will always have the greater than > sign in H1 and the right tail shaded. The left-tailed test will always have a less than < sign in H1 and the left tail shaded.

clipboard_ef6be25246ae148e1db8fb754cb98c10d.png

Figure 8-17

The label on a particular brand of cream of mushroom soup states that (on average) there is 870 mg of sodium per serving. A nutritionist would like to test if the average is actually more than the stated value. To test this, 13 servings of this soup were randomly selected and amount of sodium measured. The sample mean was found to be 882.4 mg and the sample standard deviation was 24.3 mg. Assume that the amount of sodium per serving is normally distributed. Test this claim using the traditional method of hypothesis testing. Use the \(\alpha\) = 0.05 level of significance.

Solution

Step 1: State the hypotheses and identify the claim: The statement “the average is more (>) than 870” must be in the alternative hypothesis. Therefore, the null and alternative hypotheses are:

H0: µ = 870

H1: µ > 870 (claim)

This is a right-tailed test with the claim in the alternative hypothesis.

Step 2: Compute the test statistic: We are using the t-test because we are performing a test about a population mean. We must use the t-test (instead of the z-test) because the population standard deviation σ is unknown. (Note: be sure that you know why we are using the t-test instead of the z-test in general.)

The formula for the test statistic is \(t=\frac{\bar{x}-\mu_{0}}{\left(\frac{S}{\sqrt{n}}\right)}=\frac{882.4-870}{\left(\frac{24.3}{\sqrt{13}}\right)}=1.8399\).

Note: If you were given raw data use 1-var Stats on your calculator to find the sample mean, sample size and sample standard deviation.

Step 3: Compute the critical value(s): The critical value for a right-tailed test with a level of significance \(\alpha\) = 0.05 is found in a way similar to finding the critical values from confidence intervals.

Since we are using the t-test, we must find the critical value t1–\(\alpha\) from a t-distribution with the degrees of freedom, df = n – 1 = 13 –1 = 12. Use the DISTR menu invT option. Note that if you have an older TI-84 or a TI-83 calculator you need to have the invT program installed or use Excel.

Draw and label the t-distribution curve with the critical value as in Figure 8-18.

clipboard_e2073659de08a6b4aeecacf48f025219e.png

Figure 8-18

The critical value is t1–\(\alpha\) = 1.782 and the rejection rule becomes: Reject H0 if the test statistic t ≥ t1–\(\alpha\) = 1.782.

Step 4: State the decision. Decision: Since the test statistic t =1.8399 is in the critical region, we should Reject H0.

Step 5: State the summary. Summary: At the 5% significance level, we have sufficient evidence to say that the average amount of sodium per serving of cream of mushroom soup exceeds the stated 870 mg amount.

Example 8-7 Continued:

Use the prior example, but this time use the p-value method. Again, let the significance level be \(\alpha\) = 0.05.

Solution

Step 1: The hypotheses remain the same. H0: µ = 870

H1: µ > 870 (claim)

Step 2: The test statistic remains the same, \(t=\frac{\bar{x}-\mu_{0}}{\left(\frac{S}{\sqrt{n}}\right)}=\frac{882.4-870}{\left(\frac{24.3}{\sqrt{13}}\right)}=1.8399\).

Step 3: Compute the p-value.

For a right-tailed test, the p-value is found by finding the area to the right of the test statistic t = 1.8339 under a tdistribution with 12 degrees of freedom. See Figure 8-19.

clipboard_e31493abc8d7bd3d835479d9e187566db.png

Figure 8-19

Note that exact p-values for a t-test can only be found using a computer or calculator. For the TI calculators this is in the DISTR menu. Use tcdf(lower,upper,df).

For this example, we would have p-value = tcdf(1.8399,∞,12) = 0.0453.

The p-value is the probability of observing an effect as least as extreme as in your sample data, assuming that the null hypothesis is true. The p-value is calculated based on the assumptions that the null hypothesis is true for the population and that the difference in the sample is caused entirely by random chance.

Step 4: State the decision. The rejection rule: reject the null hypothesis if the p-value ≤ \(\alpha\). Decision: Since the p-value = 0.0453 is less than \(\alpha\) = 0.05, we Reject H0. This agrees with the decision from the traditional method. (These two methods should always agree!)

Step 5: State the summary. The summary remains the same as in the previous method. At the 5% significance level, we have sufficient evidence to say that the average amount of sodium per serving of cream of mushroom soup exceeds the stated 870 mg amount.

We can use technology to get the test statistic and p-value.

TI-84: If you have raw data, enter the data into a list before you go to the test menu. Press the [STAT] key, arrow over to the [TESTS] menu, arrow down to the [2:T-Test] option and press the [ENTER] key. Arrow over to the [Stats] menu and press the [ENTER] key. Then type in the hypothesized mean (µ0), sample or population standard deviation, sample mean, sample size, arrow over to the \(\neq\), <, > sign that is the same as the problem’s alternative hypothesis statement then press the [ENTER] key, arrow down to [Calculate] and press the [ENTER] key. The calculator returns the t-test statistic and p-value.

clipboard_e1f7ce7702f1fd056f927e431cd9249a6.png

Alternatively (If you have raw data in list one) Arrow over to the [Data] menu and press the [ENTER] key. Then type in the hypothesized mean (µ0), L1, leave Freq:1 alone, arrow over to the \(\neq\), <, > sign that is the same in the problem’s alternative hypothesis statement then press the [ENTER] key, arrow down to [Calculate] and press the [ENTER] key. The calculator returns the t-test statistic and the p-value.

TI-89: Go to the [Apps] Stat/List Editor, then press [2nd] then F6 [Tests], then select 2: T-Test. Choose the input method, data is when you have entered data into a list previously or stats when you are given the mean and standard deviation already. Then type in the hypothesized mean (μ0), sample standard deviation, sample mean, sample size (or list name (list1), and Freq: 1), arrow over to the \(\neq\), <, > and select the sign that is the same as the problem’s alternative hypothesis statement then press the [ENTER] key to calculate. The calculator returns the t-test statistic and p-value.

clipboard_ee6ee07b247a8dbe3c99babfd60690494.png

The weight of the world’s smallest mammal is the bumblebee bat (also known as Kitti’s hog-nosed bat or Craseonycteris thonglongyai) is approximately normally distributed with a mean 1.9 grams. Such bats are roughly the size of a large bumblebee. A chiropterologist believes that the Kitti’s hog-nosed bats in a new geographical region under study has a different average weight than 1.9 grams. A sample of 10 bats weighed in grams in the new region are shown below. Use the confidence interval method to test the claim that mean weight for all bumblebee bats is not 1.9 g using a 10% level of significance.

clipboard_eded04400d8044b884aaddc3055aeac28.png

Solution

Step 1: State the hypotheses and identify the claim. The key phrase is “mean weight not equal to 1.9 g.” In mathematical notation, this is μ ≠ 1.9. The not equal ≠ symbol is only allowed in the alternative hypothesis so the hypotheses would be:

H0: μ = 1.9

H1: μ ≠ 1.9

Step 2: Compute the confidence interval. First, find the t critical value using df = n – 1 = 9 and 90% confidence. In Excel t\(\alpha\)/2 = T.INV(.1/2,9) = 1.833113.

Then use technology to find the sample mean and sample standard deviation and substitute in your numbers to the formula.

\(\begin{aligned}
&\bar{x} \pm t_{\alpha / 2}\left(\frac{s}{\sqrt{n}}\right) \\
&\Rightarrow 1.985 \pm 1.833113\left(\frac{0.235242}{\sqrt{10}}\right) \\
&\Rightarrow 1.985 \pm 1.833113(0.07439) \\
&\Rightarrow 1.985 \pm 0.136365 \\
&\Rightarrow(1.8486,2.1214)
\end{aligned}\)

The answer can be given as an inequality 1.8486 < µ < 2.1214

or in interval notation (1.8486, 2.1214).

Step 3: Make the decision: The rejection rule is to reject H0 when the hypothesized value found in H0 is outside the bounds of the confidence interval. The null hypothesis was μ = 1.9 g. Since 1.9 is between the lower and upper boundary of the confidence interval 1.8486 < µ < 2.1214 then we would not reject H0.

The sampling distribution, assuming the null hypothesis is true, will have a mean of μ = 1.9 and a standard error of \(\frac{0.2352}{\sqrt{10}}=0.07439\). When we calculated the confidence interval using the sample mean of 1.985 the confidence interval captured the hypothesized mean of 1.9. See Figure 8-20.

clipboard_e1ff990721e3ff41b0f17bc336941ad97.png

Figure 8-20

Step 4: State the summary: At the 10% significance level, there is not enough evidence to support the claim that the population mean weight for bumblebee bats in the new geographical region is different from 1.9 g.

This interval can also be computed using a TI calculator or Excel.

TI-84: Enter the data in a list, choose Tests > TInterval. Select and highlight Data, change the list and confidence level to match the question. Choose Calculate.

clipboard_e05f42f706148f07b9c5b947adde95723.png

Excel: Select Data Analysis > Descriptive Statistics: Note, you will need to change the cell reference numbers to where you copy and paste your data, only check the label box if you selected the label in the input range, and change the confidence level to 1 – \(\alpha\).

clipboard_ee0f9e73c4603bccdf7780454d5d6d5de.png

Below is the Excel output. Excel only calculates the descriptive statistics with the margin of error.

clipboard_efad58f58634f5aa4fb0915e33cd36ab3.png

Use Excel to find each piece of the interval \(\bar{x} \pm t_{\alpha / 2}\left(\frac{s}{\sqrt{n}}\right)\).

Excel \(t_{\alpha / 2}\) = T.INV(0.1/2,9) = 1.8311.

\(\begin{aligned}
&\bar{x} \pm t_{\alpha / 2}\left(\frac{s}{\sqrt{n}}\right) \\
&\Rightarrow 1.985 \pm 1.8311\left(\frac{0.2352}{\sqrt{10}}\right) \\
&\Rightarrow 1.985 \pm 1.8311(0.07439)
\end{aligned}\)

Can you find the mean and standard error \(\frac{s}{\sqrt{n}}=0.07439\) in the Excel output?

\(\Rightarrow 1.985 \pm 0.136365\)

Can you find the margin of error \(t_{\frac{\alpha}{2}}\left(\frac{s}{\sqrt{n}}\right)=0.136365\) in the Excel output?

Subtract and add the margin of error from the sample mean to get each confidence interval boundary (1.8486, 2.1214).

If we have raw data, Excel will do both the traditional and p-value method.

Example 8-8 Continued:

Use the prior example, but this time use the p-value method. Again, let the significance level be \(\alpha\) = 0.05.

Solution

Step 1: State the hypotheses. The hypotheses are: H0: μ = 1.9

H1: μ ≠ 1.9

Step 2: Compute the test statistic, \(t=\frac{\bar{x}-\mu_{0}}{\left(\frac{s}{\sqrt{n}}\right)}=\frac{1.985-1.9}{\left(\frac{.235242}{\sqrt{10}}\right)}=1.1426\)

Verify using Excel. Excel does not have a one-sample t-test, but it does have a twosample t-test that can be used with a dummy column of zeros as the second sample to get the results for just one sample. Copy over the data into cell A1. In column B, next to the data, type in a dummy column of zeros, and label it Dummy. (We frequently use placeholders in statistics called dummy variables.)

clipboard_e8fe4ea3b1ac9776d898cad2e7ebc32c2.png

Select the Data Analysis tool and then select t-Test: Paired Two Sample for Means, then select OK.

clipboard_e3e25b9a5546cfb3adc2524d7d021a559.png

For the Variable 1 Range select the data in cells A1:A11, including the label. For the Variable 2 Range select the dummy column of zeros in cells B1:B11, including the label. Change the hypothesized mean to 1.9. Check the Labels box and change the alpha value to 0.10, then select OK.

clipboard_e0cb7805c1a1e3d2f112ec1e586082c82.png

Excel provides the following output:

clipboard_edb0ff9a634588c2198429949099261bc.png

Step 3: Compute the p-value. Since the alternative hypothesis has a ≠ symbol, use the Excel output next two-tailed p-value = 0.2826.

Step 4: Make the decision. For the p-value method we would compare the two-tailed p-value = 0.2826 to \(\alpha\) = 0.10. The rule is to reject H0 if the p-value ≤ \(\alpha\). In this case the p-value > \(\alpha\), therefore we do not reject H0. Again, the same decision as the confidence interval method.

For the critical value method, we would compare the test statistic t = 1.142625 with the critical values for a twotailed test \(t_{\frac{\alpha}{2}}\) = ±1.833113. Since the test statistic is between –1.8331 and 1.8331 we would not reject H0, which is the same decision using the p-value method or the confidence interval method.

Step 5: State the summary. There is not enough evidence to support the claim that the population mean weight for all bumblebee bats is not equal to 1.9 g.

One-Tailed Versus Two-Tailed Tests

Most software packages do not ask which tailed test you are performing. Make sure you look at the sign in the alternative hypothesis to and determine which p-value to use. The difference is just what part of the picture you are looking at. In Excel, the critical value shown is for a one-tail test and does not specify left or right tail. The critical value in the output will always be positive, it is up to you to know if the critical value should be a negative or positive value. For example, Figures 8-21, 8-22, and 8-23 uses df = 9, \(\alpha\) = 0.10 to show all three tests comparing either the test statistic with the critical value or the p-value with \(\alpha\).

Two-Tailed Test

The test statistic can be negative or positive depending on what side of the distribution it falls; however, the p-value is a probability and will always be a positive number between 0 and 1. See Figure 8-21.

clipboard_e5ee81ac959096119cce3bfbdab8735e2.png

Figure 8-21

Right-Tailed Test

If we happened to do a right-tailed test with df = 9 and \(\alpha\) = 0.10, the critical value t1-\(\alpha\) = 1.383 will be in the right tail and usually the test statistic will be a positive number. See Figure 8-22.

clipboard_e7fe1f5f4c7373d90184ab24f2ffd3be4.png

Figure 8-22

Left-Tailed Test

If we happened to do a left-tailed test with df = 9 and \(\alpha\) = 0.10, the critical value t\(\alpha\) = –1.383 will be in the left tail and usually the test statistic will be a negative number. See Figure 8-23.

clipboard_ee09b4c2bae7bf3f01077a238ffbefc0a.png

Figure 8-23

8.04: Hypothesis Test for One Proportion

When you read a question, it is essential that you correctly identify the parameter of interest. The parameter determines which model to use. Make sure that you can recognize and distinguish between a question regarding a population mean and a question regarding a population proportion.

The z-test is a statistical test for a population proportion. It can be used when np ≥ 10 and nq ≥ 10.

Definition: z-Test

The formula for the test statistic is:

\[Z=\dfrac{\hat{p}-p_{0}}{\sqrt{\left(\dfrac{p_{0} q_{0}}{n}\right)}}.\]

where

\(n\) is the sample size

\(\hat{p}=\dfrac{x}{n}\) is the sample proportion (sometimes already given as a %) and

\(p_0\) is the hypothesized population proportion,

\[q_0 = 1 – p_0. \nonumber\]

Use the phrases in Figure 8-24 to help with setting up the hypotheses.

clipboard_eb78509765d4da42df3ffbbc7bcd7ddf2.png

Figure 8-24

Note we will not be using the t-distribution with proportions. We will use a standard normal z distribution for testing a proportion since this test uses the normal approximation to the binomial distribution (never use the t-distribution).

If you are doing a left-tailed z-test the critical value will be negative. If you are performing a right-tailed z-test the critical value will be positive. If you were performing a two-tailed z-test then your critical values would be ±critical value. The p-value will always be a positive number between 0 and 1. The most important step in any method you use is setting up your null and alternative hypotheses. The critical values and p-value can be found using a standard normal distribution the same way that we did for the one sample z-test.

It has been found that 85.6% of all enrolled college and university students in the United States are undergraduates. A random sample of 500 enrolled college students in a particular state revealed that 420 of them were undergraduates. Is there sufficient evidence to conclude that the proportion differs from the national percentage? Use \(\alpha\) = 0.05. Show that all three methods of hypothesis testing yield the same results.

Solution

At this point you should be more comfortable with the steps of a hypothesis test and not have to number each step, but know what each step means.

Critical Value Method

Step 1: State the hypotheses: The key words in this example, “proportion” and “differs,” give the hypotheses:

H0: p = 0.856

H1: p ≠ 0.856 (claim)

Step 2: Compute the test statistic. Before finding the test statistic, find the sample proportion \(\hat{p}=\dfrac{420}{500}=0.84\) and q0 = 1 – 0.856 = 0.144.

Next, compute the test statistic:

\[z=\dfrac{\hat{p}-p_{0}}{\sqrt{\left(\dfrac{p_{0} q_{0}}{n}\right)}}=\dfrac{0.84-0.856}{\sqrt{\left(\dfrac{0.856 \cdot 0.144}{500}\right)}}=-1.019.\]

Step 3: Draw and label the curve with the critical values. See Figure 8-25.

Use \(\alpha\) = 0.05 and technology to compute the critical values \(z_{\alpha / 2}\) and \(z_{1-\alpha / 2}\).

Excel: \(z_{\alpha / 2}\) =NORM.S.INV(0.025) = –1.96 and \(z_{1-\alpha / 2}\) =NORM.S.INV(0.975) = 1.96.

TI-Calculator: \(z_{\alpha / 2}\) = invNorm(0.025,0,1) = –1.96 and \(z_{1-\alpha / 2}\) = invNorm(0.975,0,1) = 1.96.

clipboard_e8f099b5a122df52e032d75a08e8ad736.png

Figure 8-25

Step 4: State the decision. Since the test statistic is not in the shaded rejection area, do not reject H0.

Step 5: State the summary. At the 5% level of significance, there is not enough evidence to conclude that the proportion of undergraduates in college for this state differs from the national average of 85.6%.

P-value Method

The hypotheses and test statistic stay the same.

H0: p = 0.856

H1: p ≠ 0.856 (claim)

\(Z=\dfrac{\hat{p}-p_{0}}{\sqrt{\left(\dfrac{p_{0} q_{0}}{n}\right)}}=\dfrac{0.84-0.856}{\sqrt{\left(\dfrac{0.856 \cdot 0.144}{500}\right)}}=-1.019\)

To find the p-value we need to find the P(Z > |1.019|) the area to the left of z = –1.019 and to the right of z = 1.019. First, find the area below (since the test statistic is negative) z = –1.019 using the normalcdf we get 0.1541. Then, double this area to get the p-value = 0.3082.

clipboard_e84277aeef92e4e279a2f67125448603e.png

Since the p-value > \(\alpha\) the decision is to not reject H0.

Summary: There is not enough evidence to conclude that the proportion of undergraduates in college for this state differs from the national average of 85.6%.

There is a shortcut for this test on the TI Calculators, which will quickly find the test statistic and p-value.

The rejection rule for the two methods are:

  • P-value method: reject H0 when the p-value ≤ \(\alpha\).
  • Critical value method: reject H0 when the test statistic is in the critical region.

TI-84: Press the [STAT] key, arrow over to the [TESTS] menu, arrow down to the option [5:1-PropZTest] and press the [ENTER] key. Type in the hypothesized proportion (p0), x, sample size, arrow over to the \(\neq\), <, > sign that is the same in the problem’s alternative hypothesis statement then press the [ENTER] key, arrow down to [Calculate] and press the [ENTER] key.

clipboard_e9057c4e0ef886e707af43c98457d8249.png

The calculator returns the z-test statistic and the p-value. Note: sometimes you are not given the x value but a percentage instead. To find the x to use in the calculator, multiply \(\hat{p}\) by the sample size and round off to the nearest integer. The calculator will give you an error message if you put in a decimal for x or n. For example, if \(\hat{p}\)= 0.22 and n = 124 then 0.22*124 = 27.28, so use x = 27.

TI-89: Go to the [Apps] Stat/List Editor, then press [2nd] then F6 [Tests], then select 5: 1-PropZ-Test. Type in the hypothesized proportion (p0), x, sample size, arrow over to the \(\neq\), <, > sign that is the same in the problem’s alternative hypothesis statement then press the [ENTER] key to calculate. The calculator returns the z-test statistic and the pvalue. Note: sometimes you are not given the x value but a percentage instead. To find the x value to use in the calculator, multiply \(\hat{p}\) by the sample size and round off to the nearest integer. The calculator will give you an error message if you put in a decimal for x or n. For example, if \(\hat{p}\) = 0.22 and n = 124 then 0.22*124 = 27.28, so use x = 27.

8.05: Chapter 8 Exercises

Chapter 8 Exercises

1. The plant-breeding department at a major university developed a new hybrid boysenberry plant called Stumptown Berry. Based on research data, the claim is made that from the time shoots are planted 90 days on average are required to obtain the first berry. A corporation that is interested in marketing the product tests 60 shoots by planting them and recording the number of days before each plant produces its first berry. The sample mean is 92.3 days. The corporation wants to know if the mean number of days is different from the 90 days claimed. Which one is the correct set of hypotheses?

a) H0: p = 90% H1: p ≠ 90%

b) H0: μ = 90 H1: μ ≠ 90

c) H0: p = 92.3% H1: p ≠ 92.3%

d) H0: μ = 92.3 H1: μ ≠ 92.3

e) H0: μ ≠ 90 H1: μ = 90

2. Match the symbol with the correct phrase.

clipboard_ee1678dd08291f01a71ee20a4b540c641.png

3. According to the February 2008 Federal Trade Commission report on consumer fraud and identity theft, 23% of all complaints in 2007 were for identity theft. In that year, Alaska had 321 complaints of identity theft out of 1,432 consumer complaints. Does this data provide enough evidence to show that Alaska had a lower proportion of identity theft than 23%? Which one is the correct set of hypotheses?

Federal Trade Commission, (2008). Consumer fraud and identity theft complaint data: January-December 2007. Retrieved from website: http://www.ftc.gov/opa/2008/02/fraud.pdf.

a) H0: p = 23% H1: p < 23%

b) H0: μ = 23 H1: μ < 23

c) H0: p < 23% H1: p ≥ 23%

d) H0: p = 0.224 H1: p < 0.224

e) H0: μ < 0.224 H1: μ ≥ 0.224

4. Compute the z critical value for a right-tailed test when \(\alpha\) = 0.01.

5. Compute the z critical value for a two-tailed test when \(\alpha\) = 0.01.

6. Compute the z critical value for a left-tailed test when \(\alpha\) = 0.05.

7. Compute the z critical value for a two-tailed test when \(\alpha\) = 0.05.

8. As of 2018, the Centers for Disease Control and Protection’s (CDC) national estimate that 1 in 68 \(\approx\) 0.0147 children have been diagnosed with autism spectrum disorder (ASD). A researcher believes that the proportion of children in their county is different from the CDC estimate. Which one is the correct set of hypotheses?

a) H0: p = 0.0147 H1: p ≠ 0.0147

b) H0: μ = 0.0147 H1: μ ≠ 0.0147

c) H0: p ≠ 0.0147 H1: p = 0.0147

d) H0: μ = 68 H1: μ ≠ 68

e) H0: = 0.0147 H1: ≠ 0.0147

9. Match the phrase with the correct symbol.

a. Sample Size i. α

b. Population Mean ii. n

c. Sample Variance iii. σ²

d. Sample Mean iv. s²

e. Population Standard Deviation v. s

f. P(Type I Error) vi. \(\bar{x}\)

g. Sample Standard Deviation vii. σ

h. Population Variance viii. μ

10. The Food & Drug Administration (FDA) regulates that fresh albacore tuna fish contains at most 0.82 ppm of mercury. A scientist at the FDA believes the mean amount of mercury in tuna fish for a new company exceeds the ppm of mercury. Which one is the correct set of hypotheses?

a) H0: p = 82% H1: p > 82%

b) H0: μ = 0.82 H1: μ > 0.82

c) H0: p > 82% H1: p ≤ 82%

d) H0: μ = 0.82 H1: μ ≠ 0.82

e) H0: μ > 0.82 H1: μ ≤ 0.82

11. Match the symbol with the correct phrase.

clipboard_eb79ddd287ba75c522a7665c1831b3609.png

12. The plant-breeding department at a major university developed a new hybrid boysenberry plant called Stumptown Berry. Based on research data, the claim is made that from the time shoots are planted 90 days on average are required to obtain the first berry. A corporation that is interested in marketing the product tests 60 shoots by planting them and recording the number of days before each plant produces its first berry. The sample mean is 92.3 days. The corporation will not market the product if the mean number of days is more than the 90 days claimed. The hypotheses are H0: μ = 90 H1: μ > 90. Which answer is the correct type I error in the context of this problem?

a) The corporation will not market the Stumptown Berry even though the berry does produce fruit within the 90 days.

b) The corporation will market the Stumptown Berry even though the berry does produce fruit within the 90 days.

c) The corporation will not market the Stumptown Berry even though the berry does produce fruit in more than 90 days.

d) The corporation will market the Stumptown Berry even though the berry does produce fruit in more than 90 days.

13. The Food & Drug Administration (FDA) regulates that fresh albacore tuna fish contains at most 0.82 ppm of mercury. A scientist at the FDA believes the mean amount of mercury in tuna fish for a new company exceeds the ppm of mercury. The hypotheses are H0: μ = 0.82 H1: μ > 0.82. Which answer is the correct type II error in the context of this problem?

a) The fish is rejected by the FDA when in fact it had less than 0.82 ppm of mercury.

b) The fish is accepted by the FDA when in fact it had less than 0.82 ppm of mercury.

c) The fish is rejected by the FDA when in fact it had more than 0.82 ppm of mercury.

d) The fish is accepted by the FDA when in fact it had more than 0.82 ppm of mercury.

14. A two-tailed z-test found a test statistic of z = 2.153. At a 1% level of significance, which would the correct decision?

a) Do not reject H0

b) Reject H0

c) Accept H0

d) Reject H1

e) Do not reject H1

15. A left-tailed z-test found a test statistic of z = -1.99. At a 5% level of significance, what would the correct decision be?

a) Do not reject H0

b) Reject H0

c) Accept H0

d) Reject H1

e) Do not reject H1

16. A right-tailed z-test found a test statistic of z = 0.05. At a 5% level of significance, what would the correct decision be?

a) Reject H0

b) Accept H0

c) Reject H1

d) Do not reject H0

e) Do not reject H1

17. A two-tailed z-test found a test statistic of z = -2.19. At a 1% level of significance, which would the correct decision?

a) Do not reject H0

b) Reject H0

c) Accept H0

d) Reject H1

e) Do not reject H1

18. According to the February 2008 Federal Trade Commission report on consumer fraud and identity theft, 23% of all complaints in 2007 were for identity theft. In that year, Alaska had 321 complaints of identity theft out of 1,432 consumer complaints. Does this data provide enough evidence to show that Alaska had a lower proportion of identity theft than 23%? The hypotheses are H0: p = 23% H1: p < 23%. Which answer is the correct type I error in the context of this problem?

Federal Trade Commission, (2008). Consumer fraud and identity theft complaint data: January-December 2007. Retrieved from website: http://www.ftc.gov/opa/2008/02/fraud.pdf.

a) It is believed that less than 23% of Alaskans had identity theft and there really was 23% or less that experienced identity theft.

b) It is believed that more than 23% of Alaskans had identity theft and there really was 23% or more that experience identity theft.

c) It is believed that less than 23% of Alaskans had identity theft even though there really was 23% or more that experienced identity theft.

d) It is believed that more than 23% of Alaskans had identity theft even though there really was less than 23% that experienced identity theft

19. A hypothesis test was conducted during a clinical trial to see if a new COVID-19 vaccination reduces the risk of contracting the virus. What is the Type I and II errors in terms of approving the vaccine for use?

20. A manufacturer of rechargeable laptop batteries claims its batteries have, on average, 500 charges. A consumer group decides to test this claim by assessing the number of times 30 of their laptop batteries can be recharged and finds a p-value is 0.1111; thus, the null hypothesis is not rejected. What is the Type II error for this situation?

21. A commonly cited standard for one-way length (duration) of school bus rides for elementary school children is 30 minutes. A local government office in a rural area conducts a study to determine if elementary schoolers in their district have a longer average one-way commute time. If they determine that the average commute time of students in their district is significantly higher than the commonly cited standard they will invest in increasing the number of school buses to help shorten commute time. What would a Type II error mean in this context?

22. The Centers for Disease Control and Prevention (CDC) 2018 national estimate that 1 in 68 \(\approx\) 0.0147 children have been diagnosed with autism spectrum disorder (ASD). A researcher believes that the proportion of children in their county is different from the CDC estimate. The hypotheses are H0: p = 0.0147 H1: p ≠ 0.0147. Which answer is the correct type II error in the context of this problem?

a) The proportion of children diagnosed with ASD in the researcher’s county is believed to be different from the national estimate, even though the proportion is the same.

b) The proportion of children diagnosed with ASD in the researcher’s county is believed to be different from the national estimate and the proportion is different.

c) The proportion of children diagnosed with ASD in the researcher’s county is believed to be the same as the national estimate, even though the proportion is different.

d) The proportion of children diagnosed with ASD in the researcher’s county is believed to be the same as the national estimate and the proportion is the same.

23. The Food & Drug Administration (FDA) regulates that fresh albacore tuna fish contains at most 0.82 ppm of mercury. A scientist at the FDA believes the mean amount of mercury in tuna fish for a new company exceeds the ppm of mercury. A test statistic was found to be 2.576 and a critical value was found to be 1.645, what is the correct decision and summary?

a) Reject H0, there is enough evidence to support the claim that the amount of mercury in the new company’s tuna fish exceeds the FDA limit of 0.82 ppm.

b) Accept H0, there is not enough evidence to reject the claim that the amount of mercury in the new company’s tuna fish exceeds the FDA limit of 0.82 ppm.

c) Reject H1, there is not enough evidence to reject the claim that the amount of mercury in the new company’s tuna fish exceeds the FDA limit of 0.82 ppm.

d) Reject H0, there is not enough evidence to support the claim that the amount of mercury in the new company’s tuna fish exceeds the FDA limit of 0.82 ppm.

e) Do not reject H0, there is not enough evidence to support the claim that the amount of mercury in the new company’s tuna fish exceeds the FDA limit of 0.82 ppm.

24. The plant-breeding department at a major university developed a new hybrid boysenberry plant called Stumptown Berry. Based on research data, the claim is made that from the time shoots are planted 90 days on average are required to obtain the first berry. A corporation that is interested in marketing the product tests 60 shoots by planting them and recording the number of days before each plant produces its first berry. The corporation wants to know if the mean number of days is different from the 90 days claimed. A random sample was taken and the following test statistic was z = -2.15 and critical values of z = ±1.96 was found. What is the correct decision and summary?

a) Do not reject H0, there is not enough evidence to support the corporation’s claim that the mean number of days until a berry is produced is different from the 90 days claimed by the university.

b) Reject H0, there is enough evidence to support the corporation’s claim that the mean number of days until a berry is produced is different from the 90 days claimed by the university.

c) Accept H0, there is enough evidence to support the corporation’s claim that the mean number of days until a berry is produced is different from the 90 days claimed by the university.

d) Reject H1, there is not enough evidence to reject the corporation’s claim that the mean number of days until a berry is produced is different from the 90 days claimed by the university.

e) Reject H0, there is not enough evidence to support the corporation’s claim that the mean number of days until a berry is produced is different from the 90 days claimed by the university.

25. You are conducting a study to see if the accuracy rate for fingerprint identification is significantly different from 0.34. Thus, you are performing a two-tailed test. Your sample data produce the test statistic z = 2.504. Use your calculator to find the p-value and state the correct decision and summary.

26. The SAT exam in previous years is normally distributed with an average score of 1,000 points and a standard deviation of 150 points. The test writers for this upcoming year want to make sure that the new test does not have a significantly different mean score. They have a random sample of 20 students take the SAT and their mean score was 1,050 points.

a) Test to see if the mean time has significantly changed using a 5% level of significance. Show all your steps using the critical value method.

b) What is a type I error for this problem?

c) What is a type II error for this problem?

27. A sample of 45 body temperatures of athletes had a mean of 98.8˚F. Assume the population standard deviation is known to be 0.62˚F. Test the claim that the mean body temperature for all athletes is more than 98.6˚F. Use a 1% level of significance. Show all your steps using the p-value method.

28. Compute the t critical value for a left-tailed test when \(\alpha\) = 0.10 and df = 10.

29. Compute the t critical value for a two-tailed test when \(\alpha\) = 0.05 with a sample size of 18.

30. Using a t-distribution with df = 25, find the P(t ≥ 2.185).

31. A student is interested in becoming an actuary. They know that becoming an actuary takes a lot of schooling and they will have to take out student loans. They want to make sure the starting salary will be higher than $55,000/year. They randomly sample 30 starting salaries for actuaries and find a p-value of 0.0392. Use \(\alpha\) = 0.05.

a) Choose the correct hypotheses.

i. H0: μ = 55,000 H1: μ < 55,000

ii. H0: μ > 55,000 H1: μ ≤ 55,000

iii. H0: μ = 55,000 H1: μ > 55,000

iv. H0: μ < 55,000 H1: μ ≥ 55,000

v. H0: μ = 55,000 H1: μ ≠ 55,000

b) Should the student pursue an actuary career?

i. Yes, since we reject the null hypothesis.

ii. Yes, since we reject the claim.

iii. No, since we reject the claim.

iv. No, since we reject the null hypothesis.

32. The workweek for adults in the United States work full-time is normally distributed with a mean of 47 hours. A newly hired engineer at a start-up company believes that employees at start-up companies work more on average then working adults in the U.S. She asks 12 engineering friends at start-ups for the lengths in hours of their workweek. Their responses are shown in the table below. Test the claim using a 5% level of significance. Show all 5 steps using the p-value method.

clipboard_e05e9179bee29036da5c04653dc60b5a9.png

33. The average number of calories from a fast food meal for adults in the United States is 842 calories. A nutritionist believes that the average is higher than reported. They sample 11 meals that adults ordered and measure the calories for each meal shown below. Test the claim using a 5% level of significance. Assume that fast food calories are normally distributed. Show all 5 steps using the p-value method.

clipboard_ec82daa00c5055fa172062bc390cd0017.png

34. Honda advertises the 2018 Honda Civic as getting 32 mpg for city driving. A skeptical consumer about to purchase this model believes the mpg is less than the advertised amount and randomly selects 35 2018 Honda Civic owners and asks them what their car’s mpg is. Use a 1% significance level. They find a p-value of 0.0436.

a) Choose the correct hypotheses.

i. H0: μ = 32 H1: μ < 32

ii. H0: μ < 32 H1: μ ≥ 32

iii. H0: μ = 32 H1: μ > 32

iv. H0: μ = 35 H1: μ ≠ 35

v. H0: μ = 32 H1: μ ≠ 32

b) Choose the correct decision based off the reported p-value.

i. Reject H0

ii. Do not reject H0

iii. Do not reject H1

iv. Reject H1

For exercises 35-40, show all 5 steps for hypothesis testing:

a) State the hypotheses.

b) Compute the test statistic.

c) Compute the critical value or p-value.

d) State the decision.

e) Write a summary.

35. The total of individual pounds of garbage discarded by 17 households in one week is shown below. The current waste removal system company has a weekly maximum weight policy of 36 pounds. Test the claim that the average weekly household garbage weight is less than the company's weekly maximum. Use a 5% level of significance.

clipboard_e5767b380e411f7a83f622cf74e2334a0.png

36. The world’s smallest mammal is the bumblebee bat (also known as Kitti’s hog-nosed bat or Craseonycteris thonglongyai). Such bats are roughly the size of a large bumblebee. A sample of 10 bats weighed in grams are shown below. Test the claim that mean weight for all bumblebee bats is not equal to 2.1 g using a 1% level of significance. Assume that the bat weights are normally distributed.

clipboard_e3842c5dfa2c15b54fd6613e46137e22d.png

37. The average age of an adult's first vacation without a parent or guardian was reported to be 23 years old. A travel agent believes that the average age is different from what was reported. They sample 28 adults and they asked their age in years when they first vacationed as an adult without a parent or guardian, data shown below. Test the claim using a 10% level of significance.

clipboard_e6795627590f26d4a82ef6e3e62c51045.png

38. Test the claim that the proportion of people who own dogs is less than 32%. A random sample of 1,000 people found that 28% owned dogs. Do the sample data provide convincing evidence to support the claim? Test the relevant hypotheses using a 10% level of significance.

39. The National Institute of Mental Health published an article stating that in any one-year period, approximately 9.3% of American adults suffer from depression or a depressive illness. Suppose that in a survey of 2,000 people in a certain city, 11.1% of them suffered from depression or a depressive illness. Conduct a hypothesis test to determine if the true proportion of people in that city suffering from depression or a depressive illness is more than the 9.3% in the general adult American population. Test the relevant hypotheses using a 5% level of significance.

40. The United States Department of Energy reported that 48% of homes were heated by natural gas. A random sample of 333 homes in Oregon found that 149 were heated by natural gas. Test the claim that the proportion of homes in Oregon that were heated by natural gas is different from what was reported. Use a 1% significance level.

41. A 2019 survey by the Bureau of Labor Statistics reported that 92% of Americans working in large companies have paid leave. In January 2021, a random survey of workers showed that 89% had paid leave. The resulting p-value is 0.009; thus, the null hypothesis is rejected. It is concluded that there has been a decrease in the proportion of people, who have paid leave from 2019 to January 2021. What type of error is possible in this situation?

a) Type I Error

b) Type II Error

c) Standard Error

d) Margin of Error

e) No error was made.

For exercises 42-44, show all 5 steps for hypothesis testing:

a) State the hypotheses.

b) Compute the test statistic.

c) Compute the critical value or p-value.

d) State the decision.

e) Write a summary.

42. Nationwide 40.1% of employed teachers are union members. A random sample of 250 Oregon teachers showed that 110 belonged to a union. At \(\alpha\) = 0.10, is there sufficient evidence to conclude that the proportion of union membership for Oregon teachers is higher than the national proportion?

43. You are conducting a study to see if the proportion of men over the age of 50 who regularly have their prostate examined is significantly less than 0.31. A random sample of 735 men over the age of 50 found that 208 have their prostate regularly examined. Do the sample data provide convincing evidence to support the claim? Test the relevant hypotheses using a 5% level of significance.

44. Nationally the percentage of adults that have their teeth cleaned by a dentist yearly is 64%. A dentist in Portland, Oregon believes that regionally the percent is higher. A sample of 2,000 Portlanders found that 1,312 had their teeth cleaned by a dentist in the last year. Test the relevant hypotheses using a 10% level of significance.

Answer to Odd Numbered Exercises

1) b

3) a

5) ±2.5758

7) ±1.96

9) a) ii. b) viii. c) iv. d) vi. e) vii. f) i. g) v. h) iii.

11) 100(1 – α)% = Confidence Level 1 – β = Power β = P(Type II Error) µ = Parameter α = Significance Level

13) d

15) b

17) a

19) The implication of a Type I error from the clinical trial is that the vaccination will be approved when it indeed does not reduce the risk of contracting the virus. The implication of a Type II error from the clinical trial is that the vaccination will not be approved when it indeed does reduce the risk of contracting the virus.

21) The local government decides that the data do not provide convincing evidence of an average commute time higher than 30 minutes, when the true average commute time is in fact higher than 30 minutes.

23) a

25) 0.0123

27) H0: µ = 98.6; H1: µ > 98.6; z = 2.1639; p-value = 0.0152; Do not reject H0. There is not enough evidence to support the claim that the mean body temperature for all athletes is more than 98.6˚F.

29) ±2.1098

31) a) iii. b) i.

33) H0: µ = 842; H1: µ > 842; t = 0.8218; p-value = 0.2152; Do not reject H0. We do not have evidence to support the claim the average calories from a fast food meal is higher than reported.

35) H0: µ = 36; H1: µ < 36; t = -1.9758; p-value = 0.0438; Reject H0. There is enough evidence to support the claim the average weekly household garbage weight is less than the company’s weekly 36 lb. maximum.

37) H0: µ = 23; H1: µ ≠ 23; t = 1.4224; p-value = 0.1664; Do not reject H0. We do not have enough evidence to support the claim that the mean age adults travel without a parent or guardian differs from 23.

39) H0: p = 0.093; H1: p > 0.093; t = 2.7116; p-value = 0.0027; Reject H0. There is enough evidence to support the claim the population proportion of American adults that suffer from depression or a depressive illness is more than 9.3%.

41) a

43) H0: p = 0.31; H1: p < 0.31; t = -1.5831; p-value = 0.0567; Do not reject H0. There is not enough evidence to support the claim the population proportion of men over the age of 50 who regularly have their prostate examined is significantly less than 0.31

8.06: Chapter 8 Formulas

Hypothesis Test for One Mean

Use z-test when σ is given.

Use t-test when s is given.

If n < 30, population needs to be normal.

Type I Error-

Reject H0 when H0 is true.

Type II Error-

Fail to reject H0 when H0 is false.

Z-Test:

H0: μ = μ0

H1: μ ≠ μ0

\(Z=\frac{\bar{x}-\mu_{0}}{\left(\frac{\sigma}{\sqrt{n}}\right)}\)

TI-84: Z-Test

t-Test:

H0: μ = μ0

H1: μ ≠ μ0

\(t=\frac{\bar{x}-\mu_{0}}{\left(\frac{s}{\sqrt{n}}\right)}\)

TI-84: T-Test

z-Critical Values

Excel:

Two-tail: \(z_{\alpha / 2}\) = NORM.INV(1–\(\alpha\)/2,0,1)

Right-tail: \(z_{1-\alpha}\) = NORM.INV(1–\(\alpha\),0,1)

Left-tail: \(z_{\alpha}\) = NORM.INV(\(\alpha\),0,1)

TI-84:

Two-tail: \(z_{\alpha / 2}\) = invNorm(1–\(\alpha\)/2,0,1)

Right-tail: \(z_{1-\alpha}\) = invNorm(1–\(\alpha\),0,1)

Left-tail: \(z_{\alpha}\) = invNorm(\(\alpha\),0,1)

t-Critical Values

Excel:

Two-tail: \(t_{\alpha / 2}\) =T.INV(1–\(\alpha\)/2,df)

Right-tail: \(t_{1-\alpha}\) = T.INV(1–\(\alpha\),df)

Left-tail: \(t_{\alpha}\)= T.INV(\(\alpha\),df)

TI-84:

Two-tail: \(t_{\alpha / 2}\) = invT(1–\(\alpha\)/2,df)

Right-tail: \(t_{1-\alpha}\) = invT(1–\(\alpha\),df)

Left-tail: \(t_{\alpha}\) = invT(\(\alpha\),df)

Hypothesis Test for One Proportion

H0: p = p0

H1: pp0

\(Z=\frac{\hat{p}-p_{0}}{\sqrt{\left(\frac{p_{0} q_{0}}{n}\right)}}\)

TI-84: 1-PropZTest

Rejection Rules:

  • P-value method: reject H0 when the p-value ≤ \(\alpha\).
  • Critical value method: reject H0 when the test statistic is in the critical region (shaded tails).

From Mostly Harmless Statistics by Rachel L. Webb, adapted from LibreTexts. Licensed CC BY-SA 4.0. XYZ Homework OER web edition, adapted with 7 verified corrections (see errata).