7 Confidence Intervals for One Population
7.01: Introduction
Statistical inference is used to draw conclusions about a population based on a sample. We can use the probability distributions and Central Limit Theorem to understand what is going on in the population. The population can be difficult to measure so we take a sample from that population and use descriptive statistics to measure the sample. We can then use those sample statistics to infer back to what is happening in our population. Although there are many types of statistical inference tools, we will only cover some of the more common techniques.
Distinguishing between a population and a sample is important in statistics. We frequently use a representative sample to generalize a population.
- A statistic is any characteristic or measure from a sample. One example is the sample mean \(\overline{ x }\).
- A parameter is any characteristic or measure from a population. One example is the population mean µ.
- A point estimate for a parameter (a characteristic from a population) is a statistic (a characteristic from a sample). For example, the point estimate for the population mean µ is the sample mean \(\overline{ x }\). The point estimate for the population standard deviation σ is the sample standard deviation s, etc.
- A 100(1 – α)% confidence interval for a population parameter (μ, σ, etc.) represents that the proportion 100(1 – α)% of times the true value of the population parameter is contained within the interval.
- The confidence level (or level of confidence) is 1 – α. The common percentages used for confidence interval levels are 90%, 95%, and 99%. Some corresponding values of alpha are: 90% would be α = 0.10 = 10%, 95% would be α = 0.05 = 5%, and 99% would be α = 0.01 = 1%. In this context, α, “alpha,” represents the complement of the confidence level, and its definition will be explained in the next chapter.
When a symmetric distribution, such as a normal distribution, is used, confidence intervals are always of the form: point estimate ± margin of error
The margin of error defines the “radius” of the interval necessary to obtain the desired confidence level. The margin of error depends on the desired confidence level. Higher levels of confidence come at a cost, namely larger margins of error, which means our estimate will be less accurate.
The margin of error formula will usually include a value from a sampling distribution called the critical value. The critical value measures the number of standard errors to be added and subtracted in order to achieve your desired confidence level based on the α level chosen.
For large sample sizes, the sampling distribution of a mean is normal. We can use the standard normal distribution values that would give the middle 95% of the distribution when α = 0.05 since 100(1 – 0.05)% = 95%.
The two critical values –zα/2 and +zα/2, as shown Figure 7-1. Note: in the notation zα/2 the α/2 represents the area in each of the tails.

Figure 7-1
Assumption: If the sample size is small (n < 30), the population we are sampling from must be normal. If the sample size is “large” (n ≥ 30), the Central Limit Theorem guarantees that the sampling distribution will be normally distributed no matter how the population distribution is distributed.
7.02: Confidence Interval for a Proportion
Suppose you want to estimate the population proportion, p. As an example, an administrator may want to know what proportion of students at your school smoke. An insurance company may want to know what proportion of accidents are caused by teenage drivers who do not have a drivers’ education class. Every time we collect data from a new sample, we would expect the estimate of the proportion to change slightly. If you were to find a range of values over an interval this would give a better estimate of where the population proportion falls. This range of values that would better predict the true population parameter is called an interval estimate or confidence interval.
The sample proportion \(\hat{p}\) is the point estimate for p, the standard error (the standard deviation of the sampling distribution) of \(\hat{p}\) is \(\sqrt{\left(\frac{\hat{p} \cdot \hat{q}}{n}\right)}\), the zα/2 is the critical value using the standard normal distribution, and the margin of error \(\mathrm{E}=Z_{\alpha / 2} \sqrt{\left(\frac{\hat{p} \cdot \hat{q}}{n}\right)}\). Some textbooks use \(\pi\) instead of p for the population proportion, and \(\bar{p}\) (pronounced “p-bar”) instead of \(\hat{p}\) for sample proportion.
Choose a simple random sample of size n from a population having unknown population proportion p. The 100(1 – \(\alpha\))% confidence interval estimate for p is given by \(\hat{p} \pm Z_{\alpha / 2} \sqrt{\left(\frac{\hat{p} \hat{q}}{n}\right)}\).
Where \(\hat{p}=\frac{x}{n}=\frac{\# \text { of successes }}{\# \text { of trials }}\) (read as “p hat”) is the sample proportion, and \(\hat{q}=1-\hat{p}\) is the complement.
The above confidence interval can be expressed as an inequality or an interval of values.
\(\hat{p}-z_{\alpha / 2} \sqrt{\left(\frac{\hat{p} \hat{q}}{n}\right)}<p<\hat{p}+z_{\alpha / 2} \sqrt{\left(\frac{\hat{p} \hat{q}}{n}\right)} \quad \text { or } \quad\left(\hat{p}-z_{\frac{\alpha}{2}} \sqrt{\left(\frac{\hat{p} \hat{q}}{n}\right)}, \hat{p}+z_{\alpha / 2} \sqrt{\left(\frac{\hat{p} \hat{q}}{n}\right)}\right)\)
Assumption: \(n \cdot \hat{p} \geq 10 \text { and } n \cdot \hat{q} \geq 10\)
*This assumption must be addressed before using these statistical inferences.
This formula is derived from the normal approximation of the binomial distribution, therefore the same conditions for a binomial need to be met, namely a set sample size of independent trials, two outcomes that have the same probability for each trial.
Steps for Calculating a Confidence Interval
1. State the random variable and the parameter in words.
x = number of successes
p = proportion of successes
2. State and check the assumptions for confidence interval.
a. A simple random sample of size n is taken.
b. The conditions for the binomial distribution are satisfied.
c. To determine the sampling distribution of \(\hat{p}\), you need to show that \(n \cdot \hat{p} \geq 10 \text { and } n \cdot \hat{q} \geq 10\), where \(\hat{q}\) = 1 − \(\hat{p}\). If this requirement is true, then the sampling distribution of \(\hat{p}\) is well approximated by a normal curve. (In reality, this is not really true, since the correct assumption deals with p. However, in a confidence interval you do not know p, so you must use \(\hat{p}\). This means you just need to show that x ≥ 10 and n – x ≥ 10.)
3. Compute the sample statistic \(\hat{p}=\frac{x}{n}\) and the confidence interval \(\hat{p} \pm z_\frac{\alpha}{2} \sqrt{\left(\frac{\hat{p} \hat{q}}{n}\right)}\).
4. Statistical Interpretation: In general, this looks like:
“We can be (1 – α)*100% confident that the interval \[\hat{p}-z_\frac{\alpha}{2} \sqrt{\left(\frac{\hat{p} \hat{q}}{n}\right)}<p<\hat{p}+z_\frac{\alpha}{2} \sqrt{\left(\frac{\hat{p} \hat{q}}{n}\right)}\]
Real World Interpretation: This is where you state what interval contains the true proportion.
A concern was raised in Australia that the percentage of deaths of indigenous Australian prisoners was higher than the percent of deaths of nonindigenous Australian prisoners, which is 0.27%. A sample of six years (1990- 1995) of data was collected, and it was found that out of 14,495 indigenous Australian prisoners, 51 died (“Indigenous deaths in,” 1996). Find a 95% confidence interval for the proportion of indigenous Australian prisoners who died.
Solution
1. State the random variable and the parameter in words.
x = number of indigenous Australian prisoners who die
p = proportion of indigenous Australian prisoners who die
2. State and check the assumptions for a confidence interval.
a. A simple random sample of 14,495 indigenous Australian prisoners was taken. However, the sample was not a random sample, since it was data from six years. It is the numbers for all prisoners in these six years, but the six years were not picked at random. Unless there was something special about the six years that were chosen, the sample is probably a representative sample. This assumption is probably met.
b. There are 14,495 prisoners in this case. The prisoners are all indigenous Australians, so you are not mixing indigenous Australian with nonindigenous Australian prisoners. There are only two outcomes, the prisoner either dies or does not. The chance that one prisoner dies over another may not be constant, but if you consider all prisoners the same, then it may be close to the same probability. Thus, the assumptions for the binomial distribution are satisfied.
c. In this case, x = 51 and n – x = 14,495 – 51 = 14,444. Both are greater than or equal to 10. The sampling distribution for \(\hat{p}\) is a normal distribution.
3. Compute the sample statistic and the confidence interval.
Sample Proportion: \(\hat{p}=\frac{x}{n}=\frac{51}{14495}=.003518\),
Critical Value: \(z_{\alpha / 2}=1.96\), since 95% confidence level
Margin of Error \(\mathrm{E}=z_{\alpha / 2} \sqrt{\left(\frac{\hat{p} \cdot \hat{q}}{n}\right)}=1.96 \sqrt{\left(\frac{0.003518(1-0.003518)}{14495}\right)}=0.000964\)
Confidence Interval: \(\hat{p}-\mathrm{E}<p<\hat{p}+\mathrm{E}\)
0.003518 – 0.000964 < p < 0.003518 + 0.000964
0.002554 < p < 0.004482 or (0.002554, 0.004482)
4. Statistical Interpretation: We can be 95% confident that 0.002554 < p < 0.004482 contains the proportion of all indigenous Australian prisoners who died.
5. Real World Interpretation: We can be 95% confident that the percentage of all indigenous Australian prisoners who died is between 0.26% and 0.45%.
Using Technology
Excel has no built-in shortcut key for finding a confidence interval for a proportion, but if you type in the following formulas shown below you can make your own Excel calculator where you just change the highlighted cells and all the numbers below will update with the relevant information.
Type in the following be cognizant of cell reference numbers.

You get the following answers where the last two numbers are your confidence interval limits.

Make sure to put your answer in interval notation (0.002555, 0.004482) or 0.26% < p < 0.45%.
You can also do the calculations for the confidence interval with the TI Calculator.
TI-84: Press the [STAT] key, arrow over to the [TESTS] menu, arrow down to the [A:1-PropZInterval] option and press the [ENTER] key. Then type in the values for x, sample size and confidence level, arrow down to [Calculate] and press the [ENTER] key. The calculator returns the answer in interval notation. Note: Sometimes you are not given the x value but a percentage instead. To find the x to use in the calculator, multiply \(\hat{p}\) by the sample size and round off to the nearest integer. The calculator will give you an error message if you put in a decimal for x or n. For example, if \(\hat{p}\) = 0.22 and n = 124 then 0.22*124 = 27.28, so use x = 27.

TI-89: Go to the [Apps] Stat/List Editor, then press [2nd] then F7 [Ints], then select 5: 1-PropZInt. Type in the values for x, sample size and confidence level, and press the [ENTER] key. The calculator returns the answer in interval notation. Note: sometimes you are not given the x value but a percentage instead. To find the x value to use in the calculator, multiply \(\hat{p}\) by the sample size and round off to the nearest integer. The calculator will give you an error message if you put in a decimal for x or n. For example, if \(\hat{p}\)= 0.22 and n = 124 then 0.22*124 = 27.28, so use x = 27.
A researcher studying the effects of income levels on new mothers breastfeeding their infants hypothesizes that those countries where the income level is lower has a higher rate of infants breastfeeding than higher income countries. It is known that in Germany, considered a high-income country by the World Bank, 22% of all babies are breastfed. In Tajikistan, considered a low-income country by the World Bank, researchers found that in a random sample of 500 new mothers that 125 were breastfeeding their infants. Find a 90% confidence interval of the proportion of mothers in low-income countries who breastfeed their infants.
Solution
1. State your random variable and the parameter in words.
x = The number of new mothers who breastfeed in a low-income country.
p = The proportion of new mothers who breastfeed in a low-income country.
2. State and check the assumptions for a confidence interval.
a. A simple random sample of 500 breastfeeding habits of new mothers in a low-income country was taken as was stated in the problem.
b. There were 500 women in the study. The women are considered identical, though they probably have some differences. There are only two outcomes - either the woman breastfeeds her baby or she does not. The probability of a woman breastfeeding her baby is probably not the same for each woman, but it is probably not that different for each woman. The assumptions for the binomial distribution are satisfied.
c. x = 125 and n – x = 500 – 125 = 375 and both are greater than or equal to 10, so the sampling distribution of \(\hat{p}\) is well approximated by a normal curve.
3. Compute the sample statistic and the confidence interval. On the TI-83/84: Go into the STAT menu. Move over to TESTS and choose 1- PropZInt, then press Calculate.

4. Statistical Interpretation: We are 90% confident that the interval 0.219 < p < 0.282 contains the population proportion of all women in low-income countries who breastfeed their infants.
5. Real World Interpretation: The proportion of women in low-income countries who breastfeed their infants is between 0.219 and 0.282 with 90% confidence.
7.03: Sample Size Calculation for a Proportion
A confidence interval for a population proportion p and q = 1 – p, with specific margin of error E is given by:
\(n=p^{*} \cdot q^{*}\left(\frac{z_{\alpha / 2}}{E}\right)^{2}\) Always round up to the next whole number.
Note: If the sample size is determined before the sample is selected, the p* and q* in the above equation are our best guesses. Often times statisticians will use p* = q* = 0.5; this takes the guesswork out of determining p* and provides the “worst case scenario” for n. In other words, if p* = 0.5 is used, then you are guaranteed that the margin of error will not exceed E but you also will have to sample the largest possible sample size. Some texts will use p or π instead of p*.
A study found that 73% of prekindergarten children ages 3 to 5 whose mothers had a bachelor’s degree or higher were enrolled in early childhood care and education programs.
- How large a sample is needed to estimate the true proportion within 3% with 95% confidence?
- How large a sample is needed if you had no prior knowledge of the proportion?
Solution
a) Use \(n=p^{*} \cdot q^{*}\left(\frac{z_{\alpha / 2}}{E}\right)^{2}=0.73 \cdot 0.27\left(\frac{1.96}{0.03}\right)^{2}=841.3104\). Since we cannot have 0.3104 of a person, we need to round up to the next whole person and use n = 842. Don’t round down since we may not get within our margin of error for a smaller sample size.
b) Since no proportion is given, use the planning value of p* = 0.5.
\[n=0.5 \cdot 0.5\left(\frac{1.96}{0.03}\right)^{2}=1067.1111 \nonumber\]
Round up and use \(n = 1,068\).
Note the sample sizes of 842 and 1,068. If you have a prior knowledge about the sample proportion then you may not have to sample as many people to get the same margin of error. The larger the sample size, the smaller the confidence interval.
7.04: Z-Interval for a Mean
Suppose you want to estimate the mean weight of newborn infants, or you want to estimate the mean salary of college graduates. A confidence interval for the mean would be the way to estimate these means.
A 100(1 - \(\alpha\) )% confidence interval for a population mean μ: (σ known) Choose a simple random sample of size n from a population having unknown mean μ. The 100(1 - \(\alpha\))% confidence interval estimate for μ is given by, \(\bar{x} \pm z_{\alpha / 2}\left(\frac{\sigma}{\sqrt{n}}\right)\).
The point estimate for μ is \(\overline{ x }\), and the margin of error is \(z_{\alpha / 2}\left(\frac{\sigma}{\sqrt{n}}\right)\).
Where \(z_\frac{\alpha}{2}\) is the value on the standard normal curve with area 1 – \(\alpha\) between the critical values –z\(\alpha\)/2 and +z\(\alpha\)/2, as shown below in Figure 7-2.

Note: In the notation, z\(\alpha\)/2 the \(\alpha\)/2 represents the area in each of the tails, see Figure 7-2.
The confidence interval can be expressed as an inequality or an interval of values.
\(\bar{x}-z_{\alpha / 2} \cdot \frac{\sigma}{\sqrt{n}}<\mu<\bar{x}+z_{\alpha / 2} \cdot \frac{\sigma}{\sqrt{n}} \quad \text { or } \quad\left(\bar{x}-z_{\alpha / 2} \cdot \frac{\sigma}{\sqrt{n}}, \bar{x}+z_{\alpha / 2} \cdot \frac{\sigma}{\sqrt{n}}\right)\)
Assumptions:
1. If the sample size is small (n < 30), the population we are sampling from must be normally distributed. If the sample size is “large” (n ≥ 30) the Central Limit Theorem guarantees that the sampling distribution of the mean will be normally distributed no matter how the population distribution is distributed.
2. The population standard deviation σ must be known. Most of the time we are using σ from a similar study or a prior year’s data. If you have a sample standard deviation then we will use a different method introduced in a later section.
These assumptions must be addressed before using these statistical inferences. In most cases, we do not know the population standard deviation so will not use the z-interval. Instead, we will use a different sampling distribution called the Student’s t-distribution or t-distribution for short.
Suppose we select a random sample of 100 pennies in circulation in order to estimate the average age of all pennies that are still in circulation. The sample average age, in years, was found to be \(\overline{ x }\) = 14.6. For the sake of this example, let us assume that the population standard deviation is 4 years old. Find a 95% confidence interval for the true average age of pennies that are still in circulation.
Solution
We can use the above (z) model because σ is known, the population distribution shape is unknown, but the sample size is over 30.
Use Excel or your calculator to find z\(\alpha\)/2 for a 95% confidence interval. In Excel use =NORM.INV(lower tail area, mean, standard deviation). It is easier to deal with the positive z-score so use the z to the right of the mean which would have 1 – \(\alpha\)/2 = 0.975 area. In Excel use =NORM.INV(0.975,0,1) or the calculator invNorm(0.975,0,1) which gives z\(\alpha\)/2 = 1.96.
\(\bar{x} \pm z_\frac{\alpha}{2} \frac{\sigma}{\sqrt{n}} \quad \Rightarrow \quad 14.6 \pm 1.96\left(\frac{4}{\sqrt{100}}\right) \quad \Rightarrow \quad 14.6 \pm 0.784 \quad \Rightarrow (13.816, 15.384) \)
The point estimate for μ is 14.6 years, and the margin of error is 0.784 years. If we were to repeat this same sampling process, we would expect 95 out of 100 such intervals to contain the true population mean age of all pennies in circulation. A shorthand way to say this is, with 95% confidence that the population mean age of all pennies in circulation is between 13.816 to 15.384 years.
The answer is expressed as an inequality so the confidence interval is 13.816 < µ < 15.384. You can also use interval notation (13.816, 15.384) which is more common and matches the notation found on most calculators.
TI-84: Press the [STAT] key, arrow over to the [TESTS] menu, arrow down to the [7:ZInterval] option and press the [ENTER] key. Arrow over to the [Stats] menu and press the [ENTER] key. Then type in the population or sample standard deviation, sample mean, sample size and confidence level, arrow down to [Calculate] and press the [ENTER] key. The calculator returns the answer in interval notation.

TI-89: Go to the [Apps] Stat/List Editor, then press [2nd] then F7 [Ints], then select 1: ZInterval. Choose the input method, data is when you have entered data into a list previously or stats when you are given the mean and standard deviation already. Type in the population standard deviation, sample mean, sample size (or list name (list1), and Freq: 1) and confidence level, and press the [ENTER] key to calculate. The calculator returns the answer in interval notation.

7.05: Interpreting a Confidence Interval
There is always a chance that the confidence interval would not contain the true parameter that we are looking for. Inferential statistics does not “prove” that the population parameter is within the boundaries of the confidence interval. If the sample we took had all outliers and the sample statistic is far away from the true population parameter, then when we subtract and add the margin of error to the point estimate, the population parameter may not be within the limits.
Both the sample size and confidence level affect how wide the interval is. The following discussion demonstrates what happens to the width of the interval as you get more confident.
Think about shooting an arrow into the target. Suppose you are really good at that and that you have a 90% chance of hitting the bull’s-eye. Now the bull’s-eye is very small. Since you hit the bull’s eye approximately 90% of the time, then you probably hit inside the next ring out 95% of the time. You have a better chance of doing this, but the circle is bigger. You probably have a 99% chance of hitting the target, but that is a much bigger circle to hit. As your confidence in hitting the target increases, the circle you hit gets bigger. The same is true for confidence intervals.
The higher level of confidence makes a wider interval. There is a tradeoff between width and confidence level. You can be really confident about your answer, but your answer will not be very precise. On the other hand, you can have a precise answer (small margin of error) but not be very confident about your answer.
When we increase the confidence level, the confidence interval becomes wider to be more confident that the population parameter is within the lower and upper boundaries. A wider margin of error means less accuracy. When one is more confident, one would have a harder time predicting the true parameter with the larger range of values. See Figure 7-3.

Figure 7-3
For instance, if we wanted to find the true mean grade for a statistics course using a 99% confidence critical value, we may get a very large margin of error, 75% ± 25%. This would say that we would be 99% confident that the average grade for all students is between 50% to 100%. This is of little help since that is anywhere between the grade range of F to an A. There are two ways to narrow this margin of error. The best way to reduce the margin of error is to increase the sample size, which decreases the standard deviation of the sampling distribution. When you take a larger sample, you will get a narrower interval. The other way to decrease the margin of error is to decrease your confidence level. When you decrease the confidence level, the critical value will be smaller. If we have a smaller margin of error then one can more accurately predict the population parameter.
Now look at how the sample size affects the size of the interval. Suppose the following Figure 7-4 represents confidence intervals calculated on a 95% interval.

Figure 7-4
A larger sample size from a representative sample makes the standard error smaller and hence the width of the interval narrower. Large samples are closer to the true population so the point estimate is pretty close to the true value.
The following website is an applet where you can simulate confidence intervals with different parameters, sample sizes and confidence levels, take a moment and play around with the applet:
http://www.rossmanchance.com/applets/ConfSim.html.
The vertical bar in Figure 7-5 represents the true population mean test score of 75 (which would be unknown in real life). If you were to compute 100 confidence intervals using a 95% confidence level, then approximately 95/100 = 95% would contain the true population mean. The figure shows the confidence intervals as horizontal lines. There are 95 confidence intervals that contain the population mean shown in green. There are 5 confidence intervals that did not capture the population mean within the interval endpoints that are shown in red.
The probability that one confidence interval contains the mean is either zero or one. However, if we were to repeat the same sampling process, the proportion of times that the confidence intervals would capture the populations parameter is (1 – \(\alpha\)), where α is the complement of the confidence level. As an example, if you have a 95% confidence interval of 0.65 < p < 0.73, then you would say, “If we were to repeat this process, then 95% of the time the interval 0.65 to 0.73 would contain the true population proportion.” This means that if you have 100 intervals, 95 of them will contain the true proportion, and 5% will not.
The incorrect interpretation is that there is a 95% probability that the true value of p will fall between 0.65 and 0.73. The reason that this interpretation is incorrect is that the true value is fixed out there somewhere. You are trying to capture it with this interval. This is the chance that your interval captures the true mean, and not that the true value falls in the interval.
In addition, a real-world interpretation depends on the situation. It is where you are telling people what numbers you found the parameter to lie between. Therefore, your real-world interpretation is where you tell what values your parameter is between. There is no probability attached to this statement. That probability is in the statistical interpretation.

Figure 7-5
In the following Figure 7-6, confidence intervals were simulated using a 90% confidence level and then again using the 99% confidence level. Each confidence level was run 100 times with sample sizes of n = 30, then again using a sample size of n = 100, holding all other variables constant.

Figure 7-6
Compare columns 1 & 2 with columns 3 & 4 in Figure 7-6. For columns 1 & 2, 90/100 = 90% of the confidence intervals contain the mean. For columns 3 & 4, 99/100 = 99% of the confidence intervals contain the mean Note the higher confidence level is wider for the same sample size.
Compare columns 1 & 3 in Figure 7-6 and you can see that the width of the confidence interval is wider for the 99% confidence level compared to the 90% confidence level. Holding all other variables constant the confidence interval captured the population mean 99% of the time. Then compare columns 2 & 4 to see similar results.
The wider confidence intervals will more likely capture the true population mean, however you will have less accuracy in predicting what the true mean is.
State the statistical and real-world interpretations of the following confidence intervals.
- Suppose you have a 95% confidence interval for the mean age a woman gets married in 2013 is 26 < μ < 28.
- Suppose a 99% confidence interval for the proportion of Americans who have tried cannabis as of 2019 is 0.55 < p < 0.61.
- Statistical Interpretation: We are 95% confident that the interval 26 < μ < 28 contains the population mean age of all women that got married in 2013.
- Real World Interpretation: We are 95% confident that the mean age of women that got married in 2013 is between 26 and 28 years of age.
- Statistical Interpretation: We are 99% confident that the interval 0.55 < p < 0.61 contains the population proportion of all Americans who have tried cannabis as of 2019.
- Real World Interpretation: We are 99% confident that the proportion of all Americans who have tried cannabis as of 2019 is between 55% and 61%.
Solution
a
b
“I'm not trying to prove anything, by the way. I'm a scientist and I know what constitutes proof. But the reason I call myself by my childhood name is to remind myself that a scientist must also be absolutely like a child. If he sees a thing, he must say that he sees it, whether it was what he thought he was going to see or not. See first, think later, then test. But always see first. Otherwise you will only see what you were expecting. Most scientists forget that.”
(Adams, 2002)
7.06: Sample Size for a Mean
Often, we need a specific confidence level, but we need our margin of error to be within a set range. We are able to accomplish this by increasing the sample size. However, taking large samples is often difficult or costly to accomplish. Thus, it is useful to be able to determine the minimum sample size necessary to achieve our confidence interval.
A confidence interval for a population mean µ with specific margin of error E and known population standard deviation σ is given by, \(n=\left(\frac{z_{\alpha / 2} \cdot \sigma}{E}\right)^{2}\).
Always round up to the next whole number.
Keep in mind that we rarely know the value of the population standard deviation. We can estimate σ by using a previous year’s standard deviation or a standard deviation from a similar study, a pilot sample, or by dividing the range by 4.
A researcher is interested in estimating the average salary of teachers. She wants to be 95% confident that her estimate is correct. In a previous study, she found the population standard deviation was $1,175. How large a sample is needed to be accurate within $100?
Solution
First find the z\(\alpha\)/2 for 95% confidence using Excel or your calculator, so z\(\alpha\)/2 = 1.96. Most of the time the margin of error = E follows the word “within” in the question, E = 100. The standard deviation σ = 1175. Replace each number into the formula: \(n=\left(\frac{1.96 \cdot 1175}{100}\right)^{2} = 530.38\). If we round down, we would not get “within” the $100 margin of error. Always round sample sizes up to the next whole number so that your margin of error will be within the specified amount. The larger the sample size, the smaller the confidence interval. The answer is n = 531.
7.07: t-Interval for a Mean
7.7.1 Student’s T-Distribution
A t-distribution is another symmetric distribution for a continuous random variable.

William Gosset was a statistician employed at Guinness and performed statistics to find the best yield of barley for their beer. Guinness prohibited its employees to publish papers so Gosset published under the name Student. Gosset’s distribution is called the Student’s t-distribution.
A t-distribution is another special type of distribution for a continuous random variable.
Properties of the t-distribution density curve:
- Symmetric, Unimodal (one mode) Bell-shaped.
- Centered at the mean μ = median = mode = 0.
- The spread of a t-distribution is determined by the degrees of freedom which are determined by the sample size.
- As the degrees of freedom increase, the t-distribution approaches the standard normal curve.
- The total area under the curve is equal to 1 or 100%.

Figure 7-7
Figure 7-7 shows examples of three different t-distributions with degrees of freedom of 1, 5 and 30. Note that as the degrees of freedom increase the distribution has a smaller standard deviation and will get closer in shape to the normal distribution.
The t-critical value that has 5% of the area in the upper tail for n = 13.
Solution
Use a t-distribution with the degrees of freedom, df = n – 1 = 13 – 1 = 12. Draw and shade the upper tail area as in Figure 7-8. Use the DISTR menu invT option. Note that if you have an older TI-84 or a TI-83 calculator you need to have the program INVT installed.
For this function, you always use the area to the left of the point. If want 5% in the upper tail, then that means there is 95% in the bottom tail area. tα = invT(area below t-score, df) = invT(0.95,12) = 1.782


You can download the INVT program to your calculator from http://MostlyHarmlessStatistics.com or use Excel =T.INV(0.95,12) = 1.7823.
Compute the probability of getting a t-score larger than 1.8399 with a sample size of 13.
Solution
To find the P(t > 1.8399) on the TI calculator, go to DISTR use tcdf(lower,upper,df). For this example, we would have tcdf(1.8399,∞,12). In Excel use =1-T.DIST(1.8399,12,TRUE) = 0.0453. P(t > 1.8399) = 0.0453.

Figure 7-9

7.7.2 T-Confidence Interval
Note that we rarely have a calculation for the population standard deviation so in most cases we would need to use the sample standard deviation as an estimate for the population standard deviation. If we have a normally distributed population with an unknown population standard deviation then the sampling distribution of the sample mean will follow a t-distribution.

Figure 7-10
A 100(1 - \(\alpha\))% Confidence Interval for a Population Mean μ: (σ unknown)
Choose a simple random sample of size n from a population having unknown mean μ.
The 100(1 - \(\alpha\))% confidence interval estimate for μ is given by \(\bar{x} \pm t_{\alpha / 2, n-1}\left(\frac{s}{\sqrt{n}}\right)\).
The df = degrees of freedom* are n – 1.
The degrees of freedom are the number of values that are free to vary after a sample statistic has been computed. For example, if you know the mean was 50 for a sample size of 4, you could pick any 3 numbers you like, but the 4th value would have to be fixed to have the mean come out to be 50. For this class we just need to know that degrees of freedom will be based on the sample size.
The sample mean \(\bar{x}\) is the point estimate for μ, and the margin of error is \(t_{\alpha / 2}\left(\frac{s}{\sqrt{n}}\right)\). Where t\(\alpha\)/2 is the positive critical value on the t-distribution curve with df = n – 1 and area 1 – \(\alpha\) between the critical values –t\(\alpha\)/2 and +t\(\alpha\)/2, as shown in Figure 7-11.

Figure 7-11
Before we compute a t-interval we will practice getting t critical values using Excel and the TI calculator’s built in tdistribution.
Compute the critical values –t\(\alpha\)/2 and +t\(\alpha\)/2 for a 90% confidence interval with a sample size of 10.
Solution
Draw and t-distribution with df = n – 1 = 9, see Figure 7-12. In Excel use =T.INV(lower tail area, df) =T.INV(0.95,9) or in the TI calculator use invT(lower tail area, df) = invT(0.95,9). The critical values are t = ±1.833

Figure 7-12

We can use Excel to find the margin of error when raw data is given in a problem. The following example is first done longhand and then done using Excel’s Data Analysis Tool and the T-Interval shortcut key on the TI calculator.
The yearly salary for mathematics assistant professors are normally distributed. A random sample of 8 math assistant professor’s salaries are listed below in thousands of dollars. Estimate the population mean salary with a 99% confidence interval.
66.0 75.8 70.9 73.9 63.4 68.5 73.3 65.9
Solution
First find the t critical value using df = n – 1 = 7 and 99% confidence, t\(\alpha\)/2 = 3.4995.

Then use technology to find the sample mean and sample standard deviation and substitute the numbers into the formula.
\(\bar{x} \pm t_{\alpha / 2, n-1}\left(\frac{s}{\sqrt{n}}\right) \Rightarrow 69.7125 \pm 3.4995\left(\frac{4.4483}{\sqrt{8}}\right) \Rightarrow 69.7125 \pm 5.5037 \Rightarrow(64.2088,75.2162)\)
The answer can be given as an inequality 64.2088 < µ < 75.2162 or in interval notation (64.2088, 75.2162).
We are 99% confident that the interval 64.2 and 75.2 contains the true population mean salary for all mathematics assistant professors.
We are 99% confident that the mean salary for mathematics assistant professors is between $64,208.80 and $75,216.20.
Assumption: The population we are sampling from must be normal* or approximately normal, and the population standard deviation σ is unknown. *This assumption must be addressed before using statistical inference for sample sizes of under 30.
TI-84: Press the [STAT] key, arrow over to the [TESTS] menu, arrow down to the [8:TInterval] option and press the [ENTER] key. Arrow over to the [Stats] menu and press the [ENTER] key. Then type in the mean, sample standard deviation, sample size and confidence level, arrow down to [Calculate] and press the [ENTER] key. The calculator returns the answer in interval notation. Be careful. If you accidentally use the [7:ZInterval] option you would get the wrong answer.
Alternatively (If you have raw data in list one) Arrow over to the [Data] menu and press the [ENTER] key. Then type in the list name, L1, leave Freq:1 alone, enter the confidence level, arrow down to [Calculate] and press the [ENTER] key.
TI-89: Go to the [Apps] Stat/List Editor, then press [2nd] then F7 [Ints], then select 2:TInterval. Choose the input method, data is when you have entered data into a list previously or stats when you are given the mean and standard deviation already. Type in the mean, standard deviation, sample size (or list name (list1), and Freq: 1) and confidence level, and press the [ENTER] key. The calculator returns the answer in interval notation. Be careful: If you accidentally use the [1:ZInterval] option you would get the wrong answer.
Excel Directions
Type the data into Excel. Select the Data Analysis Tool under the Data tab.

Select Descriptive Statistics. Select OK.

Use your mouse and click into the Input Range box, then select the cells containing the data. If you highlighted the label then check the box next to Labels in first row. In this case no label was typed in so the box is left blank. (Be very careful with this step. If you check the box and do not have a label then the first data point will become the label and all your descriptive statistics will be incorrect.)
Check the boxes next to Summary statistics and Confidence Level for Mean. Then change the confidence level to fit the question. Select OK.
The table output does not find the confidence interval. However, the output does give you the sample mean and margin of error.
The margin of error is the last entry labeled Confidence Level. To find the confidence interval subtract and add the margin of error to the sample mean to get the lower and upper limit of the interval in two separate cells.

The following screenshot shows the cell references to find the lower limit as =D3-D16 and the upper limit as =D3+D16. Make sure to put your answer in interval notation.

The answer is given as an inequality 64.2088 < µ < 75.2162 or in interval notation (64.2088, 75.2162).
We are 99% confident that the interval 64.2 and 75.2 contains the true population mean salary for all mathematics assistant professors.
Summary
A t-confidence interval is used to estimate an unknown value of the population mean for a single sample. We need to make sure that the population is normally distributed or the sample size is 30 or larger. Once this is verified we use the interval \(\bar{x}-t_{\alpha / 2, n-1}\left(\frac{s}{\sqrt{n}}\right)<\mu<\bar{x}+t_{\alpha / 2, n-1}\left(\frac{s}{\sqrt{n}}\right)\) to estimate the true population mean. Most of the time we will be using the t-interval, not the z-interval, when estimating a mean since we rarely know the population standard deviation. It is important to interpret the confidence interval correctly. A general interpretation where you would change what is in the parentheses to fit the context of the problem is: “One can be 100(1 – \(\alpha\))% confident that between (lower boundary) and (upper boundary) contains the population mean of (random variable in words using context and units from problem).”
7.08: Chapter 7 Exercises
1. Which confidence level would give the narrowest margin of error?
a) 80%
b) 90%
c) 95%
d) 99%
2. Suppose you compute a confidence interval with a sample size of 25. What will happen to the width of the confidence interval if the sample size increases to 50, assuming everything else stays the same? Choose the correct answer below.
a) Gets smaller
b) Stays the same
c) Gets larger
3. For a confidence level of 90% with a sample size of 35, find the critical z values.
4. For a confidence level of 99% with a sample size of 18, find the critical z values.
5. A researcher would like to estimate the proportion of all children that have been diagnosed with autism spectrum disorder (ASD) in their county. They are using 95% confidence level and the Centers for Disease Control and Prevention (CDC) 2018 national estimate that 1 in 68 \(\approx\) 0.0147 children are diagnosed with ASD. What sample size should the researcher use to get a margin of error to be within 2%?
6. A political candidate has asked you to conduct a poll to determine what percentage of people support her. If the candidate only wants a 9% margin of error at a 99% confidence level, what size of sample is needed?
7. A pilot study found that 72% of adult Americans would like an Internet connection in their car.
a) Use the given preliminary estimate to determine the sample size required to estimate the proportion of adult Americans who would like an Internet connection in their car to within 0.02 with 95% confidence.
b) Use the given preliminary estimate to determine the sample size required to estimate the proportion of adult Americans who would like an Internet connection in their car to within 0.02 with 99% confidence.
c) If the information in the pilot study was not given, determine the sample size required to estimate the proportion of adult Americans who would like an Internet connection in their car to within 0.02 with 99% confidence.
8. Out of a sample of 200 adults ages 18 to 30, 54 still lived with their parents. Based on this, construct a 95% confidence interval for the true population proportion of adults ages 18 to 30 that still live with their parents.
9. In a random sample of 200 people, 135 said that they watched educational TV. Find and interpret the 95% confidence interval of the true proportion of people who watched educational TV.
10. In a certain state, a survey of 600 workers showed that 35% belonged to a union. Find and interpret the 95% confidence interval of true proportion of workers who belong to a union.
11. A teacher wanted to estimate the proportion of students who take notes in her class. She used data from a random sample size of 82 and found that 50 of them took notes. The 99% confidence interval for the proportion of student that take notes is _______ < p < _________.
12. A random sample of 150 people was selected and 12% of them were left-handed. Find and interpret the 90% confidence interval for the proportion of left-handed people.
13. A survey asked people if they were aware that maintaining a healthy weight could reduce the risk of stroke. A 95% confidence interval was found using the survey results to be (0.54, 0.62). Which of the following is the correct interpretation of this interval?
a) We are 95% confident that the interval 0.54 < p < 0.62 contains the population proportion of people who are aware that maintaining a healthy weight could reduce the risk of stroke.
b) There is a 95% chance that the sample proportion of people who are aware that maintaining a healthy weight could reduce the risk of stroke is between 0.54 < p < 0.62.
c) There is a 95% chance of having a stroke if you do not maintain a healthy weight.
d) There is a 95% chance that the proportion of people who will have a stroke is between 54% and 62%.
14. Gallup tracks daily the percentage of Americans who approve or disapprove of the job Donald Trump is doing as president. Daily results are based on telephone interviews with approximately 1,500 national adults. Margin of error is ±3 percentage points. On December 15, 2017, the gallop poll using a 95% confidence level showed that 34% approved of the job Donald Trump was doing. Which of the following is the correct statistical interpretation of the confidence interval?
a) As of December 15, 2017, 34% of American adults approve of the job Donald Trump is doing as president.
b) We are 95% confident that the interval 0.31 < p < 0.37 contains the proportion of American adults who approve of the job Donald Trump is doing as president as of December 15, 2017.
c) As of December 15, 2017, 95% of American adults approve of the job Donald Trump is doing as president. d) We are 95% confident that the proportion of adult Americans who approve of the job Donald Trump is doing as president is 0.34 as of December 15, 2017.
15. A laboratory in Florida is interested in finding the mean chloride level for a healthy resident in the state. A random sample of 25 healthy residents has a mean chloride level of 80 mEq/L. If it is known that the chloride levels in healthy individuals residing in Florida is normally distributed with a population standard deviation of 27 mEq/L, find and interpret the 95% confidence interval for the true mean chloride level of all healthy Florida residents.
16. Out of 500 people sampled in early October 2020, 315 preferred Biden. Based on this, compute the 95% confidence interval for the proportion of the voting population that preferred Biden.
17. The age when smokers first start from previous studies is normally distributed with a mean of 13 years old with a population standard deviation of 2.1 years old. A survey of smokers of this generation was done to estimate if the mean age has changed. The sample of 33 smokers found that their mean starting age was 13.7 years old. Find the 99% confidence interval of the mean.
18. The scores on an examination in biology are approximately normally distributed with a known standard deviation of 20 points. The following is a random sample of scores from this year’s examination: 403, 418, 460, 482, 511, 543, 576, 421. Find and interpret the 99% confidence interval for the population mean scores.
19. The undergraduate grade point average (GPA) for students admitted to the top graduate business schools was 3.53. Assume this estimate was based on a sample of 8 students admitted to the top schools. Assume that the population is normally distributed with a standard deviation of 0.18. Find and interpret the 99% confidence interval estimate of the mean undergraduate GPA for all students admitted to the top graduate business schools.
20. The Food & Drug Administration (FDA) regulates that fresh albacore tuna fish that is consumed is allowed to contain 0.82 ppm of mercury or less. A laboratory is estimating the amount of mercury in tuna fish for a new company and needs to have a margin of error within 0.03 ppm of mercury with 95% confidence. Assume the population standard deviation is 0.138 ppm of mercury. What sample size is needed?
21. You want to obtain a sample to estimate a population mean age of the incoming fall term transfer students. Based on previous evidence, you believe the population standard deviation is approximately 5.3. You would like to be 90% confident that your estimate is within 1.9 of the true population mean. How large of a sample size is required?
22. SAT scores are distributed with a mean of 1,500 and a standard deviation of 300. You are interested in estimating the average SAT score of first year students at your college. If you would like to limit the margin of error of your 95% confidence interval to 25 points, how many students should you sample?
23. An engineer wishes to determine the width of a particular electronic component. If she knows that the standard deviation is 1.2 mm, how many of these components should she consider to be 99% sure of knowing the mean will be within 0.5 mm?
24. For a confidence level of 90% with a sample size of 30, find the critical t values.
25. For a confidence level of 99% with a sample size of 24, find the critical t values.
26. For a confidence level of 95% with a sample size of 40, find the critical t values.
27. The amount of money in the money market accounts of 26 customers is found to be approximately normally distributed with a mean of $18,240 and a sample standard deviation of $1,100. Find and interpret the 95% confidence interval for the mean amount of money in the money market accounts at this bank.
28. A professor wants to estimate how long students stay connected during two-hour online lectures. From a random sample of 25 students, the mean stay time was 93 minutes with a standard deviation of 10 minutes. Assuming the population has a normal distribution, compute a 95% confidence interval estimate for the population mean.
29. A random sample of stock prices per share (in dollars) is shown. Find and interpret the 90% confidence interval for the mean stock price. Assume the population of stock prices is normally distributed.
26.60 75.37 3.81 28.37 40.25 13.88 53.80 28.25 10.87 12.25
30. In a certain city, a random sample of executives have the following monthly personal incomes (in thousands) 35, 43, 29, 55, 63, 72, 28, 33, 36, 41, 42, 57, 38, 30. Assume the population of incomes is normally distributed. Find and interpret the 95% confidence interval for the mean income.
31. A tire manufacturer wants to estimate the average number of miles that may be driven in a tire of a certain type before the tire wears out. Assume the population is normally distributed. A random sample of tires is chosen and are driven until they wear out and the number of thousands of miles is recorded, find and interpret the 99% confidence interval for the mean using the sample data 32, 33, 28, 37, 29, 30, 22, 35, 23, 28, 30, 36.
32. Recorded here are the germination times (in days) for ten randomly chosen seeds of a new type of bean: 18, 12, 20, 17, 14, 15, 13, 11, 21, 17. Assume that the population germination time is normally distributed. Find and interpret the 99% confidence interval for the mean germination time.
33. A sample of the length in inches for newborns is given below. Assume that lengths are normally distributed. Find the 95% confidence interval of the mean length.
| Length | 20.8 | 16.9 | 21.9 | 18 | 15 | 20.8 | 15.2 | 22.4 | 19.4 | 20.5 |
34. Suppose you are a researcher in a hospital. You are experimenting with a new tranquilizer. You collect data from a random sample of 10 patients. The period of effectiveness of the tranquilizer for each patient (in hours) is as follows:
| Hours | 2 | 2.9 | 2.6 | 2.9 | 3 | 3 | 2 | 2.1 | 2.9 | 2.1 |
a) What is a point estimate for the population mean length of time?
b) What must be true in order to construct a confidence interval for the population mean length of time in this situation? Choose the correct answer below.
i. The sample size must be greater than 30.
ii. The population must be normally distributed.
iii. The population standard deviation must be known.
iv. The population mean must be known.
c) Construct a 99% confidence interval for the population mean length of time.
d) What does it mean to be "99% confident" in this problem? Choose the correct answer below.
i. 99% of all confidence intervals found using this same sampling technique will contain the population mean time.
ii. There is a 99% chance that the confidence interval contains the sample mean time.
iii. The confidence interval contains 99% of all sample times.
iv. 99% of all times will fall within this interval.
e) Suppose that the company releases a statement that the mean time for all patients is 2 hours. Is this possible? Is it likely?
35. Which of the following would result in the widest confidence interval?
a) A sample size of 100 with 99% confidence.
b) A sample size of 100 with 95% confidence.
c) A sample size of 30 with 95% confidence.
d) A sample size of 30 with 99% confidence.
36. The world’s smallest mammal is the bumblebee bat (also known as Kitti’s hog-nosed bat or Craseonycteris thonglongyai). Such bats are roughly the size of a large bumblebee. A sample of bats, weighed in grams, is given in the below. Assume that bat weights are normally distributed. Find the 99% confidence interval of the mean.
| Weight | |
| 2.11 | 1.53 |
| 2.27 | 1.98 |
| 2.27 | 2.11 |
| 1.75 | 2.06 |
| 1.92 | 2.01 |
37. The total of individual weights of garbage discarded by 20 households in one week is normally distributed with a mean of 30.2 lbs. with a sample standard deviation of 8.9 lbs. Find the 90% confidence interval of the mean.
38. A student was asked to find a 90% confidence interval for widget width using data from a random sample of size n = 29. Which of the following is a correct interpretation of the interval 14.3 < μ < 26.8? Assume the population is normally distributed.
a) There is a 90% chance that the sample mean widget width will be between 14.3 and 26.8.
b) There is a 90% chance that the widget width is between 14.3 and 26.8.
c) With 90% confidence, the width of a widget will be between 14.3 and 26.8.
d) With 90% confidence, the mean width of all widgets is between 14.3 and 26.8.
e) The sample mean width of all widgets is between 14.3 and 26.8, 90% of the time.
39. A researcher finds a 95% confidence interval for the average commute time in minutes using public transit is (15.75, 28.25). Which of the following is the correct interpretation of this interval?
a) We are 95% confident that all commute time in minutes for the population using public transit is between 15.75 and 28.25 minutes.
b) There is a 95% chance commute time in minutes using public transit is between 15.75 and 28.25 minutes.
c) We are 95% confident that the interval 15.75 < μ < 28.25 contains the sample mean commute time in minutes using public transportation.
d) We are 95% confident that the interval 15.75 < μ < 28.25 contains the population mean commute time in minutes using public transportation
7.09: Chapter 7 Formulas
|
Confidence Interval for One Proportion \(\begin{aligned} TI-84: 1-PropZInt |
Sample Size for Proportion \(n=p^{*} \cdot q^{*}\left(\frac{z_{\alpha / 2}}{E}\right)^{2}\) Always round up to whole number. If p is not given use p* = 0.5. E = Margin of Error |
|
Confidence Interval for One Mean Use z-interval when σ is given. Use t-interval when s is given. If n < 30, population needs to be normal. |
Z-Confidence Interval \(\bar{x} \pm z_{\frac{\alpha}{2}}\left(\frac{\sigma}{\sqrt{n}}\right)\) TI-84: ZInterval |
|
Z-Critical Values Excel: z\(\alpha\)/2 =NORM.INV(1–area/2,0,1) TI-84: z\(\alpha\)/2 = invNorm(1–area/2,0,1) |
t-Critical Values Excel: t\(\alpha\)/2 =T.INV(1–area/2,df) TI-84: t\(\alpha\)/2 = invT(1–area/2,df) |
|
t-Confidence Interval \(\bar{x} \pm t_{\alpha / 2}\left(\frac{s}{\sqrt{n}}\right)\) df = n – 1 TI-84: TInterval |
Sample Size for Mean \(n=\left(\frac{z_{\alpha / 2} \cdot \sigma}{E}\right)^{2}\) Always round up to whole number. E = Margin of Error |
From Mostly Harmless Statistics by Rachel L. Webb, adapted from LibreTexts. Licensed CC BY-SA 4.0. XYZ Homework OER web edition, adapted with 7 verified corrections (see errata).