Login
📚 Mostly Harmless Statistics
Chapters ▾
⇩ Download ▾

3.2 Measures of Spread

Variability is an important idea in statistics. If you were to measure the height of everyone in your classroom, every student gives you a different value. That means not every student has the same height. Thus, there is variability in people’s heights. If you were to take a sample of the income level of people in a town, every sample gives you different information. There is variability between samples too. Variability describes how the data are spread out. If the data are very close to each other, then there is low variability. If the data are very spread out, then there is high variability. How do you measure variability? It would be good to have a number that measures it. This section will describe some of the different measures of variability, also known as variation.

Numerical statistics for variation can show how spread out data is. The variation of data is relative, and is usually used when comparing two sets of similar data. When we are making inferences about an average, we can make better estimates when there is less variation in the data. The four most common measures of the “spread” of data are called the range, variance, standard deviation, and coefficient of variation.

A sample of house prices (in $1,000): 325, 375, 385, 395, 420, and 825, found the mean house price of $454,167. How much does this tell you about the price of all houses? Can you tell if most of the prices were close to the mean or were the prices really spread out? What is the highest price and the lowest price? All you know is that the center of the price is $454,167. What if you were approved for only $400,000 for a home loan, could you buy a home in this area? You need more information.

3.2.1 Range

The range of a set of data is the difference between the highest and the lowest data values (or maximum and minimum values). Note in statistics we only report a single number which represents the spread from the lowest to highest value.

3.2.2 Variance & Standard Deviation

The range does not really provide a very detailed picture of the variability. A better way to describe how the data is spread out is needed. Instead of looking at the distance as the highest value from the lowest, how about looking at the distance each value is from the mean? This spread is called the deviation.

The standard deviation is the average (mean) distance from a data point to the mean. It can be thought of as how much a typical data point differs from the mean.

Where x¯ is the sample mean, n is the sample size, and Σ means to find the sum.

The n – 1 in the denominator has to do with a concept called degrees of freedom (df). Dividing by the df makes the sample standard deviation a better approximation of the population standard deviation than dividing by n.

We rarely will find a population variance or standard deviation, but you will need to know the symbols.

The population variance formula: σ2=(xμ)2N.

The population standard deviation formula: σ=(xμ)2N.

The lower-case Greek letter σ pronounced “sigma” and σ2 represents the population variance, μ is the population mean, and N is the size of the population.

Note: the sum of the deviations should always be zero. Try not to round too much in the calculations for standard deviation since each rounding causes a slight error.

Descriptive statistics can be time-consuming to calculate by hand so use technology.

One Variable Statistics on the TI Calculator

The procedure for calculating the sample mean ( x¯ ) and the sample standard deviation (sx) for the TI calculator are shown below. Note, the TI calculator also gives you the population standard deviation (σx) because it does not know whether the data you input is a population or a sample. You need to decide which value you need to use, based on whether you have a population or sample. In almost all cases you have a sample and will be using sx. In addition, the calculator uses the notation of sx instead of just s. It is just a way for it to denote the information.

TI-84: Enter the data in a list and then press [STAT]. Use cursor keys to highlight CALC. Press 1 or [ENTER] to select 1:1-Var Stats. Press [2nd], then press the number key corresponding to your data list. Press [Enter] to calculate the statistics. Note: the calculator always defaults to L1 if you do not specify a data list.

Four TI-84 screens: data entered in list L1, the STAT CALC menu with 1:1-Var Stats highlighted, the 1-Var Stats input screen with List set to L1 and Calculate selected, and results showing x-bar = 2, Σx = 10, Σx² = 36, Sx = 2, σx = 1.788854382, and n = 5.

sx is the sample standard deviation. You can arrow down and find more statistics. Use the min and max to calculate the range by hand. To find the variance simply square the standard deviation.

TI-89: Press [APPS], select FlashApps then press [ENTER]. Highlight Stats/List Editor then press [ENTER]. Press [ENTER] again to select the main folder. To clear a previously stored list of data values, arrow up to the list name you want to clear, press [CLEAR], then press enter.

Two TI-89 Stats/List Editor screens: the F4 Calc menu with 1:1-Var Stats at the top, and the 1-Var Stats dialog with List set to list1, Freq set to 1, and Enter=OK and ESC=CANCEL buttons.

Press [F4], select 1: 1-Var Stats. To get the list name to the List box, press [2nd] [Var-Link], arrow down to list1 and press [Enter]. This will bring list1 to the List box. Press [Enter] to enter the list name and then enter again to calculate.

Use the down arrow key to see all the statistics.

TI-89 VAR-LINK screen listing list1 through list6 in the MAIN folder, with list1 highlighted ready to be selected.
Two TI-89 1-Var Stats result screens showing x-bar = 2, Σx = 10, Σx² = 36, Sx = 2, σx = 1.78885, n = 5, MinX = 0, Q1X = 0.5, MedX = 1, Q3X = 4, MaxX = 5, and sum of squared deviations 16.

Sx is the sample standard deviation. You can arrow down and find more statistics. Use the min and max to calculate the range by hand. To find the variance simply square the standard deviation or take the last sum of squares divided by n – 1.

Excel: Type in the data into one column, select the Data tab, and choose Data Analysis. Select Descriptive Statistics, and then select OK.

Excel Data ribbon with the Data Analysis dialog open over data values 5, 0, 1, 3, and 1 in column A; Descriptive Statistics is highlighted in the Analysis Tools list, with OK, Cancel, and Help buttons at the right.

Highlight the data for the Input Range, if you highlighted a label; check the Labels in first row box. Select the circle to the left of Output Range, then click into the box to the right of the Output Range and select one cell where you want the top left-hand corner of your summary table to start. Select the box next to Summary statistics, then select OK, see below.

Excel Descriptive Statistics dialog with Input Range $A$1:$A$6, Grouped By set to Columns, Labels in first row checked, Output Range $B$1, and the Summary statistics box checked.

We get the following summary statistics:

Excel summary statistics table for x: mean 2, standard error 0.8944, median 1, mode 1, standard deviation 2, sample variance 4, kurtosis −0.1875, skewness 0.9375, range 5, minimum 0, maximum 5, sum 10, and count 5.

In general, a “small” standard deviation means the data are close together (more consistent) and a “large” standard deviation means the data is spread out (less consistent). Sometimes you want consistent data and sometimes you do not. As an example, if you are making bolts, you want the lengths to be very consistent so you want a small standard deviation. If you are administering a test to see who can be a pilot, you want a large standard deviation so you can tell whom the good and bad pilots are.

What do “small” and “large” mean? To a bicyclist whose average speed is 20 mph, s = 20 mph is huge. To an airplane whose average speed is 500 mph, s = 20 mph is nothing. The “size” of the variation depends on the size of the numbers in the problem and the mean. Another situation where you can determine whether a standard deviation is small or large is when you are comparing two different samples. A sample with a smaller standard deviation is more consistent than a sample with a larger standard deviation.

If we were to compare the variability between two histograms. The standard deviation and variance measure the average spread from left to right. Take a moment and see if you can order the following histograms from the smallest to the largest standard deviation.

Histogram with unlabeled axes showing a U shape: the tallest bars are at the two ends and heights dip to the shortest bars in the middle.

Figure 3-16

Histogram with unlabeled axes showing a bell shape: the tallest bar is at the center and bar heights step down symmetrically toward short bars at both ends.

Figure 3-17

Histogram with unlabeled axes showing a roughly uniform shape: seven tall bars of nearly equal height across the whole plot.

FIgure 3-18

Histogram with unlabeled axes rising to a peak at the third bar and then steadily decreasing to the right, forming a mound with a longer right tail.

Figure 3-19

The histogram that has more of the data close to the mean will have the smallest standard deviation. The histogram that has more of the data towards the end points will have a larger standard deviation. Figure 3-16 will have the largest standard deviation since more of the data is grouped in the first and last class. Figure 3-17 will have the smallest standard deviation since more of the data is grouped in the center class which will be close to the mean in a symmetric distribution. Figures 3-18 and 3-19 are harder to compare without also having access to the mean and median to indicate skewness. However, Figure 3-19 does have smaller frequencies in the first and last three classes compared to Figure 3-18.

The correct order from smallest to largest standard deviation would be Figure 3-17, Figure 3-19, Figure 3-18, and then Figure 3-16.

One should not compare the range, standard deviation or variance of different data sets that have different units or scale.

3.2.3 Coefficient of Variation

The coefficient of variation, denoted by CVar or CV, is the standard deviation divided by the mean. The units on the numerator and denominator cancel with one another and the result is usually expressed as a percentage. The coefficient of variation allows you to compare variability among data sets when the units or scale is different.

Adapted from Mostly Harmless Statistics by Rachel Webb (Portland State University), hosted on LibreTexts (stats.libretexts.org) and licensed under CC BY-SA 4.0. Changes were made. License: CC-BY-SA-4.0.