Login
📚 Mostly Harmless Statistics
Chapters ▾

2 Organizing Data

2.01: Introduction

Once a sample is collected, we can organize and present the data in tables and graphs. These tables and graphs help summarize, interpret and recognize characteristics within the data more easily than raw data. There are many types of graphical summaries. We will concentrate mostly on the ones that we can use technology to create.

A population is a collection of all the measurements from the individuals of interest. Remember, in most cases you cannot collect data on the entire population, so you have to take a sample. Now you have a large number of data values. What can you do with them? Just looking at a large set of numbers does not answer our questions. If we organize the data into a table or graph, we can see patterns in the data. Ultimately, though, you want to be able to use that table or graph to interpret the data, to describe the distribution of the data set, explore different characteristics of the data and make inferences about the original population.

Some characteristics to look for in tables and graphs:

  1. Center: middle of the data set, also known as the average.
  2. Variation: how spread out is the data.
  3. Distribution: shape of the data.
  4. Outliers: data values that are far from the majority of the data.
  5. Time: changing characteristics of the data over time.

There is technology that will create most of the graphs you need, though it is important for you to understand the basics of how they are created.

Qualitative data are words describing a characteristic of the individual. Qualitative data is graphed using several different types of graphs, bar graphs, Pareto charts, and pie charts. Quantitative data are numbers that we count or measure. Quantitative data graphed using stem-and-leaf plots, dotplots, histograms, ogives, and time series.

The bar graph for quantitative data called a histogram looks similar to a bar graph for qualitative data, except there are some major differences. First, in a bar graph the categories can be put in any order on the horizontal axis. There is no set order for these data values. You cannot say how the data is distributed based on the shape, since the shape can change just by putting the categories in different orders. With quantitative data, the data are in specific orders since you are dealing with numbers. With quantitative data, you can talk about a distribution; the shape changes depending on how many categories you set up. This shape of the quantitative graph is called a frequency distribution.

This leads to the second difference from bar graphs. In a bar graph, the categories are determined by the name of the label. In quantitative data, the categories are numerical categories, and the frequencies are determined by how many categories (or what are called classes) you choose. There can be many different classes depending on the point of view of the author and how many classes there are. The third difference is that the categories touch with quantitative data, and there will be no gaps in the graph. The reason that bar graphs have gaps is to show that the categories do not continue on, as they do in quantitative data.

2.02: Tabular Displays

Frequency Tables for Quantitative Data

To create many of these graphs, you must first create the frequency distribution. The idea of a frequency distribution is to take the interval that the data spans and divide it into equal sized subintervals called classes. The grouped frequency distribution gives either the frequency (count) or the relative frequency (usually expressed as a percent) of individuals who fall into each class. When creating frequency distributions, it is important to note that the number of classes that are used and the value of the first class boundary will change the shape of, and hence the impression given by, the distribution. There is no one correct width to the class boundaries. It usually takes several tries to create a frequency distribution that looks just the way you want. As a reader and interpreter of such tables, you should be aware that such features are selected so that the table looks the way it does to show a particular point of view. For small samples, we usually have between four and seven classes, and as you get more data you will need more classes.

We will start with an example of a random sample of 35 ages from credit card applications given below in no particular order. Organize the data in a frequency distribution table.

46 47 49 25 46 22 42 32 39 24 46 40 39 27 25 30 31 29 33 27 46 21 29 20 26 39 26 25 25 26 35 49 33 26 30  
Solution

As you can see, the data is hard to interpret in this format. If we were to peruse through the data we could find that minimum age is 20 and the maximum age is 49. If we cut the age groups up in to 10-year intervals, we get only three classes 20-29, 30-39 and 40-49. Although this would work, if we had more classes we can sometimes see trends within the data at a more granular level. If we split the ages up in to 5-year intervals to 20-24, 25-29, 30-34, 35-39, 40-44, and 45-49 we would get six classes.

Note that the class limits should never overlap; we call these mutually exclusive classes. For example, if your class went from 20-25, 25-30 etc., the 25-year-olds would fall within both classes. Also, make sure each of the classes have the same width. A more formal way to pick your classes uses the following process.

Steps involved in making a frequency distribution table:

  1. Find the range = largest value – smallest value.
  2. Pick the number of classes to use. Usually the number of classes is between five and twenty. Five classes are used if there are a small number of data points and twenty classes if there are a large number of data points (over 1,000 data points).
  3. Class width = \(\frac{\text { range }}{\text { # of classes }}\). Always round up to the next integer (if the answer is already a whole number, go to the next integer). If you do not round up, your last class will not contain your largest data value, and you would have to add another class just for it. If you round up, then your largest data value will fall in the last class, and there are no issues.
  4. Create the classes. Each class has limits that determine which values fall in each class. To find the class limits, set the smallest value in the data set as the lower class limit for the first class. Then add the class width to the lower class limit to get the next lower class limit. Repeat until you get all the classes. The upper class limit for a class is one less than the lower limit for the next class.
  5. If your data value has decimal places, then round up the class width to the nearest value with the same number of decimal places as the original data. As an example, if your data was out to two decimal places, and you divided your range by the number of classes to get 4.8333, then the class width would round the second decimal place up and end on 4.84.
  6. The frequency for a class is the number of data values that fall in the class.

For the age data let us use 6 classes. Find the range by taking 49 – 20 = 29 and divide this by the number of classes 29/6 = 4.8333. Round this number up to 5 and use 5 for the class width. Once you determine your class width and class limits, place each of these classes in a table and then tally up the ages that fall within each class. Count the tally marks and record the number in the frequency table.

The total of the frequency column should be the number of observations in the data. You may want to total the frequencies to make sure you did not leave any of the numbers out of the table.

Class Limits Tally   Class Frequency
20-24 4   20-24 4
25-29 12   25-29 12
30-34 6   30-34 6
35-39 4   35-39 4
40-44 2   40-44 2
45-49 7   45-49 7
      Total 35

Figure 2-1

Using the frequency table, we can now see that there are more people in the 25- 29 year-old class, followed by the 45-49 year-old class.

We call this most frequent category the modal class. There may be no mode at all or more than one mode.

Frequency Tables for Qualitative Data

A frequency distribution can also be made for qualitative data.

Suppose you have the following data for which type of car students at a college drive. Make a frequency table to summarize the data.

Ford Honda Nissan Chevy Chevy Chevy Chevy Toyota Chevy Saturn Honda Toyota Nissan Honda Toyota Toyota Nissan Ford Toyota Chevy Toyota Ford Chevy Chevy Chevy Nissan Toyota Toyota Ford Nissan Kia Nissan Nissan Nissan Honda Nissan Mercedes Honda Toyota Toyota Chevy Chevy Porsche Chevy Toyota Toyota Ford Hyundai Honda Nissan
Solution

The list of data is hard to analyze, so you need to summarize it. The classes in this case are the car brands. However, several car brands only have one car in the list. In that case, it is easier to make a category called “other” for the categories with low frequencies.

Count how many of each type of cars there are: there are 12 Chevys, 5 Fords, 6 Hondas, 10 Nissans, 12 Toyotas, and 5 other brands (Hyundai, Kia, Mercedes, Porsche, and Saturn). Place the other brands into a frequency distribution table alphabetically:

Category Frequency
Chevy 12
Ford 5
Honda 6
Nissan 10
Toyota 12
Other 5
Total 50

For nominal data, either alphabetize the classes or arrange the classes from most frequent to least frequent, with the “other” category always at the end. For ordinal data put the classes in their order with the “other” category at the end.

Relative Frequency Tables

Frequencies by themselves are not as useful to tell other people what is going on in the data. If you want to know what percentage the category is of the total sample then we can use the relative frequency of each category. The relative frequency is just the frequency divided by the total. The relative frequency is the proportion in each category and may be given as a decimal, percentage, or fraction.

Using the car data’s frequency distribution, we will create a third column labeled relative frequency. Take each frequency and divide by the sample size, see Figure 2-2. The relative frequencies should add up to one (ignoring rounding error).

Type of Car Frequency Relative Frequency
Chevy 12 12/50 = 0.24
Ford 5 5/50 = 0.1
Honda 6 6/50 = 0.12
Nissan 10 10/50 = 0.2
Toyota 12 12/50 = 0.24
Other 5 5/50 = 0.1
Total 50 1

Figure 2-2

Many people understand percentages better than proportions so we may want to multiply each of these decimals by 100% to get the following relative frequency percent table.

Type of Car Percent
Chevy 24
Ford 10
Honda 12%
Nissan 20%
Toyota 24%
Other 10%
Total 100%

Figure 2-3

We can summarize the car data and see that for college students Chevy and Toyota make up 48% of the car models.

Excel

Recall the frequency table for the credit card applicants. The relative frequency table for the random sample of 35 ages from credit card applications follows.

ClassFrequencyRelative Frequency 20-24 4 4/35 = 0.1143 25-29 12 12/35 = 0.3429 30-34 6 6/35 = 0.1714 35-39 4 4/35 = 0.1143 40-44 2 2/35 = 0.0571 45-49 7 7/35 = 0.2 Total 35 1

Making a relative frequency table using Excel

Solution

In Excel, type your frequencies into a column and then in the next column type in =(cell reference number)/35. Then copy and paste the formula in cell B2 down the page. If you used the cell reference number, Excel will automatically change the copied cells to the next row down.

You get the following relative frequency table. The sum of the relative frequencies will be one (the sum may add to 0.9999 or 1.0001 if you are doing the calculations by hand due to rounding).

clipboard_e779c3f7abe885400c555d9449e6bbf59.png

To get Excel to show the percentage instead of the proportion by highlighting the relative frequencies and selecting the percent % button on the Home tab.

clipboard_e6eccc872d84d619559fdf0275ece321c.png

Class Frequency Relative Frequency Relative Frequency Percent
20-24 4 0.1143 11%
25-29 12 0.3429 34%
30-34 6 0.1714 17%
35-39 4 0.1143 11%
40-44 2 0.0571 6%
45-49 7 0.2 20%
Total 35 1 100%

The relative frequency table lets us quickly see that a little more than half 34% + 20% = 54% of the ages of the credit card holders are between the ages of 25 – 29 and 45 – 49 years-old.

Cumulative & Cumulative Relative Frequency Tables

Another useful piece of information is how many data points fall below a particular class. As an example, a teacher may want to know how many students received below a 70%, a doctor may want to know how many adults have cholesterol above 160, or a manager may want to know how many stores gross less than $2,000 per day. This calculation is known as a cumulative frequency and is used for ordinal or quantitative data. If you want to know what percent of the data falls below a certain class, then this fact would be a cumulative relative frequency.

To create a cumulative frequency distribution, count the number of data points that are below the upper class limit, starting with the first class and working up to the top class. The last upper class should have all of the data points below it.

Recall the credit card applicants. Make a cumulative frequency table.

ClassFrequencyRelative Frequency Percent 20-24 4 11% 25-29 12 34% 30-34 6 17% 35-39 4 11% 40-44 2 6% 45-49 7 20% Total 35 100%
Solution

To find the cumulative frequency, carry over the first frequency of 4 to the first row of the cumulative frequency column. Then take this 4 and add it to the next frequency of 12 to get 16 for the second cumulative frequency value. For the third cumulative frequency, take 16 and add to the third frequency of 6 to get 22. Keep doing this additive process until you finish the column.

The cumulative frequency in the last class should have the last number as the sample size. The cumulative relative frequency is just the cumulative frequency divided by the sample size. The cumulative relative frequencies should have one as the last number.

Class Frequency Relative Frequency Cumulative Frequency Cumulative Relative Frequency
20-24 4 0.1143 4 4/35 = 0.1143
25-29 12 0.3429 4 + 12 = 16 16/35 = 0.4571
30-34 6 0.1714 16 + 6 = 22 22/35 = 0.6286
35-39 4 0.1143 22 + 4 = 26 26/35 = 0.7429
40-44 2 0.0571 26 + 2 = 28 28/35 = 0.8
45-49 7 0.2 28 + 7 = 35 35/35 = 1
Total 35 1    

Use Excel to add up the values for you.

clipboard_e83dd7973320d277d3e1f83dd6c4d1d4b.png

You can also express the cumulative relative frequencies as a percentage instead of a proportion.

Class Cumulative Frequency Cumulative Relative Frequency
20-24 4 11
25-29 16 46
30-34 22 63
35-39 26 74
40-44 28 80
45-49 35 100

If a manager wanted to know how many applicants were under the age of 40 we could look across the 35-39 year-old class to see that there were 26 applicants that were under 40 (39 years old or younger), or if we used the cumulative relative frequency, 74% of the applicants were under 40 years old. If the manager wants to know how many applicants were over 44, we could subtract 100% – 80% = 20%.

Contingency Tables

A contingency table provides a way of portraying data that can facilitate calculating probabilities. A contingency table summarizes the frequency of two qualitative variables. The table displays sample values in relation to two different variables that may be dependent or contingent on one another. There are other names for contingency tables. Excel calls the contingency table a pivot table. Other common names are two-way table, cross-tabulation, or cross-tab for short.

A fitness center coach kept track of members over the last year. They recorded if the person stretched before they exercised, and whether they sustained an injury. The following contingency table shows their results. Find the relative frequency for each value in the table.

  Injury No Injury Stretched 52 270 Did Not Stretch 21 57
Solution

Each value in the table represents the number of times a particular combination of variable outcomes occurred, for example, there were 57 members that did not stretch and did not sustain an injury. It is helpful to total up the categories. The row totals provide the total counts across each row (i.e. 52 + 270 = 322), and column totals are total counts down each column (i.e. 52 + 21 = 73). See Figure 2-4.

  Injury No Injury Total
Stretched 52 270 322
Did Not Stretch 21 57 78
Total 73 327 400

Figure 2-4

We can quickly summarize the number of athletes for each category. The bottom right-hand number in Figure 2-4 is the grand total and represents the total number of 400 people. There were 322 people that stretched before exercising. There were 73 people that sustained an injury while exercising, etc.

If we find the relative frequency for each value in the table, we can find the proportion of the 400 people for each category. To find a relative frequency we divide each value in the table by the grand total. See Figure 2-5.

  Injury No Injury Total
Stretched 52/400 = =0.13 270/400 = 0.675 322/400 = 0.805 or 0.13 + 0.675 = 0.805
Did Not Stretch 21/400 = 0.0525 57/400 = 0.1425 78/400 = 0.195 or 0.0525 + 0.1425 = 0.195
Total 73/400 = 0.1825 or 0.13 + 0.0525 = 0.1825 327/400 = 0.8175 or 0.675 + 0.1425 = 0.8175 400/400 = 1 or the sum of either the row or column totals

Figure 2-5

When data is collected, it is usually presented in a spreadsheet where each row represents the responses from an individual or case.

“Of course, one never has the slightest notion what size or shape different species are going to turn out to be, but if you were to take the findings of the latest Mid‐Galactic Census report as any kind of accurate Guide to statistical averages you would probably guess that the craft would hold about six people, and you would be right. You'd probably guessed that anyway. The Census report, like most such surveys, had cost an awful lot of money and didn't tell anybody anything they didn't already know - except that every single person in the Galaxy had 2.4 legs and owned a hyena. Since this was clearly not true the whole thing had eventually to be scrapped.” (Adams, 2002)

Make a pivot table using Excel. A random sample of 500 records from the 2010 United States Census were downloaded to Excel. Below is an image of just the first 20 people.

There are seven variables:

2.03: Graphical Displays

Statistical graphs are useful in getting the audience’s attention in a publication or presentation. Data presented graphically is easier to summarize at a glance compared to frequency distributions or numerical summaries. Graphs are useful to reinforce a critical point, summarize a data set, or discover patterns or trends over a period of time. Florence Nightingale (1820-1910) was one of the first people to use graphical representations to present data. Nightingale was a nurse in the Crimean War and used a type of graph that she called polar area diagram, or coxcombs to display mortality figures for contagious diseases such as cholera and typhus.

clipboard_eb2e7c2490074c70c342069f0909a448a.png

Nightingale

clipboard_eefb1ad16c8747346826dbc9816844907.png

Nightingale-mortality.jpg. (2021, May 18). Wikimedia Commons, the free media repository. Retrieved July 2021 from https://commons.wikimedia.org/w/index.php?title=File:Nightingale-mortality.jpg&oldid=561529217.

It is hard to provide a complete overview of the most recent developments in data visualization with the onset of technology. The development of a variety of highly interactive software has accelerated the pace and variety of graphical displays across a wide range of disciplines.

2.3.1 Stem-and-Leaf Plot

Stem-and-leaf plots (or stemplots) are a useful way of getting a quick picture of the shape of a distribution by hand. Turn the graph sideways and you can see the shape of your data. You can now easily identify outliers. Each observation is divided into two pieces; the stem and the leaf. If the number is just two digits then the stem would be the tens digit and the leaf would be the ones digit. When a number is more than two digits then the cut point should split the data into enough classes that is useful to see the shape of the data.

To create a stem-and-leaf plot:

  1. Separate each observation into a stem and a leaf.
  2. Write the stems in a vertical column in ascending order (from smallest to largest). Fill in missing numbers even if there are gaps in the data. Draw a vertical line to the right of this column.
  3. Write each leaf in the row to the right of its stem, in increasing order.

Create a stem-and-leaf plot for the sample of 35 ages.

46 47 49 25 46 22 42 24 46 40 39 27 25 30 33 27 46 21 29 20 26 25 25 26 35 49 33 26 32 31 39 30 39 29 26
Solution

Divide each number so that the tens digit is the stem and the ones digit is the leaf. The smallest observation is 20. The stem = 2 and the leaf = 0. The next value is 21 and the stem = 2 and the leaf = 1, up to the last value of 49 which would have a stem = 4 and a leaf = 9. If we use the tens categories we have the stems 2, 3 and 4. Line up the stems without skipping a number even if there are no values in that stem. In other words, the stems should have equal spacing (for example, count by ones, tens, hundreds, thousands, etc.). Then place a vertical line to the right of the stems. In each row put the leaves with a space between each leaf. Sort each row from smallest to largest. In Figure 2- 6 the 2 | 0 = 20.

\begin{array}{l|llllllllllllllll}
2 & 0 & 1 & 2 & 4 & 5 & 5 & 5 & 5 & 6 & 6 & 6 & 6 & 7 & 7 & 9 & 9 \\
3 & 0 & 0 & 1 & 2 & 3 & 3 & 5 & 9 & 9 & 9 \\
4 & 0 & 2 & 6 & 6 & 6 & 6 & 7 & 9 & 9
\end{array}

Figure 2-6

It is hard to see the shape with so few classes and so many leaves in each class.

We can break each stem in half, putting leaves 0-4 in the first row and 5-9 in the second row, as in Figure 2-7.

\begin{array}{l|llllllllllll}
2 & 0 & 1 & 2 & 4 \\
2 & 5 & 5 & 5 & 5 & 6 & 6 & 6 & 6 & 7 & 7 & 9 & 9 \\
3 & 0 & 0 & 1 & 2 & 3 & 3 \\
3 & 5 & 9 & 9 & 9 \\
4 & 0 & 2 \\
4 & 6 & 6 & 6 & 6 & 7 & 9 & 9
\end{array}

Figure 2-7

Now, add labels and make sure the leaves are in ascending order. Be careful to line the leaves up in columns. You need to be able to compare the lengths of the rows when you interpret the graph.

Imagine lines around the leaves and turn the graph 90 degrees to the left. You can now see in Figure 2-8 the shape of the distribution. Note that Excel uses the upper class limit for the axis label.

clipboard_ec153816438a1cb237e9560821f1db3ae.png

Figure 2-8

If a leaf takes on more than the ones category then supply a footnote at the bottom of the plot with the units.

A small sample of house prices in thousands of dollars was collected: 375, 189, 432, 225, 305, 275. Make a stem-and-leaf plot.

Solution

If we were to split the stem and leaf between the ones and tens place, then we would need stems going from 18 up to 43. Twenty-six stems for only six data points is too many. The next break then for a stem would be between the tens and hundreds. This would give stems from 1 to 4. Then each leaf will be the ones and tens. For example, then number 375 would have a stem = 3 and a leaf = 75.

\begin{array}{l|ll}
1 & 89 \\
2 & 25 & 75 \\
3 & 05 & 75 \\
4 & 32
\end{array}

Leaf = $1000

A small sample of coffee prices: 3.75, 1.89, 4.32, 2.25, 3.05, 2.75 was collected. Make a stem-and-leaf plot.

Solution

\begin{array}{l|ll}
1 & 89 \\
2 & 25 & 75 \\
3 & 05 & 75 \\
4 & 32
\end{array}

Leaf = $0.01

Note that the last two stem-and-leaf plots look identical except for the footnote. It is important to include units to tell people what the stems and leaves mean by inserting a legend.

Back-to-back stem-and-leaf plots let us compare two data sets on the same number line. The two samples share the same set of stems. The sample on the right is written backward from largest leaf to smallest leaf, and the sample on the left has leaves from smallest to largest.

Use the following back-to-back stem-and-leaf plot to compare pulse rates before and after exercise.

Pulse Rates: Before and After Exercise stemleaf chart with 2 data series 2026-07-23T09:02:26.215430 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Back-to-back stem-and-leaf plot comparing pulse rates before exercise (n=28) and after exercise (n=26); stems are the tens digit of the pulse rate, leaves the ones digit. (regenerated with XYZChart for the XYZ Homework web edition.)

Solution

The group on the left has leaves going in descending order and represent the pulse rates before exercise. The stems are in the middle column. The group on the right has leaves going in ascending order and represent the pulse rates after exercise. The first row has pulse rates of 62, 65, 66, 67, 68, 68 and 69. The last row of pulse rates are 124, 125, and 128.

2.3.2 Histogram

A histogram is a graph for quantitative data (we call these bar graphs for qualitative data). The data is divided into a number of classes. The class limits become the horizontal axis demarcated with a number line and the vertical axis is either the frequency or the relative frequency of each class. Figure 2-9 is an example of a histogram.

The histogram for quantitative data looks similar to a bar graph, except there are some major differences.

First, in a bar graph the categories can be put in any order on the horizontal axis. There is no set order for these nominal data. You cannot say how the data is distributed based on the shape, since the shape can change just by putting the categories in different orders. With quantitative data, the data are in a specific order, since you are dealing with numbers. With quantitative data, you can talk about a distribution shape.

This leads to the second difference from bar graphs. In a bar graph, the categories that you made in the frequency table were the words used for the category name. In quantitative data, the categories are numerical categories, and the numbers are determined by how many classes you choose. If two people have the same number of categories, then they will have the same frequency distribution. Whereas in qualitative data, there can be many different categories depending on the point of view of the author.

The third difference is that the bars touch with quantitative data, and there will be no gaps in the graph. The reason that bar graphs have gaps is to show that the categories do not continue on, as they do in quantitative data. Since the graph for quantitative data is different from qualitative data, it is given a different name of histogram.

Some key features of a histogram:

clipboard_e5c43615a8306b45a4edee77c778a4993.png
Figure 2-9

To create a histogram, you must first create a frequency distribution. Software and calculators can create histograms easily when a large amount of sample data is being analyzed.

Excel

To create a histogram in Excel you will need to first install the Data Analysis tool.

If your Data Analysis is not showing in the Data tab, follow the directions for installing the free add-in here: https://support.office.com/en-us/article/Load-the-Analysis-ToolPak-in-Excel-6a63e598-cd6d-42e3-9317- 6b40ba1a66b4.

Type in the data into one blank column in any order. If you want to have class widths other than Excel’s default setting, type in a new column the endpoints of each class found in your frequency distribution, these are called the bins in Excel.

Using the sample of 35 ages, make a histogram using Excel.

46 47 49 25 46 22 42 24 46 40 39 27 25 30 33 27 46 21 29 20 26 25 25 26 35 49 33 26 32 31 39 30 39 29 26
Solution

Type the data in any order into column A and the bins in order in column B as shown below. Then select the Data tab, select Data Analysis, select Histogram, then select OK.

clipboard_e8a6e3e25cf0597e35771f4ca257c4f2f.png

In the dialogue box, click into the Input Range box, then use your mouse and highlight the ages including the label.

Then click into the Bin Range box and use your mouse to highlight the bins including the label.

Select the box for Labels only if you included the labels in your ranges. You can have your output default to a new worksheet, or select the circle to the left of Output Range, click into the box to the right of Output Range and then select one blank cell on your spreadsheet where you want the top left-hand corner of your table and graph to start. Then check the boxes next to Cumulative Percentage and Chart Output. Then select OK, and see below.

clipboard_ef3a2bf7a55d0fcda3237c4bb5a5df8d8.png

A histogram needs to have bars that touch, which is not the default in Excel. To get the bars to touch, right-click on one of the blue bars and select Format Data Series and slide the Gap Width to 0%.

clipboard_eca6723d9b151cd9094a0417593d812d1.png

Excel produces both a frequency table and a histogram. The table has the frequencies and the cumulative relative frequencies.

Bin Frequency Cumulative %
24 4 11.43%
29 12 45.71%
34 6 62.86%
39 4 74.29%
44 2 80.00%
49 7 100.00%
More 0 100.00%

The histogram has bars for the height of each frequency and then makes a line graph of the cumulative relative frequencies over the bars. This red line is a line graph of the cumulative relative frequencies, also called an ogive and is discussed in a later section.

clipboard_e6d641d5f1f3622850be570d011965ded.png

It is important to note that the number of classes that are used and the value of the first class boundary will change the shape of the histogram.

A relative frequency histogram is when the relative frequencies are used for the vertical axis instead of the frequencies and the y-axis will represent a percent instead of the number of people.

In Excel, after you create your histogram, you can manually change the frequency column to the relative frequency values by dividing each number by the sample size. Here is a screen shot just as the last number was changed, note as soon as you hit enter the bars will shrink and adjust.

clipboard_e10ddf9990799aa36139dd9af712b6ce8.png

After the last value =7/35 was entered and the label changed to Relative Frequency you get the following graph.

Histogram histogram chart with 1 data series 2026-07-23T09:02:26.264099 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Relative-frequency histogram of 35 ages (Mostly Harmless Statistics, Webb, §2.3). Six equal-width age classes (20-24, 25-29, 30-34, 35-39, 40-44, 45-49) with relative fre (regenerated with XYZChart for the XYZ Homework web edition.)

The shape of the histogram will be the same for the relative frequency distribution and the frequency distribution; the height, though, is the proportion instead of frequency.

TI-84: To make a histogram, enter the data by pressing [STAT]. The first option is already highlighted (1:Edit) so you can either press [ENTER] or [1]. Make sure the cursor is in the list, not on the list name and type the desired values pressing [ENTER] after each one.

clipboard_efb333a061afe3c407ce6b8855df18d9b.png

Press [2nd] [QUIT] to return to the home screen. To clear a previously stored list of data values, arrow up to the list name you want to clear, press [CLEAR], and then press enter. An alternative way is press [STAT], press 4 for 4:ClrList, press [2nd], then press the number key corresponding to the data list you wish to clear, for example, [2nd] [1] will clear L1, then press [ENTER]. After you enter the data, press [2nd] [STAT PLOT]. Select the first plot by hitting [Enter] or the number [1:Plot 1]. Turn the plot [On] by moving the cursor to On and selecting Enter. Select the Histogram option using the right arrow keys. Select [Zoom], then [ZoomStat].

clipboard_eb246b06885a6ad8e3b13354397bdc72e.png

You can see and change the class width by selecting [Window], then change the minimum x value Xmin=20, the maximum x value Xmax=50, the x-scale to Xscl=5 and the minimum y value Ymin=-6.5 and the maximum y value to Ymax=14. Select the [GRAPH] button. We get a similar looking Histogram compared to the stem-and-leaf plot and Excel histogram. Select the [TRACE] button to see the height of each bar and the classes.

clipboard_eacdd5632d21b885a8fcceab06046a10f.png

TI-89: First, enter the data into the Stat/List editor under list 1. Press [APP] then scroll down to Stat/List Editor, on the older style TI-89 calculators, go into the Flash/App menu, and then scroll down the list. Make sure the cursor is in the list, not on the list name, and type the desired values pressing [ENTER] after each one. To clear a previously stored list of data values, arrow up to the list name you want to clear, press [CLEAR], and then press enter. After you enter the data, select Press [F2] Plots, scroll down to [1: Plot Setup] and press [Enter].

clipboard_eed6707889170d9327de8ec8a35301a0c.png

Select [F1] Define. Use your arrow keys to select Histogram for Type, and then scroll down to the x-variable box. Press [2nd] [Var-Link] this key is above the [+] sign. Then arrow down until you find your List1 name under the Main file folder. Then press [Enter] and this will bring the name List1 back to the menu. You will now see that Plot1 has a small picture of a histogram. To view the histogram, select [F5] [Zoom Data].

clipboard_e9f29fc4f8f772d101c259e0b2e46ca3b.png

The histogram looks a little different from Excel; you can change the settings for the bucket to match your table. Press [♦] [F2:Window]. Change the minimum x value xmin=20, the maximum x value xmax=50, the x-scale to xscl=5 and the minimum y value ymin=-6.5 and the maximum y value to ymax=14. Then press the [♦] [F3:GRAPH] button. Select [F3:Trace] to see the frequency for each bar. Then use your left and right arrow keys to move to the other bars.

clipboard_e8470841df965c98e3bd3238fb36877fb.png

Make a histogram for the following random sample of student rent prices using Excel.

1500 1350 350 1200 850 900 1500 1150 1500 900 1400 1100 1250 600 610 960 890 1325 900 800 2550 495 1200 690
Solution

Start by making a relative frequency distribution table with 7 classes.

  1. Find the range: largest value – smallest value = 2550 – 350 = 2200, range = $2,200.
  2. Find the class width: width = \(\frac{\text { range }}{\text { 7 }}\) = \(\frac{\text { 2000 }}{\text { 7 }}\) ≈ 314.286. Round up to 315. Always round up to the next integer even if the width is already an integer.
  3. Find the class limits: Start at the smallest observation. This is the lower class limit for the first class. Add the class width to get the lower limit of the next class. Keep adding the class width to get all the lower limits, 350 + 315 = 665, 665 + 315 = 980, 980 + 315 = 1295, etc. The upper limit is one unit less than the next lower limit: so, for the first class the upper class limit would be 665 – 1 = 664. When you have all 7 classes, make sure the last number, in this case the 2550, is at least as large as the largest value in the data. If not, you made a mistake somewhere.

Using Excel: Type the raw data in Excel in column A, the right-hand class endpoints for the bins in column B. Select Data, Data Analysis, Histogram.

Select the Input Range, Bin Range, Labels (if you selected them), output option, Chart Output, then OK.

See finished histogram below in Figure 2-13.

Distribution of Ages histogram chart with 1 data series 2026-07-23T03:25:34.124555 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Figure 2-13 Histogram of the age distribution (regenerated with XYZChart for the XYZ Homework web edition).

Figure 2-10

By hand: Tally and find the frequency of the data.

Frequency Distribution for Monthly Rent

Class Limits Tally Frequency Relative Frequency
350-664 4 4 0.1667
665-979 8 8 0.333
980-1294 5 5 0.2083
1295-1609 6 6 0.25
1610-1924 0 0 0
1925-2239 0 0 0
2240-2554 1 1 0.0417
Total 0 24 1

Figure 2-11

Make sure the total of the frequencies is the same as the number of data points and the total of the relative frequency is one. Since we want the bars on the histogram to touch, the number line needs to use the class boundaries that are half way between the endpoints of the class limits. Start by finding the distance between the class endpoints and divide by two: (665-664)/2 = 0.5. Then subtract 0.5 from the left-hand side of each class limit and this will give you the points to use on the x-axis: 349.5, 664.5, 979.5, 1294.5, 1609.5, 1924.5, 2239.5, and 2554.5. Then draw your graph as in Figure 2-12. You can use frequencies or relative frequencies for the y-axis.

clipboard_ef61b10d754b81260a491e14b9ced98cb.png

Figure 2-12

clipboard_ee724d4ce0b96d53e4b01d8b31309d3e1.png

Figure 2-13

Reviewing the graph in Figure 2-13, you can see that most of the students pay around $750 per month for rent, with about $1,500 being the other common value. Most students pay between $600 and $1,600 per month for rent. Of course, these values are just estimates pulled from the graph.

There is a large gap between the $1,500 class and the highest data value. This seems to say that one student is paying a great deal more than everyone else is. This value may be an outlier.

An outlier is a data value that is far from the rest of the values. It may be an unusual value or a mistake. It is a data value that should be investigated. In this case, the student lives in a very expensive part of town, thus the value is not a mistake, and is just very unusual. There are other aspects that can be discussed, but first some other concepts need to be introduced.

2.3.3 Ogive

The line graph for the cumulative or cumulative relative frequency is called an ogive (oh-jyve). To create an ogive, first create a scale on both the horizontal and vertical axes that will fit the data. Then plot the points of the upper class boundary versus the cumulative (or cumulative relative) frequency. Make sure you include the point with the lowest class and the zero cumulative frequency. Then just connect the dots.

The steeper the line the more accumulation occurs across the corresponding class. If the line is flat then the frequency for that class is zero. The ogive graph will always be going uphill from left to right and should never dip below the previous point. Figure 2-14 is an example of an ogive.

Ogive comes from the uphill shape used in architecture. Here is an example of an ogive in the East Hall staircase at PSU.

clipboard_eff7eea06a1faa7ff3d14232fda37fd5a.png

clipboard_e30e7c6658ce1b30d2d0cd2e97c260ac7.png

Figure 2-14

Make an ogive for the following random sample of rent prices students pay with the corresponding cumulative frequency distribution table.

1500 1350 350 1200 850 900 1250 600 610 960 890 1325 1500 1150 1500 900 1400 1100 900 800 2550 495 1200 690
Class Limits Frequency Cumulative Frequency
350 - 664 4 4
665 - 979 8 12
980 - 1294 5 17
1295 - 1609 6 23
1610 - 1924 0 23
1925 - 2239 0 23
2240 - 2554 1 24
Solution

Find the class boundaries, 349.5, 664.5 … use these for the tick mark labels on the horizontal x-axis, the same as what was used for the histogram. The y-axis uses the cumulative frequencies. The largest cumulative frequency is 24. Every third number is marked on the y-axis units. See Figure 2-15 and Figure 2-16.

By hand:

clipboard_e12c6072fa9696e35bc747b124e87b09d.png

Figure 2-15

Using software:

clipboard_e965ee5ce6a583dd335f9d627a70de79b.png

Figure 2-16

The usefulness of an ogive is to allow the reader to find out how many students pay less than a certain value, and what amount of monthly rent a certain number of students pay.

For instance, if you want to know how many students pay less than $1,500 a month in rent, then you can go up from the $1,500 until you hit the line and then you go left to the cumulative frequency axis to see what cumulative frequency corresponds to $1,500. It appears that around 21 students pay less than $1,500. See Figure 2-17.

If you want to know the cost of rent that 15 students pay less than, then you start at 15 on the vertical axis and then go right to the line and down to the horizontal axis to the monthly rent of about $1,200. You can see that about 15 students pay less than about $1,200 a month. See Figure 2-18.

clipboard_ed42b868b914e331900f0d35b58f74af8.png

Figure 2-17

clipboard_e26a016d8dd476a05198002dbb5be3201.png

Figure 2-18

If you graph the cumulative relative frequency then you can find out what percentage is below a certain number instead of just the number of people below a certain value.

Using the sample of 35 ages, make an ogive.

46 47 49 25 46 22 42 24 46 40 39 27 25 30 33 27 46 21 29 20 26 25 25 26 35 49 33 26 32 31 39 30 39 29 26
Solution

Excel

Excel will plot an ogive over a histogram as one of its options, but the scale is harder to read.

Type the data in any order into column A and the bins in order in column B as shown below. Then select the Data tab, select Data Analysis, select Histogram, then select OK, see below.

clipboard_e365673eb4fef9d99958e354f5149b81c.png

In the dialogue box, click into the Input Range box, then use your mouse and highlight the ages including the label. Then click into the Bin Range box and use your mouse to highlight the bins including the label. Select the box for Labels only if you included the labels in your ranges. You can have your output default to a new worksheet, or select the circle to the left of Output Range, click into the box to the right of Output Range and then select one blank cell on your spreadsheet where you want the top left-hand corner of your table and graph to start. Then check the boxes next to Cumulative Percentage and Chart Output. Then select OK.

clipboard_ebdf12f2164e74bdc0639cd5d1bef2f6f.png

A histogram needs to have bars that touch, which is not the default in Excel. To get the bars to touch, right-click on one of the blue bars and select Format Data Series and slide the Gap Width to 0%.

clipboard_eb71b90c8ff3af670b35713611622ef23.png

Excel produces both a frequency table and a histogram. The table has the frequencies and the cumulative relative frequencies.

Bin Frequency Cumulative %
24 4 11.43%
29 12 45.71%
34 6 62.86%
39 4 74.29%
44 2 80.00%
49 7 100.00%
More 0 100.00%

The orange line is the ogive and the vertical axis is on the right side.

clipboard_efc9754e8e10f76332c66c8eb6f80b94e.png

2.3.4 Pie Chart

You cannot make stem-and-leaf plots, histograms, ogives or time series graphs for qualitative data. Instead, we use bar or pie charts for a qualitative variable, which lists the categories and gives either the frequency (count) or the relative frequency (percent) of individual items that fall into each category.

A pie chart or pie graph is a very common and easy-to-construct graph for qualitative data. A pie chart takes a circle and divides the circle into pie shaped wedges that are proportional to the size of the relative frequency. There are 360 degrees in a full circle. Relative frequency is just the percentage as a decimal. To find the angle for each pie wedge, multiply the relative frequency for each category by 360 degrees. Figure 2-19 is an example of a pie chart.

Soft Drink Purchases pie chart with 1 data series 2026-07-23T09:02:26.296080 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Figure 2-19 Pie chart of soft drink purchases by brand, shown as percentage of total purchases: Coke Classic 38%, Pepsi-Cola 26%, Diet Coke 16%, Dr. Pepper 10%, Sprite 10%. (regenerated with XYZChart for the XYZ Homework web edition.)

Figure 2-19

Use Excel to make a pie chart for the following frequency distribution of marital status.

Marital Status Frequency Divorced (D) 16 Married (M) 44 Single (S) 23 Widowed (W) 9
Solution

In Excel, type in the table as it appears, then use your mouse and highlight the entire table. Select the Insert tab, then select the pie graph icon, then select the first option under the 2-D Pie.

clipboard_e5b9977d8c2d7bb60f79ebb47039e6cf9.png

Once you have the pie chart you can select the Design window to get a graph to your liking.

clipboard_efd63bd4f8243d3201ac1de41bd77f2ff.png

It is good practice to include the class label and the percent. The percent should add up to 100%, although with rounding sometimes the sum can be off by 1%.

You can also click on the green plus sign to the right of the graph and add different formatting options, or the paintbrush to change colors.

clipboard_eb721cdc89367e04bca1458cdfaf8a6a0.png

Here is the finished pie graph.

Marital Status pie chart with 1 data series 2026-07-23T09:02:26.332093 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Pie chart of marital status for 92 individuals: Married 48% (44), Single 25% (23), Divorced 17% (16), Widowed 10% (9). Title "Marital Status". (regenerated with XYZChart for the XYZ Homework web edition.)

2.3.5 Bar Graph

A bar graph (column graph or bar chart) is another graph of a distribution for qualitative data. Bar graphs or charts consist of frequencies on one axis and categories on the other axis. Then you draw rectangles for each category with a height (if frequency is on the vertical axis) or length (if frequency is on the horizontal axis) that is equal to the frequency. All of the rectangles should be the same width, and there should be equally wide gaps between each bar. Figure 2-20 is an example of a bar chart.

Soft Drink Purchases bar chart with 1 data series 2026-07-23T09:02:26.364527 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Figure 2-20 Bar chart of soft drink purchase frequencies (Coke Classic 19, Diet Coke 8, Dr. Pepper 5, Pepsi-Cola 13, Sprite 5), illustrating a qualitative-data bar graph. (regenerated with XYZChart for the XYZ Homework web edition.)

Figure 2-20

Some key features of a bar graph:

You can draw a bar graph with frequency or relative frequency on the vertical axis. The relative frequency is useful when you want to compare two samples with different sample sizes. The relative frequency graph and the frequency graph should look the same, except for the scaling on the frequency axis.

Use Excel to make a bar chart for the following frequency distribution of marital status.

Marital Status Frequency Divorced (D) 16 Married (M) 44 Single (S) 23 Widowed (W) 9
Solution

In Excel, type in the table as it appears, then use your mouse and highlight the entire table.

Similar steps as the pie chart, but this time choose the column graph option we get the following bar graph for marital status.

Frequency bar chart with 1 data series 2026-07-23T09:02:26.397158 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Bar (column) graph of a marital-status frequency distribution: Divorced (D) 16, Married (M) 44, Single (S) 23, Widowed (W) 9. This is the initial (unformatted) Excel colu (regenerated with XYZChart for the XYZ Homework web edition.)

Then format the graph as needed.

clipboard_eb41c9c61cfe4f82b3115a9b62bafc739.png

The completed bar graph is below.

Marital Status bar chart with 1 data series 2026-07-23T09:02:26.426689 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Bar graph of the frequency distribution of marital status: Divorced (16), Married (44), Single (23), Widowed (9). This is the completed/formatted Excel bar chart from Exa (regenerated with XYZChart for the XYZ Homework web edition.)

Pie charts are useful for comparing sizes of categories. Bar charts show similar information. It really is a personal preference and what information you are trying to address. However, pie charts are best when you only have a few categories and the data can be expressed as a percentage.

The data does not have to be percentages to draw the pie chart, but if a data value can fit into multiple categories, you cannot use a pie chart to display the data. As an example, if you are asking people which is their favorite national park and you ask them to pick their top three choices, then the total number of answers can add up to more than 100% of the people surveyed. Therefore, you cannot use a pie chart to display the favorite national park, but a bar chart would be appropriate.

2.3.6 Pareto Chart

A Pareto (pronounced pə-RAY-toh) chart is a bar graph that starts from the most frequent class to the least frequent class. The advantage of Pareto charts is that you can visually see the more popular answer to the least popular. This is especially useful in business applications, where you want to know what services your customers like the most, what processes result in more injuries, which issues employees find more important, and other type of questions where you are interested in comparing frequency. Figure 2-21 is an example of a Pareto chart.

clipboard_eedbf9ccbcaede15d2dfc2d34d1c22815.png

Pareto

clipboard_eeeee9958874248d5ffcfc204d0207b76.png

Figure 2-21

Use Excel to make a Pareto chart for the following frequency distribution of marital status.

Marital Status Frequency Divorced (D) 16 Married (M) 44 Single (S) 23 Widowed (W) 9
Solution

In Excel, type in the table as it appears, then use your mouse and highlight the entire table. Highlight the table, then select the Home tab, then select Sort & Filter, then select Custom Sort.

clipboard_eb5191206e5f70eec72349f70d568f5a8.png

Change the Sort by to Frequency and the Order to Largest to Smallest and click OK.

This will automatically arrange the bars in your bar chart from largest to smallest.

Many Pareto charts will have the bars touching. You can right click on the bars, choose format data series, and then change the Gap Width to zero.

clipboard_e4f9a49a953cf9c80227c5bb5adcbb0a2.png

Here is the completed Pareto chart.

Marital Status bar chart with 1 data series 2026-07-23T09:02:26.455790 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Pareto chart of marital status frequency: bars ordered from most to least frequent (Married 44, Single 23, Divorced 16, Widowed 9), with the y-axis showing frequency. (regenerated with XYZChart for the XYZ Homework web edition.)

There are many other types of graphs used on qualitative data. There are software packages that will create most of them. It depends on your data as to which graph may be best to display the data.

2.3.7 Stacked Column Chart

The next example illustrates one of these types known as a stacked column chart. Stacked column (bar) charts are used when we need to show the ratio between a total and its parts. Each color shows the different series as a part of the same single bar, where the entire bar is used as a total.

In the Wii Fit game, you can do four different types of exercises: yoga, strength, aerobic, and balance. The Wii system keeps track of how many minutes you spend on each of the exercises every day. The following graph is the data for Niko over one-week time-period. Discuss any interpretations you can infer from the graph.

Wii Fit Credits bar chart with 4 data series 2026-07-23T09:02:26.498405 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Figure 2-22 Stacked column chart of Niko's daily Wii Fit exercise minutes (Yoga, Strength, Aerobic, Balance) over 15-21 Aug; daily totals range 28-44 minutes, peaking 20 Aug. (regenerated with XYZChart for the XYZ Homework web edition.)

Figure 2-22

Solution

It appears that Niko spends more time on yoga than on any other exercises on any given day. He seems to spend less time on aerobic exercises on a given day. There are several days when the amount of exercise in the different categories is almost equal. The usefulness of a stacked column chart is the ability to compare several different categories over another variable, in this case time. This allows a person to interpret the data with a little more ease.

Data scientists write programming using statistics to filter spam from incoming email messages. By noting specific characteristics of an email, a data scientist may be able to classify some emails as spam or not spam with high accuracy. One of those characteristics is whether the email contains no numbers, small numbers, or big numbers. Make a stacked column chart with the data in the table. Which type of email is more likely to be spam?

  Number     None Small Big Total Spam 149 168 50 367 Not Spam 400 2659 495 3554 Total 549 2827 545 3921

Example from OpenIntroStatistics.

Solution

Type the summarized table into Excel. Highlight just the inside of the table from the row label, column label and data (do not include the totals or Number label). Select the Insert tab, and then select the 2nd option under the column chart. Add a legend, labels and change colors for clarity.

clipboard_eae536463f0d21f1887dec88ebee16009.png

The completed stacked bar graph is shown in Figure 2-23.

Email Spam bar chart with 2 data series 2026-07-23T09:02:26.538492 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Figure 2-23 Figure 2-23: "Email Spam" — a stacked column chart showing the number of emails (y-axis, 0-3000) by whether the email contains numbers (x-axis: None, Small, Big), with ea (regenerated with XYZChart for the XYZ Homework web edition.)

Figure 2-23

Emails with no numbers have a relatively high rate of spam email (149/549 = 0.271) about 27%. On the other hand, less than 10% of email with small numbers (168/2827 = 0.059) or big numbers (50/545 = 0.092) are spam.

2.3.8 Multiple or Side-by-Side Bar Graph

A multiple bar graph, also called a side-by-side bar graph, allows comparisons of several different categories over another variable.

The percentages of people who use certain contraceptives in Central American countries are displayed in the graph below. Use the graph to find the type of contraceptive that is most used in Costa Rica and El Salvador.

clipboard_e25edadbce88ea3e00dabff572dd0da33.png

(9/21/2020) Retrieved from https://public.tableau.com/profile/prbdata#!/vizhome/AccesstoContraceptiveMethods/AccesstoContraceptiveMethods

Figure 2-24

Solution

This side-by-side bar graph allows you to quickly see the differences between the countries. For instance, the birth control pill is used most often in Costa Rica, while condoms are most used in El Salvador.

Make a side-by-side bar graph for the following medal count for the 2018 Olympics.

  GoldSilverBronzeNorway 14 14 11 Germany 14 10 7 Canada 11 8 10 United States 9 8 6
Solution

Copy the table over to Excel. Highlight the entire table, then use similar steps as the regular bar graph.

clipboard_eb50ac9fca438e623bf44b66d60591328.png

Add labels and change the color. The completed graph is shown below.

Medal Count of the 2018 Olympics bar chart with 3 data series 2026-07-23T09:02:26.578409 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Side-by-side (grouped) bar graph of the 2018 Winter Olympics medal count by country, with Gold, Silver, and Bronze medals shown as separate bars for Norway, Germany, Cana (regenerated with XYZChart for the XYZ Homework web edition.)

2.3.9 Time-Series Plot

A time-series plot is a graph showing the data measurements in chronological order, where the data is quantitative data. For example, a time-series plot is used to show profits over the last 5 years. To create a time-series plot, time always goes on the horizontal axis, and the frequency or relative frequency goes on the vertical axis. Then plot the ordered pairs and connect the dots. A time series allows you to see trends over time. Caution: You must realize that the trend may not continue. Just because you see an increase does not mean the increase will continue forever. As an example, prior to 2007, many people noticed that housing prices were increasing. The belief at the time was that housing prices would continue to increase. However, the housing bubble burst in 2007, and many houses lost value during the recession.

The New York Stock Exchange (NYSE) has a website where you can download information on the stock market. Use technology to make a time-series plot.

The daily trading volume for two weeks was downloaded at http://www.nyxdata.com/Data-Products/NYSE-Volume-Summary#summaries.

Trade Date NYSE Trades NYSE Volume NYSE Dollar Volume
2-Nov-17 3,126,422 900,365,416 $35,252,842,833
1-Nov-17 2,951,960 862,960,656 $32,609,257,159
31-Oct-17 2,712,611 944,584,968 $37,697,844,745
30-Oct-17 2,749,134 855,767,500 $33,455,644,270
27-Oct-17 2,816,612 882,117,579 $34,636,857,517
26-Oct-17 2,771,894 866,549,040 $34,105,511,597
25-Oct-17 2,823,272 895,997,132 $36,504,371,518
24-Oct-17 2,369,409 763,566,128 $31,139,796,280
23-Oct-17 2,165,477 745,070,262 $30,216,979,896
20-Oct-17 2,198,245 861,995,044 $37,745,087,429
19-Oct-17 2,211,641 692,171,878 $28,682,027,691
18-Oct-17 2,108,579 669,011,182 $28,642,623,992
17-Oct-17 2,045,857 680,893,022 $28,072,416,838
16-Oct-17 2,078,792 685,406,032 $27,524,409,199
13-Oct-17 2,151,643 757,155,836 $30,624,653,749
Solution

Using Excel, we will make a time series plot for NYSE daily trading volume. Using the Ctrl key highlight just the date column and the NYSE Volume, then select the Insert tab and the first 2-D line graph option.

clipboard_e333832b274e8ecd43ce27813e8daae62.png

You can then select different designs.

clipboard_e0210f1205130e85b4977554a51f2f314.png

One can use time-series plots to see when they want to cash out or buy a stock.

The time-series graph shows the behavior of one variable over time and does not reflect other variables that are influencing the trading volume.

2.3.10 Scatter Plot

Sometimes you have two quantitative variables and you want to see if they are related in any way. A scatter plot helps you to see what the relationship may look like. A scatter plot is just a plotting of the ordered pairs.

Is there any relationship between elevation and high temperature on a given day? The following data are the high temperatures at various cities on a single day and the elevation of the city.

Make a scatterplot to see what type of relationship exists.

Elevation (in feet) 7000 4000 6000 3000 7000 4500 5000 Temperature (°F) 50 60 48 70 55 55 60
Solution

Excel

Type the data into two columns next to each other. It is important not to have a blank column between the points or Excel may give you an error message. Once you type your data into columns A and B, use your mouse and highlight all the data including the labels. Select the Insert tab, and then select the first box under Scatter.

clipboard_e9cdb9edb417817056ec153756d425f93.png

Add appropriate labels. The completed scatter plot is shown below.

Scatter Plot scatter chart with 2 data series 2026-07-23T09:02:26.621866 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Scatter plot of daily high temperature (°F) versus city elevation (in feet) for seven cities, showing a weak negative relationship between elevation and temperature. (regenerated with XYZChart for the XYZ Homework web edition.)

TI-84: First, enter the data into lists 1 and 2. Press [STAT] the first option is already highlighted (1:Edit) so you can either press [ENTER] or 1.Type in the data pressing [ENTER] after each one. For x-y data pairs, enter all x -values in one list. Enter all corresponding y -values in a second list. Press [2nd] [QUIT] to return to the home screen. Make sure you turn off other stat plots or graphs in the y= menu. Press [2nd] then the [y=] button. Select the first plot. Highlight On and press [Enter] so that On is highlighted. Arrow down to Type and highlight the first option that looks like a scatter plot. Make sure your x and y lists are using L1 and L2.

clipboard_e73da6a22bfcbd1d290b1fa22b76a2e8d.png

Select Zoom and arrow down to ZoomStat and press [Enter]. You will get the following scatterplot. Select Trace and use your arrow keys to see the values at different points.

TI-89: Press [♦] then [F1] (to get Y=) and clear any equations that are in the y-editor. Open the Stats/List Editor. Press [APPS], select FlashApps then press [ENTER]. Highlight Stats/List Editor then press [ENTER]. Press [ENTER] again to select the main folder. Type in the data pressing [ENTER] after each one. Enter all x-values in one list. Enter all corresponding y-values in a second list. In the Stats/List Editor, select [F2] for the Plots menu. Use cursor keys to highlight 1:Plot Setup. Make sure that the other graphs are turned off by pressing [F4] button to remove the check marks. Under “Plot 1” press [F1] Define.

clipboard_ea5365abeee21b9d88eae53b76675a635.png

In the “Plot Type” menu, select “Scatter.” Move the cursor to the “x” space press [2nd] Var-Link, scroll down to list1, and then press [Enter]. This will put the list name in the dialogue box. Do the same for the y values, but this time choose list2. Press [ENTER] twice and you will be returned to the Plot Setup menu.

clipboard_e3de4e9f170d56988f7a613dcd8a078ad.png

Press F5 ZoomData to display the graph.

Press F3 Trace and use the arrow keys to scroll along the different points.

Interpreting the scatter plot.

The graph indicates a linear relationship between temperature and elevation. If you were to hold a pencil up to cover the dots, note that you would see that the dots roughly follow a fat line downhill.

It also appears to be a negative relationship, thus as elevation increases, the temperature decreases.

clipboard_e32720f9be4398365a502f0815fa7e188.png

Figure 2-25

Be careful with the vertical axis of both time-series and scatter plots. If the axis does not start at zero the slope of the line can be exaggerated to show more or less of increase than there really is. This is done in politics and advertising to manipulate the data.

For example, if we change the vertical axis of temperature to go between 45°F and 75°F we get the following scatter plot in Figure 2-26.

We have the same arrangements of dots, but the slope looks much steeper over the 30° range.

Temperature vs. Elevation scatter chart with 2 data series 2026-07-23T09:02:26.662883 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Figure 2-26 Figure 2-26: Scatter plot of high temperature (°F) versus city elevation (feet) for 7 cities on a single day, showing a negative relationship. This is the same data as Fi (regenerated with XYZChart for the XYZ Homework web edition.)

Figure 2-26

2.3.11 Misleading Graphs

One thing to be aware of as a consumer, data in the media may be represented in misleading graphs. Misleading graphs not only misrepresent the data, they can lead the reader to false conclusions. There are many ways that graphs can be misleading. One way to mislead is to use picture graphs or 3D graphs that exaggerate differences and should be used with caution. Leaving off units and labels can result in a misleading graph. Another more common example is to rescale or reverse the vertical axis to try to show a large difference between categories. Not starting the vertical axes at zero will show a more dramatic rate of change. Other ways that graphs can be misleading is to change the horizontal axis labels so that they are out of time sequence, using inappropriate graphs, not showing the base population.

What is misleading about the following graph?

An ad for a new diet pill shows the following time-series plot for someone that has lost weight over a 5-month period.

Weight over Time line chart with 1 data series 2026-07-23T09:02:26.702760 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Weight over Time — a time-series (line) plot of one person's weight over a 5-month period, dropping from 200 lb to 187 lb. Reproduced from the diet-pill-ad example whose (regenerated with XYZChart for the XYZ Homework web edition.)

Solution

If you do not start the vertical axis at zero, then a change can look much more dramatic than it really is. Notice the decrease in weight looks much larger in Figure 2-27. The graph in Figure 2-28 has the vertical axis starting at zero. Notice that over the 5 months, the weight appears to be decreasing, however, it does not look like there is a large decrease.

clipboard_e525678d554f586102f7e0df130962564.png

Figure 2-27

clipboard_e9f4995ff828371d093c1bd68311e798c.png
Figure 2-28

What is misleading about the graph in Figure 2-29?

clipboard_e7db4511fe0fb38c1bac7fa60e3a12439.png

https://www.mediamatters.org/blog/2014/03/31/dishonest-fox-charts-obamacare-enrollment-editi/198679.

Figure 2-29

Solution

The y-axis scale is different for each bar and there are no units on the axis. The first bar has each tic mark as 2 billion, the second bar has each tick as less then 1 billion.

This exaggerates the difference. If they used square scaling as in Figure 2-30, there would not be such an extreme difference between the height of the bars.

clipboard_ee3a9ca04a704504d8516825a8a52fb6d.png

https://www.mediamatters.org/blog/2014/03/31/dishonest-fox-charts-obamacare-enrollment-editi/198679.

Figure 2-30

What is misleading about the graph in Figure 2-31?

clipboard_e2be6760c97ced071dec328f9ee732642.png

https://www.livescience.com/45083-misleading-gun-death-chart.html

Figure 2-31

Solution

The graph has the y-axis reversed. What looks like an increasing trend line really is decreasing when you correct the y-axis. The red background is also an effect to raise alarm, almost like a curtain of blood.

What is misleading about the graph shown in a Lanacane commercial in May 2012, shown in Figure 2-32?

clipboard_e9dd527ec400d3930966ca394d728fc92.png

Retrieved 7/2/2021 from https://youtu.be/I0DapkQ-c1I?t=17

Figure 2-32

Solution

It appears that Lanacane is better than regular hydrocorisone cream at releiving itching. However, note that there are no units or labels to the axis.

What is misleading about the graph published Georgia’s Department of Public Health website in May 2020, shown in Figure 2-33?

 clipboard_eca214b90df14140425e050daa46c5f84.png

Retrieved 7/3/2021 from https://www.vox.com/covid-19-coronav...ning-reopening Figure 2-33

Solution

There are two misleading items for this graph. The horizontal axis is time, yet the dates are out of sequence starting with April 28, April 27, April 29, May 1, April 30, May 4, May 6, May 5, May 2, May 7, April 26, May 3, May 8, May 9. The first date of April 26 is presented almost at the end of the axis. The graph at first glance would deceive viewers in cases going down over time. A Pareto style chart should never be used for time series data.

The second misleading item is the graph’s title and no label on the y-axis. What does the height of each bar represent? Is the height the number of cases for each county, or is the height the number of deaths and hospitalizations? The website later corrected the graphic as shown in Figure 2-34.

clipboard_ea1b30686fb788964f3cb220e40859afc.png

Retrieved 7/3/2021 from https://www.vox.com/covid-19-coronav...ning-reopening Figure 2-34

Large data sets need to be summarized in order to make sense of all the information. The distribution of data can be represented with a table or a graph. It is the role of the researcher or data scientist to make accurate graphical representations that can help make sense of this in the context of the data. Tables and graphs can summarize data, but they alone are insufficient. In the next chapter we will look at describing data numerically.

2.04: Chapter 2 Exercises

Chapter 2 Exercises

1. Which types of graphs are used for quantitative data? Select all that apply.

a) Ogive

b) Pie Chart

c) Histogram

d) Stem-and-Leaf Plot

e) Bar Graph

2. Which types of graphs are used for qualitative data? Select all that apply.

a) Pareto Chart

b) Pie Chart

c) Dotplot

d) Stem-and-Leaf Plot

e) Bar Graph

f) Time Series Plot

3. The bars for a histogram should always touch, true or false?

4. A sample of rents found the smallest rent to be $600 and the largest rent was $2,500. What is the recommended class width for a frequency table with 7 classes?

5. An instructor had the following grades recorded for an exam.

96 66 65 82 85
82 87 76 80 85
83 69 79 70 83
63 81 94 71 83
99 75 73 83 86

a) Create a stem-and-leaf plot.

b) Complete the following table.

Class Frequency Cumulative Frequency Relative Frequency Cumulative Relative Frequency
60 – 69        
70 – 79        
80 – 89        
90 – 99        
Total 25      

c) What should the relative frequencies always add up to?

d) What should the last value always be in the cumulative frequency column?

e) What is the frequency for students that were in the C range of 70-79?

f) What is the relative frequency for students that were in the C range of 70-79?

g) Which is the modal class?

h) Which class has a relative frequency of 12%?

i) What is the cumulative frequency for students that were in the B range of 80-89?

j) Which class has a cumulative relative frequency of 40%?

6. Eyeglassomatic manufactures eyeglasses for different retailers. The number of lenses for different activities is in table.

Activity Grind Multi-coat Assemble Make Frames Receive Finished Unknown
Number of lenses 18,872 12,105 4,333 25,880 26,991 1,508

Grind means that they ground the lenses and put them in frames, multi-coat means that they put tinting or scratch resistance coatings on lenses and then put them in frames, assemble means that they receive frames and lenses from other sources and put them together, make frames means that they make the frames and put lenses in from other sources, receive finished means that they received glasses from another source, and unknown means they do not know where the lenses came from.

a) Make a relative frequency table for the data.

b) How many of the eyeglasses did Eyeglassomatic assemble?

c) How many of the eyeglasses did Eyeglassomatic manufacture all together?

d) What is the relative frequency for the assemble category?

e) What percent of eyeglasses did Eyeglassomatic grind?

7. The following table is from a sample of five hundred homes in Oregon asked the primary source of heating in their residential homes.

Type of Heat Percent
Electricity 33
Heating Oil 4
Natural Gas 50
Firewood 8
Other 5

a) How many of the households heat their home with firewood?

b) What percent of households heat their home with natural gas?

8. The following table is from a sample of 50 undergraduate PSU students.

Class Relative Frequency Percent
Freshman 18
Sophomore 13
Junior 23
Senior 46

a) What percent of students are below a senior class?

b) What is the cumulative frequency of the junior class?

9. A sample of heights of 20 people in cm is recorded below. Make a stem-and-leaf plot.

Height (cm)
167 201 170 185 175 162
182 186 172 173 188 154
185 178 177 184 178 165
169 171 185 178 175 176

10. The stem-and-leaf plot below is for pulse rates before and after exercise.

clipboard_e71351022aade20b90c63ae50627f950d.png

a) Was pulse rate higher on average before or after exercise?

b) What was the fastest pulse rate of the before exercise group?

c) What was the slowest pulse rate of the after-exercise group?

11. The following data represents the percent change in tuition levels at public, four-year colleges (inflation adjusted) from 2008 to 2013 (Weissmann, 2013). Below is the frequency distribution and histogram.

Class Limits Class Midpoint Frequency Relative Frequency
2.2 – 11.7 6.95 6 0.12
11.8 – 21.3 16.55 20 0.40
21.4 – 30.9 26.15 11 0.22
31.0 – 40.5 35.75 4 0.08
40.6 – 50.1 45.35 2 0.04
50.2 – 59.7 54.95 2 0.04
59.8 – 69.3 64.55 3 0.06
69.4 – 78.9 74.15 2 0.04

clipboard_edeb78e418560a68ef0a7897023ab5d61.png

a) How many colleges were sampled?

b) What was the approximate value of the highest change in tuition?

c) What was the approximate value of the most frequent change in tuition?

12. The following data and graph represent the grades in a statistics course.

Class Limits Class Midpoint Frequency Relative Frequency
40 – 49.9 45 2 0.08
50 – 59.9 55 1 0.04
60 – 69.9 65 7 0.28
70 – 79.9 75 6 0.24
80 – 89.9 85 7 0.28
90 – 99.9 95 2 0.04

clipboard_efd5dc4827ce5bc36cf9c7fd44fae009b.png

a) How many students were in the class?

b) What was the approximate lowest and highest grade in the class?

c) What percent of students had a passing grade of 70% or higher?

13. The following graph represents a random sample of car models driven by college students. What percent of college students drove a Nissan?

clipboard_ebf0cbb4357d6e9b242e7467c42005547.png

14. The following graph and data represent the percent change in tuition levels at public, four-year colleges (inflation adjusted) from 2008 to 2013 (Weissmann, 2013).

clipboard_edd39b5722899f737fae7e762133aeec5.png

Class Limits Cumulative Frequency
2.2 – 11.7 6
11.8 – 21.3 26
21.4 – 30.9 37
31.0 – 40.5 41
40.6 – 50.1 43
50.2 – 59.7 45
59.8 – 69.3 48
69.4 – 78.9 50

a) How many colleges were sampled?

b) What class of percent changes had the most colleges in that range?

c) How many colleges had a percent change below 50.2% change in tuition?

d) What is the cumulative relative frequency for the 50.2% – 59.7% change in tuition class?

15. Eyeglassomatic manufactures eyeglasses for different retailers. The number of lenses for different activities is in table.

Activity Grind Multi-coat Assemble Make Frames Receive Finished Unknown
Number of lenses 18,872 12,105 4,333 25,880 26,991 1,508

a) Make a pie chart.

b) Make a bar chart.

c) Make a Pareto chart.

16. The daily sales using different sales strategies is shown in the graph below.

clipboard_e257de91a2e4d278c2c574de81c8b0b54.png

a) Which strategy generated the most sales?

b) Was there a particular strategy that worked well for one product, but not for another product?

17. The following graph represents a random sample of car models driven by college students. What was the most common car model?

clipboard_e87d1bef01e562d50776e4338c1210898.png

18. The Australian Institute of Criminology gathered data on the number of deaths (per 100,000 people) due to firearms during the period 1983 to 1997. The data is in table below. Create a time-series plot of the data. What is the overall trend over time? (2013, September 26). Retrieved from http://www.statsci.org/data/oz/firearms.html.

Year Rate
1983 4.31
1984 4.42
1985 4.52
1986 4.35
1987 4.39
1988 4.21
1989 3.4
1990 3.61
1991 3.67
1992 3.61
1993 2.98
1994 2.95
1995 2.72
1996 2.95
1997 2.3

19. A scatter plot for a random sample of 24 countries shows the average life expectancy and the average number of births per woman (fertility rate). What is the approximate fertility rate for a country that has a life expectancy of 76 years?

clipboard_e306c40810d9dfa06351c87d30736580e.png

(2013, October 14). Retrieved from http://data.worldbank.org/indicator/SP.DYN.TFRT.IN.

20. The Australian Institute of Criminology gathered data on the number of deaths (per 100,000 people) due to firearms during the period 1983 to 1997. The time-series plot is below. What year had the highest rate of deaths?

clipboard_e2885b0e9f6535162df7c9e22add523df.png

(2013, September 26). Retrieved from http://www.statsci.org/data/oz/firearms.html.

21. A survey by the Pew Research Center, conducted in 16 countries among 20,132 respondents from April 4 to May 29, 2016, before the United Kingdom’s Brexit referendum to exit the EU. The following is a time series graph for the proportion of survey respondents by country that responded that the current economic situation is their country was good.

clipboard_ed7cf7ea53d3b2306e20c6b385681d258.png

http://www.pewglobal.org/2016/08/09/views-on-national-economies-mixed-as-many-countries-continue-to-struggle/

a) Which country had the most favorable outlook of their country’s economic situation in 2010?

b) Which country had the least favorable outlook of their country’s economic situation in 2016?

22. Why is this a misleading or poor graph?

clipboard_e7c308d1b3cd24c4769dc6f824b60cb71.png

23. Why is this a misleading or poor graph?

clipboard_e8de2549724a7c4082401f54c9ddc690f.png

24. Why is this a misleading or poor graph?

clipboard_efbc1b7d4410fb8224a4048a894af156c.png

United States unemployment. (2013, October 14). Retrieved from http://www.tradingeconomics.com/united-states/unemployment-rate

25. The Australian Institute of Criminology gathered data on the number of deaths (per 100,000 people) due to firearms during the period 1983 to 1997. Why is this a misleading or poor graph?

clipboard_ed3e74309729ba25e296bd7a372531450.png

(2013, September 26). Retrieved from http://www.statsci.org/data/oz/firearms.html.

26. Why is this a misleading or poor graph?

clipboard_e7b43151ef1362084d5791f9da49eae86.png

Answer to Odd Numbered Exercises

1) a, c,

3) True

5) a) \(\begin{array}{l|llllllllllll}
6 & 3 & 5 & 6 & 9 \\
7 & 0 & 1 & 3 & 5 & 6 & 9 \\
8 & 0 & 1 & 2 & 2 & 3 & 3 & 3 & 3 & 5 & 5 & 6 & 7 \\
9 & 4 & 6 & 9 \\
\end{array}\)

b) clipboard_ef94521a944ef882ee0843c3736da321f.png c) 1 d) The sample size n. e) 6 f) 0.24 g) 80-89 h) 90-99 i) 0.88 j) 70-79

7) a) 40 b) 50%

9) \(\begin{array}{l|lllllllllll}
15 & 4 & & & & & & & & & \\
16 & 2 & 5 & 7 & 9 & & & & & & & \\
17 & 0 & 1 & 2 & 3 & 5 & 5 & 6 & 7 & 8 & 8 & 8 \\
18 & 2 & 4 & 5 & 5 & 5 & 6 & 8 & & & \\
19 & & & & & & & & & & & \\
20 & 1 & & & & & & & & &
\end{array}\)

11) a) 50 b) 78 c) 16.55

13) 20%

15) a) clipboard_e3d2954c2ae826281a19af8cef2f98bed.png

b) clipboard_e6740c8926801682ed897cad4000b6d90.png

c) clipboard_e40d64bc1ad7b6f325550b1d870e64968.png

17) Chevy & Toyota

19) 1.5

21) a) Poland

b) Greece

23) There are no labels for both axis or categories.

25) The vertical axis is reversed, making the graph appear to increase when it is actually decreasing. There are no labels for both axes.

From Mostly Harmless Statistics by Rachel L. Webb, adapted from LibreTexts. Licensed CC BY-SA 4.0. XYZ Homework OER web edition, adapted with 7 verified corrections (see errata).