#set document(title: "3.10 Comparing Measures", author: "OpenStax") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 3.10#h(0.6em)Comparing Measures #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Prerequisites] Percentiles, Distributions, What is Central Tendency, Measures of Central Tendency, Mean and Median ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Learning Objectives] + Understand how the difference between the mean and median is affected by skew + State how the measures differ in symmetric distributions + State which measure(s) should be used to describe the center of a skewed distribution ] How do the various measures of central tendency compare with each other? For symmetric distributions, the mean, median, trimean, and trimmed mean are equal, as is the mode except in bimodal distributions. Differences among the measures occur with skewed distributions. Figure 1 shows the distribution of 642 scores on an introductory psychology test. Notice this distribution has a slight positive skew. #figure(figph[Histogram of 642 psychology test scores grouped in intervals of 10 with class boundaries labeled 39.5 through 169.5. Frequencies rise from near 0 at 39.5-49.5 to a peak of about 147 in the 79.5-89.5 interval, then fall through about 130, 78, 59, 36 down to single digits above 139.5. The distribution is unimodal and skewed slightly to the right.], alt: "Histogram of 642 psychology test scores grouped in intervals of 10 with class boundaries labeled 39.5 through 169.5. Frequencies rise from near 0 at 39.5-49.5 to a peak of about 147 in the 79.5-89.5 interval, then fall through about 130, 78, 59, 36 down to single digits above 139.5. The distribution is unimodal and skewed slightly to the right.", caption: [Figure 1. A distribution with a positive skew.]) Measures of central tendency are shown in Table 1. Notice they do not differ greatly, with the exception that the mode is considerably lower than the other measures. When distributions have a #strong[positive skew], the mean is typically higher than the median, although it may not be in bimodal distributions. For these data, the mean of 91.58 is higher than the median of 90. Typically the trimean and trimmed mean will fall between the median and the mean, although in this case, the trimmed mean is slightly lower than the median. The geometric mean is lower than all measures except the mode. Table 1. Measures of central tendency for the test scores. #figure(table( columns: 2, align: left, inset: 6pt, table.header([Measure], [Value]), [Mode #linebreak() Median #linebreak() Geometric Mean #linebreak() Trimean #linebreak() Mean trimmed 50% #linebreak() Mean #linebreak()], [84.00 #linebreak() 90.00 #linebreak() 89.70 #linebreak() 90.25 #linebreak() 89.81 #linebreak() 91.58], )) The distribution of baseball salaries (in 1994) shown in Figure 2 has a much more pronounced skew than the distribution in Figure 1. #linebreak() #figure(figph[Histogram of major league baseball salaries, x-axis Salary/1000 from 0 to 6500 in steps of 500 and y-axis Count: about 367 players in the lowest bin, about 120 in the next, then about 45, 30, 36, 36, a secondary rise to about 51 near 3000-3500, and a thin tail of about 28, 18, 13, 9 down to nearly 0 above 5500 — a very large positive skew.], alt: "Histogram of major league baseball salaries, x-axis Salary/1000 from 0 to 6500 in steps of 500 and y-axis Count: about 367 players in the lowest bin, about 120 in the next, then about 45, 30, 36, 36, a secondary rise to about 51 near 3000-3500, and a thin tail of about 28, 18, 13, 9 down to nearly 0 above 5500 — a very large positive skew.", caption: [Figure 2. A distribution with a very large positive skew. This histogram shows the salaries of major league baseball players (in thousands of dollars: 250 equals 250,000).]) Table 2 shows the measures of central tendency for these data. The large skew results in very different values for these measures. No single measure of central tendency is sufficient for data such as these. If you were asked the very general question: "So, what do baseball players make?" and answered with the mean of \$1,183,000, you would not have told the whole story since only about one third of baseball players make that much. If you answered with the mode of \$250,000 or the median of \$500,000, you would not be giving any indication that some players make many millions of dollars. Fortunately, there is no need to summarize a distribution with a single number. When the various measures differ, our opinion is that you should report the mean, median, and either the trimean or the mean trimmed 50%. Sometimes it is worth reporting the mode as well. In the media, the median is usually reported to summarize the center of skewed distributions. You will hear about median salaries and median prices of houses sold, etc. This is better than reporting only the mean, but it would be informative to hear more statistics. Table 2. Measures of central tendency for baseball salaries (in thousands of dollars). #figure(table( columns: 2, align: left, inset: 6pt, table.header([Measure], [Value]), [Mode #linebreak() Median #linebreak() Geometric Mean #linebreak() Trimean #linebreak() Mean trimmed 50% #linebreak() Mean #linebreak()], [250 #linebreak() 500 #linebreak() 555 #linebreak() 792 #linebreak() 619 #linebreak() 1,183], ))