#set document(title: "12.9 Outliers", author: "Rachel Webb") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 12.9#h(0.6em)Outliers A scatter plot should be checked for outliers. An outlier is a point that seems out of place when compared with the other points. Some of these points can affect the equation of the regression line. #examplebox("Example 1")[][ Should linear regression be used with this data set? #math.equation(block: false, alt: "x")[$x$] 1 3 8 2 1 3 2 2 3 1 #math.equation(block: false, alt: "y")[$y$] 2 3 8 2 3 1 3 1 2 1 #solutionbox[ A regression analysis for the data set was run on Excel. #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Watch one point manufacture a correlation] Preloads the ten points including the leverage point (8, 8): the panel shows r = 0.844 and R^2 = 0.7119, a 'significant' fit. Now delete the 8 from L1 and the 8 from L2, recalculate, and watch r collapse to exactly 0 - the entire relationship was one point. - r = 0.844 with (8, 8) - delete it and r = 0 ] #figure(figph[Excel-generated table of regression statistics for the given data.], alt: "Excel-generated table of regression statistics for the given data.", caption: none) If we test for a significant correlation: #math.equation(block: true, alt: "H sub 0 : ρ equals 0")[$H_(0) : ρ = 0$] #linebreak() #math.equation(block: true, alt: "H sub 1 : ρ not equal to 0")[$H_(1) : ρ ≠ 0$] The correlation is #math.equation(block: false, alt: "r equals 0.844")[$r = 0.844$] and the p-value is 0.002, which is less than #math.equation(block: false, alt: "α")[$α$] = 0.05, so we would reject #math.equation(block: false, alt: "H sub 0")[$H_(0)$] and conclude there is a significant relationship between #math.equation(block: false, alt: "x")[$x$] and #math.equation(block: false, alt: "y")[$y$]. However, if look at the scatterplot in Figure 12-18, with the regression equation we can clearly see that the point #math.equation(block: false, alt: "open parenthesis 8 , 8 close parenthesis")[$( 8 , 8 )$] is an outlier. The outlier is pulling the slope up towards the point #math.equation(block: false, alt: "open parenthesis 8 , 8 close parenthesis")[$( 8 , 8 )$]. #figure(figph[Scatterplot of the given data with a linear regression equation of y = 0.8438x + 0.4063, an R-squared value of 0.7119, and an outlier at point (8, 8).], alt: "Scatterplot of the given data with a linear regression equation of y = 0.8438x + 0.4063, an R-squared value of 0.7119, and an outlier at point (8, 8).", caption: [Figure 12-18: Scatterplot of data with an outlier.]) If we were to take out the outlier point #math.equation(block: false, alt: "open parenthesis 8 , 8 close parenthesis")[$( 8 , 8 )$] and run the regression analysis again on the modified data set we get the following Excel output. #figure(figph[Excel-generated regression statistics table of the given data with the (8, 8) outlier removed.], alt: "Excel-generated regression statistics table of the given data with the (8, 8) outlier removed.", caption: none) See Figure 12-19: note the correlation is now 0 and the p-value is 1, so there is no relationship at all between #math.equation(block: false, alt: "x")[$x$] and #math.equation(block: false, alt: "y")[$y$]. #figure(figph[Scatterplot of the given data with the outlier point (8, 8) removed. Plot now takes the form of a 3-by-3 grid of points bounded by x and y values of 1 and 3, with the linear regression line equation now being y=2 and the R-squared value being 0.], alt: "Scatterplot of the given data with the outlier point (8, 8) removed. Plot now takes the form of a 3-by-3 grid of points bounded by x and y values of 1 and 3, with the linear regression line equation now being y=2 and the R-squared value being 0.", caption: [Figure 12-19: Scatterplot of the same data as above with outlier removed.]) This type of outlier is called a #strong[leverage point]. Leverage points are positioned far away from the main cluster of data points on the #math.equation(block: false, alt: "x")[$x$]-axis. ] ] There is another type of outlier called an #strong[influential point]. Influential points are positioned far away from the main cluster of data points on the #math.equation(block: false, alt: "y")[$y$]-axis. There is an option in most software packages to get the “standardized” residuals. Standardized residuals are z-scores of the residuals. Any standardized residual that is not between #math.equation(block: false, alt: "minus 2")[$− 2$] and #math.equation(block: false, alt: "2")[$2$] may be an outlier. If it is not between #math.equation(block: false, alt: "minus 3")[$− 3$] and #math.equation(block: false, alt: "3")[$3$] then the point is an outlier. When this happens, the points are called influential points or influential observations. #examplebox("Example 2")[][ Use technology to compute the standardized residuals. Should linear regression be used with this data set? #math.equation(block: false, alt: "x")[$x$] 1 3 2 2 4 5 7 9 6 8 #math.equation(block: false, alt: "y")[$y$] 1 3 10 2 4 5 7 9 6 8 #solutionbox[ A regression analysis for the given data set was run on Excel, producing the following results: #figure(figph[Regression analysis table for the given data. Observation 3, the data point (2, 10), has a standard residual of 2.671, which is highlighted. All other observations have standard residual values between -1 and 1.], alt: "Regression analysis table for the given data. Observation 3, the data point (2, 10), has a standard residual of 2.671, which is highlighted. All other observations have standard residual values between -1 and 1.", caption: none) The point #math.equation(block: false, alt: "open parenthesis 2 , 10 close parenthesis")[$( 2 , 10 )$] shown in Figure 12-20 is pulling the left side of the line up and away from the points that form a line. This influential point changes the #math.equation(block: false, alt: "y")[$y$]-intercept and slope. #figure(figph[Scatterplot of the given data points. All points except for (2, 10) appear to line up; the regression line does not pass evenly through these points, but is tilted towards the (2, 10) point.], alt: "Scatterplot of the given data points. All points except for (2, 10) appear to line up; the regression line does not pass evenly through these points, but is tilted towards the (2, 10) point.", caption: [Figure 12-20:]) ] ]