#set document(title: "14.4 Standard Error of the Estimate", author: "OpenStax") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 14.4#h(0.6em)Standard Error of the Estimate #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Prerequisites] Measures of Variability, Introduction to Simple Linear Regression, Partitioning Sums of Squares #linebreak() #linebreak() ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Learning Objectives] + Make judgments about the size of the standard error of the estimate from a scatter plot + Compute the standard error of the estimate based on errors of prediction + Compute the standard error using Pearson's correlation + Estimate the standard error of the estimate based on a sample ] Figure 1 shows two regression examples. You can see that in Graph A, the points are closer to the line than they are in Graph B. Therefore, the predictions in Graph A are more accurate than in Graph B. #figure(figph[Two scatter plots labeled A and B, each with the same rising regression line through a cloud of circles. In A the points lie in a tight band about the line; in B the same number of points scatter much more widely around it. The predictions in A are more accurate — A has the smaller standard error of the estimate.], alt: "Two scatter plots labeled A and B, each with the same rising regression line through a cloud of circles. In A the points lie in a tight band about the line; in B the same number of points scatter much more widely around it. The predictions in A are more accurate — A has the smaller standard error of the estimate.", caption: [Figure 1. Regressions differing in accuracy of prediction.]) The standard error of the estimate is a measure of the accuracy of predictions. Recall that the regression line is the line that minimizes the sum of squared deviations of prediction (also called the #strong[sum of squares error]). The standard error of the estimate is closely related to this quantity and is defined below: #math.equation(block: true, alt: "σ sub est equals the square root of the fraction ∑ open parenthesis Y minus Y prime close parenthesis squared over N")[$σ_("est") = sqrt(frac(∑ ( Y − Y^(′) )^(2), N))$] where σ#sub[est] is the standard error of the estimate, Y is an actual score, Y' is a predicted score, and N is the number of pairs of scores. The numerator is the sum of squared differences between the actual scores and the predicted scores. Note the similarity of the formula for σ#sub[est] to the formula for σ.  It turns out that σest is the standard deviation of the errors of prediction (each Y - Y' is an error of prediction). Assume the data in Table 1 are the data from a population of five X, Y pairs. Table 1. Example data. #figure(table( columns: 6, align: left, inset: 6pt, table.header([], [X], [Y], [Y'], [Y-Y'], [(Y-Y')#super[2]]), [], [1.00], [1.00], [1.210], [-0.210], [0.044], [], [2.00], [2.00], [1.635], [0.365], [0.133], [], [3.00], [1.30], [2.060], [-0.760], [0.578], [], [4.00], [3.75], [2.485], [1.265], [1.600], [], [5.00], [2.25], [2.910], [-0.660], [0.436], [Sum], [15.00], [10.30], [10.30], [0.000], [2.791], )) The last column shows that the sum of the squared errors of prediction is 2.791. Therefore, the standard error of the estimate is #math.equation(block: true, alt: "σ sub est equals the square root of the fraction 2.791 over 5 equals 0.747")[$σ_("est") = sqrt(frac(2.791, 5)) = 0.747$] There is a version of the formula for the standard error in terms of Pearson's correlation: #math.equation(block: true, alt: "σ sub est equals the square root of the fraction open parenthesis 1 minus ρ squared close parenthesis S S Y over N")[$σ_("est") = sqrt(frac(( 1 − ρ^(2) ) S S Y, N))$] where ρ is the population value of Pearson's correlation and SSY is #math.equation(block: true, alt: "S S Y equals ∑ open parenthesis Y minus μ sub Y close parenthesis squared")[$S S Y = ∑ ( Y − μ_(Y) )^(2)$] For the data in Table 1, μ#sub[y] = 2.06, SSY = 4.597 and ρ= 0.6268. Therefore, #math.equation(block: true, alt: "σ sub est equals the square root of the fraction open parenthesis 1 minus 0.6268 squared close parenthesis open parenthesis 4.597 close parenthesis over 5 equals the square root of the fraction 2.791 over 5 equals 0.747")[$σ_("est") = sqrt(frac(( 1 − 0.6268^(2) ) ( 4.597 ), 5)) = sqrt(frac(2.791, 5)) = 0.747$] which is the same value computed previously. Similar formulas are used when the standard error of the estimate is computed from a sample rather than a population. The only difference is that the denominator is N-2 rather than N. The reason N-2 is used rather than N-1 is that two parameters (the slope and the intercept) were estimated in order to estimate the sum of squares. Formulas for a sample comparable to the ones for a population are shown below. #math.equation(block: true, alt: "s sub est equals the square root of the fraction ∑ open parenthesis Y minus Y prime close parenthesis squared over N minus 2")[$s_("est") = sqrt(frac(∑ ( Y − Y^(′) )^(2), N − 2))$] #math.equation(block: true, alt: "s sub est equals the square root of the fraction 2.791 over 3 equals 0.964")[$s_("est") = sqrt(frac(2.791, 3)) = 0.964$] #math.equation(block: true, alt: "s sub est equals the square root of the fraction open parenthesis 1 minus r squared close parenthesis S S Y over N minus 2")[$s_("est") = sqrt(frac(( 1 − r^(2) ) S S Y, N − 2))$] x=c(1,2,3,4,5) y= c(1,2,1.3,3.75,2.25) summary(lm(y~x)) Call: lm(formula = y ~ x) Residuals: 1 2 3 4 5 -0.210 0.365 -0.760 1.265 -0.660 Coefficients: Estimate Std. Error t value Pr(\>|t|) (Intercept) 0.785 1.012 0.776 0.494 x 0.425 0.305 1.393 0.258 Residual standard error: 0.9645 on 3 degrees of freedom Multiple R-squared: 0.3929, Adjusted R-squared: 0.1906 F-statistic: 1.942 on 1 and 3 DF, p-value: 0.2578