#set document(title: "9.7 Sampling Distribution of r", author: "OpenStax") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 9.7#h(0.6em)Sampling Distribution of r #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Prerequisites] Values of the Pearson Correlation, Introduction to Sampling Distributions #linebreak() #linebreak() ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Learning Objectives] + State how the shape of the sampling distribution of r deviates from normality + Transform r to z' + Compute the standard error of z' + Calculate the probability of obtaining an r above a specified value ] Assume that the correlation between quantitative and verbal SAT scores in a given population is 0.60. In other words, ρ = 0.60. If 12 students were sampled randomly, the sample correlation, r, would not be exactly equal to 0.60. Naturally different samples of 12 students would yield different values of r. The distribution of values of r after repeated samples of 12 students is the sampling distribution of r. The shape of the sampling distribution of r for the above example is shown in Figure 1. You can see that the sampling distribution is not symmetric: it is negatively #strong[skewed]. The reason for the skew is that r cannot take on values greater than 1.0 and therefore the distribution cannot extend as far in the positive direction as it can in the negative direction. The greater the value of ρ, the more pronounced the skew. #figure(figph[The sampling distribution of Pearson's r for N = 12 and a population correlation rho of 0.60, on an axis from -0.5 to 1. The curve peaks near 0.6 but is clearly asymmetric: a long tail stretches left toward 0 and below, while the right side falls away steeply and reaches zero at 1.0, because r cannot exceed 1.], alt: "The sampling distribution of Pearson's r for N = 12 and a population correlation rho of 0.60, on an axis from -0.5 to 1. The curve peaks near 0.6 but is clearly asymmetric: a long tail stretches left toward 0 and below, while the right side falls away steeply and reaches zero at 1.0, because r cannot exceed 1.", caption: [Figure 1. The sampling distribution of r for N = 12 and ρ = 0.60.]) Figure 2 shows the sampling distribution for ρ = 0.90. This distribution has a very short positive tail and a long negative tail. #figure(figph[The sampling distribution of Pearson's r for N = 12 and rho = 0.90, on an axis from 0.2 to 1. The skew is far more pronounced than at rho = 0.60: the peak sits near 0.9 with a very short tail to the right and a long tail running back past 0.4. The nearer rho is to 1, the harder the ceiling at 1 squeezes the distribution.], alt: "The sampling distribution of Pearson's r for N = 12 and rho = 0.90, on an axis from 0.2 to 1. The skew is far more pronounced than at rho = 0.60: the peak sits near 0.9 with a very short tail to the right and a long tail running back past 0.4. The nearer rho is to 1, the harder the ceiling at 1 squeezes the distribution.", caption: [Figure 2. The sampling distribution of r for N = 12 and ρ = 0.90.]) Referring back to the SAT example, suppose you wanted to know the probability that in a sample of 12 students, the sample value of r would be 0.75 or higher. You might think that all you would need to know to compute this probability is the mean and standard error of the sampling distribution of r. However, since the sampling distribution is not normal, you would still not be able to solve the problem. Fortunately, the statistician Fisher developed a way to transform r to a variable that is normally distributed with a known standard error. The variable is called z' and the formula for the transformation is given below. z' = 0.5 ln\[(1+r)/(1-r)\] The details of the formula are not important here since normally you will use either a table or #link("https://onlinestatbook.com/2/calculators/r_to_z.html")[calculator] to do the transformation. What is important is that z' is normally distributed and has a standard error of #math.equation(block: false, alt: "the fraction 1 over the square root of N minus 3")[$frac(1, sqrt(N − 3))$] where N is the number of pairs of scores. Let's return to the question of determining the probability of getting a sample correlation of 0.75 or above in a sample of 12 from a population with a correlation of 0.60. The first step is to convert both 0.60 and 0.75 to their z' values, which are 0.693 and 0.973, respectively. The standard error of z' for N = 12 is 0.333. Therefore the question is reduced to the following: given a normal distribution with a mean of 0.693 and a standard deviation of 0.333, what is the probability of obtaining a value of 0.973 or higher? The answer can be found directly from the applet "Calculate Area for a given X" to be 0.20. Alternatively, you could use the formula: z = (X - μ)/σ = (0.973 - 0.693)/0.333 = 0.841 and use a table to find that the area above 0.841 is 0.20.