#set document(title: "9.5 Sampling Distribution of the Mean", author: "OpenStax") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 9.5#h(0.6em)Sampling Distribution of the Mean #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Prerequisites] Introduction to Sampling Distributions, Variance Sum Law I #linebreak() #linebreak() ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Learning Objectives] + State the mean and variance of the sampling distribution of the mean + Compute the standard error of the mean + State the central limit theorem ] The sampling distribution of the mean was defined in the section introducing sampling distributions. This section reviews some important properties of the sampling distribution of the mean introduced in the demonstrations in this chapter. === Mean The mean of the sampling distribution of the mean is the mean of the population from which the scores were sampled. Therefore, if a population has a mean μ, then the mean of the sampling distribution of the mean is also μ. The symbol μ#sub[M] is used to refer to the mean of the sampling distribution of the mean. Therefore, the formula for the mean of the sampling distribution of the mean can be written as: μ#sub[M] = μ === Variance The variance of the sampling distribution of the mean is computed as follows: #math.equation(block: true, alt: "σ sub M squared equals the fraction σ squared over N")[$σ_(M)^(2) = frac(σ^(2), N)$] That is, the variance of the sampling distribution of the mean is the population variance divided by N, the sample size (the number of scores used to compute a mean). Thus, the larger the sample size, the smaller the variance of the sampling distribution of the mean. (optional) This expression can be derived very easily from the variance sum law. Let's begin by computing the variance of the sampling distribution of the sum of three numbers sampled from a population with variance σ#super[2]. The variance of the sum would be σ#super[2] + σ#super[2] + σ#super[2]. For N numbers, the variance would be Nσ#super[2]. Since the mean is 1/N times the sum, the variance of the sampling distribution of the mean would be 1/N#super[2] times the variance of the sum, which equals σ#super[2]/N. The standard error of the mean is the standard deviation of the sampling distribution of the mean. It is therefore the square root of the variance of the sampling distribution of the mean and can be written as: #math.equation(block: true, alt: "σ sub M equals the fraction σ over the square root of N")[$σ_(M) = frac(σ, sqrt(N))$] The standard error is represented by a σ because it is a standard deviation. The subscript (M) indicates that the standard error in question is the standard error of the mean. === Central Limit Theorem The central limit theorem states that: Given a population with a finite mean μ and a finite non-zero variance σ#super[2], the sampling distribution of the mean approaches a normal distribution with a mean of μ and a variance of σ#super[2]/N as N, the sample size, increases. The expressions for the mean and variance of the sampling distribution of the mean are not new or remarkable. What is remarkable is that regardless of the shape of the parent population, the sampling distribution of the mean approaches a normal distribution as N increases. If you have used the "Central Limit Theorem Demo," you have already seen this for yourself. As a reminder, Figure 1 shows the results of the simulation for N = 2 and N = 10. The parent population was a #strong[uniform] distribution. You can see that the distribution for N = 2 is far from a normal distribution. Nonetheless, it does show that the scores are denser in the middle than in the tails. For N = 10 the distribution is quite close to a normal distribution. Notice that the means of the two distributions are the same, but that the spread of the distribution for N = 10 is smaller. #figure(figph[Two panels on the same 0-to-32 axis showing the distribution of the sample mean for a uniform population at two sample sizes. At N = 2 the distribution is a broad triangle spanning almost the whole range. At N = 10 it is a much narrower, distinctly bell-shaped mound running from about 8 to 24. Both are centered at 16; the red interval marker beneath each is correspondingly shorter for N = 10.], alt: "Two panels on the same 0-to-32 axis showing the distribution of the sample mean for a uniform population at two sample sizes. At N = 2 the distribution is a broad triangle spanning almost the whole range. At N = 10 it is a much narrower, distinctly bell-shaped mound running from about 8 to 24. Both are centered at 16; the red interval marker beneath each is correspondingly shorter for N = 10.", caption: [Figure 1. A simulation of a sampling distribution. The parent population is uniform. The blue line under "16" indicates that 16 is the mean. The red line extends from the mean plus and minus one standard deviation.]) Figure 2 shows how closely the sampling distribution of the mean approximates a normal distribution even when the parent population is very non-normal. If you look closely you can see that the sampling distributions do have a slight positive #strong[skew]. The larger the sample size, the closer the sampling distribution of the mean would be to a normal distribution. #figure(figph[Three stacked panels on a 0-to-32 axis. The top shows a strongly non-normal population — three separate humps near 5, 13 and 28 with sparse space between. Below it, the Distribution of Sample Mean at N = 5 is already a single smooth mound centered near 16, and at N = 25 it is narrower still and clearly bell-shaped. The central limit theorem at work: the distribution of the mean becomes normal even though the population is nothing like normal.], alt: "Three stacked panels on a 0-to-32 axis. The top shows a strongly non-normal population — three separate humps near 5, 13 and 28 with sparse space between. Below it, the Distribution of Sample Mean at N = 5 is already a single smooth mound centered near 16, and at N = 25 it is narrower still and clearly bell-shaped. The central limit theorem at work: the distribution of the mean becomes normal even though the population is nothing like normal.", caption: [Figure 2. A simulation of a sampling distribution. The parent population is very non-normal.])