#set document(title: "9.8 Sampling Distribution of p", author: "OpenStax") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 9.8#h(0.6em)Sampling Distribution of p #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Prerequisites] Introduction to Sampling Distributions, Binomial Distribution, Normal Approximation to the Binomial #linebreak() #linebreak() ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Learning Objectives] + Compute the mean and standard deviation of the sampling distribution of p + State the relationship between the sampling distribution of p and the normal distribution ] Assume that in an election race between Candidate A and Candidate B, 0.60 of the voters prefer Candidate A. If a random sample of 10 voters were polled, it is unlikely that exactly 60% of them (6) would prefer Candidate A. By chance the proportion in the sample preferring Candidate A could easily be a little lower than 0.60 or a little higher than 0.60. The sampling distribution of p is the distribution that would result if you repeatedly sampled 10 voters and determined the proportion (p) that favored Candidate A. The sampling distribution of p is a special case of the sampling distribution of the mean. Table 1 shows a hypothetical random sample of 10 voters. Those who prefer Candidate A are given scores of 1 and those who prefer Candidate B are given scores of 0. Note that seven of the voters prefer candidate A so the sample proportion (p) is p = 7/10 = 0.70 As you can see, p is the mean of the 10 preference scores. Table 1. Sample of voters. #figure(table( columns: 2, align: left, inset: 6pt, table.header([Voter], [Preference]), [1], [1], [2], [0], [3], [1], [4], [1], [5], [1], [6], [0], [7], [1], [8], [0], [9], [1], [10], [1], )) The distribution of p is closely related to the binomial distribution. The binomial distribution is the distribution of the total number of successes (favoring Candidate A, for example) whereas the distribution of p is the distribution of the mean number of successes. The mean, of course, is the total divided by the sample size, N. Therefore, the sampling distribution of p and the binomial distribution differ in that p is the mean of the scores (0.70) and the binomial distribution is dealing with the total number of successes (7). The binomial distribution has a mean of μ = Nπ Dividing by N to adjust for the fact that the sampling distribution of p is dealing with means instead of totals, we find that the mean of the sampling distribution of p is: μ#sub[p] = π The standard deviation of the binomial distribution is: #math.equation(block: true, alt: "the square root of N π open parenthesis 1 minus π close parenthesis")[$sqrt(N π ( 1 − π ))$] Dividing by N because p is a mean not a total, we find the standard error of p: #math.equation(block: true, alt: "σ sub p equals the fraction the square root of N π open parenthesis 1 minus π close parenthesis over N equals the square root of the fraction π open parenthesis 1 minus π close parenthesis over N")[$σ_(p) = frac(sqrt(N π ( 1 − π )), N) = sqrt(frac(π ( 1 − π ), N))$] Returning to the voter example, π = 0.60 and N = 10. (Don't confuse π = 0.60, the population proportion and p = 0.70, the sample proportion.) Therefore, the mean of the sampling distribution of p is 0.60. The standard error is #math.equation(block: true, alt: "σ sub p equals the square root of the fraction 0.60 open parenthesis 1 minus .60 close parenthesis over 10 equals 0.155")[$σ_(p) = sqrt(frac(0.60 ( 1 − ".60" ), 10)) = 0.155$] The sampling distribution of p is a discrete rather than a continuous distribution. For example, with an N of 10, it is possible to have a p of 0.50 or a p of 0.60 but not a p of 0.55. The sampling distribution of p is approximately normally distributed if N is fairly large and π is not close to 0 or 1. A rule of thumb is that the approximation is good if both Nπ and N(1 - π) are greater than 10. The sampling distribution for the voter example is shown in Figure 1. Note that even though N(1 - π) is only 4, the approximation is quite good. #figure(figph[The sampling distribution of the sample proportion p for N = 10 and a population proportion of 0.60, drawn as blue vertical spikes at 0, .1, .2, ... 1.0 with a smooth normal curve laid over them. The spikes are tallest at .6, then .5 and .7, and the curve passes close to the top of each — the normal approximation is good here even though N(1 - pi) is only 4.], alt: "The sampling distribution of the sample proportion p for N = 10 and a population proportion of 0.60, drawn as blue vertical spikes at 0, .1, .2, ... 1.0 with a smooth normal curve laid over them. The spikes are tallest at .6, then .5 and .7, and the curve passes close to the top of each — the normal approximation is good here even though N(1 - pi) is only 4.", caption: [Figure 1. The sampling distribution of p. Vertical bars are the probabilities; the smooth curve is the normal approximation.])