#set document(title: "4.3 Complement Rule", author: "Rachel Webb") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 4.3#h(0.6em)Complement Rule #examplebox("Example 1")[][ A random sample of 500 records from the 2010 United States Census were downloaded to Excel and the following contingency table was found for biological sex and marital status. Select one member at random and find the following probabilities. Count of Marital Status Column Labels Row Labels Female Male Grand Total Divorced 21 17 38 Married/spouse absent 5 9 14 Married/spouse absent 92 100 192 Never married/single 93 129 222 Separated 1 2 3 Widowed 20 11 31 Grand Total 232 268 500 a) Compute the probability that a person is divorced. b) Compute the probability that a person is not divorced. #solutionbox[ a) Take the row total of all divorced which is 38 and then divide by the grand total of 500 to get P(Divorced) = 38/500 = 0.076. b) We can add up all the other category totals besides divorced 14 + 192 + 222 + 3 + 31 = 462, then divide by the grand total to get P(Not Divorced) = 462/500 = 0.924. ] ] There is a faster way to computer these probabilities that will be important for more complicated probabilities called the complement rule. The table contains 100% (100% = 1 as a proportion) of our data so we can assume that the probability of the divorced is the opposite (complement) to the probability of not being divorced. Notice that the P(Divorced) + P(Not Divorced) = 1. This is because these two events have no outcomes in common, and together they make up the entire sample space. Events that have this property are called #strong[complementary events]. Notice P(Not Divorced) = 1 – P(Divorced) = 1 – 0.076 = 0.924. #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ If two events are #strong[complementary events], then to find the probability of one event just subtract the probability from 1. Notation used for complement of A also called “not A” is A#super[C]. P(A) + P(A#super[C]) = 1 or P(A) = 1 – P(A#super[C]) or P(A#super[C]) = 1 – P(A) ] Some texts will use the notation for a complement as #emph[A'] or #math.equation(block: false, alt: "A bar")[$limits(A)^(―)$], instead of A#super[C]. #examplebox("Example 2")[][ Suppose you know that the probability of it raining today is 80%. What is the probability of it not raining today? #figure(figph[Cartoon of a blue rain cloud labeled 80% with raindrops falling beneath it.], alt: "Cartoon of a blue rain cloud labeled 80% with raindrops falling beneath it.", caption: none) #solutionbox[ Since not raining is the complement of raining, then P(not raining) = 1 – P(raining) = 1 – 0.8 = 0.2. ] ] “On a small obscure world somewhere in the middle of nowhere in particular - nowhere, that is, that could ever be found, since it is protected by a vast field of unprobability to which only six men in this galaxy have a key - it was raining.” (Adams, 2002) #strong[Venn Diagrams] Figure 4-5 is an example of a Venn diagram and is a visual way to represent sets and probability. The rectangle represents all the possible outcomes in the entire sample space (the population). The shapes inside the rectangle represent each event in the sample space. Usually these are ovals, but they can be any shape you want. If there are any shared elements between the events then the circles should overlap one another. #figure(figph[Three-circle Venn diagram inside a rectangle: circles labeled Statistics, Computer Science, and Business & Domain Expertise; the pairwise overlaps are labeled Machine Learning, Web Developer, and Data Analysis, and the center where all three meet is labeled Data Science.], alt: "Three-circle Venn diagram inside a rectangle: circles labeled Statistics, Computer Science, and Business & Domain Expertise; the pairwise overlaps are labeled Machine Learning, Web Developer, and Data Analysis, and the center where all three meet is labeled Data Science.", caption: none) Figure 4-5 The field of statistics includes machine learning, data analysis, and data science. The field of computer science includes machine learning, data science and web development. The field of business and domain expertise includes data analysis, data science and web development. If you know machine learning, then you will need a background in both statistics and computer science. If you are a data scientist, then you will need a background in statistics, computer science, business and domain expertise. #examplebox("Example 3")[][ Suppose you know the probability of not getting the flu is 0.24. Draw a Venn diagram and find the probability of getting the flu. #solutionbox[ Since getting the flu is the complement of not getting the flu, the P(getting the flu) = 1 – P(not getting the flu) = 1 – 0.24 = 0.76. Label each space as in Figure 4-6. #figure(figph[Venn diagram with a yellow rectangle for the sample space containing one blue circle labeled P(Flu) = 0.24; the region outside the circle is labeled P(No Flu) = 0.76.], alt: "Venn diagram with a yellow rectangle for the sample space containing one blue circle labeled P(Flu) = 0.24; the region outside the circle is labeled P(No Flu) = 0.76.", caption: none) Figure 4-6 ] ] The complement is useful when you are trying to find the probability of an event that involves the words “at least” or an event that involves the words “at most.” As an example of an “at least” event is supposing you want to find the probability of making at least \$50,000 when you graduate from college. That means you want the probability of your salary being greater than or equal to \$50,000. An example of an “at most” event is supposing you want to find the probability of rolling a die and getting at most a 4. That means that you want to get less than or equal to a 4 on the die, a 1, 2, 3, or 4. The reason to use the complement is that sometimes it is easier to find the probability of the complement and then subtract from 1.