#set document(title: "15.5 Data visualization", author: "OpenStax / XYZ Homework") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 15.5#h(0.6em)Data visualization === Learning objectives By the end of this section you should be able to - Explain why visualization has an important role in data science. - Choose appropriate visualization for a given task. - Use Python visualization libraries to create data visualization. === Why visualization? Data visualization has a crucial role in data science for understanding the data. Data visualization can be used in all steps of the data science life cycle to facilitate data exploration, identify anomalies, understand relationships and trends, and produce reports. Several visualization types are commonly used: #figure(table( columns: 3, align: left, inset: 6pt, table.header([Visualization type], [Description], [Benefits/common usage]), [Bar plot], [Rectangular bars], [Compare values across different categories.], [Line plot], [A series of data points connected by line segments], [Visualize trends and changes.], [Scatter plot], [Individual data points representing the relationship between two variables], [Identify correlations, clusters, and outliers.], [Histogram plot], [Rectangular bars representing the distribution of a continuous variable by dividing the variable's range into bins and representing the frequency or count of data within each bin], [Summarizing the distribution of the data.], [Box plot], [Rectangular box with whiskers that summarize the distribution of a continuous variable, including the median, quartiles, and outliers], [Summarizing the distribution of the data and comparing different variables.], )) #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Visualization types] #link("https://www.openstax.org/r/visualization-types")[Visualization types; ch 15, video 7] ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Comparing visualization methods] ] === Data visualization tools Many Python data visualization libraries exist that offer a range of capabilities and features to create different plot types. Some of the most commonly used frameworks are Matplotlib, Plotly, and Seaborn. Here, some useful functionalities of Matplotlib are summarized. #figure(table( columns: 2, align: left, inset: 6pt, table.header([Plot type], [Method]), [Bar plot], [The plt.bar(x, height) function takes in two inputs, x and height, and plots bars for each x value with the height given in the height variable.], [#strong[Example]], [#strong[Output]], [import matplotlib.pyplot as plt \# Data categories = \["Course A", "Course B", "Course C"\] values = \[25, 40, 30\] \# Create the bar chart fig = plt.bar(categories, values) \# Customize the chart plt.title("Number of students in each course') plt.xlabel("Courses") plt.ylabel("Number of students") \# Display the chart plt.show()], [#figure(figph[Bar chart example], alt: "Bar chart example", caption: none)], )) #figure(table( columns: 2, align: left, inset: 6pt, table.header([Plot type], [Method]), [Line plot], [The plt.plot(x, y) function takes in two inputs, x and y, and plots lines connecting pairs of (x, y) values.], [#strong[Example]], [#strong[Output]], [import matplotlib.pyplot as plt \# Data month = \["Jan", "Feb", "Mar", "Apr", "May"\] inflation = \[6.41, 6.04, 4.99, 4.93, 4.05\] \# Create the line chart plt.plot(month, inflation, marker="o", linestyle="-", color="blue") \# Customize the chart plt.title("Inflation trend in 2023") plt.xlabel("Month") plt.ylabel("Inflation") \# Display the chart plt.show()], [#figure(figph[Line plot example], alt: "Line plot example", caption: none)], )) #figure(table( columns: 2, align: left, inset: 6pt, table.header([Plot type], [Method]), [Scatter plot], [The plt.scatter(x, y) function takes in two inputs, x and y, and plots points representing (x, y) pairs.], [#strong[Example]], [#strong[Output]], [import matplotlib.pyplot as plt \# Data x = \[1, 2, 3, 4, 5, 6, 7, 8, 9, 10\] y = \[10, 8, 6, 4, 2, 5, 7, 9, 3, 1\] \# Create the scatter plot plt.scatter(x, y, marker="o", color="blue") \# Customize the chart plt.title("Scatter Plot Example") plt.xlabel("X") plt.ylabel("Y") \# Display the chart plt.show()], [#figure(figph[Scatter plot example], alt: "Scatter plot example", caption: none)], )) #figure(table( columns: 2, align: left, inset: 6pt, table.header([Plot type], [Method]), [Histogram plot], [The plt.hist(x) function takes in one input, x, and plots a histogram of values in x to show distribution or trend.], [#strong[Example]], [#strong[Output]], [import matplotlib.pyplot as plt import numpy as np \# Data: random 1000 samples data = np.random.randn(1000) \# Create the histogram plt.hist(data, bins=30, edgecolor="black") \# Customize the chart plt.title("Histogram of random values") plt.xlabel("Values") plt.ylabel("Frequency") \# Display the chart plt.show()], [#figure(figph[Histogram plot example], alt: "Histogram plot example", caption: none)], )) #figure(table( columns: 2, align: left, inset: 6pt, table.header([Plot type], [Method]), [Box plot], [The plt.boxplot(x) function takes in one input, x, and represents minimum, maximum, first, second, and third quartiles, as well as outliers in x.], [#strong[Example]], [#strong[Output]], [import matplotlib.pyplot as plt import numpy as np \# Data: random 100 samples data = \[np.random.normal(0, 5, 100)\] \# Create the box plot plt.boxplot(data) \# Customize the chart plt.title("Box Plot of random values") plt.xlabel("Data Distribution") plt.ylabel("Values") \# Display the chart plt.show()], [#figure(figph[A box plot titled 'Box Plot of Random Values' illustrates data distribution. The median is near -1, the box covers values from -4 to 2, whiskers extend from -12 to 7, and two outliers are visible above 10.], alt: "A box plot titled 'Box Plot of Random Values' illustrates data distribution. The median is near -1, the box covers values from -4 to 2, whiskers extend from -12 to 7, and two outliers are visible above 10.", caption: none)], )) #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Matplotlib methods] ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Exploring further] Please refer to the following user guide for more information about the Matplotlib, Plotly, and Seaborn libraries. - #link("https://openstax.org/r/100matplotguide")[Matplotlib User Guide] - #link("https://openstax.org/r/100plotlyguide")[Plotly User Guide] - #link("https://openstax.org/r/100seabornguide")[Seaborn User Guide] ] === Programming practice with Google Use the Google Colaboratory document below to practice a visualization task on a given dataset. #link("https://openstax.org/r/100visualization")[Google Colaboratory document]