#set document(title: "15.3 Pandas", author: "OpenStax / XYZ Homework") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 15.3#h(0.6em)Pandas === Learning objectives By the end of this section you should be able to - Describe the Pandas library. - Create a DataFrame and a Series object. - Choose appropriate Pandas functions to gain insight from heterogeneous data. === Pandas library #strong[Pandas] is an open-source Python library used for data cleaning, processing, and analysis. Pandas provides data structures and data analysis tools to analyze structured data efficiently. The name "Pandas" is derived from the term "panel data," which refers to multidimensional structured datasets. Key features of Pandas include: - Data structure: Pandas implements two main data structures: - Series: A #strong[Series] is a one-dimensional labeled array. - DataFrame: A #strong[DataFrame] is a two-dimensional labeled data structure that consists of columns and rows. A DataFrame can be thought of as a spreadsheet-like data structure where each column represents a Series. DataFrame is a heterogeneous data structure where each column can have a different data type. - Data processing functionality: Pandas provides various functionalities for data processing, such as data selection, filtering, slicing, sorting, merging, joining, and reshaping. - Integration with other libraries: Pandas integrates well with other Python libraries, such as NumPy. The integration capability allows for data exchange between different data analysis and visualization tools. The conventional alias for importing Pandas is pd. In other words, Pandas is imported as import pandas as pd. Examples of DataFrame and Series objects are shown below. #figure(table( columns: 2, align: left, inset: 6pt, table.header([DataFrame example], [Series example]), [Name Age City 0 Emma 15 Dubai 1 Gireeja 28 London 2 Sophia 22 San Jose], [0 Emma 1 Gireeja 2 Sophia dtype: object], )) === Data input and output A DataFrame can be created from a dictionary, list, NumPy array, or a CSV file. Column names and column data types can be specified at the time of DataFrame instantiation. #figure(table( columns: 4, align: left, inset: 6pt, table.header([Description], [Example], [Output], [Explanation]), [DataFrame from a dictionary], [import pandas as pd \# Create a dictionary of columns data = {   "Name": \["Emma", "Gireeja", "Sophia"\],   "Age": \[15, 28, 22\],   "City": \["Dubai", "London", "San Jose"\] } \# Create a DataFrame from the dictionary df = pd.DataFrame(data) \# Display the DataFrame df], [Name Age City 0 Emma 15 Dubai 1 Gireeja 28 London 2 Sophia 22 San Jose], [The pd.DataFrame() function takes in a dictionary and converts it into a DataFrame. Dictionary keys will be column labels and values are stored in respective columns.], [DataFrame from a list], [import pandas as pd \# Create a list of rows data = \[   \["Emma", 15, "Dubai"\],   \["Gireeja", 28, "London"\],   \["Sophia", 22, "San Jose"\] \] \# Define column labels columns = \["Name", "Age", "City"\] \# Create a DataFrame from list using column labels df = pd.DataFrame(data, columns=columns) \# Display the DataFrame df], [Name Age City 0 Emma 15 Dubai 1 Gireeja 28 London 2 Sophia 22 San Jose], [The pd.DataFrame() function takes in a list containing the records in different rows of a DataFrame, along with a list of column labels, and creates a DataFrame with the given rows and column labels.], [DataFrame from a NumPy array], [import numpy as np import pandas as pd \# Create a NumPy array data = np.array(\[   \[1, 0, 0\],   \[0, 1, 0\],   \[2, 3, 4\] \]) \# Define column labels columns = \["A", "B", "C"\] \# Create a DataFrame from the NumPy array df = pd.DataFrame(data, columns=columns) \# Display the DataFrame df], [A B C 0 1 0 0 1 0 1 0 2 2 3 4], [A NumPy array, along with column labels, are passed to the pd.DataFrame() function to create a DataFrame object.], [DataFrame from a CSV file], [import pandas as pd \# Read the CSV file into a DataFrame df = pd.read\_csv("data.csv") \# Display the DataFrame df], [The content of the CSV file will be printed in a tabular format.], [The pd.read\_csv() function reads a CSV file into a DataFrame and organizes the content in a tabular format.], [DataFrame from a Excel file], [import pandas as pd \# Read the Excel file into a DataFrame df = pd.read\_excel("data.xlsx") \# Display the DataFrame df], [The content of the Excel file will be printed in a tabular format.], [The pd.read\_excel() function reads an Excel file into a DataFrame and organizes the content in a tabular format.], )) #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Pandas basics] ] === Pandas for data manipulation and analysis The Pandas library provides functions and techniques to explore, manipulate, and gain insights from the data. Key DataFrame functions that analyze this code are described in the following table. import pandas as pd import numpy as np \# Create a sample DataFrame days = {   'Season': \['Summer', 'Summer', 'Fall', 'Winter', 'Fall', 'Winter'\],   'Month': \['July', 'June', 'September', 'January', 'October', 'February'\],   'Month-day': \[1, 12, 3, 7, 20, 28\],   'Year': \[2000, 1990, 2020, 1998, 2001, 2022\] } df = pd.DataFrame(days) #figure(table( columns: 5, align: left, inset: 6pt, table.header([], [Season], [Month], [Month-day], [Year]), [0], [Summer], [July], [1], [2000], [1], [Summer], [June], [12], [1990], [2], [Fall], [September], [3], [2020], [3], [Winter], [January], [7], [1998], [4], [Fall], [October], [20], [2001], [5], [Winter], [February], [28], [2022], )) #figure(table( columns: 4, align: left, inset: 6pt, table.header([Function name], [Explanation], [Example], [Output]), [head(n)], [Returns the first n rows. If a value is not passed, the first 5 rows will be shown.], [df.head(4)], [Season Month Month-day Year 0 Summer July 1 2000 1 Summer June 12 1990 2 Fall September 3 2020 3 Winter January 7 1998], [tail(n)], [Returns the last n rows. If a value is not passed, the last 5 rows will be shown.], [df.tail(3)], [Season Month Month-day Year 3 Winter January 7 1998 4 Fall October 20 2001 5 Winter February 28 2022], [info()], [Provides a summary of the DataFrame, including the column names, data types, and the number of non-null values. The function also returns the DataFrame's memory usage.], [df.info()], [\ RangeIndex: 6 entries, 0 to 5 Data columns (total 4 columns):  \#    Column        Non-Null Count    Dtype ---   ------        --------------    -----  0    Season        6 non-null        object  1    Month         6 non-null        object  2    Month-day     6 non-null        int64  3    Year          6 non-null        int64 dtypes: int64(2), object(2) memory usage: 320.0+ bytes], [describe()], [Generates the column count, mean, standard deviation, minimum, maximum, and quartiles.], [df.describe()], [Month-day Year count 6.000000 6.000000 mean 11.833333 2005.166667 std 10.457852 12.875040 min 1.000000 1990.000000 25% 4.000000 1998.500000 50% 9.500000 2000.500000 75% 18.000000 2015.250000 max 28.000000 2022.000000], [value\_counts()], [Counts the occurrences of unique values in a column when a column is passed as an argument and presents them in descending order.], [df.value\_counts \\('Season')], [Season Fall  2 Summer  2 Winter  2 dtype: int64], [unique()], [Returns an array of unique values in a column when called on a column.], [df\['Season'\] \\.unique()], [​​\['Summer' 'Fall' 'Winter'\]], )) #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[DataFrame operations] ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Exploring further] Please refer to the Pandas user guide for more information about the Pandas library. - #link("https://openstax.org/r/100pandaguide")[Pandas User Guide] ] === Programming practice with Google Use the Google Colaboratory document below to practice Pandas functionalities to extract insights from a dataset. #link("https://openstax.org/r/100googlecolab")[Google Colaboratory document]