#set document(title: "14.3 Files in different locations and working with CSV files", author: "OpenStax / XYZ Homework") #set page(width: 8.5in, height: auto, margin: 1in) #import "@preview/cetz:0.5.2" #set text(font: ("STIX Two Text", "Libertinus Serif", "New Computer Modern"), size: 10.5pt, lang: "en") #show math.equation: set text(font: ("STIX Two Math", "New Computer Modern Math")) #set par(justify: true, leading: 0.62em, spacing: 0.9em) #set enum(spacing: 1.1em) // room between list items so tall inline fractions don't collide #set list(spacing: 1.1em) #set table(stroke: 0.5pt + rgb("#c7ccd3")) #let BLUE = rgb("#183B6F") // brand navy — section bars + example/solution labels (white on navy 11.09:1) #let ORANGE = rgb("#A94509") // brand primary-700 — AA-safe deep orange for TEXT (5.93:1 on white; raw brand #F37021 is 2.94:1 and must never carry text) #let RED = rgb("#DC2626") // brand error-600 #let GREEN = rgb("#059669") // brand success-600 (decoration only; small green text uses green-text #007942) #show heading.where(level: 1): it => block(width: 100%, above: 0pt, below: 16pt, fill: gradient.linear(BLUE, rgb("#2C5AA0")), inset: (x: 14pt, y: 12pt), radius: 3pt, text(fill: white, weight: "bold", size: 19pt, it.body)) #show heading.where(level: 2): it => block(width: 100%, above: 18pt, below: 10pt, fill: BLUE, inset: (x: 10pt, y: 6pt), radius: 2pt, text(fill: white, weight: "bold", size: 12pt, it.body)) #show heading.where(level: 3): it => text(fill: ORANGE, weight: "bold", size: 12.5pt, it.body) #show heading.where(level: 4): it => text(fill: BLUE, weight: "bold", size: 10.5pt, it.body) #let examplebox(label, title, body) = block(width: 100%, breakable: true, fill: rgb("#EFF1F5"), stroke: 0.5pt + rgb("#CFDDF0"), radius: 4pt, inset: 10pt, above: 12pt, below: 12pt)[ #block(below: 6pt)[#box(fill: BLUE, inset: (x: 6pt, y: 2pt), radius: 2pt, text(fill: white, weight: "bold", size: 8.5pt, label)) #h(0.4em) #strong[#title]] #body] // rail = decorative left rule (raw brand token); labelcolor = AA-safe label text shade #let notebox(label, rail, labelcolor, tint, body) = block(width: 100%, breakable: true, fill: tint, stroke: (left: 3pt + rail), inset: (left: 10pt, rest: 8pt), radius: (right: 4pt), above: 11pt, below: 11pt)[ #text(fill: labelcolor, weight: "bold", size: 7.5pt, tracking: 0.5pt)[#upper(label)] #linebreak() #body] #let solutionbox(body) = block(above: 4pt, below: 8pt)[ #text(fill: BLUE, weight: "bold", size: 8.5pt)[Solution] #linebreak() #body] #let figph(msg) = block(width: 100%, height: 60pt, fill: rgb("#f6f7f9"), stroke: (paint: rgb("#c7ccd3"), dash: "dashed"), radius: 4pt, inset: 10pt)[ #align(center + horizon, text(fill: rgb("#889"), style: "italic", size: 9pt, msg))] // Standardize inlined figure sizes: measure the natural CeTZ canvas, then scale to a // consistent envelope (aspect-aware; see build_typst.py FIG_* constants). Unlike the // print preamble, dimensions are FLOORED: in an editor a user can trim a figure to a // degenerate 1-D shape (a bare line), and w/h or tw/w would then divide by zero. #let _STD_W = 3.5 #let _WIDE_W = 5.6 #let _MAX_H = 3.4 #let _ASPECT_WIDE = 2.2 #let _UPSCALE_MAX = 1.15 #let stdfig(body) = context { let m = measure(body) let w = calc.max(m.width / 1in, 0.01) let h = calc.max(m.height / 1in, 0.01) let tw = if w / h > _ASPECT_WIDE { _WIDE_W } else { _STD_W } let s = calc.min(tw / w, _MAX_H / h, _UPSCALE_MAX) align(center, box(scale(x: s * 100%, y: s * 100%, reflow: true, body))) } #show figure: set block(breakable: false) #set figure(gap: 8pt) #show figure.caption: set text(size: 8.5pt, fill: rgb("#555")) == 14.3#h(0.6em)Files in different locations and working with CSV files === Learning objectives By the end of this section you should be able to - Demonstrate how to access files within a file system. - Demonstrate how to process a CSV file. === Opening a file at any location When only the filename is used as the argument to the open() function, the file must be in the same folder as the Python file that is executing. Ex: For fileobj = open("file1.txt") in #emph[files.py] to execute successfully, the #emph[file1.txt] file should be in the same folder as #emph[files.py]. Often a programmer needs to open files from folders other than the one in which the Python file exists. A #strong[path] uniquely identifies a folder location on a computer. The path can be used along with the filename to open a file in any folder location. Ex: To open a file named #emph[logfile.log] located in /users/turtle/desktop the following can be used: fileobj = open("/users/turtle/desktop/logfile.log") #figure(table( columns: 3, align: left, inset: 6pt, table.header([Operating System], [File location], [open() function example]), [Mac], [/users/student/], [fileobj = open("/users/student/output.txt")], [Linux], [/usr/code/], [fileobj = open("/usr/code/output.txt")], [Windows], [c:\\projects\\code\\], [fileobj = open("c:/projects/code/output.txt") #linebreak() or #linebreak() fileobj = open("c:\\\\projects\\\\code\\\\output.txt")], )) #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Opening files at different locations] For each question, assume that the Python file executing the open() function is not in the same folder as the #emph[out.txt] file. Each question indicates the location of #emph[out.txt], the type of computer, and the desired mode for opening the file. Choose which option is best for opening #emph[out.txt]. ] === Working with CSV files In Python, files are read from and written to as Unicode by default. Many common file formats use Unicode such as text files (#emph[.txt]), Python code files (#emph[.py]), and other code files (#emph[.c],#emph[.java]). Comma separated value (CSV, #emph[.csv]) files are often used for storing tabular data. These files store cells of information as Unicode separated by commas. CSV files can be read using methods learned thus far, as seen in the example below. #figure(figph[A CSV file displayed in both spreadsheet software, showing data in rows and cells, and a text editor, illustrating its raw comma-separated value format.], alt: "A CSV file displayed in both spreadsheet software, showing data in rows and cells, and a text editor, illustrating its raw comma-separated value format.", caption: [#strong[CSV files.] A CSV file is simply a text file with rows separated by newline \\n characters and cells separated by commas.]) Raw text of the file: Title, Author, Pages\\n1984, George Orwell, 268\\nJane Eyre, Charlotte Bronte, 532\\nWalden, Henry David Thoreau, 156\\nMoby Dick, Herman Melville, 538 #examplebox("Example 1")[Processing a CSV file][ """Processing a CSV file.""" \# Open the CSV file for reading file\_obj = open("books.csv") \# Rows are separated by newline \\n characters, so readlines() can be used to read in all rows into a string list csv\_rows = file\_obj.readlines() list\_csv = \[\] \# Remove \\n characters from each row and split by comma and save into a 2D structure for row in csv\_rows:   \# Remove \\n character   row = row.strip("\\n")   \# Split using commas   cells = row.split(",")   list\_csv.append(cells) \# Print result print(list\_csv) The code's output is: \[\['Title', ' Author', ' Pages'\], \['1984', ' George Orwell', ' 268'\], \['Jane Eyre', ' Charlotte Bronte', ' 532'\], \['Walden', ' Henry David Thoreau', ' 156'\], \['Moby Dick', ' Herman Melville', ' 538'\]\] ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[File types and CSV files] ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Exploring further] Files such as Word documents (#emph[.docx]) and PDF documents (#emph[.pdf]), image formats such as Portable Network Graphics (PNG, #emph[.png]) and Joint Photographic Experts Group (JPEG, #emph[.jpeg] or #emph[.jpg]) as well as many other file types are encoded differently. Some types of non-Unicode files can be read using specialized libraries that support the reading and writing of different file types. #link("https://openstax.org/r/100readthedocs")[PyPDF] is a popular library that can be used to extract information from PDF files. #link("https://openstax.org/r/100software")[BeautifulSoup] can be used to extract information from XML and HTML files. XML and HTML files usually contain unicode with structure provided through the use of angled \<\> bracket tags. #link("https://openstax.org/r/100readwrtedocx")[python-docx] can be used to read and write DOCX files. Additionally, #link("https://openstax.org/r/100csvlibrary")[csv] is a built-in library that can be used to extract information from CSV files. ] #notebox("Note", rgb("#8a94a6"), rgb("#556666"), rgb("#f7f8fa"))[ #emph[Processing a CSV file] The file #emph[fe.csv] contains scores for a group of students on a final exam. Write a program to display the average score. ]