0%

JSON and CSV · practice

Read the CSV Structure

Hand-splitting a CSV row at every comma would break the product named "Dice, Set of 6". The csv already knows the format’s quoting rules. Use that knowledge instead of rebuilding part of CSV with .

Let the header name each cell

csv.reader returns each row as a . csv.DictReader goes one step further: it uses the first row as keys for every later row.

with open("catalog.csv", newline="", encoding="utf-8") as handle:
    reader = csv.DictReader(handle, strict=True)
    for row in reader:
        print(row["sku"], row["name"])

newline="" lets the CSV parser handle CSV newlines itself. The explicit encoding and context manager continue Chapter 5’s file contract. strict=True makes malformed CSV quoting raise csv.Error instead of accepting a damaged record quietly.

The returned row is still a of text:

{
    "sku": "BK-101",
    "name": "Python Field Notes",
    "quantity": "4",
    "unit_price": "12.50",
    "in_stock": "true",
}

Validate the shape before values

reader.fieldnames is the header row. It must equal FIELDNAMES exactly, including order. Rejecting a reordered or extra header makes accidental file changes visible instead of silently changing the meaning of a column.

Rows need a width check too. DictReader uses the key None for extra cells and the None for missing cells. Reject either before returning rows. That leaves the next lesson one honest assumption: every returned dictionary has exactly the five declared text cells.

Your turn

Implement validate_headers(fieldnames) and read_csv_rows(path) in catalog.py.

  • A correct header returns normally; any other value raises with a message that names the expected header.

  • Open path with the Chapter 5 file contract shown above.

  • Build a csv.DictReader with strict=True.

  • Validate its header before reading data rows.

  • Reject a row with an extra or missing cell, and include its one-based CSV row number in the error.

  • Return a list of raw row dictionaries. Do not convert values yet.

Here “row number” means the logical CSV record number: the header is row 1 and the first data record is row 2. A quoted cell can contain a newline without becoming a second record.

You do not have to match the example line for line. What matters is that a good header returns rows, a reordered header raises, and a row with the wrong number of cells raises with its row number attached. Try each case in a temporary copy of catalog.csv and pass that copy’s path to read_csv_rows. Keep the original three-row catalog.csv unchanged for submission.

DictReader gives a row the key None or the value None. What has happened?

Task

Implement validate_headers(fieldnames) and read_csv_rows(path).

Return every raw row from catalog.csv, preserving the comma inside Dice, Set of 6. Reject a wrong header and any row with too many or too few cells. Leave every cell as text for the next lesson.