0%

Create, Import, and Evolve a Schema · practice

Check the CSV's Shape

Before an import can judge IDs or loan periods, it needs to know whether it received the promised table at all. Open catalog_import.py and complete read_item_rows(csv_path). This reads text only. It does not open, create, or prepare a database.

Run calls:

read_item_rows(Path("data/items.csv"))

The visible data/items.csv has the exact ordered header id,asset_tag,name,category,loan_days and three rows. A successful call returns raw , so values such as "1" and "7" are still :

== check the CSV shape ==
Call: read_item_rows(Path("data/items.csv"))
Rows read: 3
First row: {'id': '1', 'asset_tag': 'TL-100', 'name': 'Cordless drill', 'category': 'tools', 'loan_days': '7'}
Four-cell row -> Error: CSV row 2 must contain exactly five cells.

Use Path(csv_path).open("r", encoding="utf-8", newline="") and csv.DictReader(..., strict=True). Compare reader.fieldnames directly with ITEM_HEADERS; a set comparison would wrongly accept reordered or repeated headers. Start a row_number counter at 1 and increase it once for each data row. Reject both missing cells and extra cells. DictReader represents those cases differently, so check every and the special None key.

Normalize each input failure to the exact message below:

ProblemMessage
Missing pathCSV file does not exist.
Undecodable bytesCSV must be UTF-8 text.
Malformed CSVCSV is not valid.
Wrong or reordered headersCSV headers must be exactly: id,asset_tag,name,category,loan_days.
No item rowsCSV must contain at least one item row.
Wrong number of cells in record NCSV row N must contain exactly five cells.

Return only after reading the entire file. That gives the validator a complete in-memory batch before any database can change. At this boundary, an empty text cell is structurally valid. The validator will decide whether its meaning is acceptable.

The error row number counts CSV records, treating the header as row 1. It is not always a physical line number: a quoted field can span several lines, and DictReader skips blank lines. Use the record number to locate the item in a CSV viewer rather than assuming it names one line in a text editor. Keep commas inside quoted fields and Unicode text exactly as DictReader returns them. Do not trim or convert here; those are meaning checks, and they belong in the next function.

Task

Complete read_item_rows(csv_path) in catalog_import.py.

Require UTF-8 text, the exact ordered ITEM_HEADERS, at least one item row, and exactly five cells per row. Return raw . Use the exact normalized messages shown above, including CSV file does not exist., and do not touch a database.