Create, Import, and Evolve a Schema · practice
Validate Every CSV Row Before Writing
The CSV reader has proved that every row has five cells, but "0", " ", and a repeated asset tag still have the right shape. Now turn the complete raw batch into values that are safe to offer to the database.
Complete validate_item_rows(rows) in catalog_import.py. Run composes the two
validate_item_rows(read_item_rows(Path("data/items.csv")))
A valid row becomes a five-
== validate every row ==
Call: validate_item_rows(read_item_rows(Path("data/items.csv")))
(1, 'TL-100', 'Cordless drill', 'tools', 7)
(2, 'EV-200', 'Folding table', 'events', 14)
(3, 'EL-300', 'Projector', 'electronics', 3)
Repeated TL-100 -> Error: CSV contains asset tag TL-100 more than once.
For id and loan_days, strip whitespace and accept only ASCII decimal digits whose converted integer is greater than zero. This deliberately refuses signs, decimals, zero, negatives, and look-alike Unicode digits. For asset_tag, name, and category, strip surrounding whitespace and reject a result with no characters.
Keep separate sets for converted IDs and trimmed tags. Duplicate comparison must happen after conversion and trimming: "01" and "1" own the same ID, while " TL-100 " and "TL-100" own the same tag. Mention the normalized value in the duplicate error.
Build the result locally and return it only after every row passes. This is a pure staging step; it receives
Use the displayed row number in value errors too. A message such as CSV row 4 loan days must be a positive integer. points back to the source while keeping the rule explicit. Never skip an invalid row silently: one questionable record makes the whole supplied batch unsuitable for writing.
Task
Complete validate_item_rows(rows) in catalog_import.py.
Return five-