0%

JSON and CSV · capstone

Capstone project: Project: Convert and Prove the Catalog

The converter is now a chain of narrow promises:

catalog.csv bytes
  -> decoded UTF-8 text
  -> CSV header and raw row dictionaries
  -> normalized typed item dictionaries
  -> validated JSON document
  -> catalog.json bytes
  -> parsed and validated document

Each arrow protects a different boundary. Keeping them separate is why a bad quantity can name its CSV row, a wrong JSON field can name its item, and a programming error does not get mislabeled as input damage.

Audit the complete contract

Read SCHEMA.md once more, then inspect catalog.py from top to bottom. Replace any remaining pass, copied constant result, broad handler, or temporary debug print.

The final project must provide all of these callable behaviors:

  • validate_headers(fieldnames) accepts only the exact ordered header;

  • read_csv_rows(path) uses csv.DictReader and rejects row-width damage;

  • parse_row(row, row_number) returns a new exact typed item;

  • read_catalog_csv(path) preserves row order and duplicates;

  • validate_document(document) enforces exact JSON shape and ;

  • write_catalog_json(items, path) validates and serializes deterministically;

  • read_catalog_json(path) parses standard JSON and validates it;

  • convert_catalog(input_path, output_path) connects the two file boundaries;

  • main() converts the seeded paths with narrow error handling.

Open Bash in the project workspace and run:

python catalog.py

The command should report:

Wrote 3 items to catalog.json.

Open the generated catalog.json. Confirm the product with a comma is still one name, Café is readable, quantities are unquoted integers, prices are unquoted numbers, and stock values are JSON .

Test it on something you have not seen

Your converter has only ever met one CSV file, and it was written to suit it. That is a weak position to submit from, so give it something unfamiliar first.

Open catalog.csv and change things a real file would differ in: rename a product, add a fourth row, set a quantity to 0, put a comma inside a name, and use a price with more decimal places. Run the converter again and read the JSON it produced.

Then break it deliberately, one thing at a time, and check that the message names the right layer:

  • put many in a quantity cell; the error should name the row;

  • swap two column headers; the error should name the expected header;

  • delete a closing brace in catalog.json and read it back; the error should point at the JSON, not the CSV.

Restore catalog.csv to three valid rows when you are finished, then press Check my work.

Doing this before you submit is not busywork. It is the habit that separates code you believe works from code you have watched work.

The practical rule to keep

Formats do not remove the need for decisions. CSV does not decide which cells are numbers. JSON syntax does not decide which fields your program requires. Write the schema, convert at the boundary, validate before trusting, and keep errors specific enough to repair the damaged layer.

Task

Finish the complete converter, remove every remaining pass or temporary shortcut, then run python catalog.py in Bash to create the current catalog.json.

Submit when the saved JSON matches catalog.csv, reads back through the same schema, preserves the declared on unseen rows, and rejects malformed CSV and JSON without a broad catch-all.