0%

Create, Import, and Evolve a Schema · capstone

Capstone project: Project: Prepare a Catalog from Source

The pieces now have clear jobs. catalog_setup.py decides whether a database may become the current schema. catalog_import.py reads and validates a complete CSV, then applies one idempotent . The project adds one path-owning that composes them without weakening either boundary.

Complete:

import_items_csv(database_path, csv_path)

First wrap database_path in Path and refuse a missing path with ValueError("Database does not exist. Run init first."). Only the explicit setup call may create a new file. Next call read_item_rows(csv_path) and validate_item_rows(rows) completely. This ordering matters: an invalid CSV must not migrate an otherwise valid version 1 database.

After staging succeeds, call prepare_database(database_path). Then open one configured application connection with open_database(database_path), pass it and the staged to import_staged_items, and return its two counts. Close that connection in finally, whether the import returns or raises.

Run uses a fresh temporary database_path. It calls prepare_database(database_path), then calls import_items_csv(database_path, Path("data/items.csv")) twice:

== prepare a catalog from source ==
Call: prepare_database(database_path)
Schema version: 2
First import: inserted 3, unchanged 0
Second import: inserted 0, unchanged 3
Items after reopening: 3

The first displayed call is the separate setup boundary. The wrapper intentionally refuses to hide creation inside an import command. That gives a future program an explicit init action and makes a misspelled database path safe.

Try to read this project as a sequence of ownership transfers. The CSV reader owns its file handle. prepare_database owns its setup connection. import_items_csv owns its application connection. The transaction function leaves a supplied connection open. Each owner closes only what it opened.

Keep the five-field import boundary unchanged. In particular, the new members.email column proves the migration had a visible purpose later, but item CSV data must never write member rows or email.

Run reopens the file before reporting the final count, so those three rows are durable rather than visible only through one connection.

Task

Complete import_items_csv(database_path, csv_path) in catalog_import.py.

Refuse a missing database, fully read and validate the CSV before setup, call prepare_database, open one configured connection, call import_staged_items, and close in finally. Return the inserted and unchanged counts without catching programming errors.