Create, Import, and Evolve a Schema · practice
Make an Exact Repeated Import a No-Op
Running the same trusted import twice should not create duplicates or fail. But “same” has to mean more than finding the same tag. Open your saved import_staged_items(connection, items) and improve the decision inside its existing
For each incoming
If every selected
equals the incoming tuple, make no write and increment unchanged.If the tag exists but any owned field differs, raise
ValueError("CSV item TAG disagrees with the stored item.").If the tag is new, check its incoming ID. When another tag owns that ID, raise
ValueError("CSV item TAG uses an ID that belongs to another item."); otherwise insert it.
Keep both selects and the insert parameterized. Do not update or delete an existing row. A CSV owns exactly these five fields, so an exact comparison uses all five in their declared order.
Run repeats the exact calls you just saw. Only the observed decisions change:
== import staged items ==
First call: import_staged_items(connection, staged_items)
First result: (3, 0)
1 | TL-100 | Cordless drill | tools | 7
2 | EV-200 | Folding table | events | 14
3 | EL-300 | Projector | electronics | 3
Exact-repeat call: import_staged_items(connection, staged_items)
Exact-repeat result: (0, 3)
Changed-row call: import_staged_items(connection, changed_items)
Changed-row result: ValueError: CSV item TL-100 disagrees with the stored item.
Items after reopening: 3
This property is called idempotence: once an exact batch has been imported, repeating it reaches the same stored state. The unchanged count still gives the caller useful evidence that three rows were examined and already agreed.
Keep the complete (inserted, unchanged) only after every staged row has been classified successfully.
The counters describe this invocation, not totals already stored in the table. An empty staged (0, 0), although the public CSV reader never produces one.
Task
Edit the existing import_staged_items(connection, items) in catalog_import.py.
Implement the exact by-tag and by-ID decision matrix. Count only all-five-field matches as unchanged, insert only a new tag with an unused ID, and raise the exact disagreement or ID-ownership