0%

Create, Import, and Evolve a Schema · practice

Make an Exact Repeated Import a No-Op

Running the same trusted import twice should not create duplicates or fail. But “same” has to mean more than finding the same tag. Open your saved import_staged_items(connection, items) and improve the decision inside its existing .

For each incoming , first select all five CSV-owned fields by asset tag. There are three outcomes:

  • If every selected equals the incoming tuple, make no write and increment unchanged.

  • If the tag exists but any owned field differs, raise ValueError("CSV item TAG disagrees with the stored item.").

  • If the tag is new, check its incoming ID. When another tag owns that ID, raise ValueError("CSV item TAG uses an ID that belongs to another item."); otherwise insert it.

Keep both selects and the insert parameterized. Do not update or delete an existing row. A CSV owns exactly these five fields, so an exact comparison uses all five in their declared order.

Run repeats the exact calls you just saw. Only the observed decisions change:

== import staged items ==
First call: import_staged_items(connection, staged_items)
First result: (3, 0)
1 | TL-100 | Cordless drill | tools | 7
2 | EV-200 | Folding table | events | 14
3 | EL-300 | Projector | electronics | 3
Exact-repeat call: import_staged_items(connection, staged_items)
Exact-repeat result: (0, 3)
Changed-row call: import_staged_items(connection, changed_items)
Changed-row result: ValueError: CSV item TL-100 disagrees with the stored item.
Items after reopening: 3

This property is called idempotence: once an exact batch has been imported, repeating it reaches the same stored state. The unchanged count still gives the caller useful evidence that three rows were examined and already agreed.

Keep the complete inside one connection context. If a new first row is followed by a disagreement, raising the error exits the context and rolls that new row back. Classifying retries does not weaken atomicity. Return (inserted, unchanged) only after every staged row has been classified successfully.

The counters describe this invocation, not totals already stored in the table. An empty staged would therefore return (0, 0), although the public CSV reader never produces one.

Task

Edit the existing import_staged_items(connection, items) in catalog_import.py.

Implement the exact by-tag and by-ID decision matrix. Count only all-five-field matches as unchanged, insert only a new tag with an unused ID, and raise the exact disagreement or ID-ownership . Keep one and roll back all earlier new rows on any conflict.