When a File Is No Longer Enough
When the Catalog Starts Contradicting Itself
Suppose Riley borrows a cordless drill. The lending program adds a row to its CSV file and copies Riley’s name and email address into that row.
A few weeks later, Riley changes their email address. When they borrow a projector, the new loan gets the new address. The old loan still carries the old one.
Now the file contains this:
| loan_id | item_name | member_id | member_name | member_email |
|---|---|---|---|---|
| 401 | Cordless drill | M-17 | Riley Chen | riley.old@example.com |
| 402 | Projector | M-17 | Riley Chen | riley.new@example.com |
Imagine that you need to send Riley a reminder. Which address should the program use?
The problem lives between the rows
There is nothing obviously broken inside either row. Both loan numbers are present, every cell has a
The contradiction only appears when you put the rows together. M-17 identifies the same member twice, but the two copies of that member’s email disagree.
CSV has stored exactly what the program gave it. The difficulty is that one fact about Riley was copied into every loan. Changing one copy did not change the others.
The same thing can happen to an item. If its name is copied into every past loan, renaming the item means finding every copy and hoping none are missed.
Why can both Riley rows look valid while the catalog is still untrustworthy?
Give each fact one home
What if Riley’s current details were stored once, with the member ID M-17? A loan could then refer to that member instead of carrying another copy of the name and email. Item details could have their own home too.
This is the kind of job a
That does not make CSV a bad format. A CSV file is still a good choice for a snapshot, an export, or a
You have found the reason for changing tools: the problem is no longer one malformed cell, but several facts that need to stay connected. The database we will use is SQLite. You may already be carrying hundreds of SQLite databases without knowing it.