Chapter 13 · practice
Capstone: A Practice Report Tool
A Cache Is Data with a Contract
Here is the problem this capstone solves. You have a command that fetches a page from the practice API and writes a report. Run it five times in a row and you have made five identical requests for data that did not change. That is slow, it is rude to the service, and it means your tool stops working the moment the network does.
The fix is a cache: keep the last answer on disk and reuse it while it is still good enough.
That sounds simple, and the simple version is where the trouble starts.
Derived does not mean trustworthy
Chapter 3 taught you that derived state is replaceable. .venv can be deleted and rebuilt, so losing it costs nothing. A cache is derived in exactly the same sense: throw it away and the program fetches again.
But there is a second question that .venv never made you ask. When your program reads that cache file, how does it know what it is getting?
The file is ordinary JSON sitting in a folder. Anything could have happened to it. It could be half-written from a run that was interrupted. It could be from three versions ago, when your report had different fields. It could be a cache of page 2 while you are asking for page 1. It could have been edited by hand by someone curious.
Reusing that blindly is worse than having no cache at all, because a wrong answer that arrives instantly is much harder to notice than no answer.
So: derived, yes. Disposable, yes. Trusted on sight, no. A cache is input, and input gets validated.
Give the cache an envelope
A bare copy of the API response is not enough. It tells you what the data was, but not what you need to know before reusing it. Wrap it:
{
"schema_version": 1,
"fetched_at": 1754000000,
"request": {"page": 1, "page_size": 3},
"data": {"items": [], "page": 1, "page_size": 3, "total": 0, "has_next": false, "next_page": null}
}
Exactly four keys, and each earns its place by answering a question you will actually have to ask:
schema_version answers “do I still understand this shape?” It is the integer 1. When you change the cache format later, this is what lets an old file be recognized and discarded instead of misread.
fetched_at answers “how old is this?” Epoch seconds, as a non-negative integer. Without it, freshness is unanswerable and the whole cache is a guess.
request answers “is this even the thing I asked for?” A cache of page 2 is a perfectly valid cache document and completely wrong for a request for page 1.
data is the validated page itself, in the shape Chapter 11 already taught your validator to check.
Booleans are integers, and that will bite you
You met this in Chapter 11 and it matters again here, so it is worth repeating in its new context.
In Python, bool is a subclass of int. That means isinstance(True, int) is True, and so a timestamp check written with isinstance will happily accept True as a valid fetched_at.
if type(value) is not int or value < 0:
raise ValueError("fetched_at must be a non-negative integer")
type(value) is int rejects True. Use it for every field where the schema says integer: the version, the timestamp, the page, the page size, the item ids, the values. JSON has a true literal, so this is not a hypothetical.
Validate before you open the file for writing
One ordering rule, and it is the kind that only announces itself at the worst moment.
Build the complete document, check it, and only then open the destination. Not the other way around.
Mode "w" empties a file the instant it opens. If you open first and discover a problem while writing, you have destroyed a perfectly good previous cache and replaced it with nothing. Validate first and a bad document simply never reaches the file, leaving yesterday’s good one intact.
You saw the same principle in Chapter 9 with the report writer. It generalizes: do the work that can fail before the step that destroys something.
Add the report’s stricter item rules
The Chapter 11 transport contract allows any integer for an item ID or
Add _validate_report_page(document). First call api_catalog.validate_items_page(document), then check those three report-specific rules and return the same
Your turn
In practice_report.py, implement four
validate_cache_document(document, *, page, page_size)checks the envelope, checks the page inside it, confirms the request matches what was asked for, and returns the data and its timestamp. Anything wrong raises. read_cache(path, *, page, page_size)opens the file, parses the JSON, and hands it to the validator.make_cache_document(page_document, *, fetched_at)builds the four-key envelope around a validated page.write_cache(page_document, path, *, fetched_at)builds the document, validates it, and only then writes it as UTF-8 JSON with two-space indentation and one trailing newline.
Use _validate_report_page for the page inside data. Copying the Chapter 11 validator here would give you two versions to keep in agreement.
Why validate a cache file your own program wrote?
Why check type(value) is int rather than isinstance(value, int) for fetched_at?
Task
Implement the report-specific page wrapper plus the exact cache validator, reader, constructor, and canonical writer in practice_report.py. Preserve api_catalog.py and generated-file sentinels.