Capstone: A Practice Report Tool · practice
Report the Page, and Test What You Built
Your tool can now find a page from three different places. What it cannot do yet is say anything about it. This lesson turns a validated page into the text a person actually reads.
The work splits in two, and keeping the halves apart is most of the lesson.
Counting is not printing
summarize_page(page_document) answers questions about the data. How many items? What do their values add up to? What does that look like broken down by category?
It returns data, not text:
{
"item_count": 3,
"value_total": 60,
"categories": [
{"category": "color", "item_count": 2, "value_total": 30},
{"category": "shape", "item_count": 1, "value_total": 30},
],
}
render_report(...) takes a page and produces the finished summarize_page and lays out the answer.
Why bother separating them? Because they change for different reasons. You will want to adjust the wording of the report far more often than you will want to change what “total” means, and a
Sort the categories, always
Categories come out of a dictionary you built while walking the items, so their order reflects the order they happened to appear. Two runs over the same data in a different order produce reports that differ only in the order of their lines.
That is a bad property for something people compare. Sort by category name and the report becomes a stable thing you can diff against yesterday’s.
The items themselves keep the order the API sent, because that order is data. The categories are your own grouping, so their order is yours to decide.
Make the names unambiguous
A category is text from a service. Text can contain a comma, a quote, or a newline, and any of those turn a tidy line into a confusing one.
name = json.dumps(category["category"], ensure_ascii=False)
json.dumps on a string returns that string quoted and escaped: "color" stays readable, and a name containing a newline comes out as a single unambiguous token instead of silently splitting your report across two lines. ensure_ascii=False keeps genuine non-ASCII characters legible rather than turning café into caf\u00e9.
This costs one function call and removes a whole family of confusing bugs.
The exact shape
Practice API report
source: network
fetched_at: 1754000000
page: 1
page_size: 3
dataset_total: 5
item_count: 3
value_total: 60
categories:
- "color": 2 item(s), 30 value
- "shape": 1 item(s), 30 value
source is one of network, fresh-cache, stale-cache, or offline, so the report says where its data came from. That matters: a reader deserves to know whether they are looking at a live answer or an hour-old one.
fetched_at is the timestamp for every source except offline, which has none and prints none. An offline report has no fetch time, and inventing one would be a lie in a file people keep.
One trailing newline at the end, as always.
Keep it pure
render_report takes values and returns a string. It does not open files, read the clock, or make requests.
That is what lets Chapter 13’s
Your turn
Two pieces of work, and they belong together.
First, add summarize_page(page_document) and render_report(page_document, *, source, fetched_at) to practice_report.py, following the shape above. Both use the report-specific page validation and leave the input unchanged. Pass source and fetched_at by name when calling the renderer.
The renderer rejects an unknown source with source="offline", require fetched_at=None and render the word none. For every other source, require a genuine non-negative integer timestamp; reject ValueError. An empty page has zero items, zero total categories: and the final newline.
Then write tests/test_practice_report.py, covering not just the report but everything the chapter has built so far. This is the
cache validation: a good document round-trips; a boolean
fetched_atis refused; a cache for page 2 is refused for a request for page 1.freshness: age equal to
max_ageis fresh, one second more is stale, a future timestamp raises.source selection: a fresh cache makes no request; a missing cache fetches once;
--refreshdoes not fall back; aTypeErrorfrom the fetch is not swallowed.the report: categories sorted, an empty page handled,
offlineprintingfetched_at: none, exactly one trailing newline.
Use the fake-session pattern from Chapter 11 for anything that would otherwise reach the network, inject now_epoch everywhere rather than reading the clock, and keep every file your tests create under tmp_path.
Aim for tests that would notice a realistic mistake. The question to ask about each one is not “does this pass?” but “what would have to break for this to fail?” If you cannot answer that, the test is decoration.
Why does summarize_page return a dictionary instead of formatted text?
Why sort the category rows by name?
Why must render_report avoid opening files or reading the clock?
Task
Implement summarize_page and render_report in practice_report.py, then cover both in tests/test_practice_report.py: sorted categories, an empty page, offline rendering fetched_at: none, a category name containing a comma, and exactly one trailing newline.