0%

Capstone: A Practice Report Tool · practice

Choose Network, Cache, or Offline Data

You now have three possible sources for a page: the network, the cache on disk, and an authored file that never involves either. This lesson decides which one a given run uses.

That decision is the heart of the tool, and it is where a program like this usually goes wrong. Not because any single rule is difficult, but because the rules get discovered one at a time and bolted on, until nobody can say what the program will do without running it.

So write the table first.

The table

SituationFetches?Falls back?Source
--offline-source givennoneveroffline
Cache valid and freshnonot neededfresh-cache
Cache missingoncenonetwork
Cache invalidonceno, and warnsnetwork
Cache valid but staleonceyes, on expected failurenetwork or stale-cache
--refresh givenoncenevernetwork

Six rows. Every run lands in exactly one of them. Read it twice, because everything below is just those rows in more detail.

Offline means offline

--offline-source reads the authored page and stops. It does not check the cache, does not read the clock, and does not make a request. If the file is missing or does not match the requested page, that is an error, not a reason to go and fetch.

This has to be the very first branch. Anything that runs before it, including the clock read, breaks the promise that offline mode touches nothing else.

Missing and invalid are different failures

Both end up fetching, but they are not the same event and should not look the same.

A missing cache is completely normal. It is the first run. It is after someone cleaned up. Say nothing and fetch.

An invalid cache is worth a word. The file was there and could not be used, and a person who never hears about that will wonder why their tool is slow. Emit one warning, then fetch:

Warning: ignored invalid cache practice-cache.json.

One line, on stderr, and then carry on. This is a warning rather than an error because the program recovered.

Note which belong to which case. FileNotFoundError is missing. json.JSONDecodeError and ValueError are invalid. Any other OSError, a permissions problem, say, is neither: your program could not read a file it was told to read, and pretending that is an empty cache would hide a real problem.

Stale is a fallback, not a preference

A stale cache is old but structurally fine. The plan is to fetch fresh data. The of the stale copy is that it is there if fetching fails.

So: try the fetch. If it succeeds, use the new page. If it fails in one of the ways a network is expected to fail, fall back to the stale copy and warn that you did:

Warning: refresh failed; using stale cache practice-cache.json.

“Expected” is doing real work in that sentence, and it means exactly the failures Chapter 11 taught: requests.Timeout, requests.ConnectionError, requests.HTTPError, a JSON decode failure, or a ValueError from validating the response.

It does not mean a TypeError from a mistake in your own code. If your fetch has a bug and you catch it here, the program quietly serves old data and reports success, and the bug survives to production wearing a disguise. Let it crash.

Refresh means refresh

--refresh says: I want current data, and I know I might not get it. It fetches once and never falls back, even when a perfectly good stale cache is sitting right there.

Serving cached data to someone who explicitly asked for fresh data would be the single most confusing thing this tool could do.

Never more than one request

Every row of that table makes zero or one GET. Never two.

No retry , no pagination loop, no “try again with a smaller page size”. If a request fails, this program reports it or falls back. A retry policy is a real design decision that needs backoff, a cap, and a reason, and none of that belongs in the middle of a source-selection function.

Your turn

Implement this interface in practice_report.py:

import time

DEFAULT_TIMEOUT = api_catalog.DEFAULT_TIMEOUT


def choose_page(
    page, page_size, cache_path, max_age, *,
    refresh=False, offline_source=None, timeout=DEFAULT_TIMEOUT,
    session=None, now_epoch=None,
):
    ...

Return a four-item (page_document, source, fetched_at, warnings). source is exactly "offline", "fresh-cache", "network", or "stale-cache". Use None for the offline timestamp, the cache’s stored timestamp for either cache source, and the single now_epoch value for a successful network fetch. warnings is a of the warning shown above, without trailing newlines. Return that list; the prints it later.

offline_source names a UTF-8 JSON file containing a bare page, not a cache envelope. Parse it, validate it with _validate_report_page, and require its page and page size to match the requested integers. A helper such as read_offline_page(path, *, page, page_size) can keep that branch short.

For other modes, require genuine integers for page (1 through 10), page size (1 through 3), and non-negative max_age and now_epoch. Fetch through api_catalog.fetch_items_page(page, page_size, timeout=timeout, session=session) and apply _validate_report_page to the result. Let unrelated request-programming errors escape, just as TypeError does.

This function selects data; it does not write the cache or a report. Leave those writes to the CLI lesson.

Handle the branches in the table’s order. Sample int(time.time()) once, only when no now_epoch was supplied, and only after the offline branch has been ruled out.

Test choose_page directly with a fake session and an explicit now_epoch. The command-line entry point is completed in Lesson 5, so running practice_report.py --refresh is not yet a test of this function.

A cache file exists but contains malformed JSON. What should happen?

Why does --refresh refuse to fall back to a stale cache?

Why must a TypeError raised inside the fetch not trigger the stale-cache fallback?

Task

Implement the exact network/cache/offline decision table with one injected time, one bounded fetch at most, and narrow stale fallback .