0%

Capstone: A Practice Report Tool · practice

Make the CLI Boundary Honest

Everything works. Nobody can use it yet, because the only way to run it is to import it and call by hand.

This lesson puts a on the tool. You did this in Chapter 8, so the mechanics will feel familiar. What is new is that this program now writes two files, and the order it writes them in is a decision with consequences.

The interface

python practice_report.py REPORT_PATH [--cache PATH] [--page N] [--page-size N]
                          [--max-age SECONDS] [--refresh | --offline-source PATH]

The report path is positional, because a tool that writes a report should not make you guess where it went. Everything else has a default: cache at practice-cache.json, page 1, page size 3, maximum age 86400 seconds.

Three rules the parser enforces before anything else happens:

--refresh and --offline-source are mutually exclusive. One says “get me current data”, the other says “do not touch the network”. Together they are a contradiction, and argparse can reject a contradiction more clearly than your code can untangle one.

Abbreviation is off. allow_abbrev=False, as in Chapter 8. Without it, --ref silently means --refresh, and a typo becomes a feature nobody documented.

Path collisions fail before any I/O. If the report path and the cache path are the same file, stop. Writing both to one destination produces a file that is neither, and finding that out after the first write is too late.

Render before you open

Here is the ordering rule this lesson is really about.

1. choose the page          (may fail: network, cache, validation)
2. render the whole report  (may fail: bad data, defect in your code)
3. write the cache          (only on the network path)
4. write the report
5. print one line to stdout

Steps 1 and 2 can fail without touching either destination. Steps 3 and 4 can replace whatever was there before, and either write can still fail. So all validation and in-memory rendering finish before the first destination is opened.

Get this backwards, open the report file and render into it as you go, and a failure halfway through leaves the user with a truncated report and no previous version. They lose a good file and gain a broken one, from a run that was never going to succeed anyway.

This is the third time this idea has come up: the report writer in Chapter 5, the cache writer in lesson 1, and now the tool as a whole. Same rule each time, at a larger scale.

Cache first, then report

On the network path, write the cache before the report.

The reasoning is about which failure you would rather have. If the cache write succeeds and the report write fails, you have fresh cached data and no new report, so the next run is fast and produces the report. If the report write succeeds and the cache write fails, you have a report and no cache, so the next run fetches again unnecessarily.

Neither is a disaster. The first is better. Order them accordingly.

The other three paths, fresh cache, stale cache, and offline, write only the report. There is nothing new to cache in any of them.

One line out, one line wrong

The exit contract, unchanged in spirit from Chapter 8:

OutcomestdoutstderrStatus
Successone line naming the reportwarnings, if any0
Expected failureemptyone line1
Bad emptyargparse usage2
Defect in your codenothing promisednonzero

Catch only the expected failures: Requests Timeout, ConnectionError, HTTPError, and JSON decode errors; invalid input data; and OSError from reading or writing files. An unrelated RequestException, including InvalidURL, InvalidSchema, or MissingSchema, keeps its traceback. Those request-programming errors also inherit from and OSError, so re-raise remaining RequestException values after the expected Requests handlers and before either broad built-in handler.

Note that warnings and errors share stderr but mean different things. A warning accompanies a successful run and exits 0. An error replaces the result and exits 1.

Your turn

Finish build_parser() and this function in practice_report.py:

def run_report(
    report_path, cache_path, page, page_size, max_age, *,
    refresh=False, offline_source=None, session=None, now_epoch=None,
):
    ...

On success, print Wrote N items to REPORT_PATH from SOURCE. and one newline, replacing those placeholders with the selected item count, requested path, and source. Return 0. Return 1 for the expected operational failures above; let unrelated request errors and programming defects propagate.

Wire the parser to choose_page, render, write in the documented order, print the warnings you were handed, and return the right status. Keep main(argv=None) import-safe, exactly as in Chapter 8: parse, delegate, return.

Extend tests/test_practice_report.py with these CLI boundaries now: conflicting modes and colliding paths fail before any I/O; a render failure preserves the existing report and cache; offline mode touches neither network nor cache; and a successful command prints the documented message and returns zero. Use tmp_path, fake sessions, and capsys as before. The reference test also includes final CLI checks, which can pass once this lesson’s functions are implemented.

Then try to break it. Point --cache and the report at the same path. Pass --page 0. Pass both --refresh and --offline-source. Put text in the report file and run a command that will fail, then confirm your text survived.

Why render the complete report before opening the destination file?

On a successful network run, why write the cache before the report?

The tool falls back to a stale cache and writes a report. What should it do?

Task

Complete the exact matrix, path-collision preflight, narrow messages, cache-before-report order, and sentinel-preserving pre-write behavior.