0%

JSON and CSV · practice

Read and Validate JSON

Parsing answers one question: “Is this text JSON, and which Python values does it describe?” Validation answers a different one: “Do those values satisfy our catalog schema?” A successful parse is necessary, not sufficient.

This is valid JSON:

{"message": "hello"}

It is not a catalog document. It has the wrong fields and no item array.

load reads a file; loads reads a string

The s names the versions:

Try it

json.loads(text) parses a string already in memory. Its partner json.dumps(value) returns a JSON string. For the project files, use the handle-based pair:

with open(path, encoding="utf-8") as handle:
    document = json.load(handle)

load and dump each work with one open file . The names are close; the input tells you which one belongs.

Validate containers before looking inside

Start at the top and move inward:

  1. The document’s exact type is dict, with exactly schema_version and items.

  2. The version’s exact type is int and its is 1.

  3. items has exact type list.

  4. Every item has exact type dict and exactly the five item fields.

  5. Each field has its declared type and validity rule.

The word exact matters around . In Python, bool is a subclass of int, which has a consequence people rarely believe until they see it:

Try it

The first line is why a schema check written with isinstance will happily accept true where it wanted a number. The third line is just there to make the point stick.

The schema does not permit true as the quantity or schema version, so use type(value) is int for those fields.

For unit_price, accept exact int or float, but not bool. Reject negative values and reject non-finite floats with math.isfinite.

Keep non-standard numbers out

Python’s JSON parser accepts NaN and Infinity by default for compatibility, although they are not JSON standard numbers. Pass a parse_constant that raises :

def reject_constant(value):
    raise ValueError(f"invalid JSON number: {value}")


json.load(handle, parse_constant=reject_constant)

Your turn

Implement validate_document(document) and read_catalog_json(path). Update write_catalog_json to validate the document before it writes, so the same schema guards both directions.

Return the validated document from both validation and reading. Let json.JSONDecodeError and your ValueError messages remain visible; Lesson 6 will decide where the user-facing command catches them.

json.load succeeded. What has that proved?

Why does the schema use type(value) is int for quantity?

Task

Implement validate_document(document) and read_catalog_json(path), then call validation from write_catalog_json before opening its output file.

Reject malformed JSON, non-standard NaN/Infinity, wrong or extra fields, wrong containers, empty/untrimmed text, negative or non-finite numbers, and in the integer or number fields. Return the valid document unchanged.