0%

Chapter 11 · practice

Iteration as an Interface

Iterable and Iterator

You have written hundreds of for . Here is what one actually does.

for question in questions:
    print(question.prompt)

Python does not walk questions by . It asks questions for an , then asks that iterator for one item at a time until it says there are no more. The loop is a conversation with two participants, and knowing which is which explains a whole family of confusing behavior.

Two different things

An is anything you can loop over. A , a , a , a set, a file.

An iterator is the thing that does the walking. It holds the position. It knows what comes next.

They are usually not the same :

Try it

iter() asks an iterable for an iterator. A list is not its own iterator; it hands out a new one each time it is asked. An iterator instead returns itself, keeping the same position. It is therefore iterable too:

Try it

That is the design, and it has a consequence you can see immediately.

A list is an iterable that creates a fresh iterator, while the iterator moves forward one item at a time.

An iterator is used up

Try it

The first list(walker) collected all three. The second collected nothing. The iterator did not rewind, because it has no way back: all it holds is a position, and the position is now past the end.

The list itself is untouched:

Try it

Both give three items, because each list() call asked questions for a new iterator starting at the beginning.

Why does looping over the same list twice work, when looping over the same iterator twice does not?

Where this bites

This is not trivia. It causes a specific bug that looks like data disappearing:

Try it

Same numbers, two different answers. Given a list, sum() gets an iterator, exhausts it, and then list(scores) asks the list for a fresh one: three items. Given an iterator, sum() exhausts the only one there is, and list(scores) finds nothing left.

The is not wrong so much as underspecified: it walks its twice, and never said so. When you write a function that iterates its input more than once, either document that it needs a re-iterable collection, or turn the input into a list first and be explicit about the cost.

Why have the split at all

It would be simpler if a list just walked itself. The split buys two things.

Several loops at once. Two nested loops over the same list each get their own position, so the inner loop finishing does not end the outer one.

Try it

Nine, not three. With a single shared position that could not work.

Things too big to hold. An iterator only has to know how to produce the next item. It does not need them all to exist, which is what lets a program read a file larger than memory a line at a time. Lesson 6 returns to this.

The next lesson drives the conversation by hand, so the protocol stops being a description and becomes something you have used.

Task

Write three small that tell an and an apart by what they do, not by their type name.

  • walks_twice(items) returns a of two : everything collected on a first pass, and everything collected on a second. A list gives two full lists; an iterator gives one full and one empty.

  • is_own_iterator(items) reports whether asking the for an iterator hands back the object itself. That is true of an iterator and false of a list.

  • safe_count(items) returns how many items there are, and must work for both. Nothing that iterates its input twice can do this.

Run the program to see each answer for a list and for an iterator over the same three values.