Lecture 5 (Lists and Eratosthenes Sieve)
Download the original Jupyter notebookAt a last lecture we started introducing lists. We have seen how to create a new list containing specific values:
my_list = [1,2,6,9,10]And how to access the value at a given index:
my_list[0]1
my_list[2]6
We can similarly change the value stored at a given position in the list:
my_list[2] = 20my_list[1, 2, 20, 9, 10]
Note: As a reminder, again my_list[2] = 20 changed the position at the third position at the list. The index 2 of refers to the third position, since the lists are indexed from 0 in Python.
The built-in function len outputs the current number of elements in the list:
len(my_list)5
Exercise 1
Write a function is_sorted(lst) that checks if the list lst is sorted, that is for all valid i it holds lst[i] <= lst[i + 1].
def is_sorted(lst):
for i in range(len(lst) - 1):
if lst[i] > lst[i+1]:
return False
return TrueNote that in the iteration we used the construction for i in range(...) as opposed to for x in lst:. Inside the loop, we need to have access to the index of the element we are looking at, to compare it with the next element. The for x in lst: construction would not give us that. Let us try to see if it works on some simple examples.
is_sorted([1,2,2,2,6,10])True
is_sorted([1,2,-10, 2,2,6])False
Building up variable length lists
We have seen now how to initialize an list of some constant size, access or modify an element of the list at a given position. Another crucial operation that we might care about is appending a new element at the end of the list. We do this with the syntax list_name.append(value). (We will discuss where this syntax comes from in the later lecture, when we will introduce objects, classes and methods). S far you can think of append as a function that takes two arguments: a list, and a value, and inserts the value at the end of the list.
We can start by creating an empty list, and assigning it to the variable lst.
lst = []Now we append an element 10 at the end of the list:
lst.append(10)Let us look up what is the value of the list now:
lst[10]
If we append another element, the list lst will be of length two, with those two elements in order:
lst.append(20)lst[10, 20]
A bit of unusual part of Python is that * operator is allowed for a list and a number: it will just create a new list, which consist of concatenation of the original list with itself $n$ times:
[10, 20] * 5[10, 20, 10, 20, 10, 20, 10, 20, 10, 20]
This is particularly useful if we want to initialize a list with the same value repeated over again n times. For example:
zeros = [0] * 20zeros[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
Exercise
Identity list:
Write a function identity(n) that takes a value n, and returns a list [0, 1, ..., n-1]
Solution 1.
We can just insert elements 0 to n-1 to an empty list one by one, using a for loop.
def identity(n):
result = []
for i in range(n):
result.append(i)
return resultidentity(10)[0, 1, 2, 3, 4, 5, 6, 7, 8, 9]
Solution 2.
We could have alternatively initialized an list with $n$ elements, all 0, and then in a loop set their values to consecutive numbers.
def identity(n):
result = [0] * n
for i in range(n):
result[i] = i
return resultidentity(5)[0, 1, 2, 3, 4]
Exercise
What follows below is the code from the last lecture for checking if a given number q is prime. Write a new function all_primes(n) that takes a single integer argument n, and returns a list of all primes from 2 to n. You should use the is_prime function to check if a number is prime.
def is_prime(q):
if q < 2:
return False
for i in range(2, q):
if q % i == 0:
return False
return TrueSolution.
def all_primes(n):
result = []
for i in range(2, n+1):
if is_prime(i):
result.append(i)
return resultall_primes(20)[2, 3, 5, 7, 11, 13, 17, 19]
This is better than our version of all_primes from the previous lecture, that printed all the primes in the range. Now, once we have a function that returns a list of all primes, if we want to print them, we can do this, but if we want to process it further, they are already in a convenient form.
For example, with the function product we wrote at some point before (that takes a list of integers, and returns a product of all the elements in the list), we can easily combine the two to produce the product of all primes from $2$ to $100$.
def product(lst):
result = 1
for x in lst:
result = result * x
return resultproduct(all_primes(100))2305567963945518424753102147331756070
References instead of values
Let us look at a potentially confusing behavior that Python starts to exhibiting when we operate with lists, and modify them.
Consider the following three lines of code: we first create a list [1,2,4] and assign it to variable a:
a = [1, 2, 4]Perfom another assignment operation:
b = aNow both a and b have value [1,2,4]:
a[1, 2, 4]
b[1, 2, 4]
Let us try to modify the value at the first position in the list b:
b[0] = 10The list b is now [10,2,4] as expected:
b[10, 2, 4]
But how about a?
a[10, 2, 4]
Surprisingly, it is also [10,2,4]. Even tough the assignment operation b[0] = 10 never explicitly mentioned a! What happens?
Answer
The crucial thing to keep in mind is that variables in Python contain references to specific values. You can visualize them as arrows pointing to specific value. Two different variables can now point to the same list!
In the code below, we create a new list, with two elements [10, 20], and make variable a point to this list.
a = [10, 20]Instruction b = a will not copy the list. It will make variable a and b to point to the same list!
b = aa[10, 20]
b[10, 20]
We can see this by using id function that will tell us the unique identifier of a specific object:
id(a)4500645568
id(b)4500645568
Variables a and b are pointing to the same list, with the same id. On the other hand:
c = [10,20]Here we create a new list with the same values 10,20, and make the variable c point to this new list.
c[10, 20]
a[10, 20]
id(c)4500636032
id(a)4500645568
The list pointed to by c have a different id. Now, we have two lists, and three variables a, b and c. The variables a and b are pointing to the same list, the variable c is pointing to a different list that happens to have the same values inside.
If we append a new value at the end of the list pointed to by a:
a.append(30)Clearly, value of a is now…
a[10, 20, 30]
But b is pointing to the exact same list:
b[10, 20, 30]
Whereas c is a separate list, so no element was added to it.
c[10, 20]
Important: Make sure you understand why the code above behaves the way it does. Missing the point here may lead to some very confusing behavior.
Passing a list to a function as an argument just creates another reference:
def add_star(items):
items.append("star")my_list = ["John"]print(my_list)
add_star(my_list)
print(my_list)['John']
['John', 'star']
my_list['John', 'star']
On the other hand, the code below will not change anything in the list passed by to the function by the argument lst. It will only make the variable lst inside the function to point to a new, empty list:
def replace(lst):
lst = []my_list = ["John"]
print(my_list)
replace(lst)
print(my_list)['John']
['John']
What if we want to create a new list that happens to have the same values as the list a variable a is pointing to - so that we can modify those two lists independently from each other? We can do it using the copy function:
a = [10, 20]
b = a.copy()
a.append(30)a[10, 20, 30]
b[10, 20]
We could have written a copy-function ourselves as well. It is doing basically the same thing as the function below:
def copy_list(lst):
result = []
for x in lst:
result.append(x)
return resulta = [10, 20]
b = copy_list(a)
a.append(30)a[10, 20, 30]
b[10, 20]
Exercise Look at a code below, and try to predict what is going to get printed without executing it:
a = [1,2]
b = a
c = b.copy()
b[0] = 9
c.append(3)
print(a)
print(b)
print(c)Listing all primes in range
We already wrote two times over a code that attempts to find all primes in a range $2$ to $n$. Both of those operated on the same, very basic premise: we iterate over all numbers from $2$ to $n$, and call a function checking if a number $k$ is prime. To check if a number $k$ is prime, we iterated over all potential divisiors, and checked if this is actually a divisor of $k$.
In the first implementation if is_prime(k) function, we checked the potential divsors in the range $2$ to $k-1$, leading to a final algorithm for finding all primes in the range $2, \ldots, n$ running in time $O(n^2)$. A very simple observation lead us to improve the algorithm significantly: to check if a number $k$ is prime, we only need to check all potential divisors in range $2, \ldots \sqrt{k}$. Indeed, if a number $k$ is composite, say $q$ is divisor of $k$, then either $q$ of $k/q$ is smaller than $\sqrt{k}$ — and both are divisors of $k$. Using this simple observation, we could list all primes in the range $2, \ldots n$ in time $O(n^{3/2})$ — a noticable improvement over the previous algorithm.
Sieve of Eratosthenes
In today’s lecture we will implement an even faster algorithm to achieve this bound, Erathostenes’s Sieve - an algorithm reportedly dating back to 3rd century BCE, and with documented refences going back to 2nd century CE.
In order to list all prime numbers in the range of $2, ... n$, we keep an array found_divisor of Booleans, of length $n+1$ (i.e. indexed from 0 to n). Over the course of the algorithm, found_divisor[k] for $k \geq 2$ is True if we have already found a non-trivial divisor for k. Initially the array is set to all values False.
The outer loop then iterates over the numbers i from 2 to n. If at a given step i, we have not found a divisor of i yet, we declare that i is prime (we will never find any divisor of i again), and we “cross out” all multiples of i: that is, for all numbers $2i,3i, 4i, 5i, \ldots$ we set found_divisor[k*i] to True. How long should we iterate? As long as $k*i$ is still a valid index in the array found_divisor, i.e. as long as k*i < len(found_divisor). You can see on the lecture slides a graphic visualization of the algorithm.
Let us start implementing the sieve. We focus on the high-level structure of the algorithm: we iterate over all possible i in the range, if we have not found a divisor of i yet, we declare it to be prime, append to the list result, and cross out all multiples of i. Since crossing out all multiples of i is self-contained piece of logic that would be too burdensome to write down right now, we can just postpone it until later - pretend that we have a function doing it, and just call it. We will implement this function next.
Exercise
Write a function erathostenes(n) that produces a list of all primes up to n. It should create a list with $n+1$ values False:
is_crossed = [False] * (n+1)then iterate over all numbers from 2 to n. If the number is not crossed, add it to the list of primes, and cross all multiples, assuming that there is some function doing this for us (we will implement it later):
def erathostenes(n):
is_crossed = [False] * (n+1)
is_crossed[0] = True
is_crossed[1] = True
result = []
for q in range(2, n+1):
if not is_crossed[q]:
result.append(q)
cross_all_multiples(is_crossed, q)
return resultExercise
Write a function cross_all_multiples(lst, q), that takes a list, and q, and sets lst[q*k] = False for all k from 2 until q*k goes beyond the list length
def cross_all_multiples(lst, q):
for k in range(2, len(lst)):
if q * k >= len(lst):
return
lst[q * k] = Trueerathostenes(100)[2,
3,
5,
7,
11,
13,
17,
19,
23,
29,
31,
37,
41,
43,
47,
53,
59,
61,
67,
71,
73,
79,
83,
89,
97]
We can see that it works much faster than our previous implementation. Counting the number of primes up to 10 milion with this sieve is much faster than using the previous implementation up to one milion
len(erathostenes(10000000))len(all_primes(1000000))Comment regarding the time complexity of the Eratosthenes’s Sieve.
As it turns out, the time complexity of this algorithm is a bit more difficult to analyse than for many other algorithms you have encountered so far. The function cross_out_all_multiples(arr, k) takes time $O(n/k)$ (where $n = \mathrm{len}(\mathrm{arr})$ is the size of the range we care about), and is called for all primes $k$ in a range of $1, \ldots n$. We can upper bound the total number of operations as
The sum on the right over all prime numbers $k$ in a range, is clearly smaller then the same sum, but over all numbers in the same range, that is
$$ \sum_{\substack{k \leq n \\ k\text{ prime}}} 1/k \leq \sum_{k \leq n} 1/k $$leading to
$$ T(n) \leq O(n) \sum_{k=1}^{n} 1/k. $$The sums of inverses of integers up to $n$, i.e. $H_{n} := \sum_{k=1}^n 1/k$, are called harmonic numbers. It is a well-known fact (you will probably learn this relatively soon in your mathematics course), that the $n$-th Harmonic number grows roughly as $\log n$, leading to a bound $T(n) \leq O(n \log n)$.
In fact, the sum only over the primes,
$$ \sum_{\substack{k \leq n \\ k\text{ prime}}} 1/k, $$while unbounded, grows even slower:
$$ \sum_{\substack{k \leq n \\ k\text{ prime}}} 1/k \approx \log \log n, $$leading to an improved bound $T(n) \leq O(n \log \log n)$. This is a rather deep mathematical fact, that follows from the Prime number theorem. See for example here, for more on that.