if, for, functions, and what to do
when your program explodes. Almost half of today.By the end you will have a Python program that reads the log you spent Week 1 grepping, counts what is genuinely in it, and writes the answer out as a CSV. You type in every part.
grep -c ERROR gave you 294. Then 263. The real number is 279.Same file. Same question. Nothing was broken.
What is grep counting that you never asked it to count?
There is no compile step and no build. python3 count.py starts the interpreter, which
reads your file from the first line to the last and executes each statement as it meets it.
Stop half way and everything above the stopping point has already happened.
Three ways to run it. A file is what you ship. python3 -c 'print(1 + 1)' runs a single
line straight from the shell. The REPL — plain python3 — reads one line, runs it,
prints the result, and is where you check what something does before you commit to it.
exit() or Ctrl-D leaves.
$ python3
>>> 294 - 263
31
>>> exit()
Type python3, not python. On Ubuntu and WSL a bare python often does not exist.
service = "checkout-api" binds the name on the left to the value on the right. You never
declare a type — the value knows what it is, and type(x) will tell you. = assigns,
== compares. Text is str, whole numbers int, decimals float, True/False bool.
A list holds values in order: levels[0] is the first, levels[-1] the last,
len(levels) how many, and "ERROR" in levels asks whether it is in there at all.
An f-string puts values inside text: put f before the quote and names in braces.
service = "checkout-api"
errors = 262
levels = ["FATAL", "ERROR", "WARN", "INFO"]
print(type(errors), len(levels), levels[-1])
print(f"{service}: {errors} errors")
<class 'int'> 4 INFO
checkout-api: 262 errors
if."if takes a condition, ends the line with a colon, and everything indented under it runs
only when that condition is true. There are no braces and no endif — the indent is the
block. Four spaces, consistently, and Python refuses to run a file whose indentation does not
line up rather than guess what you meant.
elif tests the next case only when everything above it failed; else catches the rest.
Conditions use == != < > <= >=, joined with and, or and not.
if errors > 100:
print("page someone")
elif errors > 0:
print("open a ticket")
else:
print("quiet night")
page someone
for loop names one thing at a time out of a listfor level in levels: runs the indented block once per item, with level holding each item
in turn. You never manage an index. Anything walkable works — a list, the characters of a
string, and later today the lines of a file.
range(1, 4) produces 1, 2, 3. The stop value is not included, which is the most common
off-by-one in the language.
Two words steer the loop from inside it: break leaves it immediately, continue
abandons this item and starts the next.
for level in ["FATAL", "ERROR", "WARN"]:
if level == "WARN":
continue
print(level, len(level))
for attempt in range(1, 4):
print("attempt", attempt)
FATAL 5
ERROR 5
attempt 1
attempt 2
attempt 3
def names a block of code. import borrows one somebody else already named.A function is a name for a piece of work. def is_bad(level): declares it, the indented body
is what it does, and return hands a value back to the caller — a function with no return
hands back None. The names in the brackets are variables that exist only inside.
You will not write most of the functions you use. Python ships with a few hundred modules,
already on every machine here: import json gives you json.loads, import time gives you
time.sleep.
Nothing today needs pip. That changes in the last block, and there is a reason.
def is_bad(level):
return level in ["ERROR", "FATAL"]
print(is_bad("ERROR"), is_bad("INFO"))
True False
A list, a function, a loop, an if and an f-string — five slides, one program. Type it, do
not read it.
mkdir -p /tmp/lab && cd /tmp/lab
cat > triage.py <<'EOF'
#!/usr/bin/env python3
levels = ["INFO", "ERROR", "INFO", "FATAL", "ERROR", "INFO", "WARN"]
def is_bad(level):
return level in ["ERROR", "FATAL"]
bad = 0
for level in levels:
if is_bad(level):
bad = bad + 1
print(f"{len(levels)} records, {bad} of them need someone")
EOF
python3 triage.py
Yours must say 7 records, 3 of them need someone. Now add "FATAL" to the list and say
the new number out loud before you run it again.
Stretch: make is_bad count WARN too, but only when a second argument asks it to.
When Python cannot do what a line says, it stops and prints a traceback: every call that
was in flight, outermost first, then a final line naming the type of failure and the
detail. Read it from the bottom. The bottom line says what went wrong; the File …, line N above it says where.
Everything before the failure already ran — starting is on screen. And the process exits
with status 1, the same non-zero exit code you read all through Week 1.
$ python3 rate.py
starting
Traceback (most recent call last):
File "/private/tmp/lab/rate.py", line 5, in <module>
print(rate(279, 0))
~~~~^^^^^^^^
File "/private/tmp/lab/rate.py", line 2, in rate
return errors / total
~~~~~~~^~~~~~~
ZeroDivisionError: division by zero
try / except is Week 1's exit code, wearing different clothesIn the shell you ran a command and looked at $?. In Python a failure is an exception —
an object with a type — and you say in advance which types you are willing to survive.
The fragile line goes in try:. except ValueError: catches only that type and runs your
recovery; as e hands you the exception, which prints as its message. else: runs only when
nothing was raised, so it holds the work that depends on success.
Name the exception you expect. A bare except: also swallows your own typos, and you will
spend an evening looking for the bug it hid.
for raw in ["262", "twelve", "17"]:
try:
n = int(raw)
except ValueError as e:
print(f"skipping {raw}: {e}")
else:
print(f"{raw} -> {n * 2}")
262 -> 524
skipping twelve: invalid literal for int() with base 10: 'twelve'
17 -> 34
for loop with a try inside it. You already have both halves.Nothing new here: loop a fixed number of attempts, put the fragile call in try, and break
the moment it works. raise is how a failure gets created — below, it stands in for a server
that is down.
Sleeping longer after each failure is backoff. It is what stops your retries becoming the outage.
Retry a timeout, never a rejection — a 401 will be a 401 next time too. Cap the attempts, and when the last one fails, say so and exit non-zero.
import time
def fetch(attempt):
if attempt < 3:
raise ConnectionError("upstream timed out")
return "the data"
for attempt in range(1, 4):
try:
print("got", fetch(attempt), "on attempt", attempt)
break
except ConnectionError as e:
print(f"attempt {attempt} failed: {e}")
time.sleep(attempt)
mkdir -p /tmp/lab && cd /tmp/lab
cat > retry.py <<'EOF'
#!/usr/bin/env python3
import time
def fetch(attempt):
if attempt < 3:
raise ConnectionError("upstream timed out")
return "the data"
for attempt in range(1, 4):
try:
print("got", fetch(attempt), "on attempt", attempt)
break
except ConnectionError as e:
print(f"attempt {attempt} failed: {e}")
time.sleep(attempt)
EOF
python3 retry.py
It takes three seconds — the sleeps are real, and watching them is the point.
Now change attempt < 3 to attempt < 9 and predict what you will see.
Stretch: print gave up after the loop when no attempt worked.
open() hands you a handle to a stream, not the contents of the fileopen("app.log") reads nothing. It returns a file object — a cursor sitting at byte zero,
plus an operating-system handle you are responsible for giving back. Forget to give it back
and a long-running program eventually runs out of them.
with open(...) as f: closes it for you at the end of the indented block, even if the code
inside raises. It is the only form worth learning.
f.read() returns the whole file as one string; f.readline() returns one line. read() on
last week's 900 KB log is fine. On the 40 GB log that less opened instantly it is not —
same judgement as cat versus less, in a different language.
with open("app.log") as f:
first = f.readline()
print(first[:58])
{"ts":"2026-08-01T00:01:10.927+05:30","level":"ERROR","ser
\nA file object is something for can walk, and it hands back one line per turn, reading
only as much as it needs — so a file larger than your memory is not a problem. What arrives
includes the newline that ended it, which is why comparing a line against "INFO" quietly
fails when it looks right.
line.strip() returns a copy with whitespace stripped off both ends, and line.upper() an
upper-case copy. String methods return a new string and never change the original.
"ERROR" in line asks whether that text appears anywhere in it — the same in you used on a
list.
lines = 0
errors = 0
with open("app.log") as f:
for line in f:
lines = lines + 1
if "ERROR" in line:
errors = errors + 1
print(f"{lines} lines, {errors} contain ERROR")
5000 lines, 294 contain ERROR
"w" empties the file before you write a byte. "a" adds to the end.The mode is open()'s second argument, and "r" for reading is the default. "w" truncates
the file to zero length the moment it is opened — before your first write, and whether or
not your program then crashes half way through. "a" appends instead. If you have ever
destroyed a file with > in the shell, this is the same trapdoor with a different spelling.
f.write() writes exactly the string you handed it and adds no newline, unlike print.
You put the \n in yourself, every time.
with open("report.txt", "w") as f:
f.write("level,count\n")
f.write("ERROR,262\n")
with open("report.txt", "a") as f:
f.write("FATAL,17\n")
level,count
ERROR,262
FATAL,17
pathlib is the module that knows the differenceGluing paths with + and "/" breaks the day somebody leaves a trailing slash in a variable,
or runs your code on Windows. from pathlib import Path gives you a Path, and / joins two
of them properly: Path("/tmp/lab") / "app.log".
A Path answers what a string cannot: .exists(), .name, .suffix, .parent, and
.stat().st_size for the byte count. .read_text() and .write_text() open, read or write,
and close in one call — right when the file is small enough to hold in memory.
from pathlib import Path
log = Path("/tmp/lab") / "app.log"
print(log, log.exists(), log.suffix, log.stat().st_size)
/tmp/lab/app.log True .log 905782
mkdir -p /tmp/lab && cd /tmp/lab
[ -s app.log ] || curl -fsS -o app.log https://dataeko-training.studiotypo.xyz/files/app.log
cat > count.py <<'EOF'
#!/usr/bin/env python3
lines = 0
errors = 0
with open("app.log") as f:
for line in f:
lines = lines + 1
if "ERROR" in line:
errors = errors + 1
print(f"{lines} lines, {errors} contain ERROR")
EOF
python3 count.py
grep -c ERROR app.log
Your script and grep must print the same 294. Hands up if they disagree.
Stretch: count the lines that do not start with {, and tell me what those lines are.
A list is indexed by position; a dict is indexed by a key you choose. record["level"]
looks up whatever is stored under "level". {} is an empty one, and assigning to a key that
does not exist creates it.
Reading a key that does not exist raises KeyError and stops your program. That is why
record.get("user_id") is the safer read: it returns None when the key is missing, or a
default you supply as a second argument.
len() counts the keys, and for key in record: walks them in the order they were
inserted.
record = {"level": "ERROR", "service": "checkout-api"}
print(record["level"], record.get("user_id", "unknown"))
record["retries"] = 3
for key in record:
print(key, "=", record[key])
ERROR unknown
level = ERROR
service = checkout-api
retries = 3
json.loads takes a string and gives back a dict. Printing a dict is not JSON.Every line of last week's log is a string that happens to be JSON-shaped. json.loads(line)
parses it into a real dict you can subscript; a line that is not valid JSON raises
json.JSONDecodeError. json.dumps(record) goes the other way and returns a string.
They are not interchangeable, and print is where people learn that the hard way. Python
prints a dict in its own notation — single quotes, False, None. JSON has none of those:
it demands double quotes, false and null. Write a printed dict to a file, call it JSON, and
nothing else will ever read it.
import json
line = '{"level":"ERROR","service":"checkout-api","ok":false}'
record = json.loads(line)
print(record["service"], type(record))
print(record)
print(json.dumps(record))
checkout-api <class 'dict'>
{'level': 'ERROR', 'service': 'checkout-api', 'ok': False}
{"level": "ERROR", "service": "checkout-api", "ok": false}
Nothing about JSON promises you a list of records. This is the shape Session 2's weather API
genuinely returns: an object holding two parallel arrays, one of timestamps and one of
temperatures, lined up by position. There is no {"time": …, "temp": …} anywhere in it.
Ask it for two days and each array holds 48 entries. hourly["temperature_2m"][3] on its own
tells you nothing — it means something only when read next to hourly["time"][3].
''' opens a string that may run over several lines.
import json
saved = '''{"hourly": {
"time": ["2026-08-10T00:00", "2026-08-10T01:00", "2026-08-10T02:00"],
"temperature_2m": [21.2, 20.7, 20.4]}}'''
hourly = json.loads(saved)["hourly"]
print(len(hourly["time"]), hourly["time"][0], hourly["temperature_2m"][0])
3 2026-08-10T00:00 21.2
This is a saved sample, trimmed to three hours. The live API returns different numbers every single call.
zip walks two lists side by side and hands you a pair each turnzip(a, b) gives you the first item of each, then the second, then the third, and stops when
the shorter list runs out. for t, temp in zip(...) unpacks that pair into two names on the
spot — one name per list, in order.
It is the tool for precisely the shape on the last slide: two columns that mean nothing apart and a record when you put them together.
rows.append(x) adds x to the end of a list. Building one dict per row is what the next
ten minutes needs, because a CSV writer wants exactly that.
rows = []
for t, temp in zip(hourly["time"], hourly["temperature_2m"]):
rows.append({"time": t, "temp_c": temp})
print(len(rows), rows[0])
3 {'time': '2026-08-10T00:00', 'temp_c': 21.2}
csv module does the quoting you will forgetcsv.DictWriter(f, fieldnames=[…]) fixes the column order. writeheader() writes the first
line; writerows(rows) writes a list, writerow(row) writes one — both quoting any value
containing a comma so it stays one column. csv.DictReader(f) reads the header and hands you
one dict per row.
Two things to hold. Pass newline="" to open for any CSV, or you get blank rows between
your data. And everything read back is a string — a CSV carries no types, so 21.2
returns as "21.2" until you call float().
import csv
with open("hourly.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=["time", "temp_c"])
w.writeheader()
w.writerows(rows)
with open("hourly.csv", newline="") as f:
for row in csv.DictReader(f):
print(row["time"], row["temp_c"], type(row["temp_c"]))
2026-08-10T00:00 21.2 <class 'str'>
2026-08-10T01:00 20.7 <class 'str'>
mkdir -p /tmp/lab && cd /tmp/lab
[ -s app.log ] || curl -fsS -o app.log https://dataeko-training.studiotypo.xyz/files/app.log
cat > levels.py <<'EOF'
import csv, json
counts = {}
skipped = 0
with open("app.log") as f:
for line in f:
try:
record = json.loads(line)
except json.JSONDecodeError:
skipped = skipped + 1
continue
level = record["level"].upper()
counts[level] = counts.get(level, 0) + 1
with open("levels.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=["level", "count"])
w.writeheader()
for level in counts:
w.writerow({"level": level, "count": counts[level]})
print("skipped", skipped, "lines that were not JSON")
print("needs someone:", counts["ERROR"] + counts["FATAL"])
EOF
python3 levels.py && cat levels.csv
279. Stretch: delete .upper() and explain the two numbers that change.
subprocess.run starts a program. Hand it a list, not a sentence.subprocess.run(["./logsum.sh", "app.log"]) starts the script you wrote in Week 1 as a child
process and waits for it to finish. The list is the argument vector: program name first, then
each argument as its own item. Pass one string instead and Python looks for a file whose name
contains a space, and you get FileNotFoundError.
There is no shell in between, and that is the point — no globbing, no $HOME, and no way for a
filename containing a semicolon to become a second command.
capture_output=True collects the child's output instead of letting it hit your terminal, and
text=True gives you str rather than bytes.
import subprocess
done = subprocess.run(["./logsum.sh", "app.log"],
capture_output=True, text=True)
print(done.returncode, type(done.stdout))
print(done.stdout.splitlines()[1])
0 <class 'str'>
lines: 5000
returncode is $?, and check=True turns a bad one into an exceptionrun hands back a CompletedProcess with three fields worth knowing: returncode, stdout
and stderr. returncode is exactly the exit status you spent Week 1 reading — logsum.sh
exits 1 when you call it with no argument, and Python sees that 1.
By default Python does not care that it failed. check=True makes a non-zero return raise
CalledProcessError, which you handle with the try/except you wrote forty minutes ago.
The exception carries returncode and stderr with it.
That is the whole pattern: a failing child becomes a failure in your program, instead of a number nobody looked at.
import subprocess
try:
subprocess.run(["./logsum.sh"], capture_output=True, text=True, check=True)
except subprocess.CalledProcessError as e:
print("failed with", e.returncode)
print(e.stderr.strip())
failed with 1
usage: ./logsum.sh <logfile>
mkdir -p /tmp/lab && cd /tmp/lab
[ -s app.log ] || curl -fsS -o app.log https://dataeko-training.studiotypo.xyz/files/app.log
[ -x logsum.sh ] || { cat > logsum.sh <<'EOF'
#!/usr/bin/env bash
if [ -z "$1" ]; then echo "usage: $0 <logfile>" >&2; exit 1; fi
echo "file: $1"
echo "lines: $(wc -l < "$1" | tr -d ' ')"
EOF
chmod +x logsum.sh; }
cat > runner.py <<'EOF'
import subprocess
ok = subprocess.run(["./logsum.sh", "app.log"], capture_output=True, text=True)
print("exit", ok.returncode, "|", ok.stdout.splitlines()[1])
try:
subprocess.run(["./logsum.sh"], capture_output=True, text=True, check=True)
except subprocess.CalledProcessError as e:
print("exit", e.returncode, "|", e.stderr.strip())
EOF
python3 runner.py
Both lines are your own script, seen from inside a Python program.
A credential typed into a .py file is committed the moment you git add ., and Git does not
forget — deleting the line later leaves it in the history for anyone who clones. Treat every key
you have ever committed as burnt, and rotate it.
Configuration belongs in the environment instead. os.environ is a dict of the variables
your shell exported, so os.environ.get("WEATHER_API_KEY") returns None when it is not set
and your program can refuse rather than send an empty key. Set it for one command:
WEATHER_API_KEY=abc python3 app.py.
.env and .venv/ go in .gitignore — the same .gitignore you were graded on last week.
import os, sys
key = os.environ.get("WEATHER_API_KEY")
if key is None:
print("set WEATHER_API_KEY first", file=sys.stderr)
sys.exit(1)
pip install requests refuses on a modern machine, and it is right toNothing today needed an install: json, csv, subprocess, pathlib and os all ship with
Python. Session 2 needs requests, which does not. Try it and Homebrew Python and current
Ubuntu both stop you:
$ python3 -m pip install requests
error: externally-managed-environment
× This environment is externally managed
╰─> To install Python packages system-wide, try brew install
xyz, where xyz is the package you are trying to
install.
That is not a broken machine. The system Python belongs to the operating system, and packages
you shove into it break the tools that depend on it. A virtual environment is a private copy
of the interpreter and its packages, in a folder inside your project. python3 -m venv .venv
builds one; source .venv/bin/activate puts its bin at the front of your PATH.
requests installed.mkdir -p ~/wk2 && cd ~/wk2
python3 -m venv .venv
source .venv/bin/activate # macOS, Linux and WSL — all the same line
which python # ~/wk2/.venv/bin/python, not /usr/bin/python3
pip install requests
python -c "import requests; print(requests.__version__)"
printf '.venv/\n.env\n' > .gitignore
Expect 2.32 or newer — which version you get depends on your Python — and a pip notice about a newer pip. Both are normal.
Hands up if which python still points outside ~/wk2. deactivate gives you your old
python3 back.
try and except, including everything you did not see today..venv you just made exists, from the people who decide.Session 2 opens inside that .venv, against a live API that needs no key.