.venv is a folder in one directory. Walk away from it and it is gone.Last session ended with a working virtual environment and requests installed into it. That
environment is a folder called .venv inside one project directory — not a setting on
your machine, and not something your terminal remembers. A brand-new terminal starts outside
it. A cd anywhere else leaves it behind.
So the first thing you do, every session: walk back into the directory and switch it on.
Re-running python3 -m venv .venv on a venv that already exists is safe and keeps what you
installed. python -c runs the code you hand it instead of running a file.
mkdir -p ~/wk2 && cd ~/wk2
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -q requests
python -c "import requests; print(requests.__version__)"
Expect 2.32 or newer — which version you get depends on your Python, so the number on your screen may be lower than the one on the projector and still be right.
Hands up if that last line is not a version number at all.
By the end you will read from a real API, write to it, survive the request that fails, and know which header a secret belongs in. Most of today is REST, because REST is what you will be asked to write.
https://api.open-meteo.com/v1/forecast
?latitude=12.97
&longitude=77.59
¤t=temperature_2m
&timezone=Asia%2FKolkata
A resource is anything the server can hand you a copy of: a forecast, an order, one user,
a list of users. REST's rule is that the path names the resource and nothing else.
/v1/forecast is a forecast. /users is the collection. /users/42 is one item out of that
collection, and /users/42/orders is a collection that hangs off it. Nouns, plural for
collections, and an id where you mean one specific thing.
What you want done to it is not in the path. /getUser?id=42, /deleteUser, /api/doThing
are the anti-pattern: the URL has swallowed the verb, so every action needs a new URL and
nothing can tell which of them are safe.
The path is the noun. The verb comes next.
| verb | what it means | safe to send again? |
|---|---|---|
GET | give me a copy. Changes nothing on the server | yes — and it can be cached |
POST | here is a new one, create it | no — twice means two orders |
PUT | make the resource look exactly like this | yes — same body, same end state |
PATCH | change these fields, leave the rest alone | usually, but only if the API says so |
DELETE | remove it | yes — gone, then already gone |
An operation is idempotent when doing it five times leaves the server in exactly the state one call would have left it in. That is not a vocabulary question. Your request will time out halfway one day, you will not know whether the server got it, and idempotence is the whole of what decides whether you may send it again.
Every HTTP response starts with a three-digit number, and the first digit carries the
headline. 2xx — it worked. 3xx — it moved, look somewhere else. 4xx — your request was
wrong: malformed, unauthenticated, or asking for something that is not there. 5xx — the
server broke while handling a request that was perfectly fine.
That is the same question you asked in Week 1 when a command failed. 127 meant the shell
never found your program; 126 meant it found it and refused to run it. Two failures, and
the number told you which half of the world to go and look at.
You fix a 4xx in your code. You retry a 5xx, and if it keeps happening you tell them.
| code | what happened | what you do about it |
|---|---|---|
200 OK | it worked and the body is your data | read the body |
400 Bad Request | the server understood you and disagreed | fix the parameters you sent |
401 Unauthorized | no credential, or the wrong one | fix the key, not the code |
404 Not Found | nothing lives at that path | check the URL, and the id inside it |
429 Too Many Requests | you are going too fast | slow down, and read Retry-After |
500 Internal Server Error | their bug, not yours | retry it, then go and tell them |
There are dozens more and you do not need them. The pair that genuinely confuses people:
401 is "I do not know who you are"; 403 is "I know exactly who you are, and no."
params= builds the URL for you. You never glue one by hand.A query parameter narrows a resource without changing which resource it is: the same
/v1/forecast, at different coordinates. They ride after a ?, joined by &, written
name=value.
Building that string yourself is where the bugs live, because plenty of characters cannot
appear raw in a URL. Hand requests a plain dict and it does the encoding: the / inside
Asia/Kolkata becomes %2F, a space becomes %20. One parameter can also carry several
values, comma separated — current here asks for two — and that comma becomes %2C.
r = requests.get(
"https://api.open-meteo.com/v1/forecast",
params={"latitude": 12.97, "longitude": 77.59,
"current": "temperature_2m,relative_humidity_2m",
"timezone": "Asia/Kolkata"},
)
print(r.url)
https://api.open-meteo.com/v1/forecast?latitude=12.97&longitude=77.59¤t=temperature_2m%2Crelative_humidity_2m&timezone=Asia%2FKolkata
requests.get(...) hands back a Response object, and the data is only one of the things
you can get out of it.
r.status_code — the number from the last two slides, as an int. 200, 404.r.text — the body exactly as it arrived, as a string. Always there, even when the
response is an error page full of HTML.r.json() — the body parsed into ordinary Python dicts and lists. Only works if the body
really is JSON.r.json() is a method call, not an attribute: the brackets are not optional, and it
re-parses the body every time you call it. Call it once and keep the result — data = r.json() — then index into data like any other dict.
requests has no default timeout. If a server accepts your connection and then never
answers, your program does not crash and does not loop — it stops, on that line, until
somebody notices. A cron job that never returns still holds its lock; a web handler that never
returns still holds its worker.
timeout=10 says give up after ten seconds without progress and raise an exception instead —
something your code can respond to. Ten seconds is a sane start for a JSON API.
r = requests.get(url, params=p, timeout=10)
Put a timeout= on every request you ever write. One keyword, and it is the difference
between a failure and a hang.
mkdir -p ~/wk2 && cd ~/wk2
python3 -m venv .venv && source .venv/bin/activate && python -m pip install -q requests
cat > weather.py <<'EOF'
import requests
r = requests.get(
"https://api.open-meteo.com/v1/forecast",
params={"latitude": 12.97, "longitude": 77.59, "timezone": "Asia/Kolkata",
"current": "temperature_2m,relative_humidity_2m,wind_speed_10m"},
timeout=10,
)
print(r.status_code)
print(r.json()["current"])
EOF
python weather.py
200
{'time': '2026-08-10T17:30', 'interval': 900, 'temperature_2m': 22.8, 'relative_humidity_2m': 88, 'wind_speed_10m': 11.5}
Your numbers will not match mine — they change every run. Hands up if line one is not
200. Stretch: swap in Delhi, 28.61 and 77.21.
400 is a sentence, not just a number. Read it before you change anything.The status code tells you whose fault it is. The body tells you what specifically was
wrong — Open-Meteo answers a bad request with JSON carrying a reason, and any API worth
using does something equivalent. Reading it is faster than guessing.
r = requests.get(
"https://api.open-meteo.com/v1/forecast",
params={"latitude": 999, "longitude": 77.59, "current": "temperature_2m"},
timeout=10,
)
print(r.status_code)
print(r.json()["reason"])
400
Latitude must be in range of -90 to 90°. Given: 999.0.
Ask instead for a path that does not exist — /v1/nosuchthing — and the same server answers
404 with Not Found. Same machine, two different mistakes of yours, two different numbers.
raise_for_status() turns a bad number into a stack trace you cannot ignoreNothing about a 400 makes requests complain. You get a Response, your program carries on,
and it falls over four lines later on r.json()["current"] with a KeyError — pointing at
the line that read the data, not at the line that asked for it. That is a bad half hour.
r.raise_for_status() closes the gap. It does nothing at all on a 2xx, and raises
requests.exceptions.HTTPError on any 4xx or 5xx, right where the mistake was made.
r = requests.get("https://api.open-meteo.com/v1/nosuchthing", timeout=10)
r.raise_for_status()
requests.exceptions.HTTPError: 404 Client Error: Not Found for url: https://api.open-meteo.com/v1/nosuchthing
Call it on the line after every get.
The server answered, and the answer is bad. You have a Response, a number and a body.
raise_for_status() turns it into HTTPError. This is the 400 and the 404 you just saw.
Nothing answered at all. Wrong hostname, no wifi, DNS down, or your timeout fired. There
is no Response, no status code, no body — requests raises ConnectionError, Timeout or
TooManyRedirects instead.
All of them inherit from requests.exceptions.RequestException, so one except clause
catches every case. Inside it, type(e).__name__ gives the class name of the one you got —
"ConnectionError", "HTTPError" — as a plain string you can print.
Content-Type is the label that says how to read them.When you POST, the body leaves your machine as a plain run of bytes. The server has no way
to look at those bytes and know whether they are JSON, a form, or a photograph. So you tell
it, in a header, and the server believes the label rather than guessing.
Get the label wrong and you do not get a parsing error. You get 415 Unsupported Media
Type — "I am not even going to try." That is the server refusing to guess, which is the
correct behaviour and the reason requests sets the label for you when you use json=.
Content-Type: application/json {"drink":"latte"}
Content-Type: application/x-www-form-urlencoded drink=latte
Content-Type: multipart/form-data; boundary=----X (bytes, in parts)
| what goes on the wire | when you use it | nesting? | |
|---|---|---|---|
JSON application/json | {"drink":"latte","sugar":2} | almost every API you will write | yes — objects inside objects |
form application/x-www-form-urlencoded | drink=latte&sugar=2 | what an HTML <form> posts by default | no — flat keys and strings only |
multipart multipart/form-data | each field in its own labelled part, separated by a boundary string | files, or files mixed with fields | fields yes, structure no |
Two things fall out of that table. Form encoding has no types — sugar=2 arrives as the
string "2", exactly like a CSV. And it has no nesting, so the moment your data has a
list or an object inside it, form encoding cannot express it and JSON is the only one left.
JSON holds strings, numbers, booleans, lists and objects. It has no type for "raw bytes", so a photograph cannot go in a JSON body as itself. The workaround is base64 — re-spelling the bytes using 64 safe characters — and it costs you: every 3 bytes becomes 4 characters, so the upload grows by a third before it leaves your machine.
multipart/form-data exists to avoid that. It splits the body into labelled parts separated
by a boundary — a random string the client invents and announces in the header — and the
file part carries its bytes untouched, with its own filename and its own content type.
Content-Type: multipart/form-data; boundary=----ab17
------ab17
Content-Disposition: form-data; name="file"; filename="notes.txt"
Content-Type: text/plain
the quick brown fox
------ab17--
Content-Type is what you sent. Accept is what you want back.They are the two halves of the same conversation and beginners collapse them into one. The
request carries both: Content-Type describes the body you are sending, and Accept
describes the representation you would like in return. A GET has no body at all, so it
carries Accept and no Content-Type.
The same URL can hand back the same data in more than one shape. That is not a trick — it is what "representation" means in Representational State Transfer.
curl "https://api.studiotypo.xyz/orders?limit=3"
curl "https://api.studiotypo.xyz/orders?limit=3" -H "Accept: text/csv"
id,seat,drink,size,sugar,name,status
1,seat-01,flat white,medium,0,Aarav,pending
2,seat-02,cappuccino,large,2,Priya,brewing
cat > bodies.py <<'EOF'
import requests
url = "https://api.studiotypo.xyz/orders"
key = {"X-API-Key": "dataeko-2026"}
o = {"drink": "latte", "size": "small", "name": "Ravi", "sugar": 1}
print("json= ", requests.post(url, headers=key, json=o, timeout=10).status_code)
print("data= ", requests.post(url, headers=key, data=o, timeout=10).status_code)
print("files=", requests.post(url, headers=key, files={"f": ("a.txt", b"x")}, timeout=10).status_code)
EOF
python bodies.py
json= 201
data= 415
files= 415
Only the first worked, and you never typed a Content-Type header. requests set it three
different ways because you used three different keywords. Hands up if you did not get
201 415 415.
401 is "I do not know who you are". 403 is "I know exactly who you are, and no."These are not two flavours of the same rejection and the difference decides what you do next. 401 Unauthorized is about identity — you sent no credential, or a broken one. Fix it by sending a valid one. 403 Forbidden is about permission — your credential is fine, the server has read it, and this thing is still not yours to have.
That matters because a 401 is worth retrying with a fresh login and a 403 never is. Log in again as hard as you like; you will get 403 again. You need different rights, not a fresh token.
$ curl -i /admin/stats -> 401 "who are you?"
$ curl -i /admin/stats -H "X-API-Key: dataeko-2026" -> 403 "I know who you are, and no"
$ curl -i /admin/stats -H "X-API-Key: <the admin key>" -> 200 {"orders": 15, ...}
It is tempting to write ?api_key=abc123 because it is easy to test in a browser. Do not, and
know why: a URL is not private. It is written to the server's access log in plain text, it
sits in your shell history and the browser's history, it is forwarded in the Referer header
when the page loads anything from another site, and it is pasted into tickets and chat by
people trying to be helpful.
A header is written to none of those by default. Same secret, same request, wildly different blast radius.
GET /orders?api_key=abc123 <- now in six log files you do not control
GET /orders
Authorization: Bearer abc123 <- travels once, logged nowhere by default
Both are encrypted in transit by HTTPS. Encryption in transit is not the problem — where the string comes to rest afterwards is.
| scheme | what you send | who it identifies | where you meet it |
|---|---|---|---|
| API key | X-API-Key: abc123 — one long-lived string | the application, not a person | weather, maps, SMS, our lab API |
| Basic | Authorization: Basic <base64 user:pass> | a person, by password, every request | old internal tools, routers, .htpasswd |
| Bearer token | Authorization: Bearer <token> | whoever the token was minted for | almost every modern API |
| OAuth 2 | a flow that ends in a Bearer token | a person, at another company, with limits | "Sign in with Google", Stripe, GitHub |
Basic is base64, and base64 is not encryption. Anyone who sees the header can decode it in one command. It is only ever safe over HTTPS, and it sends the password on every single request, which is why it is fading.
Notice the bottom two rows are the same header. OAuth 2 is not an alternative to Bearer tokens — it is the process for obtaining one.
A JSON Web Token is the most common thing to find inside Authorization: Bearer. It is
not encrypted and it is not meant to be — it is signed. Split it on the dots and the
middle chunk decodes, with no key and no permission, into ordinary JSON.
eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9 header which algorithm
.eyJzdWIiOiJzZWF0LTA3Iiwic2NvcGUiOi... payload the claims
.l0ApzEJpTMCM6Ey8UsivvpHSjDzaxGDHDAB signature the proof
{"sub":"seat-07","name":"Chandni","scope":"orders:read orders:write","exp":1786500000}
The signature is the whole point. Change one character of the payload — give yourself
admin — and the signature no longer matches what the server computes, so the server rejects
it. You can read the claims. You cannot change them.
Imagine a service that prints your photos. It needs the photos in your Google account. The obvious approach is to give it your Google password — and that is a catastrophe: it can read your mail, it can change your password, it has access forever, and you cannot take it back without changing the password everywhere.
OAuth 2 replaces "give it your password" with "let Google give it a limited token". The printing service never sees your password. What it gets is narrower on three axes at once:
photos.readonly, and nothing else in your accountThat is the entire motivation. The flow on the next slide is just the plumbing that delivers it.
1 you click "Connect Google" on printshop.com
2 printshop REDIRECTS your browser to Google, saying:
who it is (client_id), what it wants (scope), where to come back (redirect_uri)
3 Google shows ITS OWN login page. You type your password INTO GOOGLE.
4 Google asks: "printshop wants to read your photos. Allow?"
5 Google redirects your browser back to printshop with a short-lived CODE in the URL
6 printshop's SERVER swaps that code with Google for an access token,
sending its client_secret to prove it is really printshop <- back channel
7 printshop calls the Photos API with Authorization: Bearer <token>
Two things do the security work. Your password only ever reaches Google — steps 3 and 4 happen on Google's domain, not printshop's. And the code in the URL is not the token: it is useless without the client secret, which never touches your browser.
The access token from that flow typically lives about an hour. That is not an inconvenience, it is the design: a token that leaks is a problem that expires by itself. But nobody wants to click "Connect Google" every hour, so the swap in step 6 hands back two things.
So your code does something ordinary: call the API, and if it answers 401 because the token expired, use the refresh token to get a new one and try the call again — once. That is the retry loop from Tuesday, with a specific trigger.
Scopes are checked at the API, not at login. A token minted for orders:read that tries to
write gets 403, not 401 — which is exactly the distinction from three slides ago.
count is what you got. total is what exists. They are not the same number.An API that returned fifty thousand rows in one response would fall over, so almost every list endpoint hands back a page and expects you to ask for the rest. The response tells you both numbers, and reading only the list is how you confidently report a wrong answer.
curl "https://api.studiotypo.xyz/book"
{"count": 50, "total": 100, "limit": 50, "offset": 0, "orders": [ ... ]}
You asked for the book and received half of it, with no error, no warning, and a 200.
The fix is a loop: keep asking with a bigger offset until you have total rows.
Two styles exist. limit and offset is what you see here — simple, and it can skip or
repeat rows if the data changes underneath you. Cursor paging hands you an opaque
next string to send back, which is slower to explain and correct under change.
429 is the server asking you to slow down, and Retry-After says how longEvery public API limits how often you may call it. Go over and you get 429 Too Many Requests — not an error in your code, and not something a retry-immediately loop will fix. Retrying harder is how a slow morning becomes an outage and how your key gets suspended.
The server usually tells you exactly what it wants. Retry-After is a number of seconds,
and reading it is the difference between a well-behaved client and a denial of service you
wrote by accident.
for i in $(seq 1 7); do curl -s -o /dev/null -w "%{http_code} " \
https://api.studiotypo.xyz/limited; done
200 200 200 200 200 429 429
HTTP/2 429
retry-after: 29
x-ratelimit-remaining: 0
Sleep for what it tells you, not for what you guess.
// gRPC: the contract is a FILE, not a document
service Orders {
rpc GetOrder (OrderRequest) returns (Order);
}
message OrderRequest { int32 id = 1; }
# GraphQL: the CLIENT writes the shape of the answer
{ order(id: 3) { drink size customer { name } } }
gRPC compiles that .proto into real client and server code in whatever language you use,
sends compressed binary over HTTP/2, and is fast. The price is a build step and the fact that
you cannot read it with curl. It lives between your own services, rarely on a public edge.
GraphQL gives you one endpoint and lets the caller ask for exactly the fields it wants, in one round trip instead of four. The price is that caching gets hard and one innocent query can be enormously expensive to answer.
| REST | gRPC | GraphQL | |
|---|---|---|---|
| on the wire | JSON text | binary, HTTP/2 | JSON text |
| the contract | a document, or OpenAPI | a .proto file that generates code | a schema the server publishes |
debug with curl? | yes | no | awkwardly |
| who picks the fields | the server | the server | the client |
| use it when | anything public, anything you want readable | internal, high volume, low latency | many clients wanting different shapes |
If you cannot name which of those pressures you have, you want REST. It is readable in a terminal, cacheable by anything, understood by every tool, and it is what the other two get compared against.
MDN · HTTP request methods — the reference for the five verbs, with the safe and idempotent columns spelled out properly.
MDN · HTTP authentication — the
Authorization header and the schemes on it, in one page.
requests · quickstart — read
the sections on data, json and files. Those three keywords are today's second block.
The OAuth 2 simplified guide — the authorization code flow, written for people rather than for implementers.
jwt.io introduction — what is in a token and why the signature matters. Read it, do not paste real tokens into it.
Assignment 2 is out now: github.com/studio-typo-hq/dataeko-week2-assignment — the README is the whole brief, and your seat number is in it.