Not "is it installed" — is the engine running. Docker is two pieces: a command you type, and a background service that does the work. Installing gives you the first. Starting the app gives you the second, and the command is useless without it.
docker --version
docker run hello-world
The second one downloads a tiny image and runs it. You want to see Hello from Docker!
Dockerfile is a recipe, and every line becomes a layer.By the end you will build an image, run it, prove to yourself why layer order matters, and have GitHub build and publish one for you.
It is not a small computer. There is no second operating system booting inside it. A container is an ordinary Linux process, started by the kernel you already have, which has been lied to about what exists.
The kernel gives it a private view of a few things — the filesystem it can see, the process list, the network, the hostname — so from the inside it looks like it has a machine to itself. From the outside it is one process among all your others.
That is why a container starts in milliseconds and a virtual machine takes a minute. There is nothing to boot. The kernel is already running; it just hands the process a different set of blinkers.
| image | container | |
|---|---|---|
| what it is | a read-only stack of files, built once | a running process using that stack |
| how many | one image | as many containers as you like, from that one image |
| lifespan | sits there until you delete it | starts, runs, exits |
| changes you make inside | impossible — it is read-only | kept in a thin writable layer, destroyed with the container |
| the analogy | a class | an object |
The last row is the one that catches people. Files you create inside a running container are real, and they are gone the moment that container is removed. If you want to keep something, it has to leave the container — which is what ports, volumes and registries are all for.
Every image is built on top of another image, and that chain ends at a small base someone else published. You will almost never write the bottom of the stack.
python:3.13-slim is an official image: a minimal Debian with Python 3.13 already installed and
working. Starting there means you inherit a working Python and spend your effort on your own
code.
The tag after the colon matters more than it looks:
python:3.13-slim Debian, trimmed down ~120MB before you add anything
python:3.13 the full Debian toolchain ~1GB
python:latest whatever is newest today a moving target — avoid
latest is not a version. It is a label that points somewhere different next month, which
is the opposite of what you want from a build.
Dockerfile is a recipe, and every instruction becomes a layerFive instructions cover almost everything you will write:
FROM python:3.13-slim # start from this image
WORKDIR /app # cd, and create it if missing
COPY requirements.txt . # copy from your machine into the image
RUN pip install --no-cache-dir -r requirements.txt # run a command AT BUILD TIME
COPY orders.py . # now the code
CMD ["python", "orders.py"] # the default command AT RUN TIME
The distinction that matters is the last two. RUN happens once, while the image is being
built. CMD happens every time you start a container — and it is the only line here that does
not run during the build at all.
git clone https://github.com/studio-typo-hq/dataeko-week3-lab.git
cd dataeko-week3-lab
docker build -t coffee:demo .
docker images coffee:demo
docker run --rm coffee:demo
REPOSITORY TAG IMAGE ID CREATED SIZE
coffee demo a51b19372708 1 second ago 142MB
sugar: 5
-t coffee:demo names it coffee and tags it demo. The . at the end is not punctuation
— it is the build context, the folder Docker is allowed to copy files from.
Hands up when you see sugar: 5.
Each instruction in the Dockerfile adds a layer containing only what changed — the files that instruction created or modified. The final image is those layers stacked, presented to the container as one filesystem.
COPY orders.py . layer 5 a few hundred bytes
RUN pip install -r ... layer 4 your dependencies
COPY requirements.txt . layer 3 one small file
WORKDIR /app layer 2 nothing much
FROM python:3.13-slim layer 1 ~120MB, and shared with every other image using it
Two things fall out of this. Layers are shared — ten images built from python:3.13-slim
store that base once, not ten times. And layers are the unit of transfer: pushing or pulling
moves only the layers the other end does not already have.
On a rebuild, Docker walks the Dockerfile from the top and reuses each layer as long as the instruction and its inputs are unchanged. The first thing that differs invalidates that layer and every single one after it — because each layer is built on the one above.
Measured on this exact project. Build once, then change only orders.py, and rebuild:
#6 [2/5] WORKDIR /app CACHED
#7 [3/5] COPY requirements.txt . CACHED
#8 [4/5] RUN pip install ... CACHED <- the slow one, skipped
#9 [5/5] COPY orders.py test_orders.py DONE 0.0s <- only this reran
pip install took 2.6 seconds on the first build and was skipped entirely on the
rebuild — because requirements.txt was copied on an earlier line and had not changed.
# 1. build again with nothing changed
docker build -t coffee:demo .
# 2. change only the code, and rebuild
echo "# a comment" >> orders.py
docker build -t coffee:demo .
# 3. now change the dependency list instead
echo "" >> requirements.txt
docker build -t coffee:demo .
Watch the CACHED markers each time.
Step 2 should keep pip install cached. Step 3 should not. Say out loud which line caused
the difference.
.dockerignore keeps your laptop out of the imageThe . at the end of docker build . hands Docker the whole folder as the build context, and
it is sent to the engine before the build starts. Without a .dockerignore, that includes your
.git history, your venv/, every __pycache__, and any .env file sitting there.
That is three problems at once: the build is slow because it is copying rubbish, the image is
fat because a COPY . . bakes the rubbish in, and a .env file becomes part of a shareable
artifact.
.git
venv/
.venv/
__pycache__/
*.pyc
.env
Same shape as .gitignore, same instinct, different destination — and worth writing the day you
start the project rather than the day you find your credentials in a published image.
Plenty of projects need tools to build that they do not need to run — a compiler, a bundler, build headers. If those are in the image, you are shipping them to production forever.
A Dockerfile can have more than one FROM. Each starts a fresh stage, and the last stage
is the only one that becomes the image. You copy the finished output forward and leave the rest
behind.
FROM python:3.13 AS build
COPY requirements.txt .
RUN pip install --target=/deps -r requirements.txt
FROM python:3.13-slim
COPY --from=build /deps /usr/local/lib/python3.13/site-packages
COPY orders.py .
CMD ["python", "orders.py"]
COPY --from=build reaches into the earlier stage. Everything else about that stage is
discarded — its compilers, its caches, its size.
docker run| flag | what it does | when |
|---|---|---|
--rm | delete the container when it exits | always, unless you need the corpse |
-it | interactive terminal | docker run -it --rm python:3.13-slim bash |
-p 8000:80 | your port 8000 reaches its port 80 | anything that serves |
-e KEY=value | set an environment variable inside | config and secrets |
-v $(pwd):/app | mount a real folder into the container | keeping files, live editing |
-d | detached — run in the background | servers you are not watching |
-p is the one to get right, and it is host:container, in that order. The container's port
is decided by whatever is running inside it; the host port is whatever is free on your machine.
docker run --rm -d -p 8000:80 --name web nginx
curl -s localhost:8000 | head -4
docker ps
docker logs web
docker stop web
Open http://localhost:8000 in a browser while it is running.
<!DOCTYPE html>
<html>
<head>
<title>Welcome to nginx!</title>
You just ran a web server you have not installed and cannot find on your disk. Nothing was added to your machine except an image.
Everything a container writes goes into its thin writable layer, and that layer dies with the container. For a stateless process that is a feature. For a database it is a catastrophe.
A volume mounts something from outside into the container's filesystem. Writes go through to the other side and outlive everything.
# a folder on your machine, visible inside as /app
docker run --rm -v "$(pwd):/app" -w /app python:3.13-slim python orders.py
# a named volume Docker manages for you — the right answer for databases
docker run -d -v pgdata:/var/lib/postgresql/data postgres:17
The first form is how people develop against a container without rebuilding for every edit — the code is not baked in, it is mounted live. The second is how data survives the container being replaced, which is the entire reason production databases can run in containers at all.
A container that misbehaves is not a black box. Three commands cover almost every investigation, and they are the container versions of things they already do.
docker logs web # everything the process printed
docker logs -f web # follow it live, like tail -f
docker exec -it web bash # get a shell INSIDE the running container
docker inspect web # the full configuration, as JSON
docker exec is the one that changes how this feels. You are standing inside the container's
view of the world — its files, its environment, its network — while it runs. ls, env, cat
a config file, and find out that the path you were sure existed does not.
A container is a process, so debugging it is debugging a process. Nothing new is required.
You have been pulling from one all session. python:3.13-slim and nginx came from Docker
Hub, the default registry, and the full name of that first image is really
docker.io/library/python:3.13-slim — Docker just fills in the boring parts.
Pushing works the same way in reverse. Tag an image with a registry address and push it, and it becomes something anyone with access can pull and run without having your code at all.
ghcr.io/studio-typo-hq/dataeko-week3-lab:latest
└─ registry ┘└──── owner ────┘└──── name ────┘└─ tag ┘
GitHub Container Registry comes with the repository, so there is nothing to sign up for. Docker Hub, GHCR, AWS ECR and Google Artifact Registry are all the same idea with different addresses.
Session 1 gave you a robot that runs tests. This is the same robot, with a build and a push bolted on the end — and it is the whole of "CI/CD" in one file.
permissions:
packages: write
steps:
- uses: actions/checkout@v5
- run: echo "${{ secrets.GITHUB_TOKEN }}" | docker login ghcr.io -u ${{ github.actor }} --password-stdin
- run: docker build -t ghcr.io/${{ github.repository }}:latest .
- run: docker push ghcr.io/${{ github.repository }}:latest
No credential was created for this. secrets.GITHUB_TOKEN already exists in every run — the
only thing needed was permissions: packages: write, because that token starts with read-only
access and you have to ask.
In your fork of the lab, create .github/workflows/publish.yml — the same
Add file → Create new file trick as last session:
name: publish
on: workflow_dispatch
permissions:
contents: read
packages: write
jobs:
push:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- run: echo "${{ secrets.GITHUB_TOKEN }}" | docker login ghcr.io -u ${{ github.actor }} --password-stdin
- run: docker build -t ghcr.io/${{ github.repository }}:latest .
- run: docker push ghcr.io/${{ github.repository }}:latest
Commit, open Actions → publish → Run workflow.
Hands up when you see Pushed. Then find your image on your GitHub profile, under Packages.
| Yes | your code, your dependencies, the runtime, a sensible default command |
| No | secrets, credentials, .env files, API keys — in any layer, ever |
| No | data you need to survive the container — that is a volume or a database |
| No | your .git folder, your venv, your editor config |
| Careful | anything that differs per environment — pass it with -e at run time instead |
The last row is the design rule underneath all of this: one image, many environments. The same image should run in test and in production, configured differently by what you pass it, not rebuilt differently.
If you have to rebuild the image to deploy it somewhere else, something is in the wrong place.
Three weeks ago a program was a file on your laptop. Now:
That last line is what "deployment" mostly means in practice. A platform pulls your image and runs it — with more machinery around scaling, health checks and rollback, but the artifact at the centre is the thing you built today.
The gap between this and a production system is real, and it is smaller than you think.
Docker · what is a container — the official version of minute 8, with the diagram this deck deliberately did not use.
Dockerfile reference — every instruction. Skim it; you have met five of them and there are not many more that matter.
Building best practices — layer ordering, small images, and why the cache behaves as it did at minute 38.
Publishing images with Actions — the grown-up version of the workflow you just wrote.
Assignment 3 is out now — it uses both halves of this week, so start with the session you found harder.