Tuesday you asked a database what happened. Today you watch software while it is still happening.
Everything runs in containers — the same skill as last week, three services instead of one. Start Docker first if it is not already running.
git clone https://github.com/studio-typo-hq/dataeko-week4-lab.git
cd dataeko-week4-lab
docker compose up -d --build
docker compose ps
Three containers: an app, Prometheus, and Grafana.
Hands up when docker compose ps shows all three as Up.
By the end you will have watched a service die and seen exactly what your monitoring did about it.
| what it is | good at | bad at | |
|---|---|---|---|
| Logs | a line of text per event | what exactly happened to this one request | "is the system healthy" — too much text |
| Metrics | a number over time | is it getting slower, is the error rate rising | "why did THIS request fail" — no detail |
| Traces | one request's path through many services | which service in the chain was slow | expensive, and needs every service to play along |
You already have logs — docker logs, and every print() you have written for three weeks.
Today is metrics, because metrics are what answers "is it healthy" for a whole system at once, and because they are cheap enough to keep forever. Logs are expensive to store; a number every fifteen seconds is not.
demo_http_requests_total{endpoint="/work", status="500"} 47
└────────── name ──────────┘└────────── labels ─────────┘ └value┘
The name says what is being measured. The labels slice it — the same metric, broken down by endpoint and status code, without needing a separate metric for each combination. The value is a number. Prometheus records the time for you.
Labels are the powerful part and the dangerous part. endpoint has four values, status has
three, so this one metric becomes twelve time series. Add a label with a thousand values — a user
id, a request id — and you get thousands, and you have broken your monitoring.
| what it does | example | the rule | |
|---|---|---|---|
| Counter | only ever goes up, or resets to zero on restart | requests served, errors, bytes sent | never read the raw number — read its rate |
| Gauge | goes up and down | queue depth, memory used, temperature | read it directly; it is a current value |
| Histogram | counts observations into buckets | request duration, response size | gives you averages and estimated percentiles |
The counter rule is the one that catches people. demo_http_requests_total is 2876 right now.
That number is meaningless on its own — it is "since this process started", and it resets to zero
when the container restarts.
What you actually want is rate(): how fast it is climbing. That is a real, comparable number.
curl -s localhost:8081/metrics | head -30
curl -s localhost:8081/metrics | grep '^# TYPE'
Then make some traffic and watch a number move:
curl -s localhost:8081/work > /dev/null
curl -s localhost:8081/metrics | grep 'demo_http_requests_total{'
Run those last two lines twice. Hands up when you have watched a counter go up.
This is the type people misunderstand, and misunderstanding it is how dashboards lie.
demo_http_request_duration_seconds_bucket{le="0.005"} 0
demo_http_request_duration_seconds_bucket{le="0.01"} 9
demo_http_request_duration_seconds_bucket{le="0.025"} 48
demo_http_request_duration_seconds_bucket{le="0.1"} 207
demo_http_request_duration_seconds_bucket{le="0.25"} 573
demo_http_request_duration_seconds_bucket{le="0.5"} 961
demo_http_request_duration_seconds_bucket{le="+Inf"} 961
le means less than or equal. Each bucket counts every request at or below its boundary,
so they climb: 573 requests took under 0.25s, 961 took under 0.5s.
The individual durations are gone. Only the counts were kept.
This is the design decision that explains everything else about it.
┌──────────────┐ every 5s ┌────────────┐
│ Prometheus │ ────GET────> │ your app │
│ │ /metrics │ :8080 │
└──────────────┘ <───text──── └────────────┘
Your app does not know Prometheus exists. It publishes a page and forgets about it. Prometheus holds a list of targets and scrapes each one on a schedule.
Three consequences, and all of them matter. Your app cannot flood the monitoring — the
scraper sets the pace. Prometheus knows when a target is missing, because it tried and failed.
And adding monitoring requires no change to a running app if it already exposes /metrics.
Open http://localhost:9090 in a browser. You get a query box.
Type each of these and press Execute:
up
demo_http_requests_total
sum(demo_http_requests_total)
rate(demo_http_requests_total[1m])
Then click the Graph tab on the last one.
up should show two rows, both with value 1. What do you think the second one is?
rate() is the function you will use more than all the others togetherA counter's raw value is meaningless. rate() turns it into something you can compare, alert on,
and put on a dashboard.
rate(demo_http_requests_total[1m])
Read it as: over the last 1 minute, how fast is this counter climbing, per second.
The [1m] is a range. It tells Prometheus to look back one minute and work out the slope, so
the answer is smoothed rather than jumpy. A longer range is smoother and slower to react; a
shorter one is twitchier and needs more scrapes to be meaningful.
Measured here right now: 5.56 requests per second. That number means something. 2876 did not.
In a second terminal, from the lab directory:
./traffic.sh
It hits the app in a loop. Roughly one in five requests to /work returns a 500 on purpose.
Back in Prometheus, run these and switch to the Graph tab:
sum(rate(demo_http_requests_total[1m]))
sum(rate(demo_http_requests_total{status="500"}[1m]))
/ sum(rate(demo_http_requests_total[1m]))
The second one is an error rate. What number do you get, and does it match one in five?
Here is a real measurement from this exact stack, taken at one instant.
histogram_quantile(0.95, ...) -> 0.4686 s what Prometheus reported
true 95th percentile -> 0.3867 s computed from the logs
true slowest request of all -> 0.4041 s the actual maximum
The estimate is larger than the slowest request that occurred. Not slightly wrong — reporting a duration that no request ever took.
It is not a bug. The durations were thrown away, remember. Prometheus knows 573 requests were under 0.25s and 961 were under 0.5s, so it assumes they are spread evenly across that gap and draws a straight line. Reality was clustered near the bottom of the bucket. The line overshoots.
Prometheus has a query box and an ugly graph. Grafana is the part people actually look at — and it holds no data of its own. Every panel is a PromQL query, run on a timer, drawn as a picture.
http://localhost:3001 — log in with admin / admin.
The dashboard is already there. It was not clicked into existence: it is a JSON file in the repository, loaded at startup, along with a file that tells Grafana where Prometheus lives.
That matters more than it sounds. A dashboard someone built by clicking exists on one server and dies with it. A dashboard in a file is reviewed in a pull request, versioned, and rebuilt anywhere in seconds — the same argument as the workflow file in Week 3.
At http://localhost:3001, open the dashboard. Four panels are already there.
Look at each one and work out which of the three metric types it is drawing — counter, gauge, or histogram.
Then add a fifth panel of your own:
sum by (endpoint) (rate(demo_http_requests_total[1m]))
sum by (endpoint) splits the total into one line per endpoint. Hands up when you see three
separate lines.
The best thing about pull-based monitoring is that a failure is a fact it records rather than a silence it cannot interpret.
docker compose stop app
Then, in Prometheus:
up
Within one scrape interval — five seconds here — up{job="demo-app"} becomes 0. Nobody
wrote that metric. Prometheus tried to scrape, failed, and wrote down the failure.
Now try demo_http_requests_total. It returns nothing at all. Not zero — empty. The app's own
metrics simply stop existing, because there is nothing publishing them.
| what you see | what it means | |
|---|---|---|
| Typo in a metric name | HTTP 200 and an empty result | not "no such metric" — indistinguishable from a dead service |
| Dashboard blank for 15s | nothing on any panel | rate() needs two scrapes; nothing is wrong |
Broken prometheus.yml | reload returns 500, old config keeps running | your change did not apply and nothing obviously failed |
| Time range too wide | flat line, or "No data" | the default is 6 hours; your stack is 5 minutes old |
Row one is the dangerous one. A query for a metric that does not exist returns success with no rows, which looks exactly like a query for a service that has stopped. There is no error to read.
You now have everything an alert needs. This is what one actually looks like.
- alert: AppDown
expr: up{job="demo-app"} == 0
for: 2m
annotations:
summary: "demo-app has not answered a scrape for two minutes"
The expr is a PromQL query you already know how to write. for: 2m is the part that matters —
the condition has to stay true for two minutes before anyone is woken up. Without it, one slow
scrape at 3am pages a human.
That single line is the difference between monitoring people trust and monitoring people mute. An alert that cries wolf gets ignored, and then the real one gets ignored too.
Four weeks ago a program was a file on your laptop that you ran by hand.
That is genuinely most of what "running software in production" means. What is left is where it runs, who is allowed to touch it, and what happens when the machine itself disappears.
That is next week.
Prometheus — overview and concepts — the official version of today. The data model page is the important one.
Metric types, explained properly — counter, gauge, histogram, summary. Ten minutes, and it settles the summary question.
Google SRE Book — Monitoring Distributed Systems — where the four golden signals come from. Free, and the best writing on this subject anywhere.
Grafana — get started — building dashboards deliberately rather than by clicking around.
Leave the stack running if you want to play. docker compose down stops it, and pkill -f traffic.sh stops the load generator.