On Tuesday you drove a cloud by hand — commands in, resources out. It worked. It worked exactly once, and nothing about it is written down anywhere a colleague could read, review or repeat.
Today is the other half, and it is the last session of the programme. Two ideas.
First, infrastructure as code: describing a whole system in a file, so a machine can build it, check it and rebuild it. Second, the shapes that real production systems take, and why almost all of them take the same one.
Very little typing today. This is the theory session, and theory is the part that survives you changing employer, cloud and job title.
plan: the only command you own that tells you what is
about to happen before it happens.Then we close five weeks.
Console work leaves nothing behind but the result. Three things a click cannot be:
The industry word for this is click-ops, and it is said with affection by nobody. It is how almost every cloud account starts, including the good ones.
Every system starts documented and ends drifted. Someone raises a timeout during a 2am incident. Someone opens a port to unblock a demo. Someone bumps an instance size and fully intends to change it back on Monday.
None of that is malice and all of it is invisible. The wiki still shows the system as designed. The account holds the system as patched.
You find out at the worst possible moment. The disaster-recovery rebuild produces something that does not work, and the difference is four undocumented changes made by three people, two of whom have left.
Drift is not a failure of discipline. It is the default state of anything humans can touch.
The idea is one sentence: the system is described in files, the files live in Git, and a tool makes reality match the files.
Everything you learned in the last four weeks arrives at once. The description is reviewed in a
pull request, because it is a diff. It has history, so git log answers who opened port 8080 and
when and why. It can be built again — a second environment is a second run, not a second afternoon
of clicking.
And it can be deleted honestly. Nobody ever deletes a hand-built test environment, because nobody is confident about what is inside it. That is where a chunk of every cloud bill goes.
There are two ways to tell a computer to do something.
Imperative is a recipe: create the bucket, then set the tags, then check whether it already exists, then handle the case where it half-exists. You own every step, including every step that can fail.
Declarative is a description: there is a bucket called reports, and it is tagged
Owner = dataeko. That is the whole instruction. Deciding whether to create it, change it, or do
nothing is the tool's job, not yours.
You already wrote declarative code, in Week 4, and nobody called it that. SELECT ... WHERE ... ORDER BY says what rows you want back. It never once says how to scan the table.
Terraform itself is a small engine that reads files, compares things and calls APIs. Everything
cloud-specific lives in a provider — a plugin, downloaded by terraform init, that knows what
an S3 bucket is and which API calls bring one into existence.
A resource block is one thing you want to exist. It takes two labels: the type, defined by the
provider, and a name you choose so you can refer to it from your own files.
resource "aws_s3_bucket" "reports" {
bucket = "dataeko-reports-v2"
tags = { Owner = "dataeko" }
}
There is no verb anywhere in that block. Nothing says create. It asserts that this bucket exists.
| what it does | how often | |
|---|---|---|
terraform init | downloads the providers this configuration needs | once per project, then when versions change |
terraform plan | works out the difference between your files and reality, and prints it | constantly |
terraform apply | does what the plan said it would do | after you have read the plan |
terraform destroy | deletes everything this configuration manages | rarely, and deliberately |
destroy being a first-class command is the point rather than a footnote. An environment you can
delete on purpose is an environment you can afford to create.
terraform plan is a dry run, and nothing else in your toolkit does thisEvery plan compares three things: your files, which say what you asked for; the state, which says what Terraform last built; and reality, which it refreshes by calling AWS before it decides anything. Then it prints the difference and changes nothing at all.
Four symbols carry the whole language:
+ create
~ update in-place
- destroy
-/+ destroy and then create replacement
And every plan ends with the same line of arithmetic:
Plan: 1 to add, 0 to change, 0 to destroy.
Read that line before every apply. It is the last moment at which the mistake is still free.
aws_s3_bucket.reports: Refreshing state... [id=dataeko-reports-v2]
Terraform used the selected providers to generate the following execution
plan. Resource actions are indicated with the following symbols:
~ update in-place
Terraform will perform the following actions:
# aws_s3_bucket.reports will be updated in-place
~ resource "aws_s3_bucket" "reports" {
id = "dataeko-reports-v2"
~ tags = {
~ "Owner" = "someone-in-the-console" -> "dataeko"
}
~ tags_all = {
~ "Owner" = "someone-in-the-console" -> "dataeko"
}
# (12 unchanged attributes hidden)
# (3 unchanged blocks hidden)
}
Plan: 0 to add, 1 to change, 0 to destroy.
Here are three one-line edits to a working configuration. For each one, decide what terraform plan
will say — ~ update in-place, or -/+ destroy and then create replacement.
Owner = "dataeko" to Owner = "platform"dataeko-reports-v2 to dataeko-reports-prodhunter2 to hunter3For any you answered "replace", say out loud what happens to whatever was inside it.
terraform.tfstate is a JSON file that maps the names in your configuration to the real things in
the cloud: aws_s3_bucket.reports is the bucket whose id is dataeko-reports-v2. It also records
what every attribute looked like the last time Terraform saw it.
Delete that file and Terraform does not lose your bucket. It loses the mapping. The next plan says
+ create, because as far as it now knows nothing exists — and the apply either fails on a name
clash or quietly builds you a duplicate.
It is a database, not a log. Never hand-edit it; the terraform state subcommands exist for
everything you would be tempted to open a text editor for.
On your laptop, state sits next to your configuration. On a team that is immediately wrong: your colleague's Terraform has no idea what yours built, so it proposes to build it all again.
So state moves into shared storage — an S3 bucket, or Terraform Cloud. That is called a backend and it is about five lines of configuration.
Sharing creates the second problem. Two engineers apply at once. Both read the state, both decide, both write, and the version that survives is whichever finished last. The other person's resources are now alive, costing money, and tracked by nothing.
So backends take a lock. The second person waits. Same instinct as a database transaction.
plan hides your password. The state file writes it down in plain text.Measured today. A parameter declared as an encrypted SecureString, applied, then a plan for
changing its value:
~ value = (sensitive value)
And the same secret, sitting in terraform.tfstate on disk:
"value": "hunter2-correct-horse"
Terraform redacts secrets on screen because screens get shared. It cannot redact them in state, because state is how it knows what the value currently is.
Two consequences, and both are non-negotiable. terraform.tfstate goes into .gitignore on the
first commit, not the fifth. And a shared state bucket is private, encrypted and versioned — it is
exactly as sensitive as the systems it describes.
aws_s3_bucket.reports: Refreshing state... [id=dataeko-reports-v2]
No changes. Your infrastructure matches the configuration.
Terraform has compared your real infrastructure against your configuration
and found no differences, so no changes are needed.
Apply complete! Resources: 0 added, 0 changed, 0 destroyed.
That is idempotency: doing the operation again leaves you in the same place, rather than stacking up a second copy.
Third appearance in five weeks. Week 2, PUT and DELETE are idempotent and POST is not, which
is why a double-clicked checkout charges you twice. Week 3, your script had to survive being run a
second time. Now the whole cloud.
CloudFormation is AWS's own: YAML or JSON, and no state file for you to look after, because AWS
keeps it server-side as a stack. Its dry run is called a change set. Deleting a stack deletes what
it made — the same honest delete you get from destroy. It only speaks AWS, and that is the whole
difference.
Terraform is more widely used mostly because one language covers AWS, Cloudflare, GitHub and several thousand other providers, and real companies always use more than one thing.
There are others — Pulumi, AWS CDK, Kubernetes manifests. Every single one has a description, a dry run and idempotency. Learn the idea. The tool is syntax.
users ──> CDN ──────> object storage images, CSS, uploads
│
└─────> load balancer ──┬──> app server ──┐
├──> app server ──┼──> managed database
└──> app server ──┘
Static files are served by something cheap, cached, and physically close to the user. Everything dynamic goes through a load balancer to a pool of identical, interchangeable app servers. All of them talk to one managed database.
That is it. Netflix and a two-person startup draw the same picture; the difference is how many boxes are inside each box.
Boring is the achievement. Every deviation from this shape should be a company solving a problem it can name out loud.
This is the most important idea in the second half of today.
The load balancer sends each request to whichever app server is free. Your next click is a different server. So anything a server remembers about you — your session, your basket, the file you just uploaded to its local disk, the cache it spent an hour warming — is invisible to the other three and disappears when that server is replaced.
The rule: app servers keep nothing that has to survive the next request. State goes to the database, to Redis, to object storage — somewhere every server can see.
Get this right and a server becomes disposable. You can kill one, add ten, replace all of them during lunch.
Up is a bigger instance — more CPU, more memory. Simple, needs no code changes, and it has two problems: a restart to resize, and a ceiling. There is a biggest machine, and its price is not linear.
Out is more instances behind the load balancer. No meaningful ceiling, and it only works at all if the app is stateless.
An auto-scaling group does out automatically: keep between 2 and 10 of these, add one when average CPU passes 60%, remove one when it drops. It also replaces any instance that fails its health check, without anybody being woken up — which is quietly the more valuable half.
Databases are the exception. They mostly scale up, because there is one writer.
The same database, two ways to run it.
On EC2 you own all of it: installation, the operating system, patching, backups, testing that the backups actually restore, replication, failover, and the 3am page when the disk fills up. Total control, total responsibility, and a bill that looks pleasingly small.
RDS is the same Postgres with AWS doing the operations — automated backups, point-in-time restore, patching windows, and a standby in another building that takes over when the primary dies. You give up root, some extensions, and some say over versions.
The invoice is not the cost. The cost is an engineer's weekend, and the risk that the one person who understood the replica has left.
| you manage | you get | it suits | |
|---|---|---|---|
| EC2 | the OS, patching, everything on the box | total control; anything that installs, installs | legacy software, long-lived processes, odd requirements |
| Containers · ECS/EKS | the image, and how many copies run | your Week 3 image, running anywhere, restarted when it dies | most web services |
| Lambda | one function and its dependencies | no servers at all, per-millisecond billing, scales to zero | event handlers, glue, spiky or infrequent work |
Down the table you hand over more and control less. There is no top of this table — there is a right row for the workload, and plenty of real systems use all three at once.
Walk the diagram box by box and ask one question: if this dies, what still works? One app server dying is a shrug. The database dying is an outage. The load balancer dying is an outage that nobody can even reach the error page for.
That question has a name — blast radius, how much breaks when one thing breaks. It is why production and staging live in separate AWS accounts, and why one service sharing a database with three others is a decision you have to justify.
Health checks are what turn a failure into a shrug. The load balancer asks every server "are you well?" every few seconds and stops sending traffic to the ones that stop answering.
A region is a city — eu-west-2 is London. An availability zone is a separate datacentre inside it, with its own power, its own cooling and its own network, kilometres from the others and joined to them by fast private links.
Run in one zone and one building's generator is your uptime. Run in two and a rack fire, a flooded basement or a botched power maintenance costs you capacity instead of costing you the service.
Multi-AZ is not an advanced technique. It is table stakes. It is a checkbox on RDS and a list of subnets on a load balancer, and its absence is a finding in any serious design review.
Multi-region is a completely different animal: far harder, far dearer, and the answer to a much rarer question.
Not five unrelated tool weeks. One idea, gaining a layer each week: work that only you can do, on your machine, once, became work anyone can repeat.
You can take a problem, write Python for it, put it in Git, have a colleague review it, have a machine test it, build it into an image, run that image somewhere real, store its data properly, and then see whether it is healthy. That is a complete loop. A lot of people with the word engineer in their title cannot do all of it.
The gap is depth and scale. You have run one of everything. Production is many of everything, with money, security, deadlines and other people's mistakes attached to it.
So pick one and go deeper than this course could. One language, or one database, or IaC. Breadth got you through the door; depth is what you get hired for.
What infrastructure as code is for — the official version of the first half of today, in about ten minutes.
The Terraform state file — why it exists, why it must be shared, and why it is dangerous.
The Twelve-Factor App — factor VI is where "app servers keep nothing" is written down properly. Old, short, and still correct.
AWS Well-Architected: reliability — the second half of today, in AWS's own words, with the numbers attached.