Everything so far ran on hardware you can touch. From today it runs on hardware you will never see, in a building you will never enter, rented by the second.
You will not spend money learning this. LocalStack runs Amazon's API inside a Docker container
on your own machine, and the ordinary aws command talks to it. One flag — --endpoint-url — is
the entire difference between what you type today and what you would type against a real account.
The awsl() line defines a shell function, so every AWS command for the rest of the day is short.
docker run -d --name localstack -p 4566:4566 \
-e LOCALSTACK_AUTH_TOKEN=$LOCALSTACK_AUTH_TOKEN localstack/localstack
export AWS_ACCESS_KEY_ID=test
export AWS_SECRET_ACCESS_KEY=test
export AWS_DEFAULT_REGION=ap-south-1
awsl() { aws --endpoint-url=http://localhost:4566 "$@"; }
awsl sts get-caller-identity
Hands up when you see "Account": "000000000000".
By the end you will be able to read an architecture diagram and say what every box does, why it is in that subnet, and roughly what it costs to leave it running.
Before: you signed a purchase order, waited six weeks for delivery, racked it, cabled it, and depreciated it over three years. You sized it for your busiest hour of the year, so it idled at eight per cent for the other 8,759. Capacity was a budget meeting.
Now: one API call, and the machine exists in under two minutes. You pay by the second. Another call and it is gone, along with the bill.
That is the shift from capital expenditure to operating expenditure — from owning an asset to renting a service. And it cuts both ways, because the meter never stops. Nobody ever got a surprise invoice for a server they forgot about in a cupboard. People get one every month now.
A Region is a named place: ap-south-1 is Mumbai, us-east-1 is Northern Virginia,
eu-west-1 is Ireland. AWS currently runs 124 Availability Zones across 39 Regions.
An Availability Zone is one or more data centres inside that Region with its own power, cooling and network feed. AWS says zones in a Region are many kilometres apart but all within 100 km of each other — far enough that one flood does not take two, close enough that a round trip between them is a fraction of a millisecond.
Regions are isolated on purpose. Your Mumbai bucket does not exist in Ireland unless you copy it there, and that is a legal feature as much as a technical one.
awsl ec2 describe-availability-zones \
--query 'AvailabilityZones[].[ZoneName,ZoneId,State]' --output text
ap-south-1a aps1-az1 available
ap-south-1b aps1-az3 available
ap-south-1c aps1-az2 available
Look at the two columns. Why do the letters and the numbers not line up?
Things fail at five different sizes, and each needs a different answer:
DELETE is not rare at all.Multi-AZ is the answer to exactly one of those: the zone. Run in two or three zones and a building going dark becomes a capacity problem rather than an outage.
It is not a backup. A wrong delete replicates to all three copies instantly and perfectly.
who what verb on which thing
an identity -> s3:GetObject -> arn:aws:s3:::dataeko-demo/notes/day1.txt
The third part is an ARN — an Amazon Resource Name, the globally unique address of one thing.
arn:aws:s3:::dataeko-demo/notes/day1.txt reads as partition, service, region, account, resource.
S3 leaves region and account empty because bucket names are already unique on Earth.
An identity is a user (a person, or a script with keys) or a role (permissions something
borrows). IAM evaluates that sentence against every policy attached to it, and the starting
position is deny: nothing is permitted unless something explicitly permits it. An explicit
Deny anywhere beats every Allow anywhere, and there is no way to override one.
Three keys do all the work: Effect allows or denies, Action lists the verbs, and
Resource lists the ARNs. Everything else in an IAM policy is decoration on top of those.
Almost every tutorial writes "Action": "s3:*" with "Resource": "*", because it always works and
it is never the reason your example failed. That pair is how most breaches start — one leaked
credential then owns everything instead of one bucket.
awsl iam create-user --user-name intern-reader
cat > read-only.json <<'EOF'
{"Version": "2012-10-17", "Statement": [{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:ListBucket"],
"Resource": ["arn:aws:s3:::dataeko-demo",
"arn:aws:s3:::dataeko-demo/*"] }]}
EOF
awsl iam create-policy --policy-name ReadDataeko \
--policy-document file://read-only.json
awsl iam attach-user-policy --user-name intern-reader \
--policy-arn arn:aws:iam::000000000000:policy/ReadDataeko
That bucket does not exist yet. It was accepted anyway. Why?
An IAM user has an access key id and a secret. Two strings, valid from anywhere on Earth, until somebody thinks to revoke them. They end up in git history, in CI logs, in a Slack thread, and on a laptop that gets resold.
An IAM role has no credentials at all. It is a set of permissions plus a statement of who is
allowed to borrow them. Attach a role to an EC2 instance and the SDK quietly fetches credentials
from a link-local address, 169.254.169.254, gets ones that expire in about an hour, and refreshes
them without anyone writing code.
A leaked role credential is worthless by lunchtime. A leaked user key is worthless when somebody notices, and nobody ever notices quickly.
A bucket is a name that is unique across the entire planet, not just your account. Inside it there are only objects: a key, some bytes, and metadata. The key is one flat string.
notes/day1.txt is not a file in a directory. It is a single key that happens to contain a slash,
and the console draws a folder because humans like folders. There is no directory anywhere.
Two consequences that surprise everyone. You cannot append to an object, or change forty bytes in the middle — you replace the whole thing or you leave it alone. And you cannot rename; a rename is a copy followed by a delete, and it costs two requests and moves every byte.
awsl s3 mb s3://dataeko-demo
echo 'hello from week five' > note.txt
awsl s3 cp note.txt s3://dataeko-demo/notes/day1.txt
awsl s3 ls s3://dataeko-demo --recursive
awsl s3 ls s3://dataeko-demo/
2026-08-30 18:55:44 21 notes/day1.txt
PRE notes/
mb is make bucket. The two ls calls list the same single object two different ways —
PRE means prefix, not directory. The folder appeared because you asked without --recursive.
Then look at what S3 stored beside your twenty-one bytes:
awsl s3api head-object --bucket dataeko-demo --key notes/day1.txt
Read out the value of ServerSideEncryption.
| what it promises | S3 Standard | |
|---|---|---|
| Durability | the object has not been lost | 99.999999999% — eleven nines |
| Availability | your request succeeds right now | 99.99% over a year |
| how | written to a minimum of three Availability Zones | checksummed and repaired continuously |
Eleven nines means: store ten million objects and expect to lose one roughly every ten thousand years. Four nines of availability means about 53 minutes a year when a request may simply fail.
They are different promises and people conflate them constantly. Cheaper storage classes — Standard-IA, Glacier Instant, Glacier Deep Archive — keep the eleven nines and trade availability, retrieval time or retrieval cost. One Zone-IA is the exception: one AZ, so losing that zone loses the data.
Every large public S3 leak has the same shape: a bucket policy with "Principal": "*" and
"Action": "s3:GetObject". One paste, and every object is on the open internet — indexed by search
engines and by people who scan for exactly this, all day, forever.
So since 2023 every new bucket ships with Block Public Access on, all four switches true:
BlockPublicAcls true IgnorePublicAcls true
BlockPublicPolicy true RestrictPublicBuckets true
Check yours with awsl s3api get-public-access-block --bucket dataeko-demo.
The default is the seatbelt, and the only way to be exposed now is to deliberately unbuckle it. Which people still do, because a tutorial told them to.
Somewhere in Mumbai a physical server runs a hypervisor. Your instance is a portion of it — some virtual CPUs, some memory, a virtual network card, a virtual disk.
Four things define one, and you will name all four every time you launch anything:
running costs money per second. stopped costs nothing for compute — but the disk keeps
billing, because the disk is a separate thing that outlives the instance. terminate is the only
one that ends the bill.
| type | vCPU | memory | the trade |
|---|---|---|---|
t3.micro | 2 | 1 GiB | burstable — cheap, fine for anything mostly idle |
m5.large | 2 | 8 GiB | general purpose, 4 GiB per vCPU |
c5.xlarge | 4 | 8 GiB | compute optimised, 2 GiB per vCPU |
r5.large | 2 | 16 GiB | memory optimised, 8 GiB per vCPU |
m is general, c is compute, r is memory, t is burstable, g and p are GPU, i is fast
local storage. The digit is the generation — m5 is newer than m4 and usually cheaper for the
same work. After the dot is the size, and it doubles: large, xlarge, 2xlarge.
Read the middle two rows together. Both give you 8 GiB of memory. One gives two CPUs and one gives four, and they cost differently because you are buying a ratio, not a machine.
awsl ec2 describe-instance-types \
--instance-types t3.micro m5.large c5.xlarge r5.large \
--query 'InstanceTypes[].[InstanceType,VCpuInfo.DefaultVCpus,MemoryInfo.SizeInMiB]' \
--output table
Launch is four steps and nothing magical. The AMI is copied onto a fresh EBS root volume. The
machine powers on and the kernel boots. cloud-init reads your user data and runs it as root
before anything else starts. And if a role is attached, credentials appear at 169.254.169.254
without anyone writing code.
User data is how a machine configures itself with nobody logged in. That is the thread that leads to Thursday.
Now the honest part: LocalStack has no operating system behind that instance id. You get a real API, a real state machine, real addresses. You cannot SSH in. Nothing boots. Today EC2 is an API and a set of concepts — which is genuinely the part that transfers.
A VPC exists only for you, only inside one Region. You choose its address range up front as a
CIDR block. 10.0.0.0/16 means the first 16 bits are fixed, so you own 65,536 private addresses —
none of which is reachable from the internet by default.
A subnet is a slice of that range pinned to exactly one Availability Zone. 10.0.1.0/24 is 256
addresses in ap-south-1a.
AWS reserves five addresses in every subnet, so a /24 gives you 251 usable, not 256 — the
first four and the last one are taken for the network address, the router, DNS, future use, and
broadcast.
A subnet lives in one AZ. That is why "multi-AZ" always means at least two subnets.
VPC=$(awsl ec2 create-vpc --cidr-block 10.0.0.0/16 \
--query Vpc.VpcId --output text) && echo $VPC
SUB=$(awsl ec2 create-subnet --vpc-id $VPC --cidr-block 10.0.1.0/24 \
--availability-zone ap-south-1a --query Subnet.SubnetId --output text) && echo $SUB
awsl ec2 create-subnet --vpc-id $VPC --cidr-block 10.0.2.0/24 \
--availability-zone ap-south-1b \
--query 'Subnet.[SubnetId,CidrBlock,AvailableIpAddressCount]' --output text
The last number is 251, not 256. Write down your $VPC and $SUB — you need them for the rest
of the hour.
There is no checkbox called "public". Ask a brand-new VPC what it can route to and you get exactly
one line: 10.0.0.0/16 → local.
Traffic for your own network stays inside. Everything else has nowhere to go — so every subnet starts private, and it is private because no route says otherwise.
Add an internet gateway and one route for 0.0.0.0/0, and the same subnet is public. The
gateway is not a box you size or pay for; AWS runs it and attaching it is free.
RTB=$(awsl ec2 describe-route-tables --filters Name=vpc-id,Values=$VPC \
--query 'RouteTables[0].RouteTableId' --output text)
IGW=$(awsl ec2 create-internet-gateway \
--query InternetGateway.InternetGatewayId --output text)
awsl ec2 attach-internet-gateway --vpc-id $VPC --internet-gateway-id $IGW
awsl ec2 create-route --route-table-id $RTB \
--destination-cidr-block 0.0.0.0/0 --gateway-id $IGW
awsl ec2 describe-route-tables --route-table-ids $RTB \
--query 'RouteTables[0].Routes' --output table
Two rows now. Your subnet just became public.
Your database sits in a private subnet with no route to the gateway, and that is exactly right — nothing on the internet can reach it. But it still needs to download security patches.
A NAT gateway sits in a public subnet and rewrites addresses. An outbound connection from
10.0.2.7 leaves looking as though it came from the NAT's own public address, and the reply comes
back because the NAT remembered the conversation. Nothing outside can start one, because no
address out there maps to your private machine.
It is also the most expensive box people forget about: $0.045 an hour to exist, plus $0.045 per gigabyte through it. Roughly $33 a month before a single byte moves, and the textbook design puts one in every Availability Zone.
| security group | network ACL | |
|---|---|---|
| wraps | one network interface | a whole subnet |
| rules | allow only | allow and deny |
| memory | stateful — the reply to an allowed request is allowed | stateless — the reply needs its own rule |
| default | nothing in, everything out | everything in, everything out |
Stateful is the entire difference. Open 443 inbound on a security group and the reply leaves with no outbound rule, because the group remembers you asked. Do the same on a NACL and the reply is dropped — a NACL sees two unrelated packets and has an opinion about each.
Almost all real work happens in security groups.
SG=$(awsl ec2 create-security-group --group-name web --description web \
--vpc-id $VPC --query GroupId --output text)
awsl ec2 describe-security-groups --group-ids $SG \
--query 'SecurityGroups[0].{in:IpPermissions,out:IpPermissionsEgress[].IpProtocol}'
awsl ec2 authorize-security-group-ingress --group-id $SG \
--protocol tcp --port 443 --cidr 0.0.0.0/0 --query 'SecurityGroupRules[0].CidrIpv4'
Before you open 443, in is an empty list. Nothing can reach it at all.
Two addresses, and only one of them belongs to your machine.
The private IP comes out of the subnet range and is bound to the network interface. It does not change for the life of the instance.
The public IP is not on the machine at all — the OS never sees it. The internet gateway keeps a
mapping and rewrites packets in flight, which is why ifconfig inside an EC2 instance shows
10.0.1.4 and nothing else. It comes from a shared pool, and AWS takes yours back the moment the
instance stops.
A subnet only hands them out if you tell it to — that is what --map-public-ip-on-launch does.
awsl ec2 modify-subnet-attribute --subnet-id $SUB --map-public-ip-on-launch
I=$(awsl ec2 run-instances --image-id ami-760aaa0f --instance-type t3.micro \
--subnet-id $SUB --security-group-ids $SG \
--query 'Instances[0].InstanceId' --output text)
addr() { awsl ec2 describe-instances --instance-ids $I --output text \
--query 'Reservations[0].Instances[0].[State.Name,PrivateIpAddress,PublicIpAddress]'; }
addr
awsl ec2 stop-instances --instance-ids $I >/dev/null && sleep 5 && addr
awsl ec2 start-instances --instance-ids $I >/dev/null && sleep 8 && addr
Three lines of output. Read the middle column, then the right column.
An Elastic IP is a public IPv4 address allocated to your account. It is yours across stops,
starts, and replacing the instance entirely, until you explicitly give it back. Four verbs:
allocate-address takes one, associate-address points it at an instance, disassociate-address
detaches it, release-address returns it.
That detach-and-reattach is the point. When an instance dies you remap the same address to its replacement in one API call, and DNS never changes — which matters, because a DNS change takes minutes to hours to reach everyone while a remap takes a second.
The default quota is five per Region, because public IPv4 addresses are genuinely scarce. And AWS charges $0.005 an hour for every one you hold — about $3.65 a month — whether it is doing work or sitting idle.
awsl ec2 allocate-address --domain vpc --output table
ALLOC=$(awsl ec2 describe-addresses --query 'Addresses[-1].AllocationId' --output text)
awsl ec2 associate-address --instance-id $I --allocation-id $ALLOC
awsl ec2 allocate-address --domain vpc --query PublicIp --output text
awsl ec2 describe-addresses --query 'Addresses[].[PublicIp,InstanceId]' --output table
One of those two addresses is attached to nothing. Both cost exactly the same.
| what it costs | how you find it | |
|---|---|---|
| an idle Elastic IP | $0.005/hr — identical to a working one | describe-addresses, look for a blank instance column |
| a NAT gateway | $0.045/hr + $0.045/GB — about $33/month per AZ, idle or not | describe-nat-gateways |
| an unattached EBS volume | billed by provisioned GB "until you release the storage" | describe-volumes, status available |
| cross-AZ traffic | $0.01/GB in Mumbai, charged in both directions | it does not appear as a line item |
| data out to the internet | $0.1093/GB from Mumbai after the first 100 GB | the next slide |
One idea runs through all five. AWS bills for things that exist, not for things that are used. A stopped instance is free; its disk is not. A detached Elastic IP does no work and costs the same as one serving traffic. Detaching is not deleting, and nothing sends you a warning.
Data transferred into AWS from the internet: $0.00 per gigabyte. Always, everywhere.
Data transferred out to the internet from Mumbai: the first 100 GB a month is free across your whole account, then $0.1093 per gigabyte up to 10 TB. So a terabyte of downloads is about $109, ten terabytes is about a thousand dollars, and nothing warns you first.
The asymmetry is deliberate. Getting your data in costs nothing and getting it out costs real money, which is exactly the shape of a business that would prefer your data stayed.
The same instinct explains every surprise: check what exists, not what is running.
awsl ec2 describe-addresses --query 'Addresses[?!InstanceId].PublicIp' --output text
awsl ec2 describe-volumes --filters Name=status,Values=available \
--query 'Volumes[].[VolumeId,Size,VolumeType]' --output text
awsl ec2 describe-nat-gateways --query 'NatGateways[].[NatGatewayId,State]' --output text
Four commands, once a month, on any account you are responsible for.
You can read an architecture diagram and say what every box is for:
That is not AWS trivia. It is the vocabulary every infrastructure conversation is conducted in, and you now have it.
Thursday you stop typing all of this. You write it down in a file, and something else builds it.
AWS global infrastructure — the whitepaper chapter — Regions, Availability Zones and edge locations in about ten minutes, from AWS rather than a blog.
IAM security best practices — AWS's own argument for roles over long-lived keys, and the list you will be audited against.
What is Amazon VPC — subnets, route tables and gateways properly. The single most useful page for Thursday.
Elastic IP addresses — the quota, the charging rule, and the failover trick, in one short page.
Leave the container running, or start a fresh one on Thursday. Nothing you made today is needed.