Before you start · 3 min readFrom Docker to Kubernetes: A Short Story#

A very short book for developers and DevOps engineers who have a vague idea of Docker and Kubernetes and would like a clear one.

The promise

Read this in an evening or two and you will know what every core piece of Docker and Kubernetes is for, why it exists, and the twenty percent of commands that do eighty percent of the work. You will also know what Helm, Istio, Prometheus, Grafana and Argo CD add on top, and when you would reach for each. And when you inherit a repository written in 2018, you will know what every old word in it became.

It reads as one story. Maya, a developer at a tiny startup, ships an app called Pantry. Every chapter starts with something breaking or somebody asking for something the current setup cannot do. The fix is the chapter's concept. By the end of the chapter, a new crack has appeared. Concepts you learn because you needed them are the ones you remember.

Every concept comes with three memory hooks: a place in one running metaphor (a kitchen, then a city), a diagram of how it actually works, and a one-sentence Remember it as line you can repeat out loud. Every chapter ends with the commands, a hands-on section you can paste into a terminal, the traps that bite silently, five recall questions (answers in the epilogue), and exactly three takeaways.

Who this is for

You have run docker run at least once and have heard of Pods. You may have copied a YAML file without knowing what half of it did. You want to be the person on the team who knows what is going on, without reading two thousand pages of documentation. If you are already an SRE with years of cluster time, this book will be too slow for you.

How to read it

  • Straight through the first time. The plot carries the concepts.
  • Skim on a re-read: each chapter has the same shape, so jump to The idea for the concept or The commands for the recipe.
  • The Epilogue is the reference: the full metaphor table, the diagram legend, and one cheat sheet of every command in the book.

Following along

The companion app lives in examples/pantry/. It is a fifty-line Python web service that greets you and counts visits in Redis. You need:

  • Docker Desktop (or Docker Engine on Linux) with at least 8 GB of memory allocated if you plan to install Act III's tools.
  • kubectl and kind for Act II. The kind cluster config is examples/kind-config.yaml.
  • helm, istioctl and argocd CLIs for Act III, each introduced in its chapter.

Commands are written for a POSIX shell (macOS, Linux, or WSL2 on Windows). If you get lost, examples/reset.sh <chapter> puts your machine back at that chapter's starting line.

Tested on: Docker 27.5 · Docker Compose 2.32 · kind 0.33 (Kubernetes 1.37) · kubectl 1.31 · Helm 4.3 · Istio 1.31 · kube-prometheus-stack chart 80.x · Argo CD 3.5. Newer versions will almost certainly work; the commands here are the stable ones.

Contents

Act I: One kitchen (Docker)

  1. It Works On My Machine
  2. Write the Recipe Down
  3. Memory and Neighbours
  4. Ship It
  5. The Whole Band
  6. The Night It Fell Over

Act II: The city (Kubernetes)

  1. Meet the City
  2. The Smallest Thing That Runs
  3. Someone Who Never Sleeps
  4. A Number That Never Changes
  5. Sticky Notes and the Locked Drawer
  6. Are You Alive?
  7. Memory That Survives, Again
  8. Districts, Key Cards and Rush Hour
  9. The Night It Fell Over, Again

Interludes

  • Old Maps: the history of containers, orchestrators and moving YAML, told through an inherited repository
  • The Landlord's Week: emptying a node safely and putting a ceiling on a district

Act III: The city grows up (the ecosystem)

  1. Flat-Pack Furniture (Helm)
  2. A Concierge at Every Door (Istio)
  3. The Control Room (Prometheus and Grafana)
  4. Nobody Touches the City by Hand (Argo CD)

Epilogue

  1. The Map: metaphors, timeline, old words, cheat card, cheat sheet, recall answers, where to go next

Chapter 1 · Act I · 4 min readIt Works On My Machine#

The problem

Pantry is three weeks old. It is a small web service that says hello and counts how many people have said hello back, and it stores that count in Redis. Maya wrote it, Maya runs it, and on Maya's laptop it has never once failed.

Today is demo day. Sam from sales has borrowed the app to show a customer, and Sam's laptop is a different story. Python is the wrong version. Redis is not installed, and the installer wants an admin password Sam does not have. Forty minutes before the meeting, Sam sends Maya a screenshot of a stack trace and the words "it doesn't work".

Maya knows exactly what is wrong. Pantry does not just need its own code; it needs a particular Python, a particular set of libraries, a running Redis, and an environment variable telling it where Redis lives. On her laptop all of that happens to be there. On Sam's laptop none of it is. The app is fine. The surroundings are the problem.

She has a vague memory of a colleague saying "just put it in a container". She has forty minutes to find out what that means.

The idea

A container is a process running on your machine with its own private view of the filesystem, network and process list. It is not a virtual machine. There is no second operating system booting inside it. It is your kernel, running your program, with walls around it. Those walls mean the program sees exactly the files and libraries it was packaged with, and nothing from the host that could surprise it.

Where do the files come from? From an image. An image is a read-only, versioned snapshot of a filesystem plus a little metadata: which program to start, which port it listens on, which environment variables to set. You download an image once and can start as many containers from it as you like. Each container gets its own thin writable layer on top of the image, so changes inside one container never leak into another.

Docker Hub (the cookbook library)docker rundocker rundocker run --rm

Image redis:7-alpine
(recipe card, read-only)

Container 'redis'
(dish #1, running)

Container 'redis-test'
(dish #2, running)

Container (throwaway dish)

Think of the image as a recipe card. It is written down, it has a version, and it does not change while you cook. A container is a dish cooked from that card. You can cook the same dish three times for three tables, throw one away when the table leaves, and the card is untouched. This is the single most important distinction in Docker, and almost every early confusion comes from mixing the two up. docker images lists recipe cards. docker ps lists dishes.

Remember it as: an image is the recipe card, a container is a dish cooked from it, and you can cook as many as you like.

A container has a short life cycle, and it helps to know the states:

docker create / runstartprocess ends or docker stopdocker startdocker rmdocker rm -f

Created

Running

Exited

Because a stopped container keeps its writable layer, you can docker start it again and find your files still there. Because that layer is deleted with docker rm, anything you want to keep must live somewhere else, which is the subject of chapter 3.

Maya's forty-minute fix is embarrassingly small. Instead of installing Redis on Sam's laptop, Sam runs one command that downloads the Redis image and starts a container from it. Redis is now listening on port 6379 of Sam's laptop, no admin password required, and it can be thrown away after the meeting with one more command.

The commands

# Start a container in the background (-d), give it a name, publish a port host:container
docker run -d --name redis -p 6379:6379 redis:7-alpine

# Run something interactive and throw the container away when it exits
docker run -it --rm python:3.12-slim python -c 'print("hello from inside")'

# What is running? (-a also shows stopped containers)
docker ps
docker ps -a

# Read a container's stdout/stderr; -f follows like tail -f
docker logs -f redis

# Run a command inside a running container (get a shell with: docker exec -it redis sh)
docker exec -it redis redis-cli ping

# Stop (SIGTERM, then SIGKILL after 10s) and remove
docker stop redis
docker rm redis

# See what images (recipe cards) you have locally
docker images

Try it

  1. Start Redis in a container:
   docker run -d --name redis -p 6379:6379 redis:7-alpine
   

Docker prints a long container ID. The first run also pulls the image; the second time it starts in under a second.

  1. Prove it is running and talk to it:
   docker ps
   docker exec -it redis redis-cli ping
   
   CONTAINER ID   IMAGE            COMMAND                  STATUS         PORTS                    NAMES
   3f9c2a1b8e44   redis:7-alpine   "docker-entrypoint.s…"   Up 5 seconds   0.0.0.0:6379->6379/tcp   redis
   PONG
   
  1. Run a throwaway container to feel how cheap they are:
   docker run -it --rm python:3.12-slim python -c 'import sys; print(sys.version)'
   

You get a Python 3.12 prompt's worth of output on a machine that may not have Python installed at all. When the process exits, --rm deletes the container.

  1. Stop and remove Redis, then check it is gone:
   docker stop redis && docker rm redis
   docker ps -a
   

The image is still there (docker images). Only the dish was thrown away, not the recipe.

Recall

  1. What is the difference between an image and a container, in one sentence?
  2. Which command lists images, and which lists running containers?
  3. What does --rm do, and when would you not want it?
  4. After docker stop and docker rm, is the image still on your machine?
  5. In the kitchen metaphor, what is a container?
Show answers
  1. An image is a read-only recipe; a container is a running instance cooked from it.
  2. docker images lists images; docker ps lists running containers.
  3. --rm deletes the container when it exits; not when you want its logs or files afterwards.
  4. Yes, the image stays until docker rmi.
  5. A dish cooked from the recipe card.

What Maya learned

  • An image is a read-only recipe; a container is a running dish cooked from it, and the two are listed by different commands.
  • A container is just a process with walls, so starting one costs about as much as starting the program itself.
  • run, ps, logs, exec, stop and rm cover almost everything you do with a single container.

…but

Sam's demo works. Redis came from a container, but Pantry itself still runs from a folder of source files with python app.py, and it still needs the right Python and the right libraries. The customer asks a reasonable question: "Can you give us this to try on our own machines?" Maya cannot hand them her laptop. She needs a recipe card for Pantry.

Chapter 2 · Act I · 4 min readWrite the Recipe Down#

The problem

The customer wants to try Pantry on their own machines. Maya's first instinct is a README: install Python 3.12, create a virtualenv, pip install these three libraries, run Redis somehow, export REDIS_HOST, then python app.py. She writes it, reads it back, and realises it is the same list of things that broke on Sam's laptop, just written down politely.

Yesterday Redis came out of a container and it was wonderful, because somebody at Redis had already written the recipe. Nobody has written the recipe for Pantry. Maya has to write it herself.

She has one more constraint. She changes app.py twenty times a day. If every change means a full ten-minute rebuild of "download Python, install libraries, copy code", she will never use this. The recipe has to be cheap to re-cook when only the last step changed.

The idea

A Dockerfile is the recipe. It is a plain text file with one instruction per line, read top to bottom, and docker build turns it into an image. Here is Pantry's entire recipe:

FROM python:3.12-slim                       # start from someone else's recipe
WORKDIR /app                                # cd, creating the directory if needed
COPY requirements.txt .                     # copy the dependency list in first...
RUN pip install --no-cache-dir -r requirements.txt   # ...and install it
COPY app.py .                               # now the code, which changes most often
ENV PANTRY_VERSION=v1                       # default environment
EXPOSE 8080                                 # documentation: the app listens here
CMD ["python", "app.py"]                    # what to run when a container starts

Every instruction produces a layer: a snapshot of what changed on disk after that step. An image is a stack of layers, and here is the trick that makes daily use bearable. Docker caches each layer, keyed on the instruction and the files it touched. When you rebuild, Docker walks down the recipe, reusing cached layers until it hits the first step whose inputs changed. Everything after that point is rebuilt; everything before is free.

FROM python:3.12-slim
(base layers, cached)

COPY requirements.txt
(cached unless the file changes)

RUN pip install
(cached with it: the slow step)

COPY app.py
(changes every edit: rebuilt)

CMD python app.py
(metadata only)

docker run → container
(thin writable layer on top)

That is why the recipe copies requirements.txt and installs libraries before copying app.py. Editing the code invalidates only the last two layers, which take a second. Editing the dependency list invalidates the install step, which is rare. Order the recipe so the things that change most often come last.

Remember it as: a Dockerfile is a recipe, each line is a step that becomes a cached layer, so put the steps that change most often at the bottom.

A few details matter in practice. FROM picks a base image; python:3.12-slim is a trimmed Debian with Python installed, and slim is usually the right choice. RUN executes at build time, CMD at container start time; confusing the two is the second most common Dockerfile mistake. A .dockerignore file next to the Dockerfile lists paths that should never be copied in, such as .git and __pycache__, which keeps builds fast and images clean.

Finally, tags. An image name like pantry:v1 is repository:tag. The tag is free text; v1, 2024-06-01 and latest are all just labels. latest is not magic and does not mean newest, it is simply the tag used when you do not give one. Tag every build with something meaningful and treat latest with suspicion.

Two instructions look alike and are not. CMD is the default command, replaced entirely if you pass one on docker run; ENTRYPOINT is the command that always runs, with CMD or your arguments appended to it. A tool image sets ENTRYPOINT ["redis-cli"] so that docker run redis-cli ping reads naturally; a service like Pantry only needs CMD. Both come in two forms. The exec form, CMD ["python", "app.py"], runs your program as process 1. The shell form, CMD python app.py, runs /bin/sh -c "python app.py", so process 1 is a shell and your program is its child. That matters at stop time: Docker sends SIGTERM to process 1, the shell does not pass it on, and after ten seconds Docker kills everything. Use the exec form, and make sure whatever runs as process 1 actually handles the signal.

The commands

# Build an image from the Dockerfile in the current directory and tag it
docker build -t pantry:v1 .

# Build with a different value for an ARG in the Dockerfile
docker build -t pantry:v2 --build-arg PANTRY_VERSION=v2 --build-arg PANTRY_ERROR_RATE=0.3 .

# Which recipe cards do I have?
docker images

# Give an existing image a second name (same layers, no copy)
docker tag pantry:v1 pantry:latest

# Show the layers an image is made of, biggest first tells you where the weight is
docker history pantry:v1

# Delete an image you no longer need (fails if a container still uses it)
docker rmi pantry:latest

Try it

  1. From examples/pantry, build the image and time it:
   cd examples/pantry
   docker build -t pantry:v1 .
   

The first build pulls the base image and installs the libraries. It ends with naming to docker.io/library/pantry:v1.

  1. Build again without changing anything:
   docker build -t pantry:v1 .
   

Every step says CACHED. Now open app.py, change the greeting, save, and build a third time. Only the COPY app.py step and the ones after it run.

  1. Look at the layers:
   docker history pantry:v1
   
   IMAGE          CREATED BY                                      SIZE
   a33fb03b2eb8   CMD ["python" "app.py"]                         0B
   <missing>      ENV PANTRY_VERSION=v1 PANTRY_ERROR_RATE=0       0B
   <missing>      COPY app.py . # buildkit                        12.3kB
   <missing>      RUN /bin/sh -c pip install --no-cache-dir -r…   18.7MB
   <missing>      COPY requirements.txt . # buildkit              12.3kB
   <missing>      WORKDIR /app                                    8.19kB
   ...
   

The libraries sit in one 19 MB cached layer. Your code is a few kilobytes on top, and the metadata lines cost nothing.

  1. Run it. Pantry needs Redis, so start that first on a shared network (chapter 3 explains this line):
   docker network create pantry-net
   docker run -d --name redis --network pantry-net redis:7-alpine
   docker run -d --name pantry --network pantry-net -p 8000:8080 pantry:v1
   curl localhost:8000/
   
   Hello from Pantry (v1). You are visitor #1. Served by 84ba985b15bd.
   

The recipe works. Note -p 8000:8080: port 8000 on your laptop is the serving hatch, port 8080 inside the container is the kitchen door. They do not have to match.

Gotchas

  • Process 1 in a container only receives SIGTERM if it installs a handler; shells do not, and neither does Flask's dev server, so docker stop pantry waits the full ten seconds. docker run --init puts a tiny init at PID 1 and the stop takes a quarter of a second; a real server such as gunicorn handles it itself.
  • COPY . . before pip install means any code edit re-runs the install. Copy the dependency list first, install, then copy the code.
  • COPY . . also ships everything in the directory, .git and data files included; a .dockerignore keeps stray files out of the context and the image.
  • An ARG declared before FROM is not visible after it unless you declare it again.

Then and now

You will see MAINTAINER lines and ENV key value without an equals sign in older Dockerfiles; both were replaced by LABEL and ENV key=value years ago and still build, with warnings. Old Maps tells the story.

Recall

  1. Why does the Dockerfile copy requirements.txt before app.py?
  2. What is the difference between RUN and CMD?
  3. Which form of CMD lets your program receive SIGTERM, and why?
  4. What does the latest tag actually mean?
  5. In the kitchen metaphor, what is a layer?
Show answers
  1. So a code change does not invalidate the cached dependency layer.
  2. RUN executes at build time; CMD is what runs when a container starts.
  3. The exec form, and only if process 1 handles the signal; a shell at PID 1 swallows it.
  4. The tag used when none is given; nothing about freshness.
  5. One step of the recipe, cached if unchanged.

What Maya learned

  • A Dockerfile is a recipe read top to bottom; docker build cooks it into a tagged image.
  • Each instruction is a cached layer, so put slow, rarely-changing steps first and your code last.
  • Tags are just labels; latest is the default label, not a promise of freshness.

…but

The customer's trial goes well, and one of their engineers asks where the visit counter is kept. Maya says "in Redis, in the container". The engineer asks what happens to the count when the container is removed. Maya opens a terminal to check and does not like the answer.

Chapter 3 · Act I · 3 min readMemory and Neighbours#

The problem

Maya runs the experiment. Redis in a container, ten visits to Pantry, counter reads ten. She removes the Redis container, starts a fresh one from the same image, and the counter reads one. The data lived in the container's thin writable layer, and docker rm took the layer with it.

That is only half of the day's trouble. Pantry finds Redis by the hostname redis. On her laptop that works because both containers are on the network she created yesterday, but a colleague who copied the commands without the network line gets redis unreachable. And the customer's engineer wants to know why Pantry is on port 8000 on her machine but 8080 in the logs.

Two separate things have gone wrong: containers forget, and containers cannot see each other by default. Both are deliberate, and both have a proper fix.

The idea

Containers forget on purpose. The writable layer is meant to be disposable, so that a container can be replaced without ceremony. Anything that must outlive the container goes on a volume: a directory managed by Docker that lives outside any container and is mounted into one at a path. Remove the container, the volume stays. Start a new container with the same volume, the data is back.

There are two flavours. A named volume (-v pantry-data:/data) is created and stored by Docker; you rarely need to know where on disk. A bind mount (-v $(pwd):/app) mounts a directory from your laptop straight into the container. Bind mounts are for development: edit code on the host, see it change inside the container immediately. Named volumes are for data.

container pantry (dev)container redis (dish #2, later)container redis (dish #1)Your laptop

Named volume
pantry-data
(the pantry shelf)

./src
(bind mount)

/data

/data

/app

Remember it as: a volume is the pantry shelf; dishes get thrown away, the shelf stays.

Containers cannot see each other by default, and that is also on purpose. Each container gets its own network stack. On the default network, containers can only reach each other by IP address, which changes. On a user-defined network (docker network create pantry-net), Docker runs a small DNS server and every container can reach every other by its --name. That is why REDIS_HOST=redis works: redis is a hostname on pantry-net.

Ports are a separate idea. Inside its network, a container's port is just open; Pantry listens on 8080 and Redis on 6379, and other containers on the same network can connect to those directly. Nothing outside the network can, until you publish a port with -p host:container. Publishing punches a hole from your laptop's port 8000 through to the container's 8080.

network pantry-net (the kitchen corridor)published port8000 → 8080DNS: 'redis':6379

curl localhost:8000
(the street)

pantry
:8080

redis
:6379

Remember it as: a network is the corridor between kitchens, and a published port is a serving hatch to the street.

Redis is unpublished in this picture, and that is right: only Pantry needs it. Publish as little as you can.

The commands

# Volumes: create, list, inspect where it lives, remove
docker volume create pantry-data
docker volume ls
docker volume rm pantry-data

# Mount a named volume (data) or a host directory (dev) into a container
docker run -d --name redis -v pantry-data:/data redis:7-alpine redis-server --appendonly yes
docker run --rm -it -v "$(pwd)":/app -w /app python:3.12-slim python app.py

# Networks: create one so containers can find each other by name
docker network create pantry-net
docker network ls
docker run -d --name redis --network pantry-net redis:7-alpine

# Publish a container port on the host: -p host:container
docker run -d --name pantry --network pantry-net -p 8000:8080 pantry:v1

# Everything Docker knows about a container (IP, mounts, env, ports) as JSON
docker inspect pantry

# Live CPU and memory per container, and copying files in or out
docker stats --no-stream
docker cp pantry:/app/app.py ./app-from-container.py

Try it

  1. Clean up yesterday's containers and start Redis on a named volume, with persistence turned on:
   docker rm -f pantry redis
   docker run -d --name redis --network pantry-net -v pantry-data:/data redis:7-alpine redis-server --appendonly yes
   docker run -d --name pantry --network pantry-net -p 8000:8080 pantry:v1
   curl localhost:8000/; curl localhost:8000/
   

You are visitor #1, then #2.

  1. Destroy Redis and bring it back:
   docker rm -f redis
   docker run -d --name redis --network pantry-net -v pantry-data:/data redis:7-alpine redis-server --appendonly yes
   curl localhost:8000/
   
   Hello from Pantry (v1). You are visitor #3. Served by 84ba985b15bd.
   

The shelf survived the dish.

  1. See the network doing its job:
   docker exec pantry python -c "import socket; print(socket.gethostbyname('redis'))"
   docker network inspect pantry-net --format '{{range .Containers}}{{.Name}} {{.IPv4Address}}{{"\n"}}{{end}}'
   

Pantry resolved redis to the IP the network assigned to the Redis container.

  1. Try to reach Redis from your laptop:
   nc -zv localhost 6379
   

It refuses. Redis has no published port, and that is the correct state for a database.

Gotchas

  • A bind mount path must be absolute: -v app.py:/x creates a named volume called app.py. Use "$(pwd)/app.py".
  • On the default bridge network containers cannot resolve each other by name; only a network you created has DNS.
  • localhost inside a container is the container. To reach your laptop, Docker Desktop offers host.docker.internal.
  • A container running as a non-root USER may not be able to write to a bind-mounted directory owned by you; check ownership before blaming the app.

Then and now

You will see --link redis:redis and "data-only containers" in older guides; user-defined networks (Docker 1.9, 2015) and named volumes replaced both, and --link still works while everyone tells you not to use it. Old Maps tells the story.

Recall

  1. Where does a named volume live, and what happens to it on docker rm?
  2. Why does REDIS_HOST=redis resolve inside pantry-net but not on the default network?
  3. What does -p 8000:8080 open, and which side is the container?
  4. Which command shows a container's IP, mounts and environment as JSON?
  5. In the kitchen metaphor, what is a published port?
Show answers
  1. In Docker-managed storage outside any container; it survives docker rm.
  2. Only user-defined networks have DNS for container names.
  3. Laptop port 8000 forwards to container port 8080; the right-hand side is the container.
  4. docker inspect.
  5. A serving hatch to the street.

What Maya learned

  • Containers forget by design; put data on a named volume, and it survives the container being removed.
  • Containers on a user-defined network can reach each other by container name, on their real ports.
  • -p host:container opens a port to the outside world; publish only what the outside world needs.

…but

Pantry now survives restarts and finds its database by name. The customer signs. Their first request is not a feature: "Can you give us the image? We do not want to build it ourselves." Maya has an image on her laptop and no way to hand it over except, apparently, email.

Chapter 4 · Act I · 3 min readShip It#

The problem

The customer wants the image, not the source. docker save can write an image to a tar file, and Maya's is 213 MB, which is on the wrong side of the email attachment limit and the wrong side of dignity.

Then she looks at the size again. Two hundred megabytes for fifty lines of Python? She opens docker history and finds most of it is the base image, plus the libraries. The customer's engineer, who has clearly done this before, asks two questions: "Where do you publish images?" and "Do you really need a compiler in production?"

Maya does not publish images anywhere, and she had not thought about compilers at all. Both questions turn out to be the same lesson: an image is a thing you ship, and shipping has rules.

The idea

Images live in a registry. Docker Hub is the public one, and the one you have been pulling from all along; redis:7-alpine is shorthand for docker.io/library/redis:7-alpine. Companies run private registries (GitHub Container Registry, AWS ECR, Google Artifact Registry, a self-hosted Harbor) that work exactly the same way. The verbs are push to upload and pull to download, and the image name tells Docker where to go.

pushpushpull

Maya's laptop
docker build
docker push

Registry
ghcr.io/pantry/pantry:v1
(the cookbook library)

Customer's server
docker pull
docker run

CI server
docker build
docker push

A full image name has up to four parts: registry/namespace/repository:tag. Leave out the registry and Docker assumes Docker Hub. Leave out the tag and it assumes latest. To push to your own account on Docker Hub, the image must be named yourname/pantry:v1; that is what docker tag is for. Every push and pull is layer by layer, and layers already present on the other side are skipped, which is another reason to keep the slow, stable layers at the top of the Dockerfile.

Remember it as: a registry is the cookbook library; you push a recipe card in, anyone with access pulls it out.

The compiler question is about what ends up inside the image. Some Python packages need a C compiler to install. The lazy fix is to build from the full python:3.12 base image, which includes compilers, and then ship that: 1.6 GB, most of it tools nobody will run in production. The right fix is a multi-stage build: use a fat image to do the installing, then copy only the results into a slim final image. The Dockerfile has two FROM lines, and only the last stage is shipped.

FROM python:3.12 AS builder                     # stage 1: the workshop, 1.6 GB, has compilers
WORKDIR /app
COPY requirements.txt .
RUN python -m venv /venv && /venv/bin/pip install --no-cache-dir -r requirements.txt

FROM python:3.12-slim                           # stage 2: the shop floor, small
WORKDIR /app
COPY --from=builder /venv /venv                 # take only the finished goods
COPY app.py .
ENV PATH=/venv/bin:$PATH
USER nobody                                     # do not run as root
CMD ["python", "app.py"]

Pantry's multi-stage image is 217 MB. The builder stage alone is 1.63 GB. Nothing from the builder except the /venv directory made it into the image the customer gets. Smaller images pull faster, start faster, and contain fewer things an attacker could use. USER nobody is the other habit worth forming today: nothing in the image needs root, so do not give it root.

The commands

# Log in to a registry (Docker Hub by default; pass a hostname for others)
docker login
docker login ghcr.io

# Name the image for the registry, then push
docker tag pantry:v1 yourname/pantry:v1
docker push yourname/pantry:v1

# On another machine: pull and run
docker pull yourname/pantry:v1
docker run -d -p 8000:8080 yourname/pantry:v1

# Build a multi-stage Dockerfile (-f picks a file); --target stops at a stage
docker build -t pantry:slim -f Dockerfile.multistage .
docker build -t pantry:builder --target builder -f Dockerfile.multistage .

# Reclaim disk: remove stopped containers, unused networks, dangling images and build cache
docker system prune

Try it

  1. Build both stages and compare:
   docker build -t pantry:slim -f Dockerfile.multistage .
   docker build -t pantry:builder --target builder -f Dockerfile.multistage .
   docker images pantry
   
   REPOSITORY   TAG       SIZE
   pantry       builder   1.63GB
   pantry       slim      217MB
   pantry       v1        213MB
   

The workshop is seven times bigger than what ships.

  1. Run the slim image in place of v1 and confirm nothing changed for the user:
   docker rm -f pantry
   docker run -d --name pantry --network pantry-net -p 8000:8080 pantry:slim
   curl localhost:8000/
   docker exec pantry whoami
   

The greeting is the same, and the process now runs as nobody.

  1. (Optional: needs a free Docker Hub account.) Publish it:
   docker login
   docker tag pantry:slim yourname/pantry:v1
   docker push yourname/pantry:v1
   

Then on any other machine, docker run -d -p 8000:8080 yourname/pantry:v1 is the whole installation guide. Everything in this book works without a registry account, because chapter 7 loads images straight into the local cluster.

  1. See how much disk Docker is using and tidy up:
   docker system df
   docker system prune
   

Gotchas

  • docker build produces an image for the machine it runs on: arm64 on an Apple Silicon Mac. Docker Desktop runs it anyway through emulation, so you notice only when an amd64 server says exec format error. Check with docker image inspect --format '{{.Architecture}}', and build with --platform linux/amd64 or multi-arch with buildx.
  • docker push pantry:v1 is refused: without a registry and namespace it means docker.io/library/pantry, which is not yours. Tag it yourname/pantry:v1 first.
  • Anonymous pulls from Docker Hub are rate-limited (since 2020); a CI server that pulls a lot without docker login starts failing at random.

Then and now

You will see older build output with Step 3/8 lines; that was the legacy builder, replaced by BuildKit as the default in Docker 23 (2023), whose output looks quite different. Old Maps tells the story.

Recall

  1. What are the four parts of a full image name, and which two have defaults?
  2. Why is the multi-stage image 217 MB when its builder stage is 1.6 GB?
  3. Which command pushes an image, and what must its name include?
  4. Why run as USER nobody?
  5. In the kitchen metaphor, what is a registry?
Show answers
  1. Registry, namespace, repository, tag; registry defaults to Docker Hub, tag to latest.
  2. Only the final stage ships; the builder's 1.6 GB of tools stays behind.
  3. docker push; the name must include your registry namespace.
  4. Nothing in the image needs root, and a compromised process should not have it.
  5. The cookbook library.

What Maya learned

  • An image is shipped by pushing it to a registry and pulling it somewhere else; the image name says which registry.
  • A multi-stage build uses a fat image to build and a slim image to ship, so compilers never reach production.
  • Run as a non-root user and prune regularly; small, clean images are faster and safer.

…but

The customer pulls the image and runs it. Twenty minutes later they call: "It says Redis is unreachable." Of course it does. Pantry is one container, but the app is two containers, a network and a volume, started in the right order with the right flags. Maya has those flags in her shell history. The customer does not.

Chapter 5 · Act I · 3 min readThe Whole Band#

The problem

Maya writes the customer a shell script: create the network, create the volume, start Redis with persistence, wait a moment, start Pantry with the network and the port and the environment variable. It works. Then she needs to change one flag, and realises she now maintains a shell script with six docker run invocations that nobody else understands, that cannot tell whether Redis is actually ready before Pantry starts, and that has no idea how to stop everything cleanly.

She counts the moving parts. Two containers. One network. One volume. One port. One environment variable. A startup order. This is not a container problem any more; it is a "several containers that belong together" problem, and she is solving it with the wrong tool.

The idea

Docker Compose describes a set of containers that run together as one application, in one YAML file, and gives you one command to start, stop and inspect the lot. Each container is a service. Compose creates a network for the project automatically, so services reach each other by service name, and it manages named volumes declared in the file.

Here is all of Pantry:

services:
  web:
    build: .                      # build from the Dockerfile here, or use image: yourname/pantry:v1
    ports:
      - "8000:8080"
    environment:
      REDIS_HOST: redis           # the service name is the hostname
    depends_on:
      redis:
        condition: service_healthy   # wait until redis's healthcheck passes
    healthcheck:
      test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8080/healthz')"]
      interval: 10s
  redis:
    image: redis:7-alpine
    command: ["redis-server", "--appendonly", "yes"]
    volumes:
      - pantry-data:/data
    healthcheck:
      test: ["CMD", "redis-cli", "ping"]
      interval: 5s

volumes:
  pantry-data:
docker compose up (project 'pantry')network pantry_defaultREDIS_HOST=redisdepends_on:service_healthy

service web
(pantry image)
:8080

service redis
:6379

volume pantry-data

localhost:8000

Two of those lines deserve attention. A healthcheck is a command Docker runs inside the container on a timer; if it fails repeatedly, the container is marked unhealthy. On its own that only changes what docker ps shows. Combined with depends_on and condition: service_healthy, it makes Compose wait for Redis to actually answer PONG before starting Pantry. "Started" and "ready" are different things, and this is the first time in the book that difference matters. It will not be the last.

Remember it as: Compose is the dinner menu; one file lists every dish, and one command cooks them all in the right order.

The file is committed next to the code. Anyone with the repository runs docker compose up and gets the same two containers, network and volume that Maya has, with none of the shell history. It is also the honest answer to "how do I run this locally?" for most services that need a database, a queue or a cache.

Compose is a single-machine tool. It knows nothing about a second host, and if a container dies it will restart it only if you asked with restart: unless-stopped. Keep that in mind.

The commands

# Start everything in the background, building images if needed
docker compose up -d
docker compose up -d --build

# What is running in this project, and is it healthy?
docker compose ps

# Logs from all services, or one; -f follows
docker compose logs -f
docker compose logs -f web

# Run a command inside a service's container
docker compose exec redis redis-cli GET visits

# Rebuild images without starting anything
docker compose build

# Print the effective file after defaults and variables are applied; old keys warn here
docker compose config

# Stop and remove containers and the network (volumes survive); -v removes volumes too
docker compose down
docker compose down -v

Try it

  1. Tear down the hand-made containers and start the composed version:
   docker rm -f pantry redis
   docker compose up -d --build
   docker compose ps
   
   NAME             STATUS
   pantry-redis-1   Up 9 seconds (healthy)
   pantry-web-1     Up 3 seconds (health: starting)
   

Redis was healthy before web was even created. A few seconds later web is healthy too.

  1. Use it, then stop and start it, and watch the counter:
   curl localhost:8000/; curl localhost:8000/
   docker compose down
   docker compose up -d
   curl localhost:8000/
   
   Hello from Pantry (v1). You are visitor #3. Served by 30ef7b1ee3b1.
   

down removed the containers and the network but kept the volume, so the counter survived. The hostname changed because it is a new container.

  1. Peek at the data directly:
   docker compose exec redis redis-cli GET visits
   
  1. Start again from nothing:
   docker compose down -v
   

The -v deletes the volume. The next up will greet visitor #1.

Gotchas

  • A plain depends_on: [redis] orders start, not readiness: web comes up while Redis is still booting and the first requests fail. Only condition: service_healthy waits.
  • docker compose down keeps named volumes on purpose; only down -v deletes them, and there is no confirmation.
  • The project name, and so the container and volume names, comes from the directory name. Two checkouts in differently named folders are two separate projects.
  • ports: - 8000:8080 unquoted is parsed as a number by YAML in some cases; always quote port mappings.

Then and now

You will see docker-compose with a hyphen, a version: "3" line at the top, and a file called docker-compose.yml; the Go plugin docker compose (2020) ignores version and prefers compose.yaml. Old Maps tells the story.

Recall

  1. What does a service name become on the project network?
  2. What is the difference between depends_on with and without condition: service_healthy?
  3. Which command removes containers but keeps the data, and which removes both?
  4. Where does Compose get the project name from?
  5. In the kitchen metaphor, what is the Compose file?
Show answers
  1. Its hostname on the project network.
  2. Without the condition it waits for the container to start; with it, for the healthcheck to pass.
  3. docker compose down keeps volumes; down -v deletes them.
  4. The directory name.
  5. The dinner menu.

What Maya learned

  • Compose describes a multi-container app in one file and runs it with one command; service names are hostnames.
  • Healthchecks plus depends_on: condition: service_healthy make startup order mean "ready", not just "started".
  • down keeps volumes and down -v deletes them; know which one you are typing.

…but

The customer is happy. Pantry has real users now. Maya moves it from her laptop to a rented server, runs docker compose up -d there, and goes home. At 3 a.m. her phone rings.

Chapter 6 · Act I · 4 min readThe Night It Fell Over#

The problem

Pantry is down. Maya logs in to the server at 3:10 a.m. and finds the web container exited two hours ago with a Python traceback from a dependency bug that shows up once in a few thousand requests. Nothing restarted it, because nothing was told to. She adds restart: unless-stopped to the Compose file, brings it up, and goes back to bed feeling that she has fixed it.

At 9 a.m. the customer's launch email goes out and traffic triples. One Pantry container on one server cannot keep up; requests queue, then time out. Maya can run three copies of web on the same server, but they would fight over port 8000, and the server has only so much CPU anyway. She rents a second server. Now she has two Compose files on two machines, no way to send traffic to both, and no way to know when one of them dies.

At 11 a.m. she ships a fix for the traceback. docker compose up -d --build stops the old container and starts the new one. For eight seconds, Pantry is down again, on purpose, in the middle of the morning rush.

Nothing here is a bug in Pantry. Every one of these problems is a thing a single Docker host, however well configured, simply does not do.

The idea

Maya sits down with a coffee and writes the list of what she has and what she needs.

does not reachMany hosts: what 3 a.m. asked for

Self-healing:
notice a dead copy anywhere
and replace it

Scaling:
run N copies across machines
and spread traffic

Safe rollout:
replace copies one by one
with zero downtime

One address:
find the copies
wherever they are

One host: what Docker gives you

Packaged app
(image)

Run it
(container)

Run several together
(Compose)

Restart if it dies
(restart policy)

The pattern in the right-hand column is that every item is about many copies on many machines, and about something watching them. A restart policy is a bandage on one container on one host. What Maya needs is a system whose whole job is to be told "I want three copies of this image, reachable at one address, always" and to make that true, keep it true when servers die, and change it gracefully when she ships a new version.

That system is an orchestrator, and the one the industry settled on is Kubernetes. It takes a description of what you want (the desired state) and runs a loop, forever, comparing it to what exists (the actual state) and fixing the difference. Container died? The loop notices and starts another. Server died? The loop moves its containers elsewhere. New image? The loop replaces the copies one at a time and stops if the new ones fail to come up.

Remember it as: Docker cooks one meal in one kitchen; Kubernetes runs the whole city and keeps every kitchen open.

It is worth saying plainly that Kubernetes does not replace Docker. Everything from the first five chapters carries over untouched. The image Maya built is the image Kubernetes will run. The Dockerfile, the registry, the idea that containers forget and need volumes, the idea that "started" is not "ready": all of it is the foundation. Kubernetes adds a control plane, a vocabulary for describing what you want, and the loop.

The cost is real. Kubernetes has more concepts than Docker, and its YAML can be intimidating the first time. The next nine chapters introduce those concepts one at a time, each because Maya needs it, and each with a picture. By chapter 15 she will be able to diagnose a 3 a.m. outage in five minutes with four commands.

The commands

The one thing a single host can offer, so you know what it does and does not solve:

# Restart the container if it exits, unless you stopped it yourself (Compose: restart: unless-stopped)
docker run -d --restart unless-stopped --name pantry -p 8000:8080 pantry:v1

# Make the app crash from the inside (docker kill counts as "you stopped it" and is NOT restarted)
curl localhost:8000/crash

# Update means stop, then start: count the seconds of downtime
docker compose up -d --build

# What Docker saw: the container died and was started again
docker events --since 30s --until 1s --filter container=pantry --format '{{.Action}}'

Try it

  1. Start Pantry with a restart policy, then kill it and watch Docker bring it back:
   docker compose down -v
   docker network create pantry-net 2>/dev/null
   docker run -d --name redis --network pantry-net redis:7-alpine
   docker run -d --restart unless-stopped --name pantry --network pantry-net -p 8000:8080 pantry:v1
   sleep 2; curl localhost:8000/crash; sleep 3
   docker ps --filter name=pantry --format '{{.Names}} {{.Status}}'
   docker events --since 30s --until 1s --filter container=pantry --format '{{.Action}}'
   
   pantry Up 3 seconds
   create
   start
   die
   start
   

Pantry has a /crash endpoint that makes the process exit, which is the honest way to simulate a bug. (docker kill would not trigger a restart: Docker treats that as you stopping it on purpose.)

That handles the 3 a.m. crash. Now try the 9 a.m. problem.

  1. Try to run a second copy on the same port:
   docker run -d --name pantry2 --network pantry-net -p 8000:8080 pantry:v1
   
   docker: Error response from daemon: ... Bind for 0.0.0.0:8000 failed: port is already allocated.
   

One port, one container. There is no built-in way to say "spread across these" without adding a load balancer by hand.

  1. Measure the 11 a.m. problem:
   docker rm -f pantry pantry2
   docker run -d --restart unless-stopped --name pantry --network pantry-net -p 8000:8080 pantry:v1
   ( for i in $(seq 1 12); do curl -s -o /dev/null -m 1 -w '%{http_code}\n' localhost:8000/ || echo down; sleep 0.5; done ) &
   sleep 1; docker rm -f pantry && docker run -d --name pantry --network pantry-net -p 8000:8080 pantry:v2
   wait
   
   200
   200
   000
   down
   000
   down
   ...
   200
   500
   

Every down is half a second with no Pantry at all, about four seconds in total. That is downtime you chose. (The 500 at the end is v2 being deliberately flaky, which comes in useful later.)

  1. Clean up before the city:
   docker rm -f pantry redis
   docker network rm pantry-net
   

Gotchas

  • docker stop and docker kill both count as "you stopped it"; a container with --restart unless-stopped stays down afterwards. Only a crash from inside triggers the restart.
  • A restart policy restarts a container that exits; it does nothing for one that hangs while its process stays alive. That needs a health check, which Docker records but does not act on.
  • The eight seconds of downtime during compose up -d --build is not a bug to fix in Compose; it is the absence of a rolling update.

Then and now

You will see docker stack deploy and Compose files with a deploy: section in repositories from 2017 to 2019; that was Docker Swarm, the orchestrator that lost to Kubernetes, and Compose ignores deploy: today. Old Maps tells the story.

Recall

  1. Name the three things a single Docker host cannot do that the 3 a.m. story exposed.
  2. Why did docker kill not trigger the restart policy while /crash did?
  3. What is "desired state", and what does the orchestrator do with it?
  4. Does Kubernetes replace the images and Dockerfiles from chapters 1 to 5?
  5. In the metaphor, what does Docker cook and what does Kubernetes run?
Show answers
  1. Self-heal across machines, scale across machines, roll out without downtime.
  2. docker kill counts as you stopping it; /crash was the process dying on its own.
  3. What you want to be true; the orchestrator loops until reality matches it.
  4. No; Kubernetes runs the same images.
  5. Docker cooks one meal; Kubernetes runs the whole city.

What Maya learned

  • A restart policy heals one container on one host; it cannot scale, spread traffic, or roll out without downtime.
  • Every 3 a.m. problem is a "many copies on many machines, watched continuously" problem.
  • Kubernetes is that watcher: you declare the desired state and a loop keeps reality matching it, using the same images.

…but

Raj, who joined last month from a company that ran hundreds of services, listens to Maya's list and nods. "You need a cluster," he says. "Let me show you the city." He opens a terminal and types kind create cluster.

Chapter 7 · Act II · 4 min readMeet the City#

The problem

Raj's terminal says kind create cluster, and ninety seconds later Maya is looking at two Docker containers pretending to be two servers. "That's a cluster," Raj says. "Same thing the cloud gives you, just smaller and on your laptop."

Maya's questions come fast. If it is all containers, what is running Pantry? Nothing yet. Where do the requests go? Nowhere yet. What is the kubectl he keeps typing, and why does it talk to the cluster and not a server? And when he says "the control plane decides where things run", who exactly is deciding?

Raj does what good ops people do. Before running anything, he draws the map. Kubernetes has a lot of parts, but they fall into two groups, and once Maya can name them she will be able to read every kubectl output for the rest of the book.

The idea

A cluster is a set of machines, called nodes, managed as one. Some nodes run your containers; these are the worker nodes. One or more nodes run the management software; together they are the control plane. In a cloud, the provider runs the control plane for you and you only see workers. In kind, both are Docker containers on your laptop.

Worker node (a building)Worker node (a building)Control plane node (city hall)watcheswatcheswatcheswatches

API server
(the front desk:
every request goes here)

etcd
(the records room:
the only place state is stored)

Scheduler
(housing office:
picks a node for each Pod)

Controller manager
(inspectors on rounds:
make reality match the records)

kubelet
(caretaker)

Pods

kubelet
(caretaker)

Pods

kubectl / you

Think of the cluster as a city. Each node is a building. The control plane is city hall, and city hall has four departments worth knowing:

  • The API server is the front desk. Every request, from you, from kubectl, from every other component, goes through it. Nobody talks to anything else directly.
  • etcd is the records room. It is a small key-value database and it is the only place cluster state is stored. Everything else can be restarted and rebuilt from the records.
  • The scheduler is the housing office. When a new Pod needs a home, the scheduler picks a node with enough room and writes that decision back to the records.
  • The controller manager runs the inspectors. Each controller watches one kind of record ("there should be three copies of this") and takes action when reality drifts.

Every building has a kubelet, the caretaker. It watches the records for Pods assigned to its node, asks the container runtime to start them, reports their health, and restarts them if they die. The kubelet is the only part of Kubernetes that actually runs containers.

Remember it as: the cluster is a city, nodes are buildings, city hall keeps the records and everyone else watches the records and does their part.

The important word in the diagram is watches. Nothing in Kubernetes commands anything else. You write what you want into the records via the front desk, and each department notices and acts. That is what makes it self-healing: the inspectors never stop looking.

kubelet (node)ScheduleretcdAPI serverkubectlkubelet (node)ScheduleretcdAPI serverkubectlapply: "I want a Pod running pantry:v1"store the recordwatching: a Pod with no node!assign it to worker-1update the recordwatching: a Pod for my node!start the containerreport: Running

kubectl is the front-desk phone. It reads a config file (~/.kube/config) that can hold several clusters, and the currently selected one is called the context. Most of the "why is this not working" moments in a new team come from being pointed at the wrong context. The three commands for that are in the list below; learn them first.

Two kubectl verbs will carry you through the whole book. get lists things, and -o wide shows more columns while -o yaml shows everything. describe tells the story of one thing, including the events the departments logged about it, which is where the answers are when something is wrong. A third verb, explain, is the built-in manual for any field of any resource, and it means you never have to guess what a YAML key does.

The commands

# Create a local cluster from a config file (two nodes, ports mapped); delete it when done
kind create cluster --config examples/kind-config.yaml
kind delete cluster --name pantry

# Which cluster am I talking to? Switch if it is the wrong one
kubectl config current-context
kubectl config get-contexts
kubectl config use-context kind-pantry

# Where is the front desk, and is it answering?
kubectl cluster-info

# The buildings, with more columns
kubectl get nodes -o wide

# Everything running in every namespace (the control plane is just Pods too)
kubectl get pods -A

# The story of one thing, events included
kubectl describe node pantry-worker

# The built-in manual for any resource or field
kubectl explain pod.spec.containers

# Copy an image from your laptop into the cluster's nodes (no registry needed)
kind load docker-image pantry:v1 pantry:v2 --name pantry

Try it

  1. Create the cluster (this is the config the whole book uses):
   cd examples
   kind create cluster --config kind-config.yaml
   kubectl config current-context
   
   kind-pantry
   
  1. Look at the buildings:
   kubectl get nodes -o wide
   
   NAME                   STATUS   ROLES           AGE   VERSION   INTERNAL-IP   OS-IMAGE
   pantry-control-plane   Ready    control-plane   60s   v1.37.0   172.20.0.3    Debian GNU/Linux 13
   pantry-worker          Ready    <none>          45s   v1.37.0   172.20.0.2    Debian GNU/Linux 13
   

Run docker ps in another terminal: the two nodes are two containers.

  1. Look at city hall. It runs as Pods, in a namespace called kube-system:
   kubectl get pods -A
   

You will see kube-apiserver-..., etcd-..., kube-scheduler-..., kube-controller-manager-... and a kube-proxy and kindnet per node. Every box in the diagram is on that list.

  1. Read the manual for something you will write tomorrow:
   kubectl explain pod.spec.containers.image
   
  1. Load Pantry's images into the cluster so no registry is needed:
   cd pantry
   docker build -t pantry:v1 .
   docker build -t pantry:v2 --build-arg PANTRY_VERSION=v2 --build-arg PANTRY_ERROR_RATE=0.3 .
   kind load docker-image pantry:v1 pantry:v2 --name pantry
   

The images are now on both nodes' disks. On a real cluster this step is docker push to a registry the nodes can pull from.

Gotchas

  • An image tagged :latest, or with no tag, gets imagePullPolicy: Always, so the kubelet ignores the copy you loaded into kind and tries Docker Hub: ErrImagePull. Use a real tag like v1, or set imagePullPolicy: IfNotPresent.
  • kubectl talks to whichever context is current, not to the cluster you are thinking about. After creating or switching clusters, run kubectl config current-context before anything else.
  • kind load docker-image copies the image onto the nodes that exist now; a node added later, or a recreated cluster, needs it loaded again.

Then and now

You will see the control-plane node called master in older manifests and tutorials; the label went in Kubernetes 1.24 and the taint in 1.25. Old Maps tells the story, and why "Kubernetes dropped Docker" changed nothing.

Recall

  1. Name the four control-plane departments and one sentence on what each does.
  2. Which component actually starts containers on a node?
  3. Which command tells you which cluster kubectl is pointed at?
  4. What does "everything watches the API server" buy you?
  5. In the city metaphor, what is a node and what is etcd?
Show answers
  1. API server (front desk), etcd (records), scheduler (housing office), controller manager (inspectors).
  2. The kubelet on each node.
  3. kubectl config current-context.
  4. Self-healing: nothing commands anything; every part notices and acts, forever.
  5. A building, and the records room.

What Maya learned

  • A cluster is nodes plus a control plane; the control plane stores desired state in etcd and everything else watches it and acts.
  • The kubelet on each node is the only thing that runs containers; the scheduler only decides where.
  • kubectl config current-context, get, describe and explain are the first four commands to reach for in any cluster.

…but

The city is built and the images are inside it, but nothing is running. Raj hands Maya a YAML file with eleven lines in it and says, "That is the smallest thing that runs."

Chapter 8 · Act II · 4 min readThe Smallest Thing That Runs#

The problem

Maya wants to run one copy of Pantry in the cluster, the way she ran one container in chapter 1, and see it answer a request. That is all. Raj's YAML file is eleven lines, and every one of them looks like it means something.

She tries to skip the file. Surely there is a kubectl run pantry:v1 that just works? There is, sort of, and Raj lets her try it. It works. Then he asks her to run it again on Tuesday with a different environment variable, and on Wednesday to explain to a colleague exactly what she ran on Tuesday. Shell history is not a system. The file is the system.

The second surprise is the word Kubernetes uses. She asked for a container. Kubernetes gave her a Pod.

The idea

A Pod is the smallest thing Kubernetes runs. It is one or more containers that share a network address and can share storage, scheduled together onto the same node, living and dying together. Nearly every Pod you write will have exactly one container in it. The reason the concept exists is the "nearly": sometimes a program needs a helper right beside it, sharing its localhost, and Act III has a famous example.

Think of a Pod as an apartment. The containers inside share the address and the kitchen. Pantry lives alone in its apartment. Later a concierge will move in.

Node: pantry-worker (building)env REDIS_HOST=redisPod: pantry IP 10.244.1.3 (apartment)

container: web
image: pantry:v1
:8080

Service: redis
(chapter 10)

Remember it as: a Pod is an apartment: one address, usually one tenant, sometimes a helper in the spare room.

Here is the file, with the four parts every Kubernetes resource has:

apiVersion: v1          # which version of the API this kind of thing lives in
kind: Pod               # what it is
metadata:               # name, labels, namespace: how it is identified
  name: pantry
  labels:
    app: pantry
spec:                   # what you want: the desired state
  containers:
    - name: web
      image: pantry:v1
      ports:
        - containerPort: 8080
      env:
        - name: REDIS_HOST
          value: redis

apiVersion, kind, metadata, spec. Every Deployment, Service, ConfigMap and Ingress in this book has those same four top-level keys, and the only one that changes shape is spec. When Kubernetes accepts the file it adds a fifth section, status, which is what it observed. Desired state is yours; observed state is the cluster's; the gap between them is what the controllers work on.

The file goes in with kubectl apply -f. apply is idempotent: run it again with no changes and nothing happens, change one line and only that changes. It is the verb you will use most, and it is why the file is the system.

Nobody writes these files from a blank page. The honest way is to ask kubectl for a skeleton and edit it: kubectl create deployment pantry --image=pantry:v1 --dry-run=client -o yaml > deployment.yaml. --dry-run=client means "build the object but do not send it", and -o yaml prints it. The same trick works for create service, create configmap and most other kinds, and it is how experienced people produce YAML that is correct on the first try.

One more resident of the apartment is worth knowing now. An init container runs to completion before the main container starts, and the Pod stays Init:0/1 until it exits successfully. It is the helper who arrives before the tenant moves in: waiting for Redis to answer, fetching a config file, running a migration. k8s/pod-init.yaml waits for Redis with a two-line shell loop, which is the most common init container in the world.

A Pod has a life cycle, and its phase is the first thing to look at when something is off:

created, waiting for a node oran imagecontainers startedall containers exited 0 (Jobs)a container exited non-zeroand will not restartcontainer crashed, kubeletrestarts it (RESTARTS columngrows)could not be scheduled orimage never arrived

Pending

Running

Succeeded

Failed

Pending means "not yet placed or not yet started" and almost always means either no node has room or the image cannot be pulled. Running with a climbing RESTARTS count means the program keeps dying. Chapter 15 turns this diagram into a checklist.

One more thing you will do constantly: reach a Pod from your laptop without exposing it to anyone. kubectl port-forward opens a tunnel from a local port to a port in the Pod, through the API server. It is the Kubernetes equivalent of -p 8000:8080 and it is for you, not for users.

The commands

# Send a file to the front desk; apply is idempotent, run it as often as you like
kubectl apply -f k8s/pod.yaml

# List Pods; -o wide adds IP and node, -w watches for changes
kubectl get pods
kubectl get pods -o wide
kubectl get pods -w

# The story of one Pod, events at the bottom
kubectl describe pod pantry

# Everything the cluster knows about it, including status
kubectl get pod pantry -o yaml

# Logs (-f follows, -c picks a container in a multi-container Pod)
kubectl logs -f pantry

# Run a command inside; get a shell with -- sh
kubectl exec -it pantry -- sh

# Tunnel local port 8000 to the Pod's 8080 (Ctrl-C to stop)
kubectl port-forward pod/pantry 8000:8080

# A quick throwaway Pod for poking around the network, deleted on exit
kubectl run curl --rm -it --image=curlimages/curl:8.11.1 -- sh

# Generate a correct skeleton instead of typing YAML from memory
kubectl create deployment pantry --image=pantry:v1 --dry-run=client -o yaml

# Delete by file or by name
kubectl delete -f k8s/pod.yaml
kubectl delete pod pantry

Try it

  1. Pantry needs Redis. Apply Raj's Redis file first (it is a Deployment and a Service, both explained in the next two chapters; for now it is "a Redis called redis"), then the Pod:
   cd examples/pantry
   kubectl apply -f k8s/redis.yaml
   kubectl apply -f k8s/pod.yaml
   kubectl get pods -o wide
   
   NAME                     READY   STATUS    RESTARTS   AGE   IP           NODE
   pantry                   1/1     Running   0          3s    10.244.1.3   pantry-worker
   redis-5d9f8c7b6d-x2k9p   1/1     Running   0          8s    10.244.1.2   pantry-worker
   

The scheduler put both on the worker; the control-plane node is reserved for city hall.

  1. Talk to it through a tunnel:
   kubectl port-forward pod/pantry 8000:8080 &
   curl localhost:8000/
   kill %1
   
   Hello from Pantry (v1). You are visitor #1. Served by pantry.
   

The hostname is the Pod name. Pantry is running in the city.

  1. Read its story:
   kubectl describe pod pantry | tail -8
   

The Events section shows Scheduled, Pulled, Created, Started: one line per department that touched it.

  1. Delete it and watch what happens:
   kubectl delete pod pantry
   kubectl get pods
   

It is gone, and nothing brings it back. A bare Pod has no inspector watching it. That is the problem for the next chapter.

Then and now

You will see tutorials where kubectl run nginx --image=nginx creates a Deployment; it did until Kubernetes 1.18 (2020) removed the generators, and since then it creates only a Pod. Old Maps tells the story.

Recall

  1. What are the four top-level keys of every Kubernetes resource, and which one the cluster adds?
  2. Why does kubectl apply -f twice with no changes do nothing?
  3. How do you generate a Deployment manifest without writing it by hand?
  4. What does an init container guarantee about the main container?
  5. In the city metaphor, what is a Pod?
Show answers
  1. apiVersion, kind, metadata, spec; the cluster adds status.
  2. apply compares the file with the cluster and changes only differences.
  3. kubectl create deployment … --dry-run=client -o yaml.
  4. That it will not start until the init container has exited successfully.
  5. An apartment.

What Maya learned

  • A Pod is the unit Kubernetes runs: one or more containers sharing an address, almost always one.
  • Every resource is apiVersion, kind, metadata, spec; apply -f makes the file the source of truth.
  • get, describe, logs, exec and port-forward are the Pod toolkit; describe shows the events that explain problems.

…but

Maya deletes the Pod and it stays deleted. She kills the process inside it and the kubelet restarts the container, which is nice, but if the node dies the Pod dies with it and nobody notices. This is exactly the 3 a.m. problem, and a Pod on its own does not solve it. "Correct," says Raj. "Pods are not supposed to. You never run a Pod by itself. You ask someone to run it for you."

Chapter 9 · Act II · 4 min readSomeone Who Never Sleeps#

The problem

Maya's list from chapter 6 is on the whiteboard: self-healing, scaling, safe rollout. A single Pod delivers none of them. Delete it and it is gone. Ask for three and you write three files. Ship a new version and you delete the old Pod and create a new one, which is the eight seconds of downtime she was trying to escape.

"Who watches the Pods?" she asks. Raj writes one word on the whiteboard: Deployment. "You don't run Pods. You tell a Deployment how many Pods you want and what they should look like. It never sleeps."

The idea

A Deployment is a record that says: keep N copies of this Pod template running, and when the template changes, move from the old copies to the new ones gradually. A controller in city hall watches every Deployment and makes it true.

It does the job through a middle manager. The Deployment creates a ReplicaSet, whose only responsibility is "exactly N Pods matching this label selector must exist". If a Pod dies or a node vanishes, the ReplicaSet creates a replacement. When you change the Deployment's Pod template (a new image, say), the Deployment creates a second ReplicaSet for the new template and shifts the count from old to new, one Pod at a time. The old ReplicaSet is kept at zero so you can roll back.

selector:app=pantrypod-template-hash=647759c9cb

Deployment: pantry
replicas: 3
template: image pantry:v1

ReplicaSet pantry-647759c9cb
(template v1) desired 3

ReplicaSet pantry-5f67fcb5db
(template v2) desired 0
(kept for rollback)

Pod
app=pantry

Pod
app=pantry

Pod
app=pantry

The dotted lines are the important part. The ReplicaSet does not "own" Pods by holding a list of them. It selects them by labels, and labels are just key-value tags in a resource's metadata. Every Pod the Deployment creates gets app: pantry plus a hash of the template; the ReplicaSet counts Pods with those labels. Labels and selectors are how everything in Kubernetes finds everything else, and the next chapter uses the same trick for traffic.

Think of the Deployment as the building manager who guarantees N identical apartments are always occupied, and labels as the name tags on the doors the manager checks against a list.

Remember it as: a Deployment is the building manager who never sleeps: N tenants, always, and new tenants move in one at a time.

Here is the Deployment's spec, which wraps the Pod from the last chapter:

spec:
  replicas: 3
  selector:
    matchLabels:
      app: pantry            # must match the template's labels
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1            # may run one extra Pod during a rollout
      maxUnavailable: 0      # may never drop below 3 ready Pods
  template:                  # the Pod, minus apiVersion/kind/name
    metadata:
      labels:
        app: pantry
    spec:
      containers:
        - name: web
          image: pantry:v1

A rollout with those two numbers looks like this over time:

ReplicaSet v2ReplicaSet v1DeploymentReplicaSet v2ReplicaSet v1Deploymentcreate, desired 1 (surge)1 Pod readydesired 2desired 22 readydesired 1desired 33 readydesired 0 (kept for undo)

At no point are fewer than three Pods ready, so users never notice. If the new Pods fail to become ready, the rollout stops with the old ones still serving, and rollout undo shifts the counts back the other way. That is the 11 a.m. problem solved by arithmetic.

The commands

# Create or update the Deployment from the file (the normal way to change anything)
kubectl apply -f k8s/deployment.yaml

# The three layers, in one listing (-l filters by label)
kubectl get deploy,rs,pods -l app=pantry

# Change the number of copies now (the file still says 3; apply would put it back)
kubectl scale deploy/pantry --replicas=5

# Change the image without editing the file (handy in scripts and demos)
kubectl set image deploy/pantry web=pantry:v2

# Watch a rollout finish, see the history, go back one revision
kubectl rollout status deploy/pantry
kubectl rollout history deploy/pantry
kubectl rollout undo deploy/pantry

# Restart every Pod gracefully (new Pods first): picks up new ConfigMaps, clears bad state
kubectl rollout restart deploy/pantry

# Show labels, or add one
kubectl get pods --show-labels
kubectl label pod <name> tier=web

Try it

  1. Apply the Deployment and look at all three layers:
   kubectl apply -f k8s/deployment.yaml
   kubectl rollout status deploy/pantry
   kubectl get deploy,rs,pods -l app=pantry
   
   NAME                     READY   UP-TO-DATE   AVAILABLE   AGE
   deployment.apps/pantry   3/3     3            3           2s

   NAME                                DESIRED   CURRENT   READY   AGE
   replicaset.apps/pantry-647759c9cb   3         3         3       2s

   NAME                          READY   STATUS    RESTARTS   AGE
   pod/pantry-647759c9cb-64rhd   1/1     Running   0          2s
   pod/pantry-647759c9cb-7cfk6   1/1     Running   0          2s
   pod/pantry-647759c9cb-wtwht   1/1     Running   0          2s
   

Pod names are <deployment>-<template hash>-<random>. The hash is the label the ReplicaSet selects on.

  1. Kill one and watch the manager replace it:
   kubectl delete pod $(kubectl get pod -l app=pantry -o jsonpath='{.items[0].metadata.name}')
   kubectl get pods -l app=pantry
   

There are still three, and one of them is a few seconds old. The 3 a.m. problem is solved.

  1. Roll out v2 and watch it happen:
   kubectl set image deploy/pantry web=pantry:v2
   kubectl rollout status deploy/pantry
   
   Waiting for deployment "pantry" rollout to finish: 1 out of 3 new replicas have been updated...
   Waiting for deployment "pantry" rollout to finish: 2 out of 3 new replicas have been updated...
   Waiting for deployment "pantry" rollout to finish: 1 old replicas are pending termination...
   deployment "pantry" successfully rolled out
   
  1. v2 is flaky (it returns errors on purpose). Go back:
   kubectl rollout history deploy/pantry
   kubectl rollout undo deploy/pantry
   kubectl get pods -l app=pantry -o jsonpath='{.items[*].spec.containers[0].image}'
   
   pantry:v1 pantry:v1 pantry:v1
   

Both ReplicaSets still exist: kubectl get rs shows one at 3 and one at 0. The undo just moved the numbers.

Gotchas

  • A Deployment's selector cannot be changed after creation; kubectl apply fails with "field is immutable". Delete and recreate, or leave the selector alone and change labels elsewhere.
  • Without a readiness probe, a Pod counts as available the instant its container starts, so a rolling update can finish while the new version is still booting. Chapter 12 fixes this.
  • kubectl scale and set image change the cluster, not the file; the next apply -f puts the file's values back. Edit the file, or accept that you are experimenting.
  • Template labels must match the selector or the Deployment is rejected; a typo in one of them is the usual cause of "selector does not match template labels".

Then and now

You will see kind: ReplicationController and Deployments under apiVersion: extensions/v1beta1 in older manifests; the address stopped working in Kubernetes 1.16 (2019) and apps/v1 made the selector mandatory. Old Maps tells the story.

Recall

  1. What does a ReplicaSet do that a Deployment does not, and the other way round?
  2. How does a ReplicaSet know which Pods are "its" Pods?
  3. Which two numbers make a rollout zero-downtime, and what did we set them to?
  4. Which command goes back one revision?
  5. In the city metaphor, what is a Deployment?
Show answers
  1. A ReplicaSet keeps N Pods alive; a Deployment manages ReplicaSets to roll out changes.
  2. By label selector.
  3. maxSurge: 1, maxUnavailable: 0.
  4. kubectl rollout undo deploy/<name>.
  5. The building manager who never sleeps.

What Maya learned

  • Never run bare Pods; a Deployment keeps N copies alive and replaces any that die, on any node.
  • A Deployment manages ReplicaSets, which select Pods by label; changing the template creates a new ReplicaSet and shifts Pods across gradually.
  • scale, set image, rollout status, rollout undo and rollout restart are the day-to-day Deployment verbs.

…but

Three copies of Pantry are running, and each has its own Pod IP. Which one should a user talk to? Maya port-forwards to one of them, deletes that Pod to prove the manager works, and her tunnel dies with it. The new Pod has a new IP. The manager keeps the apartments full; it does not give anyone a phone number.

Chapter 10 · Act II · 4 min readA Number That Never Changes#

The problem

Three Pantry Pods, three IP addresses, and every one of them temporary. Maya's port-forward died when its Pod did. Redis has the same problem from the other side: Pantry finds it by the name redis, and in chapter 8 that worked, but she never asked why. If the Redis Pod is replaced, its IP changes, and something must be keeping the name pointed at the right place.

Then the customer asks for a URL. Not a tunnel, not a port on a node, a proper https://pantry.example.com that their users can type. Nothing Maya has so far answers to a hostname.

Two questions, then: how do Pods find each other when everything moves, and how does the outside world find the Pods at all?

The idea

A Service is a stable name and IP for a group of Pods, chosen by label selector. Traffic sent to the Service is spread across whichever Pods currently match. Pods come and go; the Service does not. Every Service also gets a DNS name inside the cluster, <name>.<namespace>.svc.cluster.local, and from the same namespace just <name> works. That is why REDIS_HOST=redis has been working: redis is a Service.

http://pantry/selector: app=pantry

client Pod
(or port-forward)

Service: pantry
ClusterIP 10.96.80.31
DNS: pantry.default.svc
port 80 → 8080

Pod 10.244.1.4
app=pantry

Pod 10.244.1.5
app=pantry

Pod 10.244.1.9 (new)
app=pantry

Think of it as a phone number for a group of apartments. Tenants move in and out; the number stays on the sign. The selector is the list of doors the number rings.

Remember it as: a Service is a phone number for a group of Pods that never changes when the Pods do.

Services come in three types, and the difference is only who can dial the number:

Type Reachable from Use it for
ClusterIP (default) Inside the cluster only Almost everything: app to database, app to app
NodePort Any node's IP on a fixed high port (30000–32767) Quick tests, kind on a laptop, the hook a load balancer attaches to
LoadBalancer A cloud load balancer with a public IP Exposing one Service directly on a cloud cluster

NodePort and LoadBalancer build on ClusterIP; each type is the one below plus a door. On a laptop, kind maps a NodePort to localhost for you. On a cloud, LoadBalancer costs money per Service, which is why nobody exposes twenty Services that way.

They do the job for "find each other". They do not answer to a hostname. For that there is Ingress: a set of rules saying "requests for host X and path Y go to Service Z". An Ingress is only a record; something has to read it and act, and that something is an Ingress controller, a reverse proxy (usually NGINX) running in the cluster, exposed once, routing to every Service by hostname. One front gate, many signposts.

read by

user
http://pantry.localtest.me

Ingress controller
(nginx Pod, exposed on :80)

Ingress rules
host pantry.localtest.me / → pantry:80
host shop.localtest.me / → shop:80

Service pantry

Service shop

Remember it as: Ingress is the city's front gate with a signpost: one entrance, routed by hostname and path.

The metaphor stretches in one place, so here it is plainly: an Ingress can also terminate TLS (the certificate lives in a Secret named in the Ingress) and rewrite paths. Those are proxy features, not gates, and the controller's documentation covers them.

The commands

# Create the Service; get shows type, cluster IP and ports
kubectl apply -f k8s/service.yaml
kubectl get svc

# Which Pod IPs is the Service currently pointing at? (empty = your selector is wrong)
kubectl get endpoints pantry

# Tunnel to the Service instead of a Pod: survives Pod restarts
kubectl port-forward svc/pantry 8000:80

# Test DNS and routing from inside the cluster with a throwaway Pod
kubectl run curl --rm -it --image=curlimages/curl:8.11.1 -- sh -c 'curl -s http://pantry/; nslookup redis'

# Install an Ingress controller (kind flavour: listens on the node's port 80)
kubectl apply -f https://raw.githubusercontent.com/kubernetes/ingress-nginx/main/deploy/static/provider/kind/deploy.yaml

# Create the Ingress and see which host it answers to
kubectl apply -f k8s/ingress.yaml
kubectl get ingress
kubectl describe ingress pantry

Try it

  1. Create the Service and look at its endpoints:
   kubectl apply -f k8s/service.yaml
   kubectl get svc pantry
   kubectl get endpoints pantry
   
   NAME     TYPE        CLUSTER-IP    EXTERNAL-IP   PORT(S)   AGE
   pantry   ClusterIP   10.96.80.31   <none>        80/TCP    2s
   NAME     ENDPOINTS                                            AGE
   pantry   10.244.1.4:8080,10.244.1.5:8080,10.244.1.6:8080      2s
   

Three Pod IPs behind one cluster IP. Delete a Pod and run the endpoints command again: one IP has changed, the Service has not.

  1. Call it from inside the city, by name:
   kubectl run curl --rm -it --restart=Never --image=curlimages/curl:8.11.1 -- -s http://pantry/
   
   Hello from Pantry (v1). You are visitor #1. Served by pantry-647759c9cb-xthw5.
   

Run it a few times: the Pod name at the end changes as the Service spreads the load.

  1. Open a door on the node. The book's kind config maps node port 30080 to your laptop:
   kubectl apply -f k8s/service-nodeport.yaml
   curl localhost:30080/
   kubectl delete -f k8s/service-nodeport.yaml
   
  1. Install the Ingress controller, wait for it, and add the front gate:
   kubectl apply -f https://raw.githubusercontent.com/kubernetes/ingress-nginx/main/deploy/static/provider/kind/deploy.yaml
   kubectl wait -n ingress-nginx --for=condition=Ready pod -l app.kubernetes.io/component=controller --timeout=180s
   kubectl apply -f k8s/ingress.yaml
   curl -H 'Host: pantry.localtest.me' localhost/
   curl -o /dev/null -w '%{http_code}\n' localhost/
   
   Hello from Pantry (v1). You are visitor #4. Served by pantry-647759c9cb-g6hp9.
   404
   

With the right hostname you reach Pantry; without it the gate has no signpost and says 404. (*.localtest.me resolves to 127.0.0.1 on most networks, in which case plain curl pantry.localtest.me/ works too; some home routers block that, so the book uses the Host header.)

Gotchas

  • Three ports, three meanings: port is the Service's own, targetPort is the container's, nodePort is the door on the node. Mixing the first two is the most common Service bug.
  • A Service whose selector matches nothing is created happily and routes nowhere; kubectl get endpoints showing <none> is the only tell.
  • From another namespace, redis does not resolve; redis.default or redis.default.svc does.
  • An Ingress without an ingressClassName, on a cluster with no default class, is accepted and ignored by every controller: no ADDRESS in kubectl get ingress, and 503 from the gate.

Then and now

You will see the kubernetes.io/ingress.class annotation and networking.k8s.io/v1beta1 Ingresses; v1 arrived in 1.19 and the old address went in 1.22 (2021), and since 1.33 get endpoints warns about EndpointSlices. Old Maps tells the story.

Recall

  1. What is a Service, in one sentence, and what makes its address stable?
  2. Which command shows the Pod IPs behind a Service, and what does an empty result mean?
  3. Who can reach a ClusterIP, a NodePort and a LoadBalancer Service?
  4. What reads an Ingress resource, and what happens without it?
  5. In the city metaphor, what are the Service and the Ingress?
Show answers
  1. A stable name and IP for whichever Pods match a selector; the Service object outlives the Pods.
  2. kubectl get endpoints <svc>; empty means the selector matches nothing.
  3. Inside the cluster; any node's IP on a high port; the internet via a cloud load balancer.
  4. An Ingress controller; without one the rules do nothing.
  5. The phone number and the front gate.

What Maya learned

  • A Service is a stable name, IP and DNS entry for whichever Pods match its selector; check endpoints when it seems broken.
  • ClusterIP is for inside, NodePort opens a fixed port on every node, LoadBalancer asks the cloud for a public IP.
  • Ingress routes by hostname and path through one controller, so a whole cluster needs one public entrance, not one per Service.

…but

Pantry has a URL. The customer's engineer opens the deployment YAML to review it and finds REDIS_HOST: redis hard-coded in the Pod template, and the greeting hard-coded in the image. "How do we change these for staging?" she asks. "Rebuild the image?" Maya opens her mouth and closes it again.

Chapter 11 · Act II · 4 min readSticky Notes and the Locked Drawer#

The problem

Staging needs a different greeting, a different Redis, and soon a different API token for the payment provider. Maya's options so far are to bake values into the image (one image per environment, which defeats the point of images) or to write them into the Deployment's Pod template (one Deployment file per environment, all nearly identical, drifting apart). Neither is right.

The token is worse. It cannot go in the image, because images get pushed to registries and pulled by whoever has access. It cannot go in the Deployment file, because the file is committed to Git. It has to be somewhere, and that somewhere has to be readable by the Pod and by nobody else.

Raj's answer is two resources that look almost identical and are treated very differently.

The idea

A ConfigMap is a named bag of key-value pairs that lives in the cluster and is injected into Pods at start-up, either as environment variables or as files. The Deployment stays the same in every environment; only the ConfigMap changes. Configuration has moved out of the image and out of the Pod template into a thing you can apply on its own.

A Secret is the same bag with a different label on it. The values are base64-encoded (that is encoding, not encryption; anyone with get secret rights can read them), and the difference is in how the cluster handles them: Secrets can be encrypted at rest in etcd, RBAC lets you grant access to ConfigMaps without granting access to Secrets, and tooling knows not to print them. Put anything you would not paste into a public chat in a Secret.

ConfigMap Secret
Holds Settings, hostnames, feature flags, config files Tokens, passwords, TLS keys
Stored as Plain text base64 (plus optional encryption at rest)
Metaphor Sticky notes on the fridge The locked drawer
Injected as env vars or files env vars or files
Rule of thumb Fine in Git Never in Git; create from a vault or CI
Pod: pantryenvFromenvFromvolume mount(key → file)

ConfigMap pantry-config
GREETING=Welcome…
REDIS_HOST=redis
motd.txt=…

Secret pantry-secret
API_TOKEN=****

environment
GREETING, REDIS_HOST,
API_TOKEN

/etc/pantry/motd.txt
(mounted file)

Remember it as: a ConfigMap is sticky notes on the fridge, a Secret is the locked drawer; both are read by the Pod, only one is meant to be seen.

There are two ways in, and they behave differently. Environment variables are simple and are what most apps read, but they are fixed when the container starts; change the ConfigMap and the Pod keeps its old values until it is restarted (kubectl rollout restart is for exactly this). Mounted files update in place a minute or so after the ConfigMap changes, which suits config files a program re-reads. Pantry uses both: settings as env vars, a message-of-the-day as a file.

Here is the piece of the Deployment that does the wiring; the ConfigMap and Secret themselves are three-line files:

containers:
  - name: web
    image: pantry:v1
    envFrom:
      - configMapRef: { name: pantry-config }    # every key becomes an env var
      - secretRef:    { name: pantry-secret }
    volumeMounts:
      - name: notes
        mountPath: /etc/pantry
        readOnly: true
volumes:
  - name: notes
    configMap:
      name: pantry-config
      items:
        - key: motd.txt
          path: motd.txt                          # one key becomes one file

The Secret file in the repository uses stringData so it is readable, and it contains a fake token. In real life the file is not committed; the Secret is created by kubectl create secret from a CI variable, or by an operator that syncs from a vault.

The commands

# Create from files (fine for config; the Secret file here holds a fake token)
kubectl apply -f k8s/configmap.yaml -f k8s/secret.yaml

# Create from the command line: the usual way for real Secrets
kubectl create configmap pantry-config --from-literal=GREETING='Hi' --from-file=motd.txt
kubectl create secret generic pantry-secret --from-literal=API_TOKEN="$TOKEN"

# See them; Secrets show sizes, not values
kubectl get configmap,secret
kubectl describe configmap pantry-config

# Read a Secret value back (base64, then decode)
kubectl get secret pantry-secret -o jsonpath='{.data.API_TOKEN}' | base64 -d

# Apply the Deployment that consumes them, then confirm from inside a Pod
kubectl apply -f k8s/deployment-config.yaml
kubectl exec deploy/pantry -- sh -c 'echo $GREETING; cat /etc/pantry/motd.txt'

# Env vars are read at start: restart Pods to pick up a changed ConfigMap
kubectl rollout restart deploy/pantry

Try it

  1. Create both and look at how they are shown:
   kubectl apply -f k8s/configmap.yaml -f k8s/secret.yaml
   kubectl get configmap pantry-config -o yaml | head -12
   kubectl get secret pantry-secret -o yaml | grep -A2 '^data'
   

The ConfigMap shows its values. The Secret shows API_TOKEN: czNjcjN0..., which is base64 and decodes with one command.

  1. Wire them into Pantry and check from inside:
   kubectl apply -f k8s/deployment-config.yaml
   kubectl rollout status deploy/pantry
   kubectl exec deploy/pantry -- sh -c 'echo GREETING=$GREETING; echo API_TOKEN=$API_TOKEN; cat /etc/pantry/motd.txt'
   
   GREETING=Welcome to the Pantry
   API_TOKEN=s3cr3t-token-do-not-commit-real-ones
   Today's special: containers, freshly orchestrated.
   

kubectl exec deploy/pantry picks a current Pod of the Deployment for you.

  1. Change a value and see which path notices:
   kubectl patch configmap pantry-config -p '{"data":{"GREETING":"Hello, staging","motd.txt":"Menu changed.\n"}}'
   sleep 60; kubectl exec deploy/pantry -- sh -c 'echo $GREETING; cat /etc/pantry/motd.txt'
   

The file says Menu changed.; the env var still says Welcome to the Pantry. Then:

   kubectl rollout restart deploy/pantry && kubectl rollout status deploy/pantry
   curl -H 'Host: pantry.localtest.me' localhost/
   
   Hello, staging (v1). You are visitor #6. Served by pantry-7c9d4b8f6-q2xkm.
   
  1. Put it back for the next chapter:
   kubectl apply -f k8s/configmap.yaml
   kubectl rollout restart deploy/pantry
   

Gotchas

  • Environment variables are read once at container start; a changed ConfigMap does nothing to running Pods until rollout restart. Mounted files do update, after up to a minute.
  • A Secret is base64, and kubectl get secret -o yaml shows it to anyone with read access. Protection is RBAC and never committing real values, not the encoding.
  • A ConfigMap is capped at about 1 MiB; a 2 MB file is refused with too large, and belongs in the image or a volume instead.
  • Keys used with envFrom must be valid environment variable names; a key like motd.txt is silently skipped as an env var and only works as a mounted file.

Recall

  1. What is the actual difference between a ConfigMap and a Secret?
  2. Which injection path picks up a changed value without a restart?
  3. How do you read a Secret's value back in plain text?
  4. Why does one image now serve staging and production?
  5. In the city metaphor, what is a Secret?
Show answers
  1. Handling: Secrets can be encrypted at rest, gated by RBAC separately, and hidden by tooling. The data is just base64.
  2. Mounted files.
  3. kubectl get secret … -o jsonpath='{.data.KEY}' | base64 -d.
  4. Configuration moved out of the image into ConfigMaps and Secrets.
  5. The locked drawer.

What Maya learned

  • Configuration lives in ConfigMaps and secrets in Secrets, injected as env vars or files, so one image and one Deployment serve every environment.
  • A Secret is base64, not encrypted; its protection is in access control and handling, so never commit real ones.
  • Env vars need a Pod restart to change; mounted files update in place.

…but

Staging is configured. During the first load test, one Pantry Pod starts answering slowly, then stops answering at all, while still showing Running. The Service keeps sending it a third of the traffic. Kubernetes cannot tell a Pod that is up from a Pod that is working, because nobody has told it how to check.

Chapter 12 · Act II · 4 min readAre You Alive?#

The problem

During the load test one Pantry Pod wedged. The process was alive, so the kubelet saw a healthy container. The Pod showed Running and 1/1, so the Service kept sending it a third of all requests, and a third of all requests timed out. Kubernetes did exactly what it had been told, which was nothing.

Two other things surfaced in the same afternoon. When Maya scaled to ten Pods, several were sent traffic in the half-second between the process starting and Redis being reachable, and returned errors. And one Pod, sharing a node with a memory-hungry neighbour, was slow for no reason of its own.

Raj sums it up: "You never told the city how to check on a tenant, and you never told it how much space a tenant needs."

The idea

A probe is a check the kubelet runs against a container on a timer. There are three, and they answer three different questions:

Probe Question On failure Metaphor
Liveness Is the process still sane? Container is restarted "Are you awake?" Ring the bell; no answer, call the locksmith
Readiness Can it take traffic right now? Pod is removed from Service endpoints (no restart) "Are you taking visitors?" No, then don't send any
Startup Has it finished starting? Other probes wait; restart if it takes too long "Are you still moving in?"

Each can be an HTTP GET (a 2xx or 3xx passes), a TCP connect, or a command run inside the container. Pantry's /healthz returns 200 only when Redis answers, so it is a good readiness check and an acceptable liveness check.

container startedstartup probe failing (withinbudget)startup ok, readiness okstartup failed too many timesreadiness fails (removed fromService)readiness passes againliveness fails N timesliveness fails N times

Starting

Ready

Restarted

NotReady

The mistake everyone makes once is using the same aggressive check for liveness and readiness. A readiness failure is gentle: stop sending traffic, keep the Pod. A liveness failure is violent: kill and restart. If Redis goes down for a minute and liveness checks Redis, every Pantry Pod is restarted in a loop for something that was never their fault. Make readiness strict and liveness lenient (longer period, more failures allowed), and give slow starters a startup probe so liveness does not kill them mid-boot.

Remember it as: liveness asks "are you awake?" and restarts you, readiness asks "are you taking visitors?" and just stops the traffic.

Resources are the other half. Every container should declare a request, the CPU and memory it needs to run properly, and a limit, the most it may ever use. The scheduler uses requests to pick a node with room; that is the only information it has. The kubelet enforces limits: a container that exceeds its memory limit is killed (OOMKilled), and one that exceeds its CPU limit is throttled.

resources:
  requests: { cpu: 50m,  memory: 64Mi }    # what the lease promises: 5% of a core, 64 MiB
  limits:   { cpu: 500m, memory: 128Mi }   # the ceiling: half a core, 128 MiB

Think of it as the lease: requests are what you are promised, limits are the ceiling you cannot exceed. Without requests, the scheduler packs Pods onto nodes blind, which is how the memory-hungry neighbour happened. m is millicores (1000m = one CPU), Mi is mebibytes. Set requests from what you measure with kubectl top, and set memory limits close to requests, because memory cannot be throttled, only killed.

The commands

# Apply the Deployment with probes and resources
kubectl apply -f k8s/deployment-probes.yaml

# READY shows containers passing readiness; RESTARTS climbs on liveness failures
kubectl get pods -l app=pantry

# See probe results and OOM kills in the events
kubectl describe pod <pod> | grep -A12 Events

# Install metrics-server on kind (cloud clusters have it already), then measure
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
kubectl patch -n kube-system deploy metrics-server --type=json -p '[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'
kubectl top pod
kubectl top node

# What has been promised on each node, and to whom
kubectl describe node pantry-worker | grep -A6 'Allocated resources'

Try it

  1. Roll out the probed version:
   kubectl apply -f k8s/deployment-probes.yaml
   kubectl rollout status deploy/pantry
   

The rollout takes a little longer now: a new Pod only counts as available once its readiness probe passes, so maxUnavailable: 0 really means zero.

  1. Make Redis disappear and watch readiness do its job:
   kubectl scale deploy/redis --replicas=0
   sleep 15; kubectl get pods -l app=pantry
   kubectl get endpoints pantry
   
   NAME                      READY   STATUS    RESTARTS   AGE
   pantry-5c8f7d9b4-2jk8x    0/1     Running   0          2m
   ...
   NAME     ENDPOINTS   AGE
   pantry   <none>      40m
   

0/1: running, not ready, no traffic. Nothing was restarted. Bring Redis back and the endpoints return within seconds:

   kubectl scale deploy/redis --replicas=1
   
  1. Break one Pod and watch liveness restart it:
   kubectl port-forward deploy/pantry 8000:8080 &
   curl localhost:8000/crash; kill %1
   sleep 5; kubectl get pods -l app=pantry
   

One Pod shows RESTARTS 1. Its events say Container web failed liveness probe, will be restarted or simply that the process exited and was restarted; either way the kubelet handled it.

  1. Measure what Pantry actually uses:
   kubectl top pod -l app=pantry
   
   NAME                     CPU(cores)   MEMORY(bytes)
   pantry-6b6ff9b46-64857   3m           28Mi
   pantry-6b6ff9b46-8wfhf   2m           27Mi
   

Two millicores and 28 MiB. The request of 50m and 64Mi is generous, and the limit of 128Mi has room. Chapter 15 shows what happens when it does not.

Gotchas

  • A liveness probe that checks a dependency (Redis) restarts every Pantry Pod whenever Redis blinks, for a fault that was never theirs. Liveness checks the process; readiness checks the dependencies.
  • CPU over the limit is throttled, memory over the limit is killed. A memory limit set from a guess is how OOMKilled happens to healthy apps.
  • Without a startup probe, a slow-booting container can fail its liveness checks before it is up and be restarted forever, looking exactly like a crash loop.
  • kubectl top shows "Metrics API not available" until metrics-server is installed and has had a minute to collect.

Then and now

You will see Heapster in older monitoring guides and dashboards; it was retired in Kubernetes 1.13 (2018) in favour of metrics-server, which is why kubectl top has a prerequisite today. Old Maps tells the story.

Recall

  1. What happens on a readiness failure, and what happens on a liveness failure?
  2. Which endpoint does Pantry use for liveness and which for readiness, and why are they different?
  3. What does the scheduler use requests for, and what does the kubelet use limits for?
  4. Which command shows real CPU and memory use per Pod?
  5. In the city metaphor, what are requests and limits?
Show answers
  1. Readiness failure removes the Pod from the Service; liveness failure restarts the container.
  2. /livez for liveness (process only), /healthz for readiness (needs Redis), so a Redis outage does not restart Pantry.
  3. Requests to place Pods; limits to throttle CPU and kill on memory.
  4. kubectl top pod.
  5. The lease: what is promised, and the ceiling.

What Maya learned

  • Readiness gates traffic and liveness triggers restarts; make readiness strict, liveness lenient, and add a startup probe for slow boots.
  • Requests are what the scheduler reserves, limits are what the kubelet enforces; memory over the limit means the container is killed.
  • Measure with kubectl top and set requests from reality, not guesses.

…but

Pantry's Pods are honest about their health now. Then Raj scales the Redis Deployment to zero and back up, and the visit counter starts again at one. Maya has seen this before, in chapter 3, and she knows the word for what is missing. She just does not know what it is called in the city.

Chapter 13 · Act II · 4 min readMemory That Survives, Again#

The problem

The Redis Pod was replaced, and the visit counter went back to one. Maya recognises the shape immediately: a Pod's filesystem is the container's writable layer, and it dies with the container, exactly as in chapter 3. Then she needed a volume. Now she needs a volume that works when the Pod might come back on a different node, where a directory on the old node's disk is no use at all.

There is a second wrinkle. Redis is a database. Databases care about identity: they want the same name and the same disk every time they start, not a random Pod name and whichever disk is free. A Deployment, which treats all its Pods as interchangeable, is the wrong manager for that.

The idea

Kubernetes separates the disk from the request for a disk. A PersistentVolume (PV) is a real piece of storage: a cloud disk, an NFS export, a directory on a node. A PersistentVolumeClaim (PVC) is a request: "I need 100 MiB, read-write by one node at a time." The cluster binds a claim to a volume that satisfies it, and the Pod mounts the claim, never the volume directly. That indirection is what lets the same Pod spec work on a laptop and in three clouds.

Nobody wants to create PVs by hand, so a StorageClass names a provisioner that creates them on demand. kind ships one called standard; clouds ship their own. A PVC that names no class gets the default.

volumeMountrequests fromprovisionsbound to

Pod redis-0
mounts /data

PVC data-redis-0
'100Mi, ReadWriteOnce'
(the rental request)

StorageClass standard
provisioner: local-path
(the storage company)

PV pvc-92ac9f2c…
100Mi on pantry-worker
(the unit)

Think of storage units: the PV is the unit, the PVC is the rental request, and the StorageClass is the storage company that builds a unit when someone asks. Delete the Pod and the claim, and the unit, stay. Delete the claim and the class's reclaim policy decides whether the unit is wiped.

Remember it as: a PVC is the rental request, the PV is the storage unit, and the Pod mounts the request so it does not care which unit it got.

One thing surprises people on kind and on most clouds: a fresh PVC shows Pending until a Pod uses it. The class's binding mode is WaitForFirstConsumer, so the unit is built on the node where the Pod is scheduled, not somewhere the Pod cannot reach. It is not an error.

For Redis's identity problem there is a second manager: the StatefulSet. It is a Deployment for things with names. Its Pods are numbered (redis-0, redis-1), start and stop in order, each get a stable DNS name through a headless Service, and each get their own PVC from a template, so redis-0 always mounts data-redis-0 even after being rescheduled. Databases, queues and anything that says "node ID" in its docs go in a StatefulSet. Everything else stays in a Deployment.

The whole StatefulSet spec adds three things to what you already know: a serviceName pointing at a headless Service, a volumeClaimTemplates list, and that is all.

kind: StatefulSet
spec:
  serviceName: redis                 # headless Service (clusterIP: None): gives redis-0.redis a DNS name
  replicas: 1
  template:
    spec:
      containers:
        - name: redis
          args: ["redis-server", "--appendonly", "yes"]
          volumeMounts:
            - name: data
              mountPath: /data
  volumeClaimTemplates:              # one PVC per Pod, named data-<pod>
    - metadata: { name: data }
      spec:
        accessModes: ["ReadWriteOnce"]
        resources: { requests: { storage: 100Mi } }

The commands

# What storage companies exist, and which is the default?
kubectl get storageclass

# Make a claim, and watch it wait for a consumer
kubectl apply -f k8s/pvc.yaml
kubectl get pvc

# Replace the throwaway Redis with the StatefulSet version
kubectl delete -f k8s/redis.yaml
kubectl apply -f k8s/redis-statefulset.yaml
kubectl rollout status sts/redis

# The claim it made for itself, and the volume that was built
kubectl get pvc,pv

# StatefulSet Pods have stable names; delete one and the same name comes back
kubectl delete pod redis-0
kubectl get pods -l app=redis -w

# Copy a file out of (or into) a Pod
kubectl cp redis-0:/data/appendonlydir ./redis-backup

Try it

  1. See the storage company and make a standalone claim:
   kubectl get storageclass
   kubectl apply -f k8s/pvc.yaml
   kubectl get pvc scratch
   
   NAME                 PROVISIONER             RECLAIMPOLICY   VOLUMEBINDINGMODE      AGE
   standard (default)   rancher.io/local-path   Delete          WaitForFirstConsumer   5m
   NAME      STATUS    VOLUME   CAPACITY   ACCESS MODES   STORAGECLASS   AGE
   scratch   Pending                                      standard       8s
   

Pending, and it will stay that way until a Pod mounts it. That is the binding mode, not a fault.

  1. Swap Redis for the version that remembers:
   kubectl delete -f k8s/redis.yaml
   kubectl apply -f k8s/redis-statefulset.yaml
   kubectl rollout status sts/redis
   kubectl get pvc,pv
   
   NAME                                 STATUS   VOLUME                                     CAPACITY   STORAGECLASS
   persistentvolumeclaim/data-redis-0   Bound    pvc-92ac9f2c-5034-485d-96f8-875737557612   100Mi      standard
   persistentvolumeclaim/scratch        Pending                                                        standard
   

data-redis-0 was claimed by the StatefulSet's template, bound, and a PV was built for it.

  1. Count some visits, destroy the Pod, and count again:
   curl -H 'Host: pantry.localtest.me' localhost/; curl -H 'Host: pantry.localtest.me' localhost/
   kubectl delete pod redis-0
   kubectl wait --for=condition=Ready pod/redis-0 --timeout=90s
   curl -H 'Host: pantry.localtest.me' localhost/
   
   Hello from Pantry (v1). You are visitor #1. ...
   Hello from Pantry (v1). You are visitor #2. ...
   pod "redis-0" deleted
   Hello from Pantry (v1). You are visitor #3. ...
   

Same name, same claim, same data. The pantry shelf has moved to the city.

  1. Clean up the unused claim:
   kubectl delete pvc scratch
   

Gotchas

  • A claim can never shrink (can not be less than status.capacity), and it can grow only if its StorageClass allows expansion; kind's standard does not, so size claims generously.
  • Deleting a StatefulSet does not delete its PVCs, on purpose; the data waits for a StatefulSet of the same name. Delete the claims yourself if you really mean it.
  • A default StorageClass with reclaim policy Delete wipes the volume when the claim is deleted; check kubectl get storageclass before deleting a claim on a real cluster.

Recall

  1. What is the relationship between a PV, a PVC and a StorageClass?
  2. Why is a fresh PVC on kind Pending, and is that a fault?
  3. What does a StatefulSet give Redis that a Deployment does not?
  4. Which command lists claims and volumes together?
  5. In the city metaphor, what is a PVC?
Show answers
  1. A PVC requests storage, a StorageClass provisions a PV to satisfy it, the Pod mounts the PVC.
  2. WaitForFirstConsumer binding; it binds when a Pod uses it. Not a fault.
  3. A stable name and its own claim that follows it.
  4. kubectl get pvc,pv.
  5. The rental request for a storage unit.

What Maya learned

  • Pods mount a PersistentVolumeClaim; a StorageClass provisions a PersistentVolume to satisfy it, so the Pod never names a disk.
  • A claim showing Pending on kind or a cloud is usually waiting for its first Pod, not failing.
  • StatefulSets give Pods stable names and their own claims; use them for databases and Deployments for everything else.

…but

Pantry is healthy, configured and durable. The company hires a second team, who want a cluster for their own service and would rather not share a default namespace with Maya's. Then the launch traffic returns, and Maya finds herself running kubectl scale by hand at lunchtime every day.

Chapter 14 · Act II · 5 min readDistricts, Key Cards and Rush Hour#

The problem

The new team's service is called Shop, and on day one one of its engineers runs kubectl delete deploy pantry in the wrong terminal. Nothing was lost, thanks to chapter 9, but everyone agrees this should not have been possible. Two teams, one cluster, one flat list of everything, and every developer holding the admin credentials Raj handed out on day one.

Meanwhile, the traffic. Pantry needs six Pods at lunchtime and two overnight. Maya has a calendar reminder to run kubectl scale at 11:45 and again at 14:30, which is a restart policy for humans.

And there is a growing list of small jobs: reset a counter once, print a report every minute, run a log shipper on every node. None of them is a long-running web service, and Deployments keep trying to restart the ones that finish.

Four problems, and it turns out to be one short chapter, because all four answers are small.

The idea

Namespaces are districts. Every namespaced resource (Pods, Services, ConfigMaps, most things) lives in exactly one, names only need to be unique within it, and kubectl commands act on one at a time (-n shop, or -A for all). Pantry gets pantry, Shop gets shop, and kubectl delete deploy pantry from the shop district finds nothing to delete. Namespaces also carry quotas and, most usefully, are the unit that permissions attach to. DNS respects them too: pantry.pantry.svc from anywhere, pantry only from inside its own district.

RBAC (role-based access control) is the key cards. A Role lists what may be done to which resources in one namespace (get, list, watch on pods). A RoleBinding hands that Role to a subject: a user, a group, or a ServiceAccount, which is the identity a Pod runs as. Cluster-wide versions (ClusterRole, ClusterRoleBinding) exist for things like nodes and for admins. Shop's engineers get a binding in shop and nothing in pantry.

Namespace: shop (district)Namespace: pantry (district)grantstocan read Pods here onlyno key card

Deployment pantry

ServiceAccount pod-reader

Role pod-reader
verbs: get,list,watch
resources: pods, pods/log

RoleBinding

Deployment shop

Key cards say who may ask city hall for what. They say nothing about which apartments may talk to each other; by default every Pod can reach every other Pod in the cluster. A NetworkPolicy closes corridors: it selects Pods and lists who may connect to them, on which ports. k8s/networkpolicy.yaml says that Redis accepts connections only from Pods labelled app: pantry, on 6379, and nothing else. A policy needs a network plugin that enforces it; kind's does, and so do all the cloud ones.

Remember it as: a Namespace is a district and RBAC is the key card that opens some doors in some districts and none in others.

The HorizontalPodAutoscaler (HPA) is hiring temps in rush hour. It watches a metric (CPU utilisation against the request from chapter 12, by default) and adjusts a Deployment's replica count between a minimum and a maximum to hold the metric near a target. It is a control loop like everything else, and it only works if the Pods declare CPU requests and metrics-server is installed.

Deployment pantryHPA (target 50% CPU, 2–6)metrics-serverDeployment pantryHPA (target 50% CPU, 2–6)metrics-serverloop[every 15 s]how busy are pantry's Pods?254% of requested CPUreplicas = ceil(3 × 254/50) → 6 (max)6 Pods runninghow busy now?40%(after cool-down) replicas = 4

Remember it as: the HPA hires temps when the queue grows and lets them go when it shrinks, within the numbers you set.

The last three are workload types for things that are not web servers, and a table is all they need:

Kind Runs Metaphor Example in k8s/
Job A Pod until it exits 0, retrying on failure, then stops A one-off contractor job.yaml resets the counter
CronJob A Job on a schedule The scheduled cleaner cronjob.yaml prints the count every minute
DaemonSet Exactly one Pod on every node (new nodes included) One caretaker per building daemonset.yaml, stand-in for a log shipper

The commands

# Districts: create, list, act inside one (-n), or everywhere (-A)
kubectl create namespace pantry-staging
kubectl get namespaces
kubectl get pods -n pantry-staging
kubectl get pods -A

# Key cards: apply a ServiceAccount + Role + RoleBinding, then ask "can X do Y?"
kubectl apply -f k8s/rbac.yaml
kubectl auth can-i list pods -n pantry-staging --as=system:serviceaccount:pantry-staging:pod-reader
kubectl auth can-i delete pods -n pantry-staging --as=system:serviceaccount:pantry-staging:pod-reader

# Corridors: who may talk to whom
kubectl apply -f k8s/networkpolicy.yaml
kubectl get networkpolicy

# Rush hour: create the autoscaler and watch it think
kubectl apply -f k8s/hpa.yaml
kubectl get hpa -w

# Contractors and cleaners
kubectl apply -f k8s/job.yaml && kubectl logs job/reset-counter
kubectl apply -f k8s/cronjob.yaml && kubectl get cronjob
kubectl apply -f k8s/daemonset.yaml && kubectl get ds

Try it

  1. Make a district and a key card, and test the card:
   kubectl apply -f k8s/namespace.yaml -f k8s/rbac.yaml
   kubectl auth can-i list pods -n pantry-staging --as=system:serviceaccount:pantry-staging:pod-reader
   kubectl auth can-i delete pods -n pantry-staging --as=system:serviceaccount:pantry-staging:pod-reader
   kubectl auth can-i list pods -n default --as=system:serviceaccount:pantry-staging:pod-reader
   
   yes
   no
   no
   

Read Pods in its district: yes. Delete them: no. Read Pods next door: no. auth can-i with --as is the fastest way to debug any permission question.

  1. Hire temps. Pantry already has CPU requests from chapter 12; add the autoscaler and some load:
   kubectl apply -f k8s/hpa.yaml
   kubectl run load --image=busybox:1.36 --restart=Never -- sh -c 'while true; do wget -qO- http://pantry/ >/dev/null 2>&1; done'
   sleep 75; kubectl get hpa pantry
   
   NAME     REFERENCE           TARGETS         MINPODS   MAXPODS   REPLICAS   AGE
   pantry   Deployment/pantry   cpu: 254%/50%   2         6         6          76s
   

One busy loop drove tiny Pods to 254% of their request, and the HPA went to the maximum. Delete the load Pod and, after a five-minute cool-down, it scales back to two.

   kubectl delete pod load
   
  1. Run a contractor and a cleaner:
   kubectl apply -f k8s/job.yaml
   kubectl wait --for=condition=Complete job/reset-counter --timeout=60s
   kubectl logs job/reset-counter
   kubectl apply -f k8s/cronjob.yaml
   sleep 70; kubectl get jobs
   

The Job printed OK and finished. The CronJob has created its first child Job, named report-visits-<timestamp>, and kubectl logs job/<that name> shows the count.

  1. Put a caretaker in every building:
   kubectl apply -f k8s/daemonset.yaml
   kubectl get ds caretaker
   
   NAME        DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   AGE
   caretaker   2         2         2       2            2           10s
   

Two nodes, two Pods. The manifest carries a toleration so it may also run on the control-plane node, which normally refuses ordinary workloads.

Gotchas

  • The HPA does nothing, and says nothing, if the Pods have no CPU requests; kubectl describe hpa shows <unknown> for the target. Requests first, autoscaler second.
  • CronJob schedules run in UTC unless you set spec.timeZone; a "7 a.m." job that fires at 8 is this.
  • kubectl delete namespace can hang in Terminating for ever when a resource inside has a finalizer nobody will honour; look for it before you start removing finalizers by hand.
  • A NetworkPolicy with an empty podSelector: {} selects every Pod in the namespace; with policyTypes: [Ingress] and no rules, that is "deny all inbound".

Then and now

You will see PodSecurityPolicy (removed in 1.25, 2022), autoscaling/v2beta2 HPAs, batch/v1beta1 CronJobs, and code that expects every ServiceAccount to have a token Secret (not since 1.24). Old Maps tells the story.

Recall

  1. What three things does a Namespace separate?
  2. Which command answers "can this ServiceAccount delete Pods here?"
  3. What two prerequisites does the HPA need before it will scale anything?
  4. Which resource controls which Pods may connect to Redis?
  5. In the city metaphor, what are a Job, a CronJob and a DaemonSet?
Show answers
  1. Names, quotas and permissions.
  2. kubectl auth can-i delete pods -n <ns> --as=system:serviceaccount:<ns>:<sa>.
  3. CPU requests on the Pods and metrics-server in the cluster.
  4. A NetworkPolicy.
  5. A one-off contractor, a scheduled cleaner, one caretaker per building.

What Maya learned

  • Namespaces split a cluster into districts for naming, quotas and permissions; -n picks one and -A sees all.
  • RBAC grants verbs on resources in a namespace to users or ServiceAccounts; auth can-i --as answers any permission question.
  • The HPA scales a Deployment on CPU against its requests; Jobs run to completion, CronJobs on a schedule, DaemonSets once per node.

…but

The city is well run. At 3 a.m., for the first time since Act I, Maya's phone rings. Four different things are wrong at once, and the dashboard says only Pending, ImagePullBackOff, CrashLoopBackOff and OOMKilled. This time she has the tools. She just needs the order.

Chapter 15 · Act II · 3 min readThe Night It Fell Over, Again#

The problem

3:04 a.m. Shop's team pushed four changes before going home, and the cluster has four unhappy Pods. Maya looks at kubectl get pods and sees the four words every Kubernetes engineer learns to read at a glance:

NAME                              READY   STATUS             RESTARTS      AGE
broken-crash-864749c4c6-d447b     0/1     CrashLoopBackOff   3 (43s ago)   60s
broken-image-85b97ddc-9wmxd       0/1     ImagePullBackOff   0             60s
broken-oom-7cd56c996d-xk8lb       0/1     OOMKilled          3 (40s ago)   60s
broken-pending-7866dd6949-7j2t8   0/1     Pending            0             60s

In Act I, one dead container took her two hours to find. Tonight there are four, and she has Raj's rule taped to her monitor: the status tells you which question to ask, and there are only four questions.

The idea

Every broken Pod is stuck at one of four points in its life: it could not be placed, its image could not be fetched, its process keeps dying, or it is being killed by the kubelet for breaking its lease. The STATUS column tells you which, and each has one command that reveals the cause.

PendingImagePullBackOff /ErrImagePullCrashLoopBackOff / ErrorOOMKilledRunning but not READY

kubectl get pods
what does STATUS say?

kubectl describe pod
read Events at the bottom

'Insufficient cpu/memory'
→ requests too big, or cluster full:
lower requests or add nodes

'didn't match node selector /
taint'
→ fix selector, tolerations, or labels

'pod has unbound PVC'
→ check kubectl get pvc

kubectl describe pod
'Failed to pull image …'

typo in name or tag,
private registry needs imagePullSecrets,
or image not loaded into kind

kubectl logs pod --previous
(the log of the crashed attempt)

app error: wrong env / config,
missing dependency, bad command;
fix the manifest

kubectl describe pod
Last State: Terminated, Reason: OOMKilled

raise memory limit,
or fix the leak

kubectl describe pod
'Readiness probe failed'

dependency down or probe path wrong;
kubectl exec + curl localhost to test

Remember it as: Pending asks the scheduler, ImagePull asks the registry, CrashLoop asks the app's previous log, OOMKilled asks the lease.

The tools are ones you already own, with two additions. describe shows the Pod's events, which is where the scheduler and kubelet explain themselves. logs --previous shows the output of the last crashed container rather than the one currently restarting, which is where the stack trace lives. get events --sort-by shows the whole namespace's recent history in time order, which is how you notice that four things broke at 2:58 a.m. and not four separate times. exec lets you test from inside (is Redis reachable from here? does /healthz answer on localhost?). And debug attaches a throwaway container with real tools to a Pod whose image has no shell at all.

Two habits make this fast. First, always read the bottom of describe; the events are last and the newest is at the bottom. Second, when a Deployment's Pods are wrong, fix the Deployment (kubectl apply a corrected file or kubectl set image), never the Pod; the manager will just replace whatever you edit by hand.

The reverse case, Running and 1/1 but users see errors, is a chapter 12 problem: the probes are not checking the right thing. Running and 0/1 is a readiness failure, and the events say why.

These four have a STATUS word. The failures that do not, the Pod that is Running and 1/1 while users see errors, or the Service that routes nowhere, are the silent kind, and each chapter's Gotchas section lists the ones its concept produces.

The commands

# Triage: statuses, then the recent history of the whole district in order
kubectl get pods
kubectl get events --sort-by=.lastTimestamp | tail -20

# The story of one Pod; events are at the bottom
kubectl describe pod <pod>

# Logs of the crashed attempt, not the current retry
kubectl logs <pod> --previous

# Test from inside: is the dependency reachable? does the probe path work?
kubectl exec -it <pod> -- sh
kubectl exec <pod> -- python -c "import socket; print(socket.gethostbyname('redis'))"

# A tools container attached to a Pod that has no shell of its own
kubectl debug -it <pod> --image=busybox:1.36 --target=web

# The Pod's last exit reason in one line
kubectl get pod <pod> -o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}'

# Fix the manager, not the Pod
kubectl set image deploy/<name> web=pantry:v1
kubectl apply -f <corrected>.yaml

Try it

Reproduce the night. The four broken Deployments are in k8s/broken/:

kubectl apply -f k8s/broken/
sleep 60; kubectl get pods -l 'app in (broken-pending,broken-image,broken-crash,broken-oom)'
  1. Pending. Ask the scheduler:
   kubectl describe pod -l app=broken-pending | tail -5
   
   Warning  FailedScheduling  ...  0/2 nodes are available: 1 Insufficient memory, 1 node(s) had untolerated taint ...
   

The manifest requests 64 GiB of memory. Nobody has a flat that big. Fix: k8s/broken/pending.yaml, lower the request, apply.

  1. ImagePullBackOff. Ask the registry:
   kubectl describe pod -l app=broken-image | grep -E 'Failed|BackOff' | tail -2
   
   Warning  Failed   ...  Failed to pull image "pantry:v1-typo": ... not found
   Normal   BackOff  ...  Back-off pulling image "pantry:v1-typo"
   

A tag that does not exist. Fix: kubectl set image deploy/broken-image web=pantry:v1.

  1. CrashLoopBackOff. Ask the previous log:
   kubectl logs -l app=broken-crash --previous
   
   cannot find REDIS_HOST=rediss
   

A typo in an environment variable, printed by the app as it died. Fix the manifest, apply.

  1. OOMKilled. Ask the lease:
   kubectl get pod -l app=broken-oom -o jsonpath='{.items[0].status.containerStatuses[0].lastState.terminated.reason}'
   
   OOMKilled
   

The Pod allocates 200 MiB and its limit is 64 MiB. Fix: raise the limit or stop allocating. Then look at the whole night in order:

   kubectl get events --sort-by=.lastTimestamp | grep -i broken | tail
   
  1. Clean up:
   kubectl delete -f k8s/broken/
   

Four failures, four questions, four commands. Maya is back in bed by 3:20.

Recall

  1. For each of Pending, ImagePullBackOff, CrashLoopBackOff and OOMKilled, which command answers "why"?
  2. Where in kubectl describe output are the events, and in what order?
  3. Why fix the Deployment rather than the Pod?
  4. Which command shows the whole namespace's recent history in time order?
  5. In the metaphor, what do the four status words ask?
Show answers
  1. describe for Pending, ImagePullBackOff and OOMKilled; logs --previous for CrashLoopBackOff.
  2. At the bottom, oldest first.
  3. The manager replaces anything you patch on a Pod.
  4. kubectl get events --sort-by=.lastTimestamp.
  5. The scheduler, the registry, the previous log, the lease.

What Maya learned

  • The STATUS column names the failure and picks the command: describe for Pending, ImagePull and OOM, logs --previous for crashes.
  • Events, read bottom-up, are where the scheduler and kubelet explain themselves; get events --sort-by shows the whole timeline.
  • Fix the Deployment, not the Pod; the manager will replace anything you patch by hand.

…but

Pantry is a solved problem. Then the company grows. Shop's team arrives with a service, a deadline, and a repository left behind by a contractor who has since moved on. Maya opens it on a quiet Friday afternoon and finds docker-compose.yml files with a version: line, a script that calls docker stack deploy, a Helm chart whose README says to run helm init, and a warning in bold that Kubernetes is about to drop Docker. Every word is one she almost knows.

Interlude · 7 min readOld Maps#

The problem

Shop's previous contractor has left, and left a repository. Maya opens it on a quiet Friday afternoon, expecting to read it in an hour, and finds a language she almost speaks.

There is a docker-compose.yml that starts with version: "2" and wires the services together with links:. There is a deploy.sh containing one line, docker stack deploy -c stack.yml pantry, and a stack.yml full of a deploy: key she has never seen. There is a Helm chart whose README says "run helm init first", with its dependencies in a separate file called requirements.yaml. There is a Deployment with apiVersion: extensions/v1beta1 and no selector. There is an Ingress whose class is set by an annotation. There is something called a PodSecurityPolicy. And there is a note at the top of the README, in bold, warning that "Kubernetes is dropping Docker support, we need to migrate off Docker before it breaks."

Some of these will apply to the cluster. Some will fail. Some will apply and silently do nothing. Maya cannot tell which is which, and she cannot tell whether the bold warning is true.

Raj reads it over her shoulder and laughs, not unkindly. "None of this is wrong," he says. "It's just old. Let me tell you when each of these was true."

The idea

Everything in that repository was correct when it was written. Understanding it means knowing three stories, each of which ends with a sentence about what survived.

Arc one: how containers escaped Docker. Docker appeared in 2013 and for two years "container" and "Docker" meant the same thing. In 2015 Docker and others founded the Open Container Initiative (OCI), which wrote down two standards: what an image is, and how a runtime starts a container from one. Once those were standards, anyone could build a runtime, and the runtime that mattered was containerd, the engine underneath Docker itself. Kubernetes talks to runtimes through its Container Runtime Interface (CRI). Docker did not speak CRI, so Kubernetes carried an adapter called the dockershim. In Kubernetes 1.20 (December 2020) the project announced the shim would go, and the internet read it as "Kubernetes drops Docker". In 1.24 (May 2022) it went. Nothing about images changed: a docker build still produces an OCI image, and every cluster still runs it, through containerd directly instead of through Docker. The README's bold warning was true, and it did not matter.

Remember it as: the recipe card became a standard, so the caretaker stopped borrowing Docker's oven and bought his own.

kubelet usedremoved

2013 Docker
(image + runtime, one product)

2015 OCI
image spec + runtime spec

containerd + runc
(the runtime under Docker)

dockershim
kubelet → Docker → containerd

1.24 (2022)
kubelet → containerd via CRI

Arc two: how the orchestrator war ended. Kubernetes was announced in 2014 and reached 1.0 in July 2015. Docker answered with its own orchestrator: Swarm mode, built into Docker 1.12 (June 2016), and a version 3 of the Compose file format with a deploy: key (Compose 1.10 and Docker 1.13, January 2017) so that one file could describe both a laptop and a cluster. docker stack deploy read that file and ran it on a swarm. For about two years teams had to choose. Then Docker itself shipped Kubernetes inside Docker Desktop (version 18.06, July 2018), which was the end of the argument. Compose survived, as the best way to run several containers on one machine, which is exactly how this book used it. A stack.yml is a Compose file with a deploy: section that Compose ignores; Swarm still exists and some small shops still run it happily.

Remember it as: two cities were built; one kept its kitchens and gave up on being a city.

Arc three: how the YAML kept moving. Kubernetes promises that apiVersion, kind, metadata, spec never change shape, and it has kept that promise. What changes is the addresses. Deployments started life in extensions/v1beta1; that address stopped working in 1.16 (September 2019), and apps/v1 requires the selector the old one let you omit. kubectl run created a Deployment until 1.18 (2020), when its generators were removed and it began creating only Pods, which is why old tutorials' kubectl run nginx --image=nginx did something different. The node that runs city hall was labelled master until the label was removed in 1.24 and the matching taint in 1.25; today it is control-plane. Ingress moved from networking.k8s.io/v1beta1 to v1 in 1.19, the old address was removed in 1.22 (2021), and with it the kubernetes.io/ingress.class annotation gave way to ingressClassName. PodSecurityPolicy was removed in 1.25 (2022) in favour of Pod Security Admission, which is a namespace label rather than a resource. Under the hood, kube-dns became CoreDNS by default in 1.13 (2018), Heapster gave way to metrics-server the same year, ServiceAccounts stopped getting a token Secret created automatically in 1.24, and in 1.33 (2025) the Endpoints API you met in chapter 10 began warning that EndpointSlices are its replacement.

The tools moved too. Helm 2 ran a server inside the cluster called Tiller, installed with helm init, with chart dependencies in requirements.yaml; Helm 3 (November 2019) removed Tiller and moved dependencies into Chart.yaml, and Helm 4 (November 2025) is the version in this book. Istio's telemetry ran through a component called Mixer, deprecated in 1.5 (2020) and removed in 1.8; its sidecars became optional when ambient mode reached GA in 1.24 (2024). Docker's own builder was replaced by BuildKit by default in Docker 23 (2023), which is why old build output looks different, and Docker Hub began rate-limiting anonymous pulls in November 2020.

Remember it as: the four-key shape never changed; only the street addresses did.

It becameYou will see

docker-compose · version: '2' · links:

extensions/v1beta1 · master · kubectl run → Deployment

ingress.class annotation · PodSecurityPolicy

helm init · Tiller · requirements.yaml

docker compose · no version · one network

apps/v1 · control-plane · kubectl create deployment

ingressClassName · Pod Security Admission label

Helm 3/4 · no server · Chart.yaml dependencies

Alive, just not here. Some names you will meet are not old, only absent from this book. Podman runs OCI containers without a daemon and accepts most docker commands unchanged. Colima and Rancher Desktop are alternatives to Docker Desktop on a Mac, which became a paid product for larger companies in 2021. minikube is kind's older sibling, a one-node cluster in a VM. Kompose translates a Compose file into Kubernetes manifests, which is a reasonable first draft and a poor final one. Swarm, as said, still runs in small shops. All of them use the same images.

The commands

# Which API addresses does this cluster actually serve? (the old ones will be missing)
kubectl api-resources --api-group=extensions
kubectl api-resources | grep -i ingress

# Validate a manifest against the live cluster without creating anything
kubectl apply --dry-run=server -f old-maps/deployment-v1beta1.yaml

# The manual for the current address of any kind
kubectl explain deployment
kubectl explain deployment --api-version=apps/v1 --recursive | head -40

# Compose: parse and print the effective file; old keys produce warnings here
docker compose -f old-maps/docker-compose.yml config

# Helm: lint an inherited chart; Helm 2 charts and their READMEs show their age
helm lint old-maps/chart-helm2
helm version --short

Try it

The inherited files are in examples/pantry/old-maps/, each with a *.modern.* neighbour that is what you would write today.

  1. Start with the Compose file. Nothing is broken, but the tool tells you it is old:
   cd examples/pantry
   docker compose -f old-maps/docker-compose.yml config | head -3
   
   level=warning msg="old-maps/docker-compose.yml: the attribute `version` is obsolete, it will be ignored, please remove it to avoid potential confusion"
   name: old-maps
   services:
   

version: is ignored, links: still works, and compose.modern.yaml is the same file without either.

  1. The Swarm script fails on the first line, because your Docker is not a swarm:
   old-maps/deploy.sh
   
   Error response from daemon: This node is not a swarm manager. Use "docker swarm init" or "docker swarm join" ...
   

Read stack.yml anyway. It is chapter 5's file with a deploy: block: replicas, update policy, restart policy. Every one of those ideas is a Deployment now.

  1. Ask the cluster about the old addresses:
   kubectl apply --dry-run=server -f old-maps/deployment-v1beta1.yaml
   kubectl apply --dry-run=server -f old-maps/ingress-v1beta1.yaml
   kubectl apply --dry-run=server -f old-maps/psp.yaml
   
   error: resource mapping not found for name: "pantry-old" ... no matches for kind "Deployment" in version "extensions/v1beta1"
   error: resource mapping not found for name: "pantry-old" ... no matches for kind "Ingress" in version "networking.k8s.io/v1beta1"
   error: resource mapping not found for name: "restricted" ... no matches for kind "PodSecurityPolicy" in version "policy/v1beta1"
   

"No matches for kind X in version Y" is the sentence that means this file is from before Y was removed. Now the modern neighbours:

   kubectl apply --dry-run=server -f old-maps/deployment.modern.yaml -f old-maps/ingress.modern.yaml -f old-maps/psa.modern.yaml
   

All three are accepted. Open the two Deployments side by side: the new one has apps/v1 and a selector, and nothing else changed.

  1. Lint the Helm 2 chart and try the command its README asks for:
   helm lint old-maps/chart-helm2
   helm init
   
   [ERROR] templates/deployment.yaml: a Deployment must contain matchLabels or matchExpressions ...
   [WARNING] chart directory is missing these dependencies: redis
   Error: unknown command "init" for "helm"
   

Three eras in three lines: the template is extensions/v1beta1 without a selector, the dependencies live in requirements.yaml where Helm 3 no longer looks, and helm init installed a server that no longer exists.

  1. Finally the bold warning. Ask the cluster what runs its containers:
   kubectl get nodes -o wide | awk '{print $1, $NF}'
   
   pantry-control-plane containerd://2.3.4
   pantry-worker containerd://2.3.4
   

Not Docker. And pantry:v1, built with Docker in chapter 2, is running on it right now.

Gotchas

  • docker compose silently ignores version:; the Python docker-compose (with a hyphen) honoured it and still exists on some machines, so which docker-compose explains a lot of "works for me" arguments.
  • A Helm 2 chart with apiVersion: v1 still installs with Helm 3 and 4, but its requirements.yaml is read only if you run helm dependency update; move the dependencies into Chart.yaml and the confusion ends.
  • kubectl convert is no longer built in; it is a plugin, and for most files it is faster to change the apiVersion by hand and add the selector.

Recall

  1. What did the OCI standardise, and why did that let Kubernetes drop the dockershim without breaking any image?
  2. What is a stack.yml, and which single Kubernetes resource replaces its deploy: block?
  3. Which Kubernetes version stopped serving extensions/v1beta1 Deployments, and what field did apps/v1 make mandatory?
  4. What was Tiller, and which Helm version removed it?
  5. In the city metaphor, what is the dockershim?
Show answers
  1. The image and runtime formats; the image never depended on Docker, only the adapter did.
  2. A Compose file with a deploy: block for Swarm; a Deployment.
  3. 1.16; the selector.
  4. Helm 2's in-cluster server; Helm 3 removed it.
  5. The adapter that let the caretaker use Docker's oven, now retired.

…but

By five o'clock Maya has modernised every file in Shop's repository by hand: new addresses, a selector here, a label there, dependencies moved into Chart.yaml. She closes the laptop feeling like an archaeologist. Then the cloud provider's email arrives: the worker node's host needs a kernel patch on Thursday, and Raj forwards it with one line. "Your building. You're the landlord this week."

Interlude · 5 min readThe Landlord's Week#

The problem

The cloud provider sends an email: the worker node's underlying host needs a kernel patch, and it will be rebooted on Thursday between two and four. Raj forwards it to Maya with one line: "Your building. You're the landlord this week."

Maya's first instinct is to do nothing. Pantry has three replicas and a Deployment that never sleeps; if the node reboots, the manager will notice the missing Pods and start new ones. That is true, and it is also a plan for three minutes of downtime and a Redis restart at an hour of the provider's choosing. What she wants is to move every tenant out before the reboot, at a time she picks, without a single failed request, and to stop anyone moving in while the paint is wet.

Then Shop's team asks a second question in the same week: their new intern deployed a load test into the staging district with a replicas: 200, and the whole cluster slowed to a crawl. Could the district have a ceiling?

Both are landlord problems: how to empty a building safely, and how to stop one tenant taking the whole block.

The idea

Emptying a building is two verbs. kubectl cordon marks a node unschedulable: existing Pods stay, new ones go elsewhere. "No new tenants." kubectl drain cordons and then evicts every Pod on the node, one at a time, so that their managers recreate them on other nodes. "Everyone out, politely." Eviction is a proper request through the API server, not a kill, so the Pod's grace period is honoured and, crucially, a PodDisruptionBudget can say no.

A PodDisruptionBudget (PDB) is a rule attached to a label selector: "at least N of these Pods must be available at all times" (minAvailable), or "at most N may be down" (maxUnavailable). Drain asks the budget before every eviction. If evicting the next Pod would break it, drain waits and retries until a replacement is running somewhere else. With three replicas and minAvailable: 2, the building empties one apartment at a time, and Pantry never drops below two serving Pods.

DeploymentPDB pantry (minAvailable 2)API serverkubectl drainDeploymentPDB pantry (minAvailable 2)API serverkubectl draindrain waits and retriescordon pantry-workerevict pantry-a3 available, may I?yes (2 remain)pantry-a gonenew Pod on another node… Pending (no room!)evict pantry-b2 available, may I?no

That last exchange is the interesting one, and on this book's two-node cluster it is exactly what happens: the only other building is the control plane, which refuses ordinary tenants with a taint, so the replacement Pod goes Pending and the budget blocks the second eviction. Drain is doing its job. The landlord's mistake was owning two buildings and keeping one locked.

Remember it as: cordon is "no new tenants", drain is "everyone out, politely", and the PodDisruptionBudget is the minimum who must stay in during works.

When the work is done, kubectl uncordon reopens the building. Pods do not move back on their own; they are only rebalanced when something replaces them, so a rollout restart afterwards is how the landlord gets tenants back in.

The ceiling on a district is two more resources, both namespaced. A ResourceQuota caps what a namespace may consume in total: Pods, requested CPU and memory, limits, even object counts. A LimitRange sets per-container defaults and bounds, so that a Pod created without any resources: gets a sensible lease instead of none. Together they are the district's building code: the quota is the ceiling for the whole district, the LimitRange is the default lease for any tenant who does not name one. Quotas count requests, which is why the LimitRange matters: without a default, a Pod with no requests counts as zero and the ceiling means nothing.

Namespace pantry-staging (district)fills incounts requests, refuses atthe ceiling

ResourceQuota staging-ceiling
pods 10 · requests.cpu 1 · requests.memory 1Gi

LimitRange default-lease
default: 200m / 128Mi
defaultRequest: 50m / 64Mi

Pod with no resources:
gets the default lease

Pod #11:
refused by the quota

Remember it as: the quota is the district's ceiling and the LimitRange is the default lease; the ceiling only works if every tenant has a lease.

Neither protects against a bad neighbour on the same floor; that is the requests-and-limits story from chapter 12. These protect the block from the district.

The commands

# No new tenants; everyone out politely; reopen
kubectl cordon pantry-worker
kubectl drain pantry-worker --ignore-daemonsets --delete-emptydir-data
kubectl uncordon pantry-worker

# The budget that drain must respect, and how many evictions it currently allows
kubectl apply -f k8s/pdb.yaml
kubectl get pdb

# Where is everyone right now?
kubectl get pods -o wide

# Let ordinary Pods onto the control-plane node (remove its taint), and put the taint back
kubectl taint nodes pantry-control-plane node-role.kubernetes.io/control-plane:NoSchedule-
kubectl taint nodes pantry-control-plane node-role.kubernetes.io/control-plane:NoSchedule

# The district's ceiling and default lease, and how much of the ceiling is used
kubectl apply -f k8s/quota.yaml -f k8s/limitrange.yaml
kubectl describe quota staging-ceiling -n pantry-staging
kubectl describe limitrange default-lease -n pantry-staging

Try it

Pantry is running with three replicas and probes from chapter 12. All three Pods are on pantry-worker, because the control plane refuses them.

  1. Add the budget, close the building, and try to empty it:
   cd examples/pantry
   kubectl apply -f k8s/pdb.yaml
   kubectl cordon pantry-worker
   kubectl drain pantry-worker --ignore-daemonsets --delete-emptydir-data --timeout=45s
   kubectl get pdb pantry
   kubectl get pods -l app=pantry -o wide
   
   evicting pod default/pantry-ff65dd7b-7sw9g
   error when evicting pods/"pantry-ff65dd7b-hhkf9" -n "default": global timeout reached: 45s
   NAME     MIN AVAILABLE   MAX UNAVAILABLE   ALLOWED DISRUPTIONS   AGE
   pantry   2               N/A               0                     20s
   pantry-ff65dd7b-4qvjz   Running   pantry-worker
   pantry-ff65dd7b-hhkf9   Running   pantry-worker
   pantry-ff65dd7b-xh79p   Pending   <none>
   

One Pod was evicted, its replacement has nowhere to go, two are still serving, and the budget's ALLOWED DISRUPTIONS is 0. Drain gave up at your timeout instead of breaking the promise. Pantry never dropped below two.

  1. Open the second building and drain again:
   kubectl taint nodes pantry-control-plane node-role.kubernetes.io/control-plane:NoSchedule-
   kubectl drain pantry-worker --ignore-daemonsets --delete-emptydir-data
   kubectl get pods -o wide
   
   node/pantry-worker drained
   pantry-ff65dd7b-mgft9   Running   pantry-control-plane
   pantry-ff65dd7b-s6p6p   Running   pantry-control-plane
   pantry-ff65dd7b-xh79p   Running   pantry-control-plane
   redis-5d85cc97f8-5fw8v  Running   pantry-control-plane
   

Every tenant moved, one at a time, with two of them always serving. This is the moment the provider may reboot the host.

  1. Repairs done. Reopen the worker, lock the control plane again, and bring the tenants home:
   kubectl uncordon pantry-worker
   kubectl taint nodes pantry-control-plane node-role.kubernetes.io/control-plane:NoSchedule
   kubectl rollout restart deploy/pantry deploy/redis
   kubectl get pods -l app=pantry -o wide
   

All three back on pantry-worker. Note that Redis, still a plain Deployment in this cluster, lost its counter in the move; chapter 13's StatefulSet with a claim would not have.

  1. Now the intern. Give the staging district a ceiling and a default lease, then watch a Pod with no resources: get one:
   kubectl apply -f k8s/namespace.yaml -f k8s/quota.yaml -f k8s/limitrange.yaml
   kubectl run lease-test -n pantry-staging --image=busybox:1.36 --restart=Never -- sleep 300
   kubectl get pod lease-test -n pantry-staging -o jsonpath='{.spec.containers[0].resources}{"\n"}'
   kubectl describe quota staging-ceiling -n pantry-staging | tail -5
   
   {"limits":{"cpu":"200m","memory":"128Mi"},"requests":{"cpu":"50m","memory":"64Mi"}}
   Resource         Used   Hard
   pods             1      10
   requests.cpu     50m    1
   requests.memory  64Mi   1Gi
   

The Pod asked for nothing and was given the default lease, and the district's ceiling now counts it. An eleventh Pod, or a request that would push CPU past one core, is refused at admission with a message naming the quota. Clean up with kubectl delete pod lease-test -n pantry-staging.

Gotchas

  • drain refuses to start if the node runs Pods that nothing manages (bare Pods) or Pods with local emptyDir data; --force and --delete-emptydir-data are you accepting that loss, so read the list first.
  • A PDB with minAvailable equal to the replica count can never be satisfied during a drain; the drain simply hangs. Leave room for at least one eviction.
  • A ResourceQuota on CPU or memory silently blocks every Pod that has no requests for them, with an admission error rather than a Pending Pod; pair every such quota with a LimitRange.

Recall

  1. What is the difference between cordon and drain?
  2. Why did the first drain stop after one eviction, and which resource stopped it?
  3. After uncordon, why are the Pods still on the other node, and what moves them back?
  4. What does a ResourceQuota count, and why is a LimitRange needed for it to mean anything?
  5. In the city metaphor, what are cordon, drain and the PodDisruptionBudget?
Show answers
  1. cordon blocks new Pods; drain cordons and evicts the existing ones.
  2. The replacement had nowhere to go, so the PodDisruptionBudget refused the second eviction.
  3. Nothing moves Pods on its own; a rollout restart replaces them where there is room.
  4. Requests; a Pod without requests counts as zero, so the LimitRange supplies defaults.
  5. "No new tenants", "everyone out politely", and the minimum who must stay in.

…but

Thursday's reboot comes and goes and nobody notices, which is the highest compliment an operation can get. Shop's staging district has a ceiling. Maya has, without quite meaning to, become the person who runs the city rather than the person who lives in it. And Shop's repository, modernised by hand in one afternoon, is one of twelve that need installing, upgrading and rolling back, plus a forty-file metrics stack the security team wants. Raj asks how she plans to do that by hand, and whether she would like a box that a whole application comes in.

Chapter 16 · Act III · 4 min readFlat-Pack Furniture#

The problem

Shop needs Redis. Raj copies Pantry's k8s/ directory, renames eleven things, and misses one, and Shop's Redis Service ends up selecting Pantry's Pods for an afternoon. Then the security team asks for a metrics stack, whose upstream install is forty YAML files and four thousand lines, and asks whether anyone plans to hand-edit that when the next version comes out.

Kubernetes has no idea that those forty files are one thing. It knows Deployments and Services; it does not know "a Redis" or "a monitoring stack" or "Pantry, the staging flavour". Maya has been treating YAML files as the unit of deployment, and they are too small. She needs a box that a whole application comes in.

The idea

Helm is the package manager for Kubernetes. A chart is a directory of YAML templates plus a values.yaml of defaults; a release is a chart installed into a cluster under a name, with your values filled in. helm install renders the templates with the values, applies the result, and records the release so it can be upgraded, rolled back, or removed as one unit. Public charts for Redis, PostgreSQL, NGINX, Prometheus and most other things live in repositories you add once and search.

helm rollback → rev 2

Chart (the flat-pack box)
Chart.yaml
values.yaml (defaults)
templates/*.yaml

Your values
--set image.tag=v2
-f staging-values.yaml
(the options sheet)

helm template / install
renders templates + values

Release 'pantry-helm' rev 3
(the assembled piece in your flat)
Deployment, Service, ConfigMap…

Think of flat-pack furniture. The chart is the box on the shelf. The values are the options sheet (colour, size). The release is the assembled piece standing in your flat, and Helm remembers every version you assembled so you can put the previous one back.

Remember it as: a chart is the flat-pack box, values are the options sheet, and a release is the piece assembled in your flat.

Pantry's chart is the manifests from chapters 8 to 12 with the parts that vary replaced by {{ .Values.… }}:

# chart/templates/deployment.yaml (excerpt)
spec:
  replicas: {{ .Values.replicas }}
  template:
    spec:
      containers:
        - name: web
          image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
          resources:
            {{- toYaml .Values.resources | nindent 12 }}
# chart/values.yaml (the defaults, the whole options sheet)
image: { repository: pantry, tag: v1 }
replicas: 3
greeting: "Welcome to the Pantry"
resources:
  requests: { cpu: 50m, memory: 64Mi }
  limits: { cpu: 500m, memory: 128Mi }
redis: { enabled: true }

Staging is now helm install pantry-staging ./chart -f staging.yaml, where staging.yaml overrides three lines. Shop's Redis is helm install shop-redis bitnami/redis, no copying. The forty-file metrics stack is one helm install, and chapter 18 does exactly that.

Two commands are worth learning before the rest: helm template renders a chart to plain YAML without installing anything, so you can read what a public chart is about to do to your cluster, and helm get values shows what a running release was actually installed with.

The alternative to know is Kustomize, built into kubectl. Instead of templates it takes plain YAML as a base and applies overlays (a patch for replicas here, a different image tag there) per environment. No templating language, no release history, no public repositories; just kubectl apply -k overlays/prod. Use it when you own all the YAML and want environments to differ by a few lines; use Helm when you install other people's software or need rollback and packaging.

Helm Kustomize
Unit Chart + values → release Base + overlays → rendered YAML
Variation Template variables Patches on plain YAML
History and rollback Yes (helm rollback) No (it is just apply)
Third-party software Thousands of public charts Bring your own YAML
Best for Installing things; packaging your app for others Your own manifests across dev/staging/prod

The commands

# Add a public repository and search it
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm search repo prometheus-community/kube-prometheus-stack

# See what a chart would create, without creating it
helm template pantry-helm ./chart

# Install; upgrade --install does both, which is what scripts should use
helm install pantry-helm ./chart
helm upgrade --install pantry-helm ./chart --set image.tag=v2 --set errorRate=0.3

# What is installed, in what state, with which values?
helm list
helm status pantry-helm
helm get values pantry-helm

# Every revision, and how to go back
helm history pantry-helm
helm rollback pantry-helm 1

# Remove the release and everything it created
helm uninstall pantry-helm

# Kustomize: render, diff, apply an overlay
kubectl kustomize kustomize/overlays/prod
kubectl diff -k kustomize/overlays/dev
kubectl apply -k kustomize/overlays/dev

Try it

  1. Render first, then install the Pantry chart under a new release name:
   cd examples/pantry
   helm template pantry-helm ./chart | grep -E '^kind:|name: pantry-helm' | head
   helm install pantry-helm ./chart
   kubectl rollout status deploy/pantry-helm
   helm list
   
   NAME         NAMESPACE  REVISION  STATUS    CHART         APP VERSION
   pantry-helm  default    1         deployed  pantry-0.1.0  v1
   

The release created its own Deployment, Service, ConfigMap and Redis, all prefixed with the release name, next to the hand-made ones.

  1. Change an option and upgrade, then look at the history:
   helm upgrade pantry-helm ./chart --set image.tag=v2 --set errorRate=0.3
   kubectl rollout status deploy/pantry-helm
   helm history pantry-helm
   helm get values pantry-helm
   
   REVISION  STATUS      CHART         DESCRIPTION
   1         superseded  pantry-0.1.0  Install complete
   2         deployed    pantry-0.1.0  Upgrade complete
   USER-SUPPLIED VALUES:
   errorRate: "0.3"
   image:
     tag: v2
   
  1. Regret it and roll back in one line:
   helm rollback pantry-helm 1
   kubectl get deploy pantry-helm -o jsonpath='{.spec.template.spec.containers[0].image}'
   
   Rollback was a success! Happy Helming!
   pantry:v1
   

That is revision 3, whose contents equal revision 1. Helm never deletes history.

  1. Try the same variation with Kustomize, then clean both up:
   kubectl create namespace pantry-dev; kubectl create namespace pantry-prod
   kubectl apply -k kustomize/overlays/dev
   kubectl apply -k kustomize/overlays/prod
   kubectl get deploy pantry -n pantry-dev -o jsonpath='dev: {.spec.replicas} × {.spec.template.spec.containers[0].image}{"\n"}'
   kubectl get deploy pantry -n pantry-prod -o jsonpath='prod: {.spec.replicas} × {.spec.template.spec.containers[0].image}{"\n"}'
   helm uninstall pantry-helm
   kubectl delete namespace pantry-dev pantry-prod
   
   dev: 1 × pantry:v1
   prod: 5 × pantry:v2
   

Same base, two overlays, two districts.

Gotchas

  • --set splits on commas, so --set greeting="Hello, world" fails with key " world" has no value; escape it as Hello\, world or use a values file. And --set-string enabled=false is the string "false", which templates treat as true.
  • Helm installs a chart's CRDs on first install and never upgrades or deletes them; a new chart version with changed CRDs needs them applied by hand.
  • A release name must be unique within a namespace; helm install pantry-helm twice fails with "cannot re-use a name that is still in use". Use upgrade --install in scripts.
  • helm uninstall removes what the release created, not what you created next to it by hand, and not PVCs made from volumeClaimTemplates.

Then and now

You will see helm init, a server called Tiller, helm install --name, requirements.yaml and the stable repository in guides from before Helm 3 (2019); none of them exist in Helm 3 or 4. Old Maps tells the story.

Recall

  1. What are a chart, a release and values, in the flat-pack metaphor?
  2. Which command shows what a chart would create without installing anything?
  3. How do you go back to the previous revision of a release?
  4. When would you pick Kustomize over Helm?
  5. Which command shows the values a running release was installed with?
Show answers
  1. The box, the assembled piece in your flat, the options sheet.
  2. helm template.
  3. helm rollback <release> <revision>.
  4. When you own all the YAML and environments differ by a few lines.
  5. helm get values <release>.

What Maya learned

  • Helm packages an application as a chart, installs it as a release with your values, and keeps every revision for rollback.
  • helm template before helm install shows exactly what a chart will do; helm get values shows what a release was installed with.
  • Kustomize varies your own YAML by overlay without templates; Helm is for packaging and for other people's software.

…but

Shop is installed from a chart, Pantry has one too, and the two services now call each other across the cluster in plain HTTP. The security review says every internal call must be encrypted and authenticated. Pantry's v2, the flaky one, needs to go out to some users before all of them. Maya has a Deployment that replaces Pods one by one, which is not the same as sending ten percent of requests somewhere. She needs something that sits in front of every Pod and understands traffic.

Chapter 17 · Act III · 5 min readA Concierge at Every Door#

The problem

The security review lists two findings. Internal traffic between Shop and Pantry is unencrypted, and nothing checks that the caller is really Shop. Maya could add TLS to every service, which means certificates in every image, rotation in every team, and a retry-and-timeout library in three languages. That is a year of work across the company to solve a problem every service has identically.

And the canary. Pantry v2 is ready enough to try on real users, but not all of them. A Deployment rollout is all-or-nothing over a few minutes; what Maya wants is ten percent of requests to v2, held there for a day while she watches, then a dial she can turn. Kubernetes Services spread traffic evenly across Pods and have no dial.

Both problems are about what happens between services. Raj says the word "mesh" and Maya braces herself.

The idea

A service mesh puts a small proxy next to every Pod and routes all of that Pod's traffic, in and out, through it. Because every request passes through a proxy the mesh controls, the mesh can encrypt it (mTLS, mutual TLS, where both sides present certificates), verify who is calling, retry it, time it out, record how long it took, and send it to a different destination than the one asked for. None of that touches application code. Pantry is unchanged.

Istio is the most widely used mesh. In its classic mode the proxy (Envoy) is injected as a second container into each Pod, the sidecar, which is the "helper in the spare room" chapter 8 promised. Istio's control plane (istiod) issues certificates and pushes routing configuration to every sidecar. A newer ambient mode moves the proxy out of the Pod and onto the node, which saves resources; the concepts are the same and this chapter uses sidecars because they are easier to see.

redis containersidecar (redis Pod)pantry containersidecar (pantry Pod)Istio ingress gatewayredis containersidecar (redis Pod)pantry containersidecar (pantry Pod)Istio ingress gatewayverify cert, log, apply rulesHTTP request (mTLS)plain HTTP on localhostconnect to redis:6379mTLS (encrypted, authenticated)plain on localhostreply (back through both sidecars)

Think of a concierge at every apartment door. Nothing reaches the tenant except through the concierge, who checks ID at the door (mTLS), keeps a visitor log (telemetry), and can, when told, send one visitor in ten to the new tenant across the hall.

Remember it as: a mesh puts a concierge at every door who checks ID, keeps the log, and can send one visitor in ten to the new tenant.

Traffic rules use three resources. A Gateway is the mesh's own front gate, like Ingress but Istio-aware. A VirtualService is the signpost: for this host, send requests to these destinations with these weights. A DestinationRule names subsets of a Service's Pods by label, so "v1" and "v2" are things the signpost can point at. The dial Maya wanted is two numbers in the VirtualService.

uses subsets from9010selects both

curl localhost:30080

Gateway pantry-gateway
(istio ingressgateway :80)

VirtualService pantry
90% → subset v1
10% → subset v2

DestinationRule pantry
v1: version=v1
v2: version=v2

Service pantry

Pods version=v1 (×3)

Pod version=v2 (×1)

# istio/virtualservice.yaml: the dial
http:
  - route:
      - destination: { host: pantry, subset: v1 }
        weight: 90
      - destination: { host: pantry, subset: v2 }
        weight: 10

The security finding is one more resource. A PeerAuthentication with mode: STRICT in a namespace means sidecars refuse any connection that is not mTLS. Applied, every internal call in the district is encrypted and authenticated, and nobody changed a line of application code. One consequence to know: anything without a sidecar can no longer talk to Pantry, and that includes the ingress-nginx controller from chapter 10, which starts returning 502. While the mesh is strict, traffic enters through Istio's own gateway (port 30080 in this book) or you inject the ingress controller's namespace too.

Meshes cost something: a proxy per Pod (memory, a little latency), and one more control plane to run. If you only need mTLS and simple retries, Linkerd is the lighter alternative: smaller proxies, fewer concepts, less configurability. The sidebar table is the honest comparison. Kiali is the dashboard that draws the mesh as a graph, and istioctl dashboard kiali opens it if you install the add-on.

Istio Linkerd
Proxy Envoy (feature-rich, heavier) Purpose-built Rust proxy (light)
Traffic rules Gateway, VirtualService, DestinationRule, and more Simpler; uses Gateway API resources
mTLS by default After you ask for STRICT On by default
Choose when You need rich routing, multi-cluster, or an ecosystem You want mTLS and metrics with the least surface area

The commands

# Install Istio into the cluster (demo profile: everything on, small footprint)
istioctl install --set profile=demo -y

# On kind: expose the ingress gateway on the node port the book maps to localhost
kubectl patch svc istio-ingressgateway -n istio-system --type=json -p '[{"op":"replace","path":"/spec/type","value":"NodePort"},{"op":"replace","path":"/spec/ports/1/nodePort","value":30080}]'

# Tell Istio to inject a sidecar into every new Pod in this namespace, then restart the Pods
kubectl label namespace default istio-injection=enabled
kubectl rollout restart deploy/pantry deploy/redis

# Apply the routing (v2 Deployment, Gateway, DestinationRule, VirtualService) and mTLS
kubectl apply -f istio/

# Is my configuration sane? Which Pods have proxies, and are they in sync?
istioctl analyze
istioctl proxy-status

# The mesh graph (needs the Kiali add-on from Istio's samples)
istioctl dashboard kiali

# Undo it all
kubectl delete -f istio/
kubectl label namespace default istio-injection-
istioctl uninstall --purge -y

Try it

  1. Install Istio and open the gate on port 30080:
   istioctl install --set profile=demo -y
   kubectl patch svc istio-ingressgateway -n istio-system --type=json -p '[{"op":"replace","path":"/spec/type","value":"NodePort"},{"op":"replace","path":"/spec/ports/1/nodePort","value":30080}]'
   
  1. Move concierges in. Label the district and restart the tenants:
   kubectl label namespace default istio-injection=enabled
   kubectl rollout restart deploy/pantry deploy/redis
   kubectl rollout status deploy/pantry
   kubectl get pods -l app=pantry
   
   NAME                      READY   STATUS    RESTARTS   AGE
   pantry-5d8d76f9f5-4vqnh   2/2     Running   0          40s
   

2/2: the app and its sidecar. Nothing in the Deployment changed.

  1. Add the new tenant and the dial, and check the mesh agrees with you:
   kubectl apply -f istio/
   kubectl rollout status deploy/pantry-v2
   istioctl analyze
   istioctl proxy-status | head -4
   

analyze may print an Info about Service port names; anything at Warning or Error is a real problem.

  1. Send a hundred visitors through the gate and count where they went:
   for i in $(seq 1 100); do curl -s localhost:30080/; done | grep -o '(v[12])' | sort | uniq -c
   for i in $(seq 1 100); do curl -s -o /dev/null -w '%{http_code}\n' localhost:30080/; done | sort | uniq -c
   
     88 (v1)
     10 (v2)
     98 200
      2 500
   

Roughly ten percent reached v2, and only those saw its errors. Turn the dial: edit the weights to 50/50, kubectl apply, and run the loop again. No Pods restart.

  1. Confirm the concierge is checking ID, then remove the mesh so the next chapter starts clean:
   kubectl get peerauthentication
   kubectl delete -f istio/
   kubectl label namespace default istio-injection-
   istioctl uninstall --purge -y
   kubectl rollout restart deploy/pantry deploy/redis
   

Gotchas

  • With PeerAuthentication STRICT, anything without a sidecar cannot reach meshed Pods, including ingress-nginx, which starts returning 502. Enter through Istio's gateway, or inject the ingress namespace too.
  • Istio wants Service ports named (http, grpc, tcp-redis); istioctl analyze reports unnamed ports as IST0118, and protocol detection may guess wrong without them.
  • Labelling a namespace istio-injection=enabled affects only Pods created afterwards; existing Pods have no sidecar until rollout restart.
  • A DestinationRule subset selects on labels; a Deployment whose Pods lack version: v2 gets zero traffic and no error.

Then and now

You will see Mixer, istio-telemetry and istio-policy Deployments in older Istio installs; Mixer was deprecated in 1.5 (2020) and removed in 1.8, and sidecars themselves became optional with ambient mode (GA in 1.24, 2024). Old Maps tells the story.

Recall

  1. What does the mesh give you without changing application code? Name three things.
  2. What do Gateway, VirtualService and DestinationRule each do?
  3. Which two numbers make the canary, and where do they live?
  4. Why does ingress-nginx return 502 once mTLS is strict?
  5. In the city metaphor, what is the sidecar?
Show answers
  1. mTLS, retries and timeouts, telemetry, traffic splitting; none in application code.
  2. Gateway is the door, VirtualService the signpost with weights, DestinationRule the named subsets.
  3. The weight values, in the VirtualService.
  4. It has no sidecar, and STRICT mode refuses plaintext.
  5. The concierge at the door.

What Maya learned

  • A mesh routes every Pod's traffic through a proxy beside it, which gives mTLS, retries, telemetry and traffic splitting with no application changes.
  • Gateway is the door, VirtualService is the signpost with weights, DestinationRule names the subsets the signpost can point at.
  • Istio is the full-featured choice and Linkerd the lighter one; both cost a proxy per Pod and a control plane to run.

…but

Ten percent of users are on v2, and Maya sees 500 in her curl loop. How many users saw one? Is v2's error rate two percent or twenty? Is it slower? kubectl top shows CPU. The curl loop shows her own hundred requests. Nobody in the company can answer "is the canary healthy?" with a number, and without a number the dial is a guess.

Chapter 18 · Act III · 4 min readThe Control Room#

The problem

"Is the canary healthy?" Maya's honest answer is a shrug. She has kubectl top, which says how much CPU Pantry uses and nothing about whether it is working. She has kubectl logs, which is a firehose across four Pods. She has her own curl loop, a hundred requests from her laptop. The customer's question is about every request from every user over the last hour, split by version, and nobody has that number.

Raj has an old sticker on his laptop that says if you can't measure it, you can't roll it out. Shop's team has been asking for dashboards since the day they joined. Both are the same request: the city needs a control room.

The idea

Prometheus is a metrics database that works by scraping: on a timer, it fetches a /metrics page from every target it knows about and stores the numbers it finds there, with labels. Applications expose counters and histograms; Prometheus does the polling. Nothing is pushed, so a dead target simply stops showing up, which is itself a signal.

Pantry has exposed /metrics since chapter 2 without anyone reading it. Two series matter: pantry_requests_total{version, status}, a counter of requests, and pantry_request_seconds_bucket, a histogram of latency. From those two, everything the customer asked for can be computed.

Grafana is the wall of screens. It queries Prometheus and draws the results, and a dashboard is a saved set of those queries. The two ship together in the kube-prometheus-stack Helm chart, along with the Prometheus Operator, which lets you tell Prometheus about a new target with a Kubernetes resource instead of editing its config file. That resource is the ServiceMonitor: "scrape every Service with this label, on this port, at this path".

tells Prometheus what toscrapescrape /metricsscrapescrapePromQL query

pantry Pod
/metrics

pantry Pod
/metrics

pantry-v2 Pod
/metrics

ServiceMonitor pantry
select Service app=pantry
port http, path /metrics, every 15s

Prometheus
(polls every building on a timer,
stores series with labels)

Grafana dashboard 'Pantry'
rate · error rate · p95

Think of the sensor network: Prometheus is the system that polls every building on a timer and files the readings, and Grafana is the control-room wall where the readings become pictures.

Remember it as: Prometheus polls every building on a timer and files the numbers; Grafana is the wall of screens that draws them.

The query language is PromQL, and it has a reputation. The truth is that three queries answer ninety percent of "is it healthy?", and they are the same three for every HTTP service in the world:

# 1. Request rate, per version (requests per second over the last minute)
sum by (version) (rate(pantry_requests_total[1m]))

# 2. Error rate, per version (fraction of requests that were 5xx)
sum by (version) (rate(pantry_requests_total{status=~"5.."}[1m]))
  / sum by (version) (rate(pantry_requests_total[1m]))

# 3. p95 latency, per version (95% of requests were faster than this, in seconds)
histogram_quantile(0.95, sum by (le, version) (rate(pantry_request_seconds_bucket[5m])))

Read them once, slowly. rate(counter[1m]) turns an ever-growing counter into "per second, recently". sum by (version) collapses the Pods into one line per version, which is the whole point for a canary. The error rate divides a filtered rate by the total rate. The histogram query is the one people copy and paste, and that is fine: it says "take the latency buckets, find the value below which 95% of requests fall".

Those three queries are also the three panels on Pantry's dashboard, in monitoring/grafana-dashboard.json. kubectl top is still useful for "which Pod is eating memory"; Prometheus is for "how is the service behaving over time, and compared to yesterday".

The commands

# Install Prometheus, Grafana and the operator in one release (takes a few minutes on first pull)
helm install monitoring prometheus-community/kube-prometheus-stack --version 91.4.1 \
  --namespace monitoring --create-namespace \
  --set grafana.adminPassword=pantry --set alertmanager.enabled=false --set nodeExporter.enabled=false

# Tell Prometheus to scrape Pantry
kubectl apply -f monitoring/servicemonitor.yaml

# Open Prometheus (targets page: Status → Targets) and Grafana (admin / pantry)
kubectl port-forward -n monitoring svc/monitoring-kube-prometheus-prometheus 9090:9090
kubectl port-forward -n monitoring svc/monitoring-grafana 3000:80

# Ask Prometheus a question from the terminal
curl -s --data-urlencode 'query=sum by (version) (rate(pantry_requests_total[1m]))' http://localhost:9090/api/v1/query

# Still the right tool for "which Pod is heavy right now?"
kubectl top pod -l app=pantry

# Remove the stack
helm uninstall monitoring -n monitoring

Try it

  1. Install the stack and wait for the Pods:
   helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
   helm repo update
   helm install monitoring prometheus-community/kube-prometheus-stack --version 91.4.1 --namespace monitoring --create-namespace --set grafana.adminPassword=pantry --set alertmanager.enabled=false --set nodeExporter.enabled=false
   kubectl get pods -n monitoring -w
   

Wait for prometheus-monitoring-kube-prometheus-prometheus-0 to show 2/2 Running. First-time image pulls can take several minutes.

  1. Point the sensor network at Pantry, generate some traffic, and check the target is up:
   kubectl apply -f monitoring/servicemonitor.yaml
   for i in $(seq 1 80); do curl -s -o /dev/null -H 'Host: pantry.localtest.me' localhost/; done
   kubectl port-forward -n monitoring svc/monitoring-kube-prometheus-prometheus 9090:9090 &
   sleep 40; curl -s 'http://localhost:9090/api/v1/targets?state=active' | grep -o '"job":"pantry"[^}]*"health":"up"' | head -3
   

Three up matches, one per Pantry Pod. If the list is empty, check that the Service carries the app: pantry label and a port named http; ServiceMonitors select on both.

  1. Ask the three questions:
   curl -s --data-urlencode 'query=sum by (version) (rate(pantry_requests_total[1m]))' http://localhost:9090/api/v1/query
   curl -s --data-urlencode 'query=histogram_quantile(0.95, sum by (le, version) (rate(pantry_request_seconds_bucket[5m])))' http://localhost:9090/api/v1/query
   kill %1
   
   {"status":"success","data":{"result":[{"metric":{"version":"v1"},"value":[...,"0.3643"]}]}}
   {"status":"success","data":{"result":[{"metric":{"version":"v1"},"value":[...,"0.0047"]}]}}
   

About a third of a request per second, and 95% of them under five milliseconds. The error-rate query returns nothing yet, because v1 has produced no 5xx responses and Prometheus does not invent zero series. Run v2 (chapter 17's canary, or kubectl set image deploy/pantry web=pantry:v2) and it appears.

  1. Open the wall of screens and import Pantry's dashboard:
   kubectl port-forward -n monitoring svc/monitoring-grafana 3000:80 &
   

Browse to http://localhost:3000, log in as admin / pantry, choose Dashboards → New → Import, and upload monitoring/grafana-dashboard.json. Three panels: request rate, error rate, p95, each with one line per version. Leave it open. The next chapter is going to make the error-rate line move.

Gotchas

  • The Prometheus installed by kube-prometheus-stack only picks up ServiceMonitors carrying the label release: <release name>; without it your target never appears and nothing complains.
  • A rate() window shorter than two scrape intervals returns nothing at all: with 15-second scrapes, [10s] is empty and [30s] is the floor; [1m] is the sensible default.
  • Prometheus does not invent zero: a series for status="500" exists only after the first 500, so an error-rate panel is empty, not zero, until something fails.
  • A ServiceMonitor selects a Service port by name; an unnamed port means no scrape.

Then and now

You will see "Prometheus Operator", "kube-prometheus" and "kube-prometheus-stack" used as if they were three products; they are the controller, the manifests built on it, and the Helm chart that installs both. Old Maps tells the story.

Recall

  1. How does Prometheus get its data, and what does that mean when a target dies?
  2. What are the three queries, in words?
  3. What does a ServiceMonitor need to match, on the Service and on itself?
  4. Which tool answers "which Pod is heavy right now" and which answers "how is the service behaving over time"?
  5. In the city metaphor, what are Prometheus and Grafana?
Show answers
  1. It scrapes /metrics on a timer; a dead target stops appearing.
  2. Request rate, error rate, p95 latency, each per version.
  3. The Service's app label and a named port; the ServiceMonitor's release label.
  4. kubectl top for now; Prometheus for over time.
  5. The sensor network and the wall of screens.

What Maya learned

  • Prometheus scrapes /metrics from targets on a timer; a ServiceMonitor tells it which Services to scrape.
  • Request rate, error rate and p95 latency, each sum by (version), answer "is it healthy?" for any HTTP service.
  • Grafana turns those queries into a dashboard; kubectl top remains the tool for "which Pod is heavy right now".

…but

Maya can finally see. What she sees, when she looks at the deployment history, is that the cluster does not match the repository. Someone raised Pantry's replicas by hand on Tuesday. Shop's ConfigMap has a value nobody committed. Three engineers have kubectl applyed from three laptops and the only record is their shell history. The dashboard shows what is happening; nothing shows what is supposed to be happening.

Chapter 19 · Act III · 5 min readNobody Touches the City by Hand#

The problem

The repository says Pantry has three replicas. The cluster has five. The repository says v1; one Deployment in staging says v2. Shop's ConfigMap has a feature flag that exists nowhere in Git. Every one of these was a sensible change made by a sensible person with kubectl apply or kubectl edit from a laptop, and every one of them is invisible to everybody else.

Maya's own release process is a Slack message saying "deploying now". If the cluster died tomorrow, rebuilding it would mean reconstructing what was running from memory.

The dashboard from chapter 18 shows what the city is. What she wants is a way to say what the city should be, in one place, and have something enforce it. She has said those words before, in chapter 6, about Pods. Raj points that out.

The idea

GitOps is the practice of keeping the desired state of a cluster in a Git repository and running a controller that continuously makes the cluster match it. Git becomes the only way to change production: a pull request is the change, its review is the approval, the merge is the deployment, and git log is the audit trail. Rolling back is git revert. Nobody runs kubectl apply against production, because anything applied by hand is drift, and the controller puts it back.

Argo CD is the most widely used of those controllers. You install it and give it an Application: a Git repository, a path, a target revision, and a destination cluster and namespace. It clones the repository, renders what it finds (plain YAML, a Helm chart, or a Kustomize overlay), compares it with the cluster, and reports Synced or OutOfSync. With automated sync it applies the difference itself; selfHeal also reverts drift, and prune deletes what Git no longer has.

ClusterArgo CDGit repositoryMayaClusterArgo CDGit repositoryMayaloop[every 3 minutes (or on webhook)]someone runs kubectl scale --replicas=5commit: image tag v1 → v2fetch examples/pantry/gitopsread live Deploymentdiff desired vs liveapply (auto-sync)drift detected → self-heal back to 3git revert (v2 was bad)apply: back to v1

Think of the city plan in the archive. The plan is the law; inspectors walk the city continuously and make every building match it, and if someone knocks a wall through by hand, they rebuild it. To change a building, you change the plan, which records every change and who made it.

Remember it as: Git is the city plan, Argo CD is the inspector who makes the city match it, and nobody changes a building by hand.

The Application is one resource:

apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: pantry
  namespace: argocd
spec:
  project: default
  source:
    repoURL: https://github.com/<you>/<this-repo>.git
    targetRevision: HEAD
    path: examples/pantry/gitops          # plain YAML; could be a chart or a kustomize overlay
  destination:
    server: https://kubernetes.default.svc
    namespace: default
  syncPolicy:
    automated:
      prune: true                          # delete what Git no longer has
      selfHeal: true                       # undo changes made by hand

The alternative to know is Flux: the same job as a set of smaller controllers configured entirely by Kubernetes resources, with no UI. Argo CD's UI draws every Application as a tree, which is a good way to learn what a Deployment owns; Flux is leaner.

Argo CD Flux
Shape One control plane with UI, CLI and API Composable controllers, no UI (community dashboards exist)
Config Application resources (or the UI) GitRepository + Kustomization / HelmRelease
Drift Reports OutOfSync; self-heal optional Reconciles on an interval
Choose when You want to see the app tree and hand it to a wider team You want the smallest footprint and everything as YAML

Here the three previous chapters meet. The v2 canary goes out through Istio at ten percent; Prometheus's error-rate panel for v2 climbs; the fix is git revert of the commit that changed the tag, and within three minutes Argo CD has put v1 back. The dial, the sensor and the plan are one system.

The commands

# Install Argo CD (takes a minute or two); --server-side because one of its CRDs is too big for client-side apply
kubectl create namespace argocd
kubectl apply -n argocd --server-side -f https://raw.githubusercontent.com/argoproj/argo-cd/v3.5.3/manifests/install.yaml
kubectl port-forward -n argocd svc/argocd-server 8443:443
kubectl get secret -n argocd argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d

# Log the CLI in through the tunnel
argocd login localhost:8443 --username admin --insecure

# Register the Application (from the file, or with the CLI) and look at it
kubectl apply -f argocd/application.yaml
argocd app create pantry --repo https://github.com/<you>/<this-repo>.git --path examples/pantry/gitops --dest-server https://kubernetes.default.svc --dest-namespace default --sync-policy automated --self-heal
argocd app list
argocd app get pantry

# The kubectl-native way to see what applying a file would change (what Argo does continuously)
kubectl diff -f gitops/

# Sync now, show what differs, see the history, go back
argocd app sync pantry
argocd app diff pantry
argocd app history pantry
argocd app rollback pantry <id>

# Remove
argocd app delete pantry
kubectl delete -n argocd -f https://raw.githubusercontent.com/argoproj/argo-cd/v3.5.3/manifests/install.yaml

Try it

This chapter needs a Git repository Argo CD can reach: fork this book's repository on GitHub and put its URL in argocd/application.yaml.

  1. Install Argo CD and log in:
   kubectl create namespace argocd
   kubectl apply -n argocd --server-side -f https://raw.githubusercontent.com/argoproj/argo-cd/v3.5.3/manifests/install.yaml
   kubectl wait -n argocd --for=condition=Available deploy/argocd-server --timeout=300s
   kubectl port-forward -n argocd svc/argocd-server 8443:443 &
   argocd login localhost:8443 --username admin --insecure --password "$(kubectl get secret -n argocd argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d)"
   

Open https://localhost:8443 in a browser for the UI (same credentials).

  1. Hand Argo CD the plan and watch it sync:
   kubectl apply -f argocd/application.yaml
   sleep 30; argocd app get pantry | grep -E 'Sync Status|Health Status'
   
   Sync Status:        Synced to  (81c6092)
   Health Status:      Healthy
   

argocd app get pantry also lists what it owns: both Deployments, both Services and the ConfigMap, each Synced and Healthy.

  1. Knock a wall through by hand and watch the inspector fix it:
   kubectl scale deploy/pantry --replicas=5
   sleep 3; kubectl get deploy pantry
   
   NAME     READY   UP-TO-DATE   AVAILABLE   AGE
   pantry   3/3     3            3           4m
   

With selfHeal on, the extra Pods were gone before the command returned. argocd app history pantry records nothing, because nothing in Git changed. The only way to have five replicas is to commit five replicas.

  1. Ship the canary the proper way, then revert it. Edit examples/pantry/gitops/deployment.yaml, change pantry:v1 to pantry:v2, commit and push, then watch Argo CD notice and act:
   git commit -am "pantry: try v2" && git push
   argocd app get pantry --refresh | grep 'Sync Status'
   sleep 60; argocd app get pantry | grep 'Sync Status'
   kubectl get deploy pantry -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
   curl -s -H 'Host: pantry.localtest.me' localhost/
   
   Sync Status:        OutOfSync from  (551fb85)
   Sync Status:        Synced to  (551fb85)
   pantry:v2
   Welcome to the Pantry (v2). You are visitor #15. Served by pantry-759dbdd485-5gfgg.
   

v2 is live, and a few of those curls will come back as a 500, because v2 is flaky on purpose. The Grafana panel from chapter 18 shows v2's error rate climbing toward 30%. That is the number the customer asked for, and it says no:

   git revert --no-edit HEAD && git push
   sleep 60; argocd app get pantry | grep 'Sync Status'
   curl -s -H 'Host: pantry.localtest.me' localhost/
   argocd app history pantry
   
   Sync Status:        Synced to  (c1c2573)
   Welcome to the Pantry (v1). You are visitor #18. Served by pantry-5784c9bc56-fcqt8.
   ID      DATE                           REVISION
   0       2026-09-19 09:07:09 +0100 BST   (81c6092)
   1       2026-09-19 09:09:21 +0100 BST   (551fb85)
   2       2026-09-19 09:09:51 +0100 BST   (c1c2573)
   

Nobody ran kubectl. The commit is the rollback, the history is the audit trail, and everyone can see both. (Argo CD polls Git every three minutes; argocd app sync pantry skips the wait, and a webhook makes it instant.)

Gotchas

  • Argo CD's install manifest must be applied with --server-side; client-side apply rejects the ApplicationSet CRD as too large and the applicationset controller crash-loops.
  • argocd app delete cascades by default and removes every resource the Application manages; use --cascade=false to forget the Application and keep the workload.
  • With selfHeal on, a kubectl hotfix during an incident is reverted within seconds; the fix has to go through Git, which is the point, and also the thing to remember at 3 a.m.
  • A private repository needs a credential registered in Argo CD before the Application is created, or creation fails with "repository not accessible".

Recall

  1. What does "GitOps" change about how production is modified?
  2. What five things does an Application name?
  3. What do prune and selfHeal each do?
  4. How is a bad release rolled back under GitOps?
  5. In the city metaphor, what are Git and Argo CD?
Show answers
  1. Git becomes the only way; a merge deploys and a revert rolls back.
  2. Repository, path, revision, destination cluster, destination namespace.
  3. prune deletes what Git removed; selfHeal reverts changes made by hand.
  4. git revert and let Argo CD sync.
  5. The city plan in the archive and the inspector who enforces it.

What Maya learned

  • GitOps keeps the cluster's desired state in Git and lets a controller make the cluster match it, so a merge is a deployment and a revert is a rollback.
  • An Argo CD Application names a repo, a path and a destination; automated sync with self-heal reverts anything changed by hand.
  • Istio limits the blast radius, Prometheus measures it, and Git plus Argo CD is how the decision is made and recorded.

…but

There is no but. The city runs itself. Maya's phone has not rung at 3 a.m. in a month, and when Shop's next canary goes wrong, the person on call reads a dashboard, reverts a commit, and goes back to sleep. Everything she needed was one concept at a time, each one because the last one had a crack in it. The Epilogue has the whole map.

Epilogue · 14 min readThe Map#

Everything in the book, on a few pages. Use this as the reference after the story has done its work.

The whole city

Act I, one kitchen:

docker builddocker push / pulldocker rundocker compose up

Dockerfile
(recipe)

Image pantry:v1
(recipe card, layers)

Registry
(cookbook library)

Container
(dish)

Volume
(pantry shelf)

Network + published port
(corridor + serving hatch)

compose.yaml
(dinner menu)

Act II, the city:

Cluster (city) · Namespace (district)labelsselector

user

Ingress
(front gate)

Service
(phone number)

Deployment / HPA
(building manager,
temps in rush hour)

Pod
(apartment)

Pod

ConfigMap / Secret
(sticky notes / locked drawer)

PVC → PV
(storage unit)

Probes + requests/limits
(doorbell + lease)

Act III, running the city:

Argo CD syncs(inspector)helm install/metrics scrapederror rate says no

Git
(city plan)

Cluster

Helm chart
(flat-pack box)

Istio sidecars
(concierge at every door)
mTLS · 90/10 split

Prometheus
(sensor network)

Grafana
(wall of screens)

Diagram legend

Mark Meaning
Solid arrow Traffic, or a command
Dashed arrow Selects, watches, or "is told by"
Box inside a box Runs inside
Yellow (image) Images, layers, registries: read-only things you build and ship
Orange (container) Containers and Compose: running things on one host
Blue (pod) Pods
Grey (node) Nodes and machines
Purple (control) Control plane parts and controllers: Deployments, probes, RBAC, ConfigMaps
Green (storage) Volumes, PVs, PVCs
Red (net) Services, Ingress, networks, ports
Pink (tool) Act III tools: Helm, Istio, Prometheus, Grafana, Argo CD

The metaphor table

Concept Remember it as Chapter
Image A recipe card: read-only, versioned 1
Container A dish cooked from the recipe; disposable, many per recipe 1
Layer / build cache A step in the recipe; unchanged steps are cached 2
Tag A label on the card; latest is just the default label 2
Volume / bind mount The pantry shelf that survives the dish; a bind mount is your own counter-top 3
Network / published port Corridor between kitchens; a published port is a serving hatch to the street 3
Registry The cookbook library you push to and pull from 4
Multi-stage build Build in the workshop, ship from the shop floor 4
Compose The dinner menu: every dish in one file, cooked in order 5
Orchestrator Docker cooks one meal; Kubernetes runs the whole city 6
Cluster / Node The city / a building 7
Control plane City hall: API server = front desk, etcd = records room, scheduler = housing office, controllers = inspectors 7
kubelet The building's caretaker: the only one who runs containers 7
Pod An apartment: one address, usually one tenant 8
Deployment / ReplicaSet The building manager who keeps N apartments occupied 9
Label / selector Name tags on doors, and the manager's list 9
Service A phone number for a group of apartments that never changes 10
ClusterIP / NodePort / LoadBalancer Inside only / a door on every building / the cloud's public number 10
Ingress The city's front gate with a signpost by hostname 10
ConfigMap Sticky notes on the fridge 11
Secret The locked drawer (base64, not encryption) 11
Liveness / readiness / startup probe Are you awake? / taking visitors? / still moving in? 12
Requests / limits The lease: what you are promised, the ceiling you cannot exceed 12
PV / PVC / StorageClass The storage unit / the rental request / the storage company 13
StatefulSet A Deployment for tenants with names and their own units 13
Namespace A district 14
RBAC (Role, RoleBinding, ServiceAccount) Key cards 14
HPA Hiring temps in rush hour 14
Job / CronJob / DaemonSet One-off contractor / scheduled cleaner / one caretaker per building 14
Pending / ImagePull / CrashLoop / OOMKilled Ask the scheduler / the registry / the previous log / the lease 15
Helm chart / values / release Flat-pack box / options sheet / the assembled piece in your flat 16
Kustomize The same blueprint with an overlay per district 16
Service mesh / sidecar A concierge at every door: checks ID, keeps the log, sends one in ten next door 17
Gateway / VirtualService / DestinationRule The mesh's gate / the signpost with weights / the named groups it points at 17
mTLS Every visitor shows ID, both ways 17
Prometheus / Grafana The sensor network that polls on a timer / the wall of screens 18
PromQL's three queries Request rate, error rate, p95 latency, each by (version) 18
GitOps / Argo CD The city plan in the archive is the law; the inspector makes the city match it 19
ENTRYPOINT vs CMD The dish the kitchen always makes, and the default order if nobody asks 2
Init container The helper who arrives before the tenant moves in 8
NetworkPolicy Which corridors connect which districts 14
cordon / drain "No new tenants" / "everyone out, politely, one at a time" Landlord's Week
PodDisruptionBudget The minimum number of tenants who must stay in during works Landlord's Week
ResourceQuota / LimitRange The district's building code: a ceiling for the district, a default lease per tenant Landlord's Week
OCI The standard for recipe cards and ovens that any kitchen can use Old Maps
CRI / containerd The oven the caretaker actually uses Old Maps
dockershim The adapter that let the caretaker use Docker's oven; retired Old Maps
Docker Swarm The other city, the one that lost Old Maps
Tiller Helm 2's butler who lived in city hall and was let go Old Maps
Gateway API The new standard front gate Old Maps

The timeline

Everything Old Maps tells, on one line.

2013Docker ships2014Kubernetes announced2015OCI foundedKubernetes 1.02016Docker Swarm mode2017Compose v3 withdeployHelm 2 and Tiller2018Kubernetes in DockerDesktopCoreDNS becomesdefault2019extensions/v1beta1removed (1.16)Helm 3 drops Tiller2013 to 2019: containers become a standard, one city wins
2020dockershimdeprecated (1.20)Docker Hub rate limits2022dockershim removed(1.24)PodSecurityPolicyremoved (1.25)2023BuildKit default(Docker 23)Gateway API GA2024Istio ambient mode GA2025Endpoints deprecated(1.33)Helm 42020 to 2025: the addresses move, the shape stays

Old words

What you will meet in inherited repositories and older tutorials, and what it became.

You will see It became When Where in the book
docker-compose (with a hyphen) docker compose, the Go plugin 2020 onward; v1 retired 2023 Ch 5, Old Maps
version: "2" / "3" at the top of a Compose file Nothing; the key is obsolete and ignored Compose spec, 2020 Ch 5, Old Maps
--link redis:redis A user-defined network; names resolve on it Docker 1.9, 2015 Ch 3
MAINTAINER, ENV key value LABEL, ENV key=value Docker 1.13, 2017 Ch 2
Legacy docker build output (Step 3/8) BuildKit, docker buildx Default in Docker 23, 2023 Ch 4
docker stack deploy, deploy: in Compose Docker Swarm; Kubernetes won Swarm mode 2016; Compose ignores deploy: Ch 6, Old Maps
"Kubernetes is dropping Docker" dockershim removed; images unchanged, containerd runs them Deprecated 1.20 (2020), removed 1.24 (2022) Ch 7, Old Maps
master node, node-role.kubernetes.io/master control-plane Label removed 1.24, taint 1.25 Ch 7, Old Maps
kubectl get componentstatuses (cs) Nothing; deprecated 1.19, 2020 Old Maps
kind: ReplicationController Deployment (with a ReplicaSet) Deployments GA 1.9, 2017 Ch 9
apiVersion: extensions/v1beta1 for Deployments, DaemonSets apps/v1, with a mandatory selector Removed 1.16, 2019 Ch 9, Old Maps
kubectl run nginx --image=nginx creating a Deployment Creates a Pod; use kubectl create deployment Generators removed 1.18, 2020 Ch 8
kubernetes.io/ingress.class annotation, serviceName/servicePort ingressClassName, backend.service, networking.k8s.io/v1 v1 in 1.19, v1beta1 removed 1.22 (2021) Ch 10, Old Maps
Endpoints EndpointSlices Deprecation warning since 1.33, 2025 Ch 10
Heapster metrics-server Retired 1.13, 2018 Ch 12
kube-dns CoreDNS Default since 1.13, 2018 Ch 10, Old Maps
PodSecurityPolicy Pod Security Admission (a namespace label) Removed 1.25, 2022 Ch 14, Old Maps
autoscaling/v2beta2 HPA, batch/v1beta1 CronJob autoscaling/v2, batch/v1 Removed 1.26 and 1.25 Ch 14
A token Secret for every ServiceAccount Created on request only (TokenRequest) 1.24, 2022 Ch 14
Tiller, helm init, helm install --name Helm 3: no server, helm install <name> Helm 3, 2019; Helm 4, 2025 Ch 16, Old Maps
requirements.yaml in a chart dependencies: in Chart.yaml Helm 3, 2019 Old Maps
The stable Helm repository Per-project repositories; Artifact Hub to find them Archived 2020 Ch 16
Istio Mixer, istio-telemetry, istio-policy Telemetry in the proxy; istiod Deprecated 1.5 (2020), removed 1.8 Ch 17, Old Maps
Sidecar-only Istio Ambient mode, sidecars optional GA in 1.24, 2024 Ch 17
kube-prometheus vs Prometheus Operator vs kube-prometheus-stack One Helm chart installing the operator and the manifests Chart, 2020 onward Ch 18

The cheat card

Fifteen commands for a normal day. Everything else is in the cheat sheet below.

kubectl config current-context                 # which city am I in?
kubectl get pods -o wide                       # who is running, where
kubectl get deploy,svc,ingress                 # the shape of the district
kubectl describe pod <pod>                     # the story; events at the bottom
kubectl logs -f <pod>                          # what it says
kubectl logs <pod> --previous                  # what it said before it crashed
kubectl exec -it <pod> -- sh                   # get inside
kubectl apply -f <file>                        # make the file true
kubectl rollout status deploy/<name>           # is the change finished?
kubectl rollout undo deploy/<name>             # put it back
kubectl port-forward svc/<name> 8000:80        # private tunnel
kubectl get events --sort-by=.lastTimestamp    # what just happened here
kubectl top pod                                # who is heavy right now
docker compose up -d                           # the whole menu, locally
docker build -t <name>:<tag> .                 # cook the recipe

The cheat sheet

Every command in the book, once, grouped by what you are trying to do. The number at the end is the chapter that introduced it.

See what is running

docker ps                                          # running containers                                1
docker ps -a                                       # ...including stopped ones                         1
docker images                                      # local images                                      1
docker compose ps                                  # services in this project, with health            5
kubectl config current-context                     # which cluster am I on?                            7
kubectl config get-contexts                        # all clusters in my kubeconfig                     7
kubectl config use-context kind-pantry             # switch cluster                                    7
kubectl cluster-info                               # is the API server answering?                      7
kubectl get nodes -o wide                          # the buildings                                     7
kubectl get pods                                   # Pods in this namespace                            8
kubectl get pods -o wide                           # ...with IP and node                               8
kubectl get pods -w                                # ...and watch changes                              8
kubectl get pods -A                                # ...in every namespace                             7
kubectl get pods -n pantry-staging                 # ...in one namespace                               14
kubectl get pods -l app=pantry                     # ...matching a label                               9
kubectl get pods --show-labels                     # labels on everything                              9
kubectl get deploy,rs,pods -l app=pantry           # the three layers of a Deployment                  9
kubectl get svc                                    # Services                                          10
kubectl get endpoints pantry                       # the Pod IPs behind a Service                      10
kubectl get ingress                                # front gates                                       10
kubectl get configmap,secret                       # sticky notes and drawers                          11
kubectl get storageclass                           # storage companies                                 13
kubectl get pvc,pv                                 # claims and volumes                                13
kubectl get namespaces                             # districts                                         14
kubectl get hpa -w                                 # autoscalers, live                                 14
kubectl get cronjob                                # scheduled cleaners                                14
kubectl get ds                                     # caretakers                                        14
kubectl top pod                                    # CPU/memory per Pod now (needs metrics-server)     12
kubectl top node                                   # CPU/memory per node now                           12
helm list                                          # installed releases                                16
helm status pantry-helm                            # one release                                       16
helm get values pantry-helm                        # what a release was installed with                 16
istioctl proxy-status                              # which Pods have sidecars, and are they in sync    17
argocd app list                                    # Applications                                      19
argocd app get pantry                              # one Application: sync and health                  19

Build and ship

docker build -t pantry:v1 .                                        # cook the recipe                    2
docker build -t pantry:v2 --build-arg PANTRY_VERSION=v2 .          # ...with a build argument           2
docker build -t pantry:slim -f Dockerfile.multistage .             # ...from a different file           4
docker build -t pantry:builder --target builder -f Dockerfile.multistage .   # ...stopping at a stage   4
docker tag pantry:v1 yourname/pantry:v1                            # name it for a registry             2
docker history pantry:v1                                           # layers and their sizes             2
docker login                                                       # sign in to Docker Hub (or ghcr.io) 4
docker push yourname/pantry:v1                                     # upload                             4
docker pull yourname/pantry:v1                                     # download                           4
kind load docker-image pantry:v1 pantry:v2 --name pantry           # copy images into the local cluster 7
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts   # add a chart repo   16
helm repo update                                                   # refresh chart lists                16
helm search repo prometheus-community/kube-prometheus-stack        # find a chart                       16
helm template pantry-helm ./chart                                  # render a chart without installing  16
kubectl kustomize kustomize/overlays/prod                          # render an overlay                  16

Run and change

docker run -d --name redis -p 6379:6379 redis:7-alpine             # background, named, port published  1
docker run -it --rm python:3.12-slim python -c 'print(1)'          # interactive, thrown away on exit   1
docker run -d --name redis --network pantry-net -v pantry-data:/data redis:7-alpine   # on a network, with a volume   3
docker run -d --restart unless-stopped --name pantry -p 8000:8080 pantry:v1   # restart if it dies      6
docker stop redis                                                  # SIGTERM, then SIGKILL              1
docker rm redis                                                    # throw the dish away                1
docker rmi pantry:latest                                           # discard a recipe card              2
docker volume create pantry-data                                   # a shelf                            3
docker volume ls                                                   # shelves                            3
docker volume rm pantry-data                                       # remove a shelf                     3
docker network create pantry-net                                   # a corridor                         3
docker network ls                                                  # corridors                          3
docker compose up -d                                               # the whole menu, in the background  5
docker compose up -d --build                                       # ...rebuilding images first         5
docker compose build                                               # only build                         5
docker compose down                                                # stop and remove (keeps volumes)    5
docker compose down -v                                             # ...and delete volumes              5
kubectl apply -f k8s/pod.yaml                                      # make the file true (idempotent)    8
kubectl apply -k kustomize/overlays/dev                            # same, from a Kustomize overlay     16
kubectl delete -f k8s/pod.yaml                                     # remove what the file describes     8
kubectl delete pod pantry                                          # remove by name                     8
kubectl run curl --rm -it --image=curlimages/curl:8.11.1 -- sh     # a throwaway Pod                    8
kubectl create deployment pantry --image=pantry:v1 --dry-run=client -o yaml   # a correct skeleton to edit   8
kubectl create namespace pantry-staging                            # a district                         14
kubectl create configmap pantry-config --from-literal=GREETING=Hi  # sticky notes from the CLI          11
kubectl create secret generic pantry-secret --from-literal=API_TOKEN=x   # a drawer from the CLI        11
kubectl scale deploy/pantry --replicas=5                           # more or fewer copies now           9
kubectl set image deploy/pantry web=pantry:v2                      # new version now                    9
kubectl rollout status deploy/pantry                               # watch a rollout finish             9
kubectl rollout history deploy/pantry                              # revisions                          9
kubectl rollout undo deploy/pantry                                 # back one revision                  9
kubectl rollout restart deploy/pantry                              # bounce every Pod gracefully        9
kubectl label pod <name> tier=web                                  # tag it                             9
kubectl patch configmap pantry-config -p '{"data":{"GREETING":"Hi"}}'   # change one field in place     11
kubectl port-forward pod/pantry 8000:8080                          # tunnel to a Pod                    8
kubectl port-forward svc/pantry 8000:80                            # tunnel to a Service                10
helm install pantry-helm ./chart                                   # assemble                           16
helm upgrade --install pantry-helm ./chart --set image.tag=v2      # reassemble (or install if absent)  16
helm rollback pantry-helm 1                                        # put the previous revision back     16
helm uninstall pantry-helm                                         # remove the release                 16
istioctl install --set profile=demo -y                             # install the mesh                   17
kubectl label namespace default istio-injection=enabled            # concierges for every new Pod here  17
kubectl apply -f istio/                                            # gateway, subsets, weights, mTLS    17
kubectl apply -n argocd --server-side -f <argo-cd install.yaml>   # install Argo CD (CRD too big for client-side)   19
kubectl apply -f argocd/application.yaml                           # hand Argo CD the plan              19
argocd login localhost:8443 --username admin --insecure            # CLI sign-in through a tunnel       19
argocd app create pantry --repo <url> --path examples/pantry/gitops --dest-server https://kubernetes.default.svc --dest-namespace default   # same, from the CLI   19
argocd app sync pantry                                             # sync now                           19
argocd app rollback pantry <id>                                    # back to an earlier sync            19
argocd app delete pantry                                           # forget the Application             19
kubectl diff -f gitops/                                            # what would change if I applied this?   19

Get inside and find out why it broke

docker logs -f redis                                               # follow a container's output        1
docker exec -it redis redis-cli ping                               # run a command inside               1
docker compose logs -f web                                         # follow one service's output        5
docker compose exec redis redis-cli GET visits                     # run a command in a service         5
docker inspect pantry                                              # everything about a container       3
docker stats --no-stream                                           # CPU and memory per container now   3
docker cp pantry:/app/app.py ./app-from-container.py               # copy files out (or in)             3
docker compose config                                              # the effective Compose file         5
docker events --since 30s --filter container=pantry                # what happened to it                6
kubectl describe pod pantry                                        # the story; Events at the bottom    8
kubectl describe node pantry-worker                                # a building, and what it has left   7
kubectl describe ingress pantry                                    # which hosts route where            10
kubectl get networkpolicy                                          # which corridors are open           14
kubectl describe quota staging-ceiling -n pantry-staging           # how much of the ceiling is used    Landlord's Week
kubectl get pdb                                                    # budgets, and disruptions allowed   Landlord's Week
kubectl apply --dry-run=server -f old-maps/deployment-v1beta1.yaml # does the cluster accept this file? Old Maps
kubectl api-resources --api-group=extensions                       # which addresses does it serve?     Old Maps
kubectl logs -f pantry                                             # follow a Pod's output              8
kubectl logs <pod> --previous                                      # output of the crashed attempt      15
kubectl exec -it pantry -- sh                                      # a shell inside                     8
kubectl exec deploy/pantry -- sh -c 'echo $GREETING'               # a command in a current Pod         11
kubectl get pod pantry -o yaml                                     # observed state, every field        8
kubectl get events --sort-by=.lastTimestamp                        # the district's timeline            15
kubectl explain pod.spec.containers                                # the manual                         7
kubectl debug -it <pod> --image=busybox:1.36 --target=web          # tools beside a bare Pod            15
kubectl cp redis-0:/data/appendonlydir ./redis-backup              # copy files out (or in)             13
kubectl auth can-i list pods -n pantry-staging --as=system:serviceaccount:pantry-staging:pod-reader   # permission question   14
istioctl analyze                                                   # is my mesh config sane?            17
argocd app diff pantry                                             # what differs from Git              19
argocd app history pantry                                          # what was synced, when              19
curl -s --data-urlencode 'query=<promql>' http://localhost:9090/api/v1/query   # ask Prometheus         18

Maintain the buildings

kubectl cordon pantry-worker                                       # no new tenants                     Landlord's Week
kubectl drain pantry-worker --ignore-daemonsets --delete-emptydir-data   # everyone out, politely       Landlord's Week
kubectl uncordon pantry-worker                                     # reopen                             Landlord's Week
kubectl taint nodes pantry-control-plane node-role.kubernetes.io/control-plane:NoSchedule-   # open city hall to tenants   Landlord's Week
kubectl apply -f k8s/pdb.yaml                                      # the minimum who must stay in       Landlord's Week
kubectl apply -f k8s/quota.yaml -f k8s/limitrange.yaml             # the district's building code       Landlord's Week
helm lint old-maps/chart-helm2                                     # how old is this chart?             Old Maps

Watch it, and clean up

kubectl port-forward -n monitoring svc/monitoring-kube-prometheus-prometheus 9090:9090   # Prometheus UI   18
kubectl port-forward -n monitoring svc/monitoring-grafana 3000:80                        # Grafana UI      18
istioctl dashboard kiali                                           # the mesh as a graph                17
docker system prune                                                # disk back                          4
kind create cluster --config examples/kind-config.yaml             # a new city                         7
kind delete cluster --name pantry                                  # the whole city, gone               7
helm uninstall monitoring -n monitoring                            # remove the monitoring stack        18
kubectl delete -f istio/ && istioctl uninstall --purge -y          # remove the mesh                    17
kubectl delete -n argocd -f https://raw.githubusercontent.com/argoproj/argo-cd/v3.5.3/manifests/install.yaml   # remove Argo CD   19

Answers

One line per Recall question, by chapter. Attempt the questions first.

It Works On My Machine

  1. An image is a read-only recipe; a container is a running instance cooked from it.
  2. docker images lists images; docker ps lists running containers.
  3. --rm deletes the container when it exits; not when you want its logs or files afterwards.
  4. Yes, the image stays until docker rmi.
  5. A dish cooked from the recipe card.

Write the Recipe Down

  1. So a code change does not invalidate the cached dependency layer.
  2. RUN executes at build time; CMD is what runs when a container starts.
  3. The exec form, and only if process 1 handles the signal; a shell at PID 1 swallows it.
  4. The tag used when none is given; nothing about freshness.
  5. One step of the recipe, cached if unchanged.

Memory and Neighbours

  1. In Docker-managed storage outside any container; it survives docker rm.
  2. Only user-defined networks have DNS for container names.
  3. Laptop port 8000 forwards to container port 8080; the right-hand side is the container.
  4. docker inspect.
  5. A serving hatch to the street.

Ship It

  1. Registry, namespace, repository, tag; registry defaults to Docker Hub, tag to latest.
  2. Only the final stage ships; the builder's 1.6 GB of tools stays behind.
  3. docker push; the name must include your registry namespace.
  4. Nothing in the image needs root, and a compromised process should not have it.
  5. The cookbook library.

The Whole Band

  1. Its hostname on the project network.
  2. Without the condition it waits for the container to start; with it, for the healthcheck to pass.
  3. docker compose down keeps volumes; down -v deletes them.
  4. The directory name.
  5. The dinner menu.

The Night It Fell Over

  1. Self-heal across machines, scale across machines, roll out without downtime.
  2. docker kill counts as you stopping it; /crash was the process dying on its own.
  3. What you want to be true; the orchestrator loops until reality matches it.
  4. No; Kubernetes runs the same images.
  5. Docker cooks one meal; Kubernetes runs the whole city.

Meet the City

  1. API server (front desk), etcd (records), scheduler (housing office), controller manager (inspectors).
  2. The kubelet on each node.
  3. kubectl config current-context.
  4. Self-healing: nothing commands anything; every part notices and acts, forever.
  5. A building, and the records room.

The Smallest Thing That Runs

  1. apiVersion, kind, metadata, spec; the cluster adds status.
  2. apply compares the file with the cluster and changes only differences.
  3. kubectl create deployment … --dry-run=client -o yaml.
  4. That it will not start until the init container has exited successfully.
  5. An apartment.

Someone Who Never Sleeps

  1. A ReplicaSet keeps N Pods alive; a Deployment manages ReplicaSets to roll out changes.
  2. By label selector.
  3. maxSurge: 1, maxUnavailable: 0.
  4. kubectl rollout undo deploy/<name>.
  5. The building manager who never sleeps.

A Number That Never Changes

  1. A stable name and IP for whichever Pods match a selector; the Service object outlives the Pods.
  2. kubectl get endpoints <svc>; empty means the selector matches nothing.
  3. Inside the cluster; any node's IP on a high port; the internet via a cloud load balancer.
  4. An Ingress controller; without one the rules do nothing.
  5. The phone number and the front gate.

Sticky Notes and the Locked Drawer

  1. Handling: Secrets can be encrypted at rest, gated by RBAC separately, and hidden by tooling. The data is just base64.
  2. Mounted files.
  3. kubectl get secret … -o jsonpath='{.data.KEY}' | base64 -d.
  4. Configuration moved out of the image into ConfigMaps and Secrets.
  5. The locked drawer.

Are You Alive?

  1. Readiness failure removes the Pod from the Service; liveness failure restarts the container.
  2. /livez for liveness (process only), /healthz for readiness (needs Redis), so a Redis outage does not restart Pantry.
  3. Requests to place Pods; limits to throttle CPU and kill on memory.
  4. kubectl top pod.
  5. The lease: what is promised, and the ceiling.

Memory That Survives, Again

  1. A PVC requests storage, a StorageClass provisions a PV to satisfy it, the Pod mounts the PVC.
  2. WaitForFirstConsumer binding; it binds when a Pod uses it. Not a fault.
  3. A stable name and its own claim that follows it.
  4. kubectl get pvc,pv.
  5. The rental request for a storage unit.

Districts, Key Cards and Rush Hour

  1. Names, quotas and permissions.
  2. kubectl auth can-i delete pods -n <ns> --as=system:serviceaccount:<ns>:<sa>.
  3. CPU requests on the Pods and metrics-server in the cluster.
  4. A NetworkPolicy.
  5. A one-off contractor, a scheduled cleaner, one caretaker per building.

The Night It Fell Over, Again

  1. describe for Pending, ImagePullBackOff and OOMKilled; logs --previous for CrashLoopBackOff.
  2. At the bottom, oldest first.
  3. The manager replaces anything you patch on a Pod.
  4. kubectl get events --sort-by=.lastTimestamp.
  5. The scheduler, the registry, the previous log, the lease.

Old Maps

  1. The image and runtime formats; the image never depended on Docker, only the adapter did.
  2. A Compose file with a deploy: block for Swarm; a Deployment.
  3. 1.16; the selector.
  4. Helm 2's in-cluster server; Helm 3 removed it.
  5. The adapter that let the caretaker use Docker's oven, now retired.

The Landlord's Week

  1. cordon blocks new Pods; drain cordons and evicts the existing ones.
  2. The replacement had nowhere to go, so the PodDisruptionBudget refused the second eviction.
  3. Nothing moves Pods on its own; a rollout restart replaces them where there is room.
  4. Requests; a Pod without requests counts as zero, so the LimitRange supplies defaults.
  5. "No new tenants", "everyone out politely", and the minimum who must stay in.

Flat-Pack Furniture

  1. The box, the assembled piece in your flat, the options sheet.
  2. helm template.
  3. helm rollback <release> <revision>.
  4. When you own all the YAML and environments differ by a few lines.
  5. helm get values <release>.

A Concierge at Every Door

  1. mTLS, retries and timeouts, telemetry, traffic splitting; none in application code.
  2. Gateway is the door, VirtualService the signpost with weights, DestinationRule the named subsets.
  3. The weight values, in the VirtualService.
  4. It has no sidecar, and STRICT mode refuses plaintext.
  5. The concierge at the door.

The Control Room

  1. It scrapes /metrics on a timer; a dead target stops appearing.
  2. Request rate, error rate, p95 latency, each per version.
  3. The Service's app label and a named port; the ServiceMonitor's release label.
  4. kubectl top for now; Prometheus for over time.
  5. The sensor network and the wall of screens.

Nobody Touches the City by Hand

  1. Git becomes the only way; a merge deploys and a revert rolls back.
  2. Repository, path, revision, destination cluster, destination namespace.
  3. prune deletes what Git removed; selfHeal reverts changes made by hand.
  4. git revert and let Argo CD sync.
  5. The city plan in the archive and the inspector who enforces it.

Where to go next

The book stopped where a laptop stops. The next steps, one pointer each:

  • A real cluster. EKS, GKE and AKS run the control plane for you; the concepts are identical, the Ingress and StorageClass names change. Start with your cloud's "create cluster" quickstart, then kubectl config use-context into it and re-run chapter 8.
  • Operators. A controller you install that understands a specific application (a database, a certificate authority) and manages it from a custom resource. Prometheus Operator in chapter 18 was one. Try cert-manager next: it turns "I want a TLS certificate for this Ingress" into a resource.
  • CI that builds images. The book pushed images by hand. GitHub Actions or GitLab CI should docker build, docker push, and bump the tag in gitops/ on every merge; Argo CD does the rest. Search for "GitOps image updater".
  • Deeper Istio. Ambient mode, retries and timeouts, authorization policies, multi-cluster. The Istio docs' "Tasks" section is task-shaped like this book.
  • Alerting. Prometheus's alertmanager (switched off in chapter 18 to save memory) turns the error-rate query into a page. Turn it on and write one rule: v2 error rate above 5% for 5 minutes.
  • Reading the source of truth. kubectl explain covers every field; the Kubernetes docs' "Concepts" section is the long version of Act II, and now you know the map.

Thank you for reading. Delete the cluster when you are done:

kind delete cluster --name pantry