Aryan Tripathi — Writing
← All writing

October 2, 2026 · 56 min read

How a Kubernetes Cluster Actually Works: From kubectl apply to Packets on the Wire

Control loops, etcd, watches, the scheduler, the kubelet, Services and kube-proxy, traced end to end through one Deployment, with failure timelines for what happens when a node, the control plane or etcd breaks.

#kubernetes#architecture#system-design#controllers#networking#etcd

Most explanations of Kubernetes stop at three sentences: pods are ephemeral and have no permanent IP, nodes are machines that run pods, and you declare what you want and Kubernetes makes it so.

Those sentences are true enough to pass a quiz. They are not enough to debug a rollout that drops requests, explain why a pod on a dead node keeps receiving traffic for almost a minute, or decide whether a controller you are about to write is safe. For that you need the machinery under the sentences: what actually stores the declaration, who notices it, who changes what, and how a request finally reaches a process.

This article traces one Deployment through the whole cluster. The question behind every section is:

What actually happens between kubectl apply and a request reaching a pod, and what keeps it true when things break?

A note on versions. This describes current upstream Kubernetes (1.3x). Defaults mentioned here are upstream defaults. Managed services (EKS, GKE, AKS) and distributions change many of them, and they hide most of the control plane from you. Where behavior depends on version or implementation, it is marked.

Three claims to replace

The common claimWhat is actually true
"A pod has no permanent IP."Correct, and it is the least interesting consequence. A pod IP lives exactly as long as one pod sandbox's network namespace. The real question is what gives clients a stable target, and the answer is a virtual IP that is not assigned to any machine (section 11).
"A node is a VM with a permanent IP that runs the pods."A node is any machine running a kubelet. Its identity is a Node object and a name, not an IP, and its IP can change. The node doesn't decide to run a pod. A kubelet runs a pod because a scheduler wrote that node's name into the pod's spec and the kubelet noticed (sections 8 and 9).
"Kubernetes reconciles declared state, like Terraform."Both are declarative, but the mechanisms differ in ways that change how systems fail. Terraform runs to completion when you ask. Kubernetes runs continuous, level-triggered control loops against a shared database (section 6).

A Contents button sits in the bottom corner the whole way down: use it to jump to any of the 15 sections below, or to see how far you are.


1. The mental model: a database with a lot of workers

The one idea

Kubernetes is a database of desired and observed state, plus many independent processes that read it and write it. Everything else follows from that.

  • The database holds objects (Pods, Deployments, Nodes, Services) as JSON or protobuf documents with a spec (what someone wants) and a status (what a controller observed).
  • A single front door, the API server, is the only component that talks to the storage layer (etcd).
  • Every other component is a client of the API server. The scheduler, the controllers, the kubelets, kube-proxy and kubectl are all clients. To coordinate, they never call each other. (The API server itself does call out in a few places: to admission webhooks, and to kubelets for kubectl logs, exec and port-forward. None of that is coordinating desired state.)
  • Coordination happens through the data. The scheduler doesn't tell a kubelet to start a pod. It writes a node name into a Pod object, and the kubelet, which is watching for Pods assigned to its node, sees the change.
§1

Every line between the control plane and the outside world ends at kube-apiserver. There are no lines between the scheduler, controllers and kubelets, and that absence is the design. Any of those components can crash, restart or be replaced, and all that happens is that it re-reads the data and continues.

Control plane vs data plane

Control planeData plane
What it isAPI server, etcd, scheduler, controller-managerkubelet, container runtime, CNI plugin, kube-proxy, and your application's pods
JobDecide, record and reconcile desired stateRun processes and move packets
If it stopsNo changes are possible: no new pods, no rollouts, no scaling, no replacement of failed podsExisting pods keep running and existing Service rules keep forwarding

This distinction explains most outage behavior (section 13). An API server outage is not an application outage, as long as nothing else breaks while it lasts.

Under the Hood In kubeadm-style clusters the control plane components run as static pods: the kubelet on a control-plane machine reads their manifests from a directory on disk, with no API server needed, which solves the bootstrapping paradox of "the API server must run before pods can be scheduled." On managed services (EKS, GKE, AKS) the provider runs the control plane for you, and you never see the machines it runs on.

Common Misconception "The master/control node runs my workloads." Control-plane machines can run pods, but they are usually tainted to repel them. Neither the scheduler nor the controllers are in the data path of any request your application serves.

What you should remember

  • Kubernetes is a database (etcd) behind one API, with independent clients that coordinate only through data.
  • Components don't call each other. They watch objects and write objects.
  • Control plane = decisions. Data plane = running processes and packets. They fail independently.

2. The lab: container-lab on Kubernetes

We continue with the same application as in the container internals article: container-lab, the small Flask service. One small change for Kubernetes: a health endpoint, GET /healthz, that returns ok without touching the visit log. Build and push it as username/container-lab:v2.

@app.get("/healthz")
def healthz():
    return "ok"

Everything we trace in this article comes from two objects.

deployment.yaml:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: container-lab
spec:
  replicas: 3
  selector:
    matchLabels:
      app: container-lab
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  template:
    metadata:
      labels:
        app: container-lab
    spec:
      terminationGracePeriodSeconds: 30
      containers:
        - name: app
          image: username/container-lab:v2
          ports:
            - containerPort: 8000
          readinessProbe:
            httpGet: { path: /healthz, port: 8000 }
            periodSeconds: 2
          resources:
            requests: { cpu: 100m, memory: 128Mi }
            limits: { memory: 256Mi }
          lifecycle:
            preStop:
              exec: { command: ["sleep", "5"] }

service.yaml:

apiVersion: v1
kind: Service
metadata:
  name: container-lab
spec:
  selector:
    app: container-lab
  ports:
    - port: 80
      targetPort: 8000
kubectl apply -f deployment.yaml -f service.yaml

A word on that command. kubectl apply doesn't create pods, schedule anything or talk to a node. It sends the objects to the API server, which stores them. Everything after that is other components reacting to what was stored. The following sections follow those reactions in order.

What you should remember

  • kubectl apply writes desired state and returns. It performs no work itself.
  • Our lab has two objects: a Deployment (what to run) and a Service (how to reach it).
  • Every later step is a different component noticing a change and writing a different object.

3. The API server: the only door

Why it exists

If every component read and wrote etcd directly, every component would need etcd credentials, would have to understand storage details, and could corrupt shared state. The API server centralizes four jobs: authentication and authorization, validation, a consistent versioned API, and change notification (watches). It is also horizontally scalable, because it keeps no state of its own. You can run several replicas behind a load balancer, all active.

What happens to one request

§3

The order is fixed: authentication, then authorization, then mutating admission, then schema validation, then validating admission, then the write. Only a request that passes all of them reaches etcd.

  • Authentication establishes an identity (client certificates, bearer tokens, OIDC). Kubelets authenticate with their own certificates.
  • Authorization asks whether that identity may perform this verb on this resource. RBAC is the usual mechanism. A special Node authorizer plus the NodeRestriction admission plugin limits each kubelet to reading and modifying only objects related to its own node, so a compromised node can't rewrite the cluster.
  • Admission can change objects (defaulting, sidecar injection) or reject them (policy, quotas, Pod Security). Validating policies can be written in CEL without running a webhook server.
  • The response includes the stored object with a new resourceVersion.

spec, status, and who owns what

Every object follows a contract:

  • spec is desired state, written by users or higher-level controllers.
  • status is observed state, written by the controller responsible for the object. For a Deployment that is the replica counts and conditions. For a Pod it is mostly written by the kubelet.
  • metadata.generation increments when spec changes, and controllers record the generation they have processed in status.observedGeneration, so you can tell whether a status reflects your latest spec.

status is usually a separate subresource with separate permissions. A user changing spec and a controller reporting status don't collide.

Optimistic concurrency

Many writers update the same objects, and nothing takes a lock. Instead, every object carries a resourceVersion. An update that includes a resourceVersion succeeds only if it still matches the stored one. Otherwise the API server returns 409 Conflict, and the writer re-reads and retries. This is optimistic concurrency control, the same idea as compare-and-swap, and nearly every controller is written around it.

Clients must pass resourceVersion back to the server unmodified. For optimistic concurrency, equality is all you need. Since Kubernetes 1.35, conformance also requires that versions of the same resource type be orderable as monotonically increasing integers, but that is a guarantee for comparing versions, not an invitation to parse them for anything else.

Other pieces of the object model you will meet

  • ownerReferences: a child records its parent (a Pod lists its ReplicaSet). They drive ownership queries and cascading deletion by the garbage collector (section 7).
  • finalizers: keys that block actual deletion until a controller removes them. Deleting an object with finalizers sets deletionTimestamp and waits, which is how controllers get to clean up external resources first.
  • labels and selectors: loose, query-based coupling. A Service doesn't contain pods. It selects them.

Try It Yourself

kubectl get deploy container-lab -o yaml | grep -E 'resourceVersion|generation|observedGeneration'
kubectl get --raw /readyz?verbose | head          # API server health checks
kubectl api-resources | head                      # the object kinds the API serves
kubectl auth can-i create pods --as=system:serviceaccount:default:default

Interview Insight "How does Kubernetes avoid two controllers overwriting each other?" Optimistic concurrency with resourceVersion plus single ownership: each field of each object is meant to have one responsible writer (user for spec, controller for status). Server-side apply formalizes this with field managers, which track which writer owns which field and flag conflicts.

What you should remember

  • The API server is a stateless front door: authenticate, authorize, mutate, validate, then store.
  • spec is desired, status is observed, and different writers own them.
  • resourceVersion gives lock-free updates: a stale write gets a 409 and retries.
  • Ownership references, finalizers and label selectors are how objects relate without direct calls.

4. etcd: where the truth lives

What it is

etcd is a distributed, strongly consistent key-value store built on the Raft consensus protocol. The API server stores each object under a key (conceptually /registry/<kind>/<namespace>/<name>) and nothing else in a healthy cluster keeps authoritative copies.

Raft in one paragraph

An etcd cluster elects one leader. Writes go to the leader, which appends them to its log and replicates the entry to followers. A write is committed once a majority (quorum) has stored it. That is why etcd clusters have odd sizes:

MembersQuorumFailures tolerated
110
321
532
431 (no better than 3)

If the leader fails, the followers elect a new one, and writes pause during that election. If a majority is lost, etcd cannot commit writes, which means the API server can no longer accept changes. Keeping the majority healthy is the first rule of operating a cluster, and the reason managed services run etcd for you.

Properties that shape everything above it

  • Every change gets a monotonically increasing revision. The API server derives resourceVersion from it. That is what makes consistent lists, resumable watches and optimistic concurrency possible.
  • etcd supports watches natively. A client can ask for "every change after revision N", and the API server builds the Kubernetes watch API on that.
  • History is compacted. Old revisions are discarded periodically. A client that asks to resume from a revision that has been compacted away gets 410 Gone and must re-list (section 5).
  • Writes go through disk and the network. etcd is latency-sensitive, particularly to disk fsync latency. That is why it prefers fast local SSDs and why overloaded etcd shows up as a slow, flaky whole cluster.

The watch cache: why the API server isn't just a proxy

If every controller and kubelet opened its own watch against etcd, etcd would drown. The API server keeps an in-memory watch cache per resource type, fed by a single watch on etcd, and serves most client watches and many reads from it. Thousands of clients watch the API server. A few watches reach etcd.

Under the Hood This is also why the API server's consistency modes matter. A list with no resourceVersion is a quorum read from etcd, which is fully up to date and expensive. A list with resourceVersion=0 may be served from the watch cache, which is fast but possibly slightly stale. Client libraries have traditionally used the cheaper path for their initial sync and then relied on watches. Newer versions can instead stream the initial state through a watch.

Common Misconception "Kubernetes keeps the state of my pods in etcd, so the pods survive an etcd outage because etcd holds them." The reverse: etcd holds the records. The processes live on nodes. If etcd goes down, running pods continue running, because kubelets don't need etcd (or even the API server) to keep existing containers alive. What you lose is the ability to change anything.

What you should remember

  • etcd is a Raft-replicated KV store. Writes commit with a majority, so use 3 or 5 members and protect quorum.
  • Revisions in etcd become resourceVersion values and make watches and optimistic concurrency possible.
  • The API server's watch cache shields etcd from thousands of watching clients.
  • Lose etcd quorum and the cluster is frozen: no change can be committed, though running workloads continue.

5. Watches, informers and work queues

How a component notices a change

Polling the API server every second from thousands of kubelets and controllers would be wasteful and slow. Instead, clients use list-then-watch:

  1. List all objects of a kind once and remember the list's resourceVersion.
  2. Watch from that resourceVersion: the server holds an HTTP connection open and streams ADDED, MODIFIED, DELETED events (plus periodic bookmarks, which advance the client's position without carrying data).
  3. If the connection drops, resume from the last seen resourceVersion. If that revision has been compacted, the server returns 410 Gone, and the client re-lists and starts over.

Controllers don't implement this by hand. The standard client libraries wrap it in an informer.

§5

The reflector keeps a local copy of the world, so a controller's reads cost nothing and don't touch the API server. Event handlers don't hand the worker the change. They enqueue only the key of the object that changed. The worker then reads the current state from the cache. Reads are local and writes are rare, which is how a controller manager scales to tens of thousands of objects.

Why the queue holds keys, not events

Suppose a Deployment changes five times in 100 ms. The queue has one entry for that key, because duplicate keys are collapsed. The worker runs reconcile("default/container-lab") once, against the latest state. Intermediate states are skipped, and nothing is lost, because the worker doesn't need the history, only the present. That leads to the central idea.

Level-triggered, not edge-triggered

  • An edge-triggered system reacts to events: "a pod died." Miss one event (a crash, a dropped connection, a restart) and the system is wrong forever.
  • A level-triggered system reacts to state: "I want 3 pods and I count 2." It doesn't matter how the state came to be. If the controller was down for an hour, it wakes up, counts, and corrects.

Kubernetes controllers are built to be level-triggered. Events from watches are only a hint that something about an object changed. The decision always comes from comparing the current desired state to the current observed state. On startup, after a network partition, or after a 410 Gone, the controller just re-lists and reconciles everything. There is no "replay the missed events" logic, because there is nothing to replay.

Deep Dive: resync Informers can also periodically re-deliver every cached object to handlers (a resync), even without changes. Level-triggered logic should make this harmless, and it acts as a safety net for controllers whose conclusion depends on something outside the watched objects.

Try It Yourself

kubectl get pods -w &                       # a watch: a long-lived HTTP stream
kubectl delete pod -l app=container-lab --wait=false
kubectl get --raw '/api/v1/namespaces/default/pods?watch=1' | head -c 400

-w is the same list-then-watch protocol used by every controller. Notice the stream of events, including the replacement pod appearing without you doing anything.

What you should remember

  • Components learn about change through list-then-watch streams, not polling.
  • An informer keeps a local cache, so controllers read for free and write only when needed.
  • Work queues hold keys and deduplicate them, so a burst of changes becomes one reconcile against the latest state.
  • Level-triggered logic compares desired and observed state every time, so missed events and restarts are harmless.

6. Reconciliation, and why it is not Terraform

The control loop

A controller is a loop that moves observed state toward desired state:

# simplified pseudocode, not actual controller source
reconcile(key):
    obj = cache.get(key)
    if obj is missing:                 # deleted: nothing left to do (or run cleanup)
        return OK
    observed = observe(obj)            # e.g. list owned ReplicaSets and their pods
    if observed == obj.spec:           # already converged
        update obj.status; return OK
    make ONE step toward obj.spec      # create a pod, delete a pod, change a count
    update obj.status
    return REQUEUE or ERROR            # on failure, retry later with backoff

Three properties make this robust:

  • Idempotent: running it twice with no change in the world does nothing the second time.
  • Convergent: each pass reduces the difference between desired and observed state. It doesn't need to finish in one pass.
  • Stateless where it matters: its correctness doesn't depend on private memory of what it did last time. The cluster's objects are the memory.

Error messages like ImagePullBackOff and CrashLoopBackOff are this loop at work: the controller (here the kubelet) retries with exponential backoff and shows the state it is in. Nothing is "stuck". The loop is running, and the world isn't cooperating yet.

Versus Terraform

If you have used Terraform, the word declarative feels familiar, and the resemblance is real but shallow:

TerraformKubernetes
DeclarationHCL files, plus a state fileObjects stored in etcd. The spec is desired state
Who actsYou, when you run plan and applyControllers, continuously
Lifetime of the processStarts, converges, exitsRuns forever
Drift detectionAt the next planImmediately, through the watch
What reality is compared withA state file you must keep consistent with reality (and refresh)Live observation, every time
Plan stepExplicit: a diff you reviewNone in the loop. kubectl diff and server-side dry run exist, but the controllers just act
Concurrency controlState lockingOptimistic concurrency: no locks, retry on conflict
When it can't convergeThe run fails and you re-runThe loop retries with backoff indefinitely
Failure mode of "someone changed it by hand"Overwritten at next apply, or an errorReverted in seconds if a controller owns that field
Blast radius of a mistakeOne applyImmediate and automatic, including mistakes: a bad spec is converged on as faithfully as a good one

Two practical consequences:

  1. kubectl apply is not the reconciliation. It only records intent. Reconciliation is the controllers' ongoing work, and it is why deleting a pod "doesn't work": the ReplicaSet controller sees the count drop and recreates it. You declared three replicas, and that declaration still holds.
  2. Mixing the two layers is common. Many teams use Terraform to create the cluster and its cloud resources, and then Argo CD or Flux to apply Git-stored manifests to it. A GitOps controller is itself a Kubernetes-style reconcile loop (compare Git to the cluster, converge), layered on top of the built-in loops.

Common Misconception "Imperative commands bypass the declarative model." kubectl scale, kubectl set image and kubectl create all just edit objects in the same database. The cluster is declarative internally whichever way you type the change. The difference is only whether your intent is kept in a file in Git.

Interview Insight "Why do Kubernetes controllers use level-triggered logic instead of reacting to events?" Because networks drop messages and processes restart. If correctness depended on seeing every event, any missed event would cause permanent divergence. By always recomputing from current state, a controller is correct after any outage without needing event replay.

What you should remember

  • A controller is an idempotent loop that compares observed to desired state and takes one step toward convergence.
  • Unlike Terraform, Kubernetes doesn't run to completion. It runs continuously and notices drift immediately.
  • apply records intent. Controllers do the work. A bad declaration is enforced as faithfully as a good one.
  • "Backoff" states (CrashLoopBackOff, ImagePullBackOff) mean the loop is running and retrying, not that it gave up.

7. The controller chain: Deployment to ReplicaSet to Pod

Controllers that each do one small thing

You created one Deployment. Three replicas of a pod came out. No single component did that. A chain of controllers did, each responsible for one relationship and each observing only the object one level below it.

§7

Every arrow labeled "watch event" is a controller noticing, not being told. Every arrow into the API server is a write. Read the chain once and notice that no component names the next one. The Deployment controller doesn't know a scheduler exists.

Ownership

The objects form a tree, linked by ownerReferences:

§7
  • The Deployment controller manages ReplicaSets. A change to the pod template (a new image, an env var) produces a new ReplicaSet, identified by a hash of the template (pod-template-hash). A change to just replicas scales the existing one.
  • The ReplicaSet controller has the simplest job in Kubernetes: keep N pods matching its selector. Count, create or delete.
  • Pods are the leaves. A ReplicaSet doesn't care which node they run on.
  • Deleting the Deployment deletes the tree beneath it, through the garbage collector, a controller that follows ownerReferences.

A rolling update is two ReplicaSets trading replicas

With maxSurge: 1 and maxUnavailable: 0 and three replicas, changing the image to v3 creates a second ReplicaSet and shifts replicas between the two:

StepNew RS (v3)Old RS (v2)Total podsReady pods
Start0333
Surge1 (starting)343
New pod Ready12 (scaling down)3–43
Repeat…22→13–4≥ 3
End3033

At no point do fewer than three pods count as available, because the Deployment controller only scales the old ReplicaSet down when enough new pods are available (Ready, and for at least minReadySeconds, which defaults to 0). "Ready" is the condition reported by the kubelet from the readiness probe. That makes the readiness probe a load-bearing part of rollout safety, not decoration.

Common Misconception "Deleting a pod restarts it." Deleting a Pod object ends that pod forever. The ReplicaSet controller then notices that it has two pods instead of three and creates a new pod, with a new name, a new IP, and possibly on a different node. Compare that with a container crashing inside a pod, which the kubelet restarts in place with the same pod IP and the same sandbox (section 9).

Try It Yourself

kubectl get rs,pods -l app=container-lab -o wide
kubectl get pod <pod> -o jsonpath='{.metadata.ownerReferences[0].kind}{"\n"}'   # ReplicaSet
kubectl set image deploy/container-lab app=username/container-lab:v3
kubectl get rs -l app=container-lab -w                                          # watch the two RS trade replicas
kubectl rollout history deploy/container-lab

What you should remember

  • One Deployment becomes pods through a chain: Deployment → ReplicaSet → Pod, each created by a different controller.
  • Controllers coordinate through objects and ownerReferences, not calls.
  • A rolling update is two ReplicaSets exchanging replicas, gated by pod Readiness.
  • Pod identity is disposable: a deleted pod is replaced by a new pod. A crashed container is restarted inside the same pod.

8. The scheduler: a controller that writes one field

What it does, and what it doesn't

The scheduler watches for Pods whose spec.nodeName is empty and, for each, chooses a node. That decision is recorded by creating a Binding, which sets spec.nodeName on the pod. That is all. The scheduler never contacts the node, never pulls an image, and never starts anything. A pod is "scheduled" the moment a field has a value.

The scheduling cycle

§8
  • Filter: hard constraints. A node that fails any is out.
  • Score: soft preferences among the survivors. The highest total wins, with ties broken randomly.
  • Reserve then bind: the scheduler assumes the placement in its own cache immediately so the next pod's decision sees it. It then writes the binding asynchronously, which is why it can place many pods per second.
  • Pods that fit nowhere stay Pending with a FailedScheduling event, and the scheduler retries when something relevant changes (a node added, a pod removed).

The inputs to the decision are the pod's requests, not its limits and not actual usage. A node with 4 CPUs of allocatable capacity can accept pods whose CPU requests sum to 4, even if they are idle, and will refuse a pod whose request doesn't fit even if the node is nearly idle. Requests are a scheduling reservation (CPU requests also become cgroup CPU weights, as in the container article). Limits are enforced by the kernel at run time.

Under the Hood The scheduler is a framework with plugin extension points (queue sort, pre-filter, filter, score, reserve, permit, bind, and so on). Features such as topology spread, affinity and preemption are plugins. You can run additional schedulers and select one per pod with schedulerName, because "scheduler" is just a client that writes nodeName. Cluster autoscalers use the same signal: pods stuck Pending with FailedScheduling are the trigger to add nodes.

Interview Insight "A pod is Pending. Walk me through it." Is nodeName set? No: the scheduler hasn't placed it. Read the FailedScheduling event (insufficient CPU or memory requests, untolerated taint, unsatisfiable affinity or topology spread, unbound PVC). Yes: the scheduler is done, and the problem is on the node. Look at ContainerCreating, ImagePullBackOff, CNI errors or volume mounts, which are the kubelet's side of the story.

What you should remember

  • The scheduler's entire output is one field: spec.nodeName.
  • Filter removes impossible nodes, score ranks the rest, and the placement is assumed in cache, then bound.
  • Scheduling uses requests, not limits or live usage.
  • "Scheduled" doesn't mean "running". The kubelet acts only after it notices the binding.

9. The node and the kubelet

What a node actually is

A node is any machine, virtual or physical, that runs:

  • a kubelet, the agent that makes pods real,
  • a container runtime that the kubelet drives through CRI (containerd or CRI-O, as in the container article),
  • a CNI plugin that gives pods network interfaces and IPs,
  • and typically a service dataplane, kube-proxy or a CNI that replaces it (for example eBPF-based networking).

That is the whole definition. It says nothing about a virtual machine, an IP address or permanence.

The cluster's view of a node is a Node object, a record in the same database as everything else. The kubelet creates it itself when it starts (self-registration), reports addresses, capacity and conditions in its status, and keeps it up to date.

  • Identity is the name, usually the hostname, not an IP. status.addresses lists InternalIP, ExternalIP and Hostname as observations. Cloud providers and autoscalers routinely replace machines, and addresses change. Nothing in Kubernetes should depend on a node IP staying fixed.
  • Nodes are replaceable. Autoscalers add and remove them. Upgrades commonly replace nodes with fresh ones instead of patching in place.
  • A Node object is not the machine. If a VM vanishes, its Node object lingers, marked NotReady, until something removes it. On cloud platforms the cloud-controller-manager checks with the provider and deletes the Node object for machines that no longer exist.

How a node proves it is alive

Each kubelet renews a Lease object in the kube-node-lease namespace. By default the lease lasts 40 seconds and is renewed every quarter of that, so every 10 seconds. A Lease is a tiny object (just a holder identity and a renew timestamp), so a heartbeat is cheap. Full Node status updates are sent when something changes, plus a slower periodic report (every 5 minutes by default). The node lifecycle controller in the controller-manager watches those leases. If a node's lease isn't renewed within the node-monitor-grace-period (50 seconds by default in current upstream), it marks the node's Ready condition Unknown and begins the failure handling in section 13.

§9

What the kubelet does

The kubelet watches the API server for Pods whose spec.nodeName equals its own node's name. That filtered watch is the way scheduled pods reach a node. (Static pods, read from manifest files on the node, are the exception.) Its sync loop then, for each pod:

  1. Admits the pod (is there room? do its requirements fit?).
  2. Calls the runtime over CRI: RunPodSandbox. The runtime creates the pod's network and IPC namespaces, with the CNI plugin attached to set up networking and assign the pod IP.
  3. Pulls images (CRI PullImage), as the pull process in the container article.
  4. Creates and starts each container (CreateContainer, StartContainer), which means containerd, a shim and runc, then namespaces, cgroups, overlay root and execve. The kubelet translates resources.limits into cgroup values.
  5. Mounts volumes (through CSI for persistent storage) before containers start.
  6. Runs probes (startup, liveness, readiness) on a timer.
  7. Reports status back to the API server: phase, pod IP, container states, the Ready condition.
  8. On deletion (deletionTimestamp set): runs the preStop hook, sends SIGTERM to containers, waits up to terminationGracePeriodSeconds, then SIGKILL.
§9

The kubelet is a controller like any other: observe actual state through the runtime, compare it to desired state from the API, converge. It is also the one component that works without the control plane. Once it has a pod's spec, it keeps the containers running, restarting them according to restartPolicy, even if the API server is unreachable.

Why the pod's IP lasts only as long as the pod

The IP belongs to the pod sandbox, which holds the network namespace (see the namespaces section of the container internals article). It is assigned by the CNI plugin from an address range (in many clusters, a podCIDR slice given to each node, though some CNIs allocate directly from the VPC network). The IP disappears with the sandbox.

EventSame pod object?Same pod IP?
Container crashes and the kubelet restarts itYesYes (same sandbox)
Pod deleted and replaced by its ReplicaSetNo (new pod)No
Pod evicted from a node and rescheduledNo (new pod)No
Rolling updateNo (new pods)No

So "pods have no permanent IP" is a consequence of the design, not its centerpiece. Pods are disposable by construction, and everything that needs a stable address is built one layer up.

Common Misconception "Kubernetes moves a pod to another node when the node fails." Kubernetes never moves a pod. Pods are bound to a node for life. When a node fails, the pod is deleted and a controller creates a different pod that the scheduler places elsewhere. This matters for anything with identity or state, and it is why StatefulSets and volumes exist.

Try It Yourself

kubectl get nodes -o wide
kubectl get node <node> -o jsonpath='{.status.addresses}{"\n"}'
kubectl get lease -n kube-node-lease
kubectl get lease -n kube-node-lease <node> -o jsonpath='{.spec.renewTime}{"\n"}'
kubectl get pods -o wide                        # pod IPs and nodes
kubectl describe pod <pod> | sed -n '/Events/,$p'

What you should remember

  • A node is any machine running a kubelet, a runtime and a CNI. Its identity is its Node object and name, not an IP.
  • The kubelet gets work through a watch on pods bound to its node, drives the runtime through CRI, and reports status back.
  • Liveness is a Lease renewal every ~10 s. About 50 s of silence marks the node unhealthy.
  • A pod IP lives as long as its sandbox. A restarted container keeps it, and a replaced pod doesn't.
  • Pods are never moved, only deleted and recreated elsewhere.

10. Three kinds of identity: pod, service, node

Putting the previous sections together, there are three different answers to "how do I refer to the thing that runs my code?", each with a different lifetime.

IdentityStable forBacked byUse it to…
Pod IPOne pod sandboxCNI-assigned addressDebug. Never hard-code
Service ClusterIP + DNS nameThe life of the Service objectA virtual IP programmed on every node, plus a CoreDNS recordReach a changing set of pods
StatefulSet pod DNS name (via a headless Service)The pod's ordinal, e.g. db-0, across rescheduling, while the IP changesDNS record updated to the new pod IPReach a specific replica
Node nameThe Node objectThe kubelet's registrationTarget scheduling, drain, taint. Don't rely on node IPs

A Service is the object that turns a changing set of pod IPs into a stable endpoint. Its machinery is a pair of controllers and a per-node agent:

§10

The EndpointSlice controller keeps the list of pod IPs for each Service up to date, and records whether each is ready. kube-proxy on each node watches Services and EndpointSlices and programs the kernel to forward traffic to ready endpoints. CoreDNS answers name lookups with the Service's virtual IP. None of them are on the request path of an individual packet except the kernel rules they installed.

This is why the readiness probe matters twice: it gates rollout progress (section 7) and it gates whether a pod receives traffic at all.

What you should remember

  • Three identities, three lifetimes: pod IP (one pod), Service virtual IP and DNS name (the Service), node name (the Node object).
  • A Service is an object plus two reconcilers: the EndpointSlice controller maintains the pod list, and kube-proxy turns it into kernel rules.
  • Readiness decides which pods are in the list that gets programmed.

11. Services at Layer 4: a virtual IP implemented in the kernel

This is the transport-layer part of the story. It explains how a packet addressed to a thing that doesn't exist ends up at a process that does.

The ClusterIP is not an address of anything

Run ip addr on any node: you will not find 10.96.0.50 in the usual iptables or nftables modes. No interface owns it and no process listens on it. In iptables mode the rules match only the Service's own protocol and port, so other traffic to that address, such as a ping, is generally not answered. It is a virtual IP: a value that exists only inside forwarding rules installed on every node. The Service's port is likewise virtual. Port 80 on the Service maps to port 8000 on the pods, and nothing listens on port 80 anywhere.

Why not DNS round robin?

Kubernetes could have given the Service name several A records, one per pod. It doesn't, because clients and resolvers cache DNS records beyond their TTL and some resolve only once at startup, so changes in the pod set would take effect unevenly and slowly. A virtual IP puts the decision in the packet path, where it can change instantly and for every client at once.

How the rules work

kube-proxy is a controller, not a proxy in the data path. (Despite its name, in current modes it doesn't forward packets itself.) It watches Services and EndpointSlices and programs the kernel's packet-processing machinery to do destination NAT (DNAT): rewrite the destination of packets sent to the virtual IP to one of the real pod IPs.

Mode (Linux)MechanismNotes
iptablesChains of netfilter rules, one set per Service and endpointLong the default, and still the upstream default at the time of writing. Rule lookup grows linearly with the number of Services and endpoints, and rewriting large rule sets slows down at scale
nftablesMaps and sets in nftablesGA in Kubernetes 1.33. More efficient updates and lookups at scale. Needs a recent kernel (5.13+)
ipvsKernel L4 load balancerDeprecated in 1.35, disabled by default from 1.40 and planned for removal in 1.43. Don't choose it for new clusters
CNI-provided (e.g. eBPF)The CNI replaces kube-proxyCilium and others implement Services in eBPF programs

For iptables and nftables, backend selection is random (by default) among the ready endpoints.

A packet's journey

Client pod A (on node 1) calls http://container-lab (port 80). The target pods run on node 2.

§11

Key observations:

  • The DNAT happens on the client's node, before the packet leaves it. No central load balancer is involved. Every node has all the rules, so the "load balancer" is distributed across the cluster.
  • conntrack is the stateful part. The first packet of a connection picks a backend, and the kernel records the translation. Every later packet in that connection (and the replies) follows the recorded entry, without re-evaluating the rules. A connection is therefore pinned to one backend for its lifetime.
  • Pod-to-pod routing across nodes is the CNI's job (overlay tunnels, routed networks or cloud VPC routing). Services sit above it and only rewrite addresses.
  • Traffic from a process on the node itself hits the same rules through the output path.

Illustrative commands to see it (Linux node, iptables mode; names and hashes will differ):

sudo iptables -t nat -S KUBE-SERVICES | grep container-lab
# -A KUBE-SERVICES -d 10.96.0.50/32 -p tcp -m tcp --dport 80 -j KUBE-SVC-XXXXXXXX

sudo iptables -t nat -S KUBE-SVC-XXXXXXXX
# -A KUBE-SVC-XXXXXXXX -m statistic --mode random --probability 0.33333 -j KUBE-SEP-AAAA
# -A KUBE-SVC-XXXXXXXX -m statistic --mode random --probability 0.50000 -j KUBE-SEP-BBBB
# -A KUBE-SVC-XXXXXXXX -j KUBE-SEP-CCCC

sudo iptables -t nat -S KUBE-SEP-BBBB
# -A KUBE-SEP-BBBB -p tcp -m tcp -j DNAT --to-destination 10.244.2.4:8000

sudo conntrack -L -d 10.96.0.50     # live connection mappings

The probabilities (1/3, then 1/2, then the remainder) are how a chain of random choices produces an equal split. In nftables mode the same logic is stored in maps, not in a linear list.

From outside the cluster

  • NodePort exposes a port (30000–32767 by default) on every node and DNATs it to the Service, so any node can accept traffic for any Service.
  • LoadBalancer adds a cloud load balancer in front, provisioned by the cloud-controller-manager, that targets the nodes (or the pods directly, depending on the provider).
  • externalTrafficPolicy: Local keeps traffic on the node that received it. This preserves the client's source IP and avoids an extra hop, but nodes without a local ready pod fail the load balancer's health check, so the LB stops sending them traffic.

The L4 consequences that bite in production

  • A Service balances connections, not requests. The decision is made once per connection, at its first packet. A client that opens one long-lived connection (gRPC, HTTP/2, database pools, keep-alive HTTP clients) will send all its traffic to one pod. Adding replicas doesn't help the connections that already exist. Solutions live above L4: client-side balancing (headless Services plus a smart client), a service mesh sidecar or an L7 proxy, or periodic connection recycling.
  • Nothing in the Service layer understands HTTP. No path routing, no retries, no header-based decisions. Those are the job of an Ingress or Gateway (L7), which is itself just pods behind a Service.
  • conntrack is finite. Each tracked connection uses table space. Heavy connection churn, and UDP in particular, can exhaust it, producing intermittent drops that look like random network failures. Watch nf_conntrack usage on busy nodes.
  • Source IPs change. After NAT, a backend sees the client pod's IP (pod-to-Service) but may see a node's IP for external traffic unless externalTrafficPolicy: Local or an LB preserving client addresses is in use.
  • Updates are not instantaneous. Between a pod becoming not ready (or terminating) and every node's rules updating there is a propagation delay of EndpointSlice change → watch → rule rewrite. Section 12 shows what that does to rollouts.

Under the Hood kube-proxy's own failure doesn't immediately break traffic. The rules it installed stay in the kernel, so existing and new connections to already-known endpoints keep working. What stops is updates: pods that appear or disappear aren't reflected, and traffic can be sent to dead endpoints. This is the control-plane / data-plane split again, one level down.

Interview Insight "A Service has three pods but one is getting 90% of the load. Why?" Almost always long-lived connections: the backend is chosen per connection, not per request. Check whether clients reuse connections (gRPC/HTTP2, pools). The fix is above L4 (client-side or L7 balancing), not "more replicas" or "a different kube-proxy mode".

What you should remember

  • A ClusterIP is a virtual IP that exists only as forwarding rules on every node, with no process and no interface owning it (in iptables/nftables modes). Only the Service's own protocol and port are matched.
  • kube-proxy is a controller that programs the kernel to DNAT virtual IPs to ready pod IPs. It's not on the data path.
  • Backend choice is made per connection, then remembered by conntrack. A Service is an L4 construct and balances connections, not requests.
  • The "load balancer" is distributed: every node holds every rule.
  • Rule updates lag behind pod changes, and that lag shapes rollout design.

12. Zero-downtime rollouts: where the loops meet the packets

A rolling update is the moment where the control-plane loops (section 7) and the packet rules (section 11) have to agree. When they don't, you see a burst of connection refused or connection reset during every deploy, even with maxUnavailable: 0.

The termination race

When an old pod is deleted, two independent things start at once:

  1. The kubelet runs the pod's preStop hook, then sends SIGTERM to the container.
  2. The EndpointSlice controller marks the endpoint not ready, and every node's kube-proxy (plus any Ingress controller or cloud load balancer) sees the change and updates its rules.

Nothing orders these. If the application exits immediately on SIGTERM, there is a window in which some nodes still send new connections to a pod that is no longer listening.

§12

The EndpointSlice records the pod's state with three conditions: serving (mirrors the pod's Ready condition), terminating (set as soon as the pod gets a deletion timestamp, usually before its containers exit), and ready (serving and not terminating). Proxies normally ignore terminating endpoints. If every endpoint of a Service is terminating, they may still send traffic to ones that are serving and terminating, so that traffic isn't lost mid-rollout.

Our lab Deployment has the standard mitigations built in. Each one maps to a mechanism you've now seen:

SettingWhat it doesMechanism
readinessProbe on /healthzA new pod receives traffic only after it is genuinely ready, and the old ReplicaSet scales down only after new pods are ReadyPod Ready condition → EndpointSlice readiness → kube-proxy rules (sections 7, 10)
maxUnavailable: 0, maxSurge: 1Capacity never dips below 3 ready pods during the rolloutDeployment controller's ReplicaSet arithmetic (section 7)
preStop: sleep 5Keeps the old pod serving while endpoint removals propagate to every nodeCloses the race: kubelet path is delayed so the endpoints path wins
terminationGracePeriodSeconds: 30Total time between the start of termination and SIGKILL, including preStopKubelet's termination sequence (section 9)

Two more things belong with these:

  • The application must handle SIGTERM. Our Flask service runs as PID 1 inside its container, and as the container article showed, the kernel doesn't apply the default SIGTERM action to PID 1. Without a handler (or an init such as tini in the image), the signal is ignored and the pod runs out the full grace period before SIGKILL: slow rollouts and dropped in-flight requests. A correct handler stops accepting new work, finishes in-flight requests and exits.
  • Anything in front of the Service has its own propagation delay. Cloud load balancers deregister targets with their own drain timers, and Ingress controllers watch endpoints like kube-proxy does. Budget for the slowest one.

Common Misconception "maxUnavailable: 0 means zero downtime." It only guarantees enough Ready pods exist. It says nothing about whether requests that were already routed to a terminating pod finish, or whether new requests are still being sent to it for a few seconds. Zero downtime comes from readiness, graceful termination and propagation delay working together.

What you should remember

  • Pod termination starts two unordered paths: the kubelet's (preStop, SIGTERM) and the endpoint propagation to every node's rules.
  • Close the race with a short preStop delay, graceful SIGTERM handling, and a terminationGracePeriodSeconds that covers both.
  • Readiness gates both rollout progress and traffic. A probe that passes too early causes errors on startup.
  • Account for load balancers and ingress controllers in front of the Service, since each has its own lag.

13. Failure timelines

Control loops explain how the cluster converges. The most useful test of whether you understand them is predicting what happens when a piece breaks and how long each reaction takes.

A container crashes or is OOM-killed

The kubelet notices through the runtime and restarts the container in the same pod (same IP, same node), subject to restartPolicy. Repeated failures trigger exponential backoff, historically starting at ten seconds and doubling to a five-minute cap (a lower cap can be configured in recent versions, behind a feature gate), and the pod shows CrashLoopBackOff while it waits. An out-of-memory kill happens in the kernel when the container exceeds its cgroup memory.max (the container article's exit code 137), and the kubelet reports OOMKilled as the reason. No control-plane component is involved, which means the control plane can be down and this still works.

A node dies

The most instructive case. The machine loses power at t = 0. The numbers below are upstream defaults. Managed services and distributions often change them.

§13

Read the timeline as three separate regimes:

WindowWhat happensWhat clients see
0 → ~50 sThe control plane doesn't know. Nobody is probing the pods: kube-proxy does not health-check endpoints, and the only prober (the kubelet) is deadConnections routed to the dead pod hang or time out, because rules still point at it
~50 s → ~6 minPods are marked NotReady and removed from endpoints, but not replaced. They still count as existing replicas, so the ReplicaSet doesn't reactThe Service works with reduced capacity (for example 2 of 3 pods, if one lived on the dead node)
~6 min onwardEviction deletes the pods, the ReplicaSet creates replacements, the scheduler places them and a kubelet starts themFull capacity returns

Two design lessons fall out:

  • Cluster-internal traffic has no active health checking. Detection of a silently dead node is bounded by the node monitor grace period. If that matters, put health checks (and retries) at a layer that has them: client-side, mesh, or L7 load balancer.
  • The five-minute toleration is a trade-off, not a mistake. A short delay reacts quickly but risks needless churn when a node just has a brief network hiccup, and, worse, if the node is actually alive but partitioned, the kubelet there would keep running the old pods while replacements start. Latency-sensitive workloads commonly shorten tolerationSeconds on node.kubernetes.io/unreachable and not-ready in the pod spec, accepting that trade.

What about stateful workloads? A Deployment replaces pods once the old ones are marked for deletion. A StatefulSet deliberately does not: its guarantee is at most one pod per identity, so it waits until the old pod is confirmed gone, and a pod on a dead, unreachable node can remain Terminating indefinitely. Kubernetes added a mechanism for declaring a node out of service after a non-graceful shutdown (the node.kubernetes.io/out-of-service taint) so that its pods and volume attachments can be cleaned up. Use it only when you are sure the machine is really off.

Planned maintenance is different. kubectl drain cordons the node (no new pods) and removes pods through the Eviction API, which respects PodDisruptionBudgets. A PDB protects against voluntary disruptions like drains. It doesn't help during a crash.

The control plane breaks

What failsWhat stopsWhat keeps working
One API server replica (of several)Nothing, if a load balancer routes around itEverything
All API serverskubectl, controllers, scheduling, status updates, all reconciliation, including node failure handlingRunning pods, kubelet container restarts, existing Service rules
Scheduler (one instance of a leader-elected set)After the lease expires, a standby takes over. Without a standby, new pods stay PendingEverything already scheduled
Controller-manager leaderA standby takes over after the lease expires (about 15 s with default leader election timings). Meanwhile, no scaling, rollouts or replacementRunning pods, existing endpoints
etcd quorum lostThe API server can't commit writes, so the cluster is frozen: no changes, no status updatesRunning pods, existing Service rules
kube-proxy on a nodeRule updates on that node. Existing rules keep forwardingTraffic to already-programmed endpoints
CoreDNSName resolution for Services. Connections by IP continueExisting connections, traffic by ClusterIP

The scheduler and controller-manager are leader-elected: several replicas may run, but only one is active at a time. The active one continuously renews a Lease object in the API server. If it stops, a standby acquires the lease and takes over. Liveness again is expressed as "keep writing a timestamp", the same trick the kubelet uses for nodes.

A control-plane outage has one more subtlety. When the API servers return after a long outage, every node's lease looks stale. The node lifecycle controller includes safeguards (rate-limited evictions, and special handling when a large fraction of a zone or cluster appears unhealthy at once) so that a control plane problem isn't mistaken for a mass node failure and doesn't cause a storm of evictions.

Interview Insight "If the API server goes down, do my pods die?" No: the kubelets keep running what they were given. But you can't deploy, scale or heal anything, and node failures go unhandled while it's down, because the controller that reacts to them is a client of the API server too. Recovery of the control plane restores the loop, not the missed history, which is again level-triggered design paying off.

What you should remember

  • A crashed container is restarted by the kubelet in place. Node failures take ~50 s to detect and ~6 min to replace with default timings.
  • During detection, nothing health-checks endpoints inside the cluster, and during replacement, capacity is reduced.
  • The control plane can fail without taking the data plane down, but while it's down, nothing heals.
  • Leader election with Leases lets standbys take over the scheduler and controller-manager.
  • StatefulSets trade availability for at-most-one safety. Don't force-delete pods or declare nodes dead unless you're sure.

14. The master architecture diagram

§14

Read it as three layers that touch only through the API server and the kernel:

  1. Declare and decide. Intent enters through the API server and lives in etcd. Controllers and the scheduler turn it into more objects, ending in a Pod bound to a node.
  2. Run. The kubelet notices the binding and drives the runtime and CNI, which end in the kernel primitives from the container article: namespaces, cgroups, overlay roots. It reports status back.
  3. Address. The EndpointSlice controller and kube-proxy convert "which pods are Ready" into kernel rules, so a virtual IP becomes a per-connection choice among real pod IPs.

Every box on the control-plane side can fail and recover independently, because all of them rebuild their view by watching the same database.


15. Design lessons Kubernetes teaches

Most of what makes Kubernetes work is a set of design choices you can reuse in your own systems:

  1. One source of truth, many stateless reconcilers. State lives in one consistent store. Everything else is a replaceable process that can be restarted at any time.
  2. Level-triggered over edge-triggered. Act on what is, not on what happened. Missed events then cost nothing.
  3. Coordinate through data, not calls. Components that don't know each other can't break each other. The cost is eventual consistency and propagation delays you must design for (section 12).
  4. Optimistic concurrency instead of locks. Version every object, reject stale writes, retry.
  5. Separate desired state from observed state. Different writers, different fields, and an explicit generation to tell whether the observation is up to date.
  6. Keep the data plane independent of the control plane. The part that serves traffic should keep working when the part that makes decisions is down.
  7. Liveness through leases. "Prove you're alive by renewing a timestamp" is cheap, and it is used for nodes and for leader election alike.
  8. Make detection time an explicit, tunable number. 50 seconds to notice, 300 seconds to evict: both are knobs, and both are a trade-off between speed and stability.
  9. Retry forever, with backoff. A loop that can't converge shouldn't die. It should wait and try again, and report why.
  10. Address by label and virtual identity, never by instance. Pods are cattle, Services are the stable handle, and selectors connect them.

Interview Insight If an interviewer asks you to design a system that "keeps N workers running across unreliable machines," you can answer with this skeleton: a durable store of desired state, an observer of actual state, an idempotent reconciler with backoff, heartbeats with leases for liveness, optimistic concurrency for writers, and a stable virtual endpoint in front. That is Kubernetes in six components, and you now know how each piece behaves under failure.


The whole journey in one paragraph

kubectl apply authenticates to the API server, which authorizes the request, runs admission and validation, and stores the Deployment in etcd through Raft consensus, producing a new resourceVersion. Informers in the controller-manager receive the change through a watch and enqueue its key. The Deployment controller creates a ReplicaSet, the ReplicaSet controller creates Pods, and each controller works from its cache against the current state, so a missed event or a restart changes nothing. The scheduler filters and scores nodes and writes spec.nodeName. The kubelet on that node, watching for pods bound to it, drives the runtime through CRI and the network through CNI. The runtime creates namespaces, cgroups and an overlay root, then execves your process, and the kubelet reports the pod's IP and readiness back. The EndpointSlice controller lists the ready pod IPs. kube-proxy, on every node, turns that list into kernel rules that DNAT a virtual Service IP to one pod per connection. If a node dies, the same loops notice through an expired Lease, remove its pods from endpoints, evict them, and recreate them elsewhere, and while the control plane is down, the data plane keeps running what it already has.

Further reading