Aryan Tripathi — Writing
← All writing

September 30, 2026 · 100 min read

Containers: From Source Code to Image Layers, SHA-256 Digests, OCI Artifacts, and Running Processes

What actually happens between docker build and a running container: BuildKit, OCI manifests, content-addressable blobs, registries, snapshotters, OverlayFS, namespaces, cgroups and runc, traced through one small application — with live, interactive diagrams for the parts you normally can't see.

#containers#docker#oci#linux#kubernetes#containerd

Containers: From Source Code to Image Layers, SHA-256 Digests, OCI Artifacts, and Running Processes

You type docker build -t container-lab:v1 ., then docker run, and a few seconds later a Python process is answering HTTP requests.

In between, your source code is turned into tar archives. Those archives are hashed into content identities and described by JSON documents that refer to each other by hash. They may be shipped to a registry and back, then unpacked into a union filesystem. Finally, the Linux kernel starts an ordinary process with a restricted view of the machine.

This article follows one small application through every one of those steps. The question behind every section is:

What actually happens between docker build and a running container?

By the end you should be able to explain what each part of the stack does, why it exists, and which of the many "image hashes" refers to what.

A note on versions. Container tooling changes quickly. This article describes the architecture current in late 2026: Docker Engine 29 (BuildKit is the default builder, and fresh installs use the containerd image store), the OCI Image and Distribution Specifications v1.1, and Linux with cgroup v2. Behavior that depends on the implementation, platform or version is marked as such. When in doubt, the specifications and your own docker version output are the source of truth.


A Contents button sits in the bottom corner the whole way down — use it to jump to any of the 24 sections below, or to see how far you are.


1. The mental model: what a container actually is

The definition

A container is one or more ordinary Linux processes that the kernel runs in a restricted execution context:

  • Namespaces control what the process can see: its own process tree, network stack, mount table, hostname and so on.
  • cgroups control how much the process can use: memory, CPU time, number of processes and I/O.
  • A root filesystem, usually assembled from read-only image layers plus a private writable layer, controls which files the process sees at /.
  • Privilege restrictions control what the process may do: a reduced set of Linux capabilities, a seccomp syscall filter, and an AppArmor or SELinux profile.

The kernel has no object called a "container". It has processes, and it has features that isolate and constrain them. A container runtime combines those features.

What a container is not

It is not…Because…
A virtual machineThere is no virtual hardware and no guest kernel. The process makes system calls directly into the host's kernel.
An imageAn image is immutable content at rest. A container is running (or stopped) state created from an image.
A file you can copyYou can export a container's filesystem, but its processes, namespaces and cgroup exist only in the kernel's memory.
Inherently a security sandboxIsolation is real but shares a kernel. Its strength depends on configuration and on the kernel's attack surface.

Six words people mix up

TermWhat it isWhere it lives
ImageAn immutable, content-addressed bundle: a manifest, a config, and an ordered list of filesystem layersRegistry, or a local content store
LayerOne filesystem changeset, stored as a (usually compressed) tar archiveA blob in a content store
ArtifactAny build output. In registries specifically, any content stored with the OCI manifest and blob model, including things that are not runnableCI systems, registries
ContainerRuntime state: metadata, a writable layer, and, while it is running, processes in namespaces and cgroupsThe host running it
Root filesystemThe merged directory tree the process sees at /A mount (usually OverlayFS) on the host
RuntimeThe software that turns an image plus configuration into isolated processes. Covered in section 14Host daemons and tools

Why "a lightweight VM" is useful but wrong

The analogy helps beginners because, from the inside, a container looks like a separate machine. It has its own hostname, IP address, filesystem and process list, and ps shows only its own processes.

That appearance is produced by the kernel limiting what the process can see. Nothing is emulated. Here is where the analogy breaks:

§1

The VM column shows a hosted arrangement. With a type-1 hypervisor (ESXi, Xen, Hyper-V) the hypervisor runs beneath any operating system, and KVM turns the Linux kernel itself into the hypervisor. Either way, each VM runs its own kernel, and that is the point that matters here.

In the VM stack, the application talks to a guest kernel, which drives virtual hardware provided by a hypervisor. In the container stack, the application's system calls go straight into the host kernel. The runtime does not sit on the execution path. It configures the kernel before the process starts and then gets out of the way. That is why the runtime is drawn off to the side and not between the kernel and the process.

The concrete consequences:

  • One kernel. uname -r inside a container prints the host's kernel version. The image supplies userspace files such as libc, Python and /etc, but never a kernel.
  • Visible from the host. On a Linux host, ps -ef shows container processes as ordinary processes.
  • Kernel compatibility matters. A Linux container needs a Linux kernel. An amd64 binary needs an amd64 CPU, or emulation.
  • A shared attack surface. A kernel vulnerability reachable from inside a container can affect the whole host. A VM puts a hypervisor between the guest and the host kernel.

Common Misconception "Containers always share the host kernel." That is true of standard Linux containers run by runc. The OCI Runtime Specification also allows other runtimes. gVisor intercepts system calls with a user-space kernel, and Kata Containers runs each pod inside a lightweight VM. Both still consume OCI images. The image format and the isolation technology are independent decisions.

An image is not a compressed container

An image contains no processes, memory, open sockets or runtime state. It is a template: an ordered set of filesystem changesets plus a JSON document that says "run python app.py in /app with these environment variables". You can create a thousand containers from one image. Each gets its own writable layer and processes, and the image never changes.

Try It Yourself (on a Linux host)

uname -r
docker run --rm alpine uname -r      # same kernel version
docker run -d --name sleeper alpine sleep 1000
ps -ef | grep "sleep 1000"           # the "container" is just a host process
docker rm -f sleeper

On macOS or Windows, both uname -r results inside containers show the kernel of Docker Desktop's Linux VM, not your host OS. Section 9 explains why.

What you should remember

  • A container is a normal process, or group of processes, that the kernel isolates with namespaces, limits with cgroups, and gives a root filesystem assembled from image layers.
  • An image is immutable content at rest. A container is runtime state created from it.
  • Standard Linux containers share the host kernel. The image never contains a kernel.
  • The runtime sets up isolation and then steps aside. It is not a layer the application's system calls pass through.

2. The lab application we will follow

Every section in this article traces the same image: container-lab:v1. The application is deliberately small, but it exposes the details we want to observe: its PID, hostname, user ID, and a file written to persistent storage.

Source code

app.py:

import os
import socket
from datetime import datetime, timezone

from flask import Flask

app = Flask(__name__)
DATA_DIR = os.environ.get("DATA_DIR", "/data")


@app.get("/")
def index():
    os.makedirs(DATA_DIR, exist_ok=True)
    visits_file = os.path.join(DATA_DIR, "visits.log")
    with open(visits_file, "a") as f:
        f.write(datetime.now(timezone.utc).isoformat() + "\n")
    with open(visits_file) as f:
        visits = sum(1 for _ in f)
    return {
        "message": "hello from container-lab",
        "hostname": socket.gethostname(),
        "pid": os.getpid(),
        "uid": os.getuid(),
        "visits": visits,
    }


if __name__ == "__main__":
    app.run(host="0.0.0.0", port=8000)

requirements.txt:

flask==3.0.3

Dockerfile

# syntax=docker/dockerfile:1
FROM python:3.12-slim
ENV PYTHONUNBUFFERED=1
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
EXPOSE 8000
CMD ["python", "app.py"]

.dockerignore

.git
.venv/
__pycache__/
*.pyc
.env
Dockerfile*
.dockerignore

Excluding Dockerfile* keeps the Dockerfile out of the COPY . . layer. BuildKit reads the Dockerfile through a separate channel, so ignoring it does not break the build.

Build and run

docker build -t container-lab:v1 .

docker run -d --name lab \
  -p 8080:8000 \
  -v lab-data:/data \
  --memory=512m --cpus=1 \
  container-lab:v1

curl http://localhost:8080/

Illustrative response (your hostname will differ):

{"hostname":"3f2c9a1b7d4e","message":"hello from container-lab","pid":1,"uid":0,"visits":1}

Three details in that output will matter later:

  • "pid": 1: the Python process believes it is the first process on the system (section 12).
  • "hostname": Docker sets the container's hostname to the short container ID (section 12).
  • "uid": 0: the process runs as root. Whether that is root on the host depends on user namespaces (sections 12 and 21).

What we will trace

StageWhat happens to container-lab:v1Section
BuildDockerfile becomes a build graph, then layers, a config and a manifest3, 4
IdentityEvery piece is hashed into a SHA-256 digest5, 6, 18
DistributionPushed to and pulled from a registry, blob by blob8
StorageKept as blobs plus unpacked snapshots9
ExecutionMounted with OverlayFS, isolated with namespaces, limited with cgroups, started by runc10–15
LifecycleStopped, restarted and removed, while its volume survives16, 17

What you should remember

  • One image, container-lab:v1, is used throughout. Every command in this article applies to it.
  • The Dockerfile has three kinds of instructions: base selection (FROM), filesystem changes (COPY, RUN) and metadata (ENV, EXPOSE, CMD).
  • The app prints its PID, hostname and UID so you can watch namespaces at work.

3. Inside docker build

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

Concept

docker build does not "run the Dockerfile top to bottom in a container". Since Docker Engine 23.0, the default builder on Linux is BuildKit (Docker Desktop switched earlier). BuildKit compiles the Dockerfile into a dependency graph, resolves as much of that graph as possible from cache, executes only what changed, and exports the result as an image.

The components

1. Docker CLI and Buildx. docker build is now an alias for docker buildx build. Buildx is the client. It connects to a builder, which by default is the BuildKit instance embedded in the Docker daemon. Buildx can also target a standalone buildkitd running in a container, in Kubernetes, or remotely.

2. The build context. The . in docker build . is the build context: the set of files the build may read. BuildKit and the legacy builder handle it very differently:

  • The legacy builder packed the entire context directory into a tarball and uploaded it to the daemon before doing anything else. That is the source of the familiar "Sending build context to Docker daemon 1.2GB" message.
  • BuildKit opens a session with the client and requests files on demand. It transfers only the paths that instructions such as COPY actually reference, and on later builds it transfers only files that changed. In both cases, .dockerignore removes paths from what the builder is allowed to see.

3. The frontend. The line # syntax=docker/dockerfile:1 tells BuildKit to fetch the Dockerfile frontend, the parser, as an image. That means you get current Dockerfile features without upgrading the daemon. Without the directive, the frontend built into your BuildKit version is used. The frontend converts the Dockerfile into LLB (low-level build definition). LLB is a directed acyclic graph of operations:

  • source ops: a base image, the local context, a Git repository, an HTTP URL
  • exec ops: RUN
  • file ops: COPY, ADD, and the directory creation performed by WORKDIR
  • metadata that will become the image config: ENV, CMD, EXPOSE, and so on

4. The solver. BuildKit walks the graph and computes a cache key for each vertex from its definition and the cache keys of its inputs. If a matching result already exists in the build cache, the step is marked CACHED and not executed. Independent branches, such as separate stages in a multi-stage build, run in parallel, and stages the final target doesn't need are skipped entirely.

5. Execution. A cache miss on RUN executes the command in a temporary container. BuildKit uses an OCI runtime (runc) or containerd as its worker, on top of a snapshot of the previous step's filesystem. The filesystem difference produced by the step is captured as a new snapshot.

6. The exporter. When the graph is solved, an exporter turns the final state into output. The default for docker build stores the image in Docker's image store and tags it. Other exporters push straight to a registry (--push), write an OCI image layout tarball (--output type=oci,dest=lab.tar), or dump the plain filesystem (--output type=local,dest=out). For an image, the exporter:

  1. Serializes each changed snapshot as a layer tar, compresses it (gzip by default, zstd optional) and stores it as a blob.
  2. Writes the image config JSON: environment, command, working directory, the list of uncompressed layer hashes, and history.
  3. Writes the image manifest JSON, which points to the config and the layers by digest.
  4. Optionally wraps the manifest in an image index. This is used for multi-platform builds and for attaching attestations such as provenance, which recent Buildx versions add by default in some configurations.
§3

Reading from top to bottom: the CLI gives the frontend the Dockerfile and gives the solver access to the context. The frontend produces a graph, and the solver checks each node against the cache and executes only the misses. The exporter then serializes the result into three kinds of blob that refer to each other by digest: layers and config are referenced from the manifest, and the manifest is optionally referenced from an index. The image store records a name pointing at the top-level blob.

That's the shape of the pipeline. Here's what "checks each node against the cache" actually does to our six-instruction Dockerfile, run three different ways:

Interactivebuildkit graph solving — same dockerfile, three builds

FROM

python:3.12-slim

queued

ENV

PYTHONUNBUFFERED=1

queued

WORKDIR

/app

queued

COPY

requirements.txt .

queued

RUN

pip install -r requirements.txt

queued

COPY

. .

queued

FINALIZE

EXPOSE 8000 · CMD

queued

Nothing exists yet, so every instruction runs and produces a new layer. Total time ≈ 8.3s of animation.

Legacy builder vs BuildKit

AspectLegacy builderBuildKit
Context transferFull tarball up frontOn demand, incremental, filtered
Execution modelLinear, one instruction at a timeGraph, with parallel independent branches
Intermediate resultsIntermediate images and containers (<none> images, "Removing intermediate container")Cache records and snapshots in BuildKit's own store, not images
Unused stagesBuilt anywaySkipped
Metadata instructionsEach one committed a new imageFolded into the final image config
ExtrasnoneCache mounts, secret mounts, SSH forwarding, remote cache import/export, multi-platform builds, attestations

Under the Hood The build cache and the image store are separate things. docker builder prune deletes cache records without touching any image, and docker image rm deletes an image without necessarily clearing the build cache. The cache can also hold content that never appears in any image, such as a pip download cache stored in RUN --mount=type=cache.

Interview Insight "Why is my build slow in CI but fast locally?" Your laptop has a warm build cache and a fresh CI runner has none. BuildKit can export and import cache to and from a registry or the CI provider's cache storage (--cache-to / --cache-from), which is how CI pipelines get cache hits.

What you should remember

  • BuildKit compiles the Dockerfile into a graph (LLB), computes cache keys, executes only cache misses, and exports the result.
  • The build context is what the builder is allowed to read. BuildKit fetches only what it needs. .dockerignore limits what it can see.
  • The output is not "a tarball". It is a set of content-addressed blobs (layers, config, manifest, and optionally an index) plus a name.
  • Build cache ≠ image layers. They are stored and pruned separately.

4. Image layers, instruction by instruction

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

What a layer is

A layer is a filesystem changeset: a tar archive containing every file and directory that was added or modified relative to the filesystem below it, plus special marker entries for deletions. In an image it is stored as a blob, usually gzip- or zstd-compressed.

Each layer has two identities:

  • The layer digest is the SHA-256 of the compressed blob. It appears in the manifest and is how registries address the blob.
  • The diff ID is the SHA-256 of the uncompressed tar. It appears in the image config's rootfs.diff_ids and is how the runtime verifies what it unpacked.

Layers are immutable because they are content-addressed: changing a single byte produces a different digest, which makes it a different layer. They are ordered, and a layer only makes sense on top of the layers below it. A deletion entry, for example, refers to a file that exists in a lower layer.

Deep Dive: chain IDs Because a layer's meaning depends on everything beneath it, runtimes identify an unpacked stack by a chain ID. The OCI Image Specification defines it recursively: ChainID(L0) = DiffID(L0) and ChainID(L0…Ln) = SHA256(ChainID(L0…Ln-1) + " " + DiffID(Ln)). Two images that share the same bottom three layers share the same chain ID for that three-layer stack, so a snapshotter can reuse one unpacked copy. containerd keys its snapshots this way.

Walking through our Dockerfile

For each instruction we answer five questions:

  1. Does it change the filesystem?
  2. Does it contribute metadata?
  3. What can be cached?
  4. What invalidates the cache?
  5. What ends up in the OCI image?

FROM python:3.12-slim

  1. It adds no new layer of its own. It inherits the base image's layers: a Debian slim root filesystem plus layers that install Python. The exact count varies between releases of the tag.
  2. It inherits the base config's Env (such as PATH, LANG and PYTHON_VERSION), its default Cmd (["python3"]) and more.
  3. BuildKit resolves the tag to a digest and caches that resolution.
  4. The cache is invalidated when the tag resolves to a different digest, for example when you build with --pull after the Python maintainers publish a patched image, or when you change the tag.
  5. The base layers' descriptors are copied into our manifest's layers array, and the base config fields are merged into our config.

ENV PYTHONUNBUFFERED=1

  1. No filesystem change.
  2. It adds PYTHONUNBUFFERED=1 to config.Env.
  3. Its cache key is the instruction plus its parent's key.
  4. Changing the value invalidates this step and every step after it, because later RUN commands execute with this environment.
  5. It becomes an entry in config.Env and a history entry marked "empty_layer": true.

WORKDIR /app

  1. It sets the working directory and, if /app doesn't exist, creates it. In BuildKit that creation can produce a tiny layer containing just the directory entry. Whether a separate layer appears is an implementation detail, so check docker history.
  2. It sets config.WorkingDir to /app.
  3. Its cache key is the instruction plus its parent's key.
  4. The cache is invalidated by changing the path or any earlier step.
  5. It becomes config.WorkingDir and possibly a small layer.

COPY requirements.txt .

  1. Yes. It adds a layer containing /app/requirements.txt.
  2. It adds only a history entry.
  3. The cache key is a checksum of the copied file's contents and relevant metadata (such as permissions), plus the destination, flags such as --chown/--chmod, and the parent's key.
  4. Editing requirements.txt invalidates it. Merely touching it (changing only its modification time) does not, because BuildKit's checksum ignores timestamps.
  5. It becomes one layer blob of a few kilobytes.

RUN pip install --no-cache-dir -r requirements.txt

  1. Yes. It adds a layer containing Flask and its dependencies under /usr/local/lib/python3.12/site-packages/ plus entry-point scripts such as /usr/local/bin/flask.
  2. It adds only a history entry.
  3. The cache key is the command string, the environment, the mounts, and the parent's key. BuildKit does not look at what the command downloads.
  4. Changing the command or any earlier step invalidates it. A new release of an unpinned dependency does not, because the cache is keyed on inputs, not on the outside world. This is why RUN apt-get update can stay stale for months in a warm cache.
  5. It becomes one layer blob, typically the largest one we add.

COPY . .

  1. Yes. It adds a layer with every context file that .dockerignore doesn't exclude: app.py and requirements.txt (again).
  2. It adds only a history entry.
  3. The cache key is checksums of all copied files.
  4. Any content change in any copied file invalidates it.
  5. It becomes one small layer blob.

EXPOSE 8000

  1. No filesystem change.
  2. It adds config.ExposedPorts: {"8000/tcp": {}}. This is documentation for tools and humans. It does not publish the port.
  3. and 4. Like ENV: the instruction plus its parent's key.
  4. It becomes a config field.

CMD ["python", "app.py"]

  1. No filesystem change.
  2. It sets config.Cmd, which replaces the base image's ["python3"].
  3. and 4. Like ENV: the instruction plus its parent's key.
  4. It becomes a config field.

The resulting image, from bottom to top:

§4

The arrows point upward because each layer is applied on top of the one below it. The config sits beside the stack, not in it. It is a separate JSON blob that lists the layers' diff IDs in order and carries all the metadata instructions.

Why the stack is a useful fiction

The "stack of boxes" picture is the right mental model for how the final filesystem is computed. It is not how bytes sit on disk:

  • In a registry or content store, layers are compressed tar blobs, not directories. Nothing is stacked. The manifest's ordered layers array is the only thing that defines the order.
  • On a host, each layer is unpacked once into a snapshot directory, identified by chain ID. OverlayFS then presents several snapshot directories as one merged tree (section 11).
  • The same unpacked base layers serve every container and image that shares them.
  • Some snapshotters (btrfs, ZFS, devmapper) don't use directories stacked by OverlayFS at all. Others, such as lazy-pulling snapshotters, fetch file contents on first access.

Reuse, deduplication and storage implications

  • Reuse: a hundred images built on the same python:3.12-slim digest share its base layers. They are stored once locally and once per registry backend, which usually deduplicates across repositories.
  • Deletions don't shrink images. RUN rm big-file in a later layer adds a deletion marker. The bytes still ship in the lower layer. Download and clean up in the same RUN, or use a multi-stage build and copy only what you need into the final stage.
  • Secrets leak through layers. A file copied in one step and deleted in the next remains retrievable from the earlier layer blob. Use RUN --mount=type=secret instead.
  • Layer order is cache order. Put the things that change least at the bottom (section 20).

Common Misconception "Every Dockerfile instruction creates a layer." With the legacy builder, every instruction created a new image in the parent chain, although metadata instructions produced empty layers. With BuildKit, only filesystem-changing operations produce layer blobs. Metadata instructions become config fields and history entries marked empty_layer.

What you should remember

  • A layer is a tar changeset of additions, modifications and deletion markers, stored as a compressed blob.
  • Each layer has two hashes: the compressed digest (in the manifest) and the uncompressed diff ID (in the config).
  • Only COPY, ADD, RUN and directory-creating WORKDIR change the filesystem. ENV, EXPOSE, CMD and similar only change the config.
  • Cache keys are computed from inputs: file checksums, command strings and parent keys. They never reflect what a command would download today.
  • Deleting a file in a later layer never removes its bytes from the image.

5. SHA-256 and content-addressable storage

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

This is the concept that holds the whole system together.

What SHA-256 is

SHA-256 is a cryptographic hash function. It takes any sequence of bytes and deterministically produces a 256-bit (32-byte) value, usually written as 64 hexadecimal characters. It is designed to have these properties:

  • Deterministic: the same bytes always produce the same hash.
  • Preimage resistant: given a hash, it is computationally infeasible to find input that produces it.
  • Collision resistant: it is computationally infeasible to find two different inputs with the same hash. No practical SHA-256 collision is known. SHA-1, by contrast, has had practical collisions since 2017.
  • Avalanche effect: changing one bit of input changes about half the output bits, in a pattern that can't be predicted.

Here are two real values computed with shasum -a 256. The inputs differ by one capital letter:

$ printf 'hello world\n' | shasum -a 256
a948904f2f0f479b8f8197694b30184b0d2ed1c1cd2a1ec0fb85d299a192a447  -

$ printf 'hello World\n' | shasum -a 256
0c23d0ceae909c42439cbb3069887888cb829f6c9d2c93966c944c65a6b6ed59  -
Interactiveone byte, new digest — sha-256, computed live in your browser

sha256(text)

sha256:

Change the text above and watch which characters move.

The two outputs are unrelated. Nothing about the second hash hints that the input was nearly identical.

Hash vs encryption

Hashing (SHA-256)Encryption (e.g. AES)
PurposeIdentity and integrityConfidentiality
Uses a key?NoYes
Reversible?No: you cannot recover the input from the hashYes, with the key
Output sizeFixed (256 bits)Proportional to the input
In containersEvery blob's identityOnly if you add it: TLS in transit, optional encrypted layers (ocicrypt)

A digest hides nothing. Anyone who can fetch the blob can read it. The digest proves which bytes you have, not that anyone was prevented from seeing them.

What exactly gets hashed

Always the exact bytes of a blob, as stored. That produces four distinct hashes for our image:

HashInput bytesWhere it appears
Layer digestThe compressed layer tarball (.tar.gz bytes)Manifest layers[].digest, registry blob URL
Diff IDThe uncompressed layer tarConfig rootfs.diff_ids[]
Config digestThe config JSON, byte for byteManifest config.digest
Manifest digestThe manifest JSON, byte for byteIndex entries, image@sha256:… references

Two consequences surprise people:

  • The same filesystem content can have different layer digests. Compressing an identical tar with a different gzip implementation or compression level produces different compressed bytes, and therefore a different layer digest, while the diff ID stays the same. The diff ID exists for exactly this reason.
  • JSON is hashed as bytes, not as data. Re-indenting a manifest or reordering its keys changes its digest even though it "means" the same thing. Registries must therefore serve manifests byte for byte as they were pushed, and clients must hash exactly what they received.

Digests form a hash tree

The manifest contains the config digest and every layer digest. So the manifest digest commits to everything underneath it: change one byte in any layer and that layer's digest changes, which changes the manifest bytes, which changes the manifest digest. An index commits to its manifests in the same way. This is a Merkle DAG, the same structure Git uses for commits and trees.

§5

Every arrow labeled "contains digest of" means the parent's bytes include the child's hash. So one digest at the top is enough to verify, transitively, every byte of the image.

The general pattern is simple:

§5

In an ordinary filesystem, a name (a path) points to content that can change. In a content-addressable store, the name is derived from the content, so the content behind a given name can never change.

What digests enable

  • Integrity verification. After downloading a blob, the client recomputes its SHA-256 and compares it with the digest it asked for. A mismatch, whether from corruption, a truncated download or a tampering proxy, is rejected.
  • Deduplication. If the local store already has sha256:X, it doesn't need to download sha256:X again, whichever image or repository refers to it.
  • Caching. Build cache keys and snapshot keys are also digests, so "have I done this before?" becomes a key lookup.
  • Addressing. Registries serve blobs at URLs such as GET /v2/<repository>/blobs/sha256:<hex>. The URL is the hash, so a registry can't serve different bytes under the same address without clients detecting it.

The OCI digest grammar is algorithm:encoded. sha256 is the algorithm used almost everywhere, and the specification also registers sha512.

Tags vs digests

ubuntu:latest                       # a tag: a mutable pointer
ubuntu@sha256:<64 hex characters>   # a digest: an immutable content identity

A tag is a name in a repository that the registry maps to a manifest or index digest. Whoever has push access can move it at any time. A digest reference names the content itself, so it can only ever resolve to those exact bytes, or to nothing if the content has been deleted.

Tags move routinely. Official images such as python:3.12-slim are rebuilt whenever the Debian base receives security fixes or a new Python patch release ships:

§5

The digests are illustrative. On Monday the tag points to one index. By Thursday the same tag points to a new one. A build that says FROM python:3.12-slim silently gets different bytes, while a build that says FROM python:3.12-slim@sha256:1111… still gets the old ones.

Common Misconception "The digest is a hash of the image." There is no single "the image" to hash. ubuntu@sha256:… is the digest of a specific manifest or index document. For multi-platform images that is usually the index, so the same digest resolves to different platform-specific manifests on an ARM laptop and on an x86 server, each fully verified.

Interview Insight If an interviewer asks "how does Docker know a layer is already present?", answer with content addressing. The client knows the digest from the manifest before downloading anything, and checks the local content store for that digest. When pushing, it asks the registry the same question with an HTTP HEAD request (section 8).

What you should remember

  • A digest is algorithm:hex of exact bytes. One changed byte gives an entirely different digest.
  • Hashing gives identity and integrity, not secrecy.
  • Manifests contain the digests of their config and layers, and indexes contain the digests of their manifests, so one top-level digest verifies everything.
  • A layer has a compressed digest (manifest) and an uncompressed diff ID (config).
  • Tags are mutable pointers. Digests are immutable identities.

6. OCI terminology: blobs, descriptors, manifests, indexes

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

The Open Container Initiative

The Open Container Initiative (OCI) is a Linux Foundation project founded in 2015 to standardize container formats so that images and runtimes are not tied to one vendor. Docker donated its image format (Docker Image Manifest V2, Schema 2) and its runtime, runc, as starting points. OCI maintains three specifications:

SpecificationStandardizesOur image's journey
Image SpecificationThe on-disk and in-registry format: manifests, indexes, configs, layers, descriptors, media typesWhat docker build produces
Distribution SpecificationThe HTTP API for pushing and pulling content to and from registries (/v2/...), including the referrers APIWhat docker push and docker pull speak
Runtime SpecificationThe bundle (a root filesystem plus config.json) and the lifecycle operations (create, start, kill, delete) an OCI runtime must implementWhat runc consumes

Versions 1.1 of the Image and Distribution specs (2024) added first-class support for non-image content: the artifactType and subject fields and the referrers API. These are covered in section 7.

Deep Dive: Docker vs OCI formats Docker's own media types (for example application/vnd.docker.distribution.manifest.v2+json) and the OCI media types are structurally almost identical, and every mainstream tool reads both. Which one your build emits depends on your builder, exporter and image store. BuildKit can emit either, and OCI media types are increasingly the default. When you see "Docker image", it is almost always either a Docker v2 schema 2 image or an OCI image, and runtimes treat them the same way.

The object vocabulary

Blob. An opaque sequence of bytes stored and retrieved by its digest: a compressed layer, a config JSON, a manifest JSON. The registry doesn't need to understand a blob's contents to store it.

Media type. A string such as application/vnd.oci.image.layer.v1.tar+gzip that tells a client how to interpret a blob's bytes. The same bytes with a different media type mean something different.

Digest. The content identity of a blob (see section 5).

Size. The blob's exact length in bytes. Clients use it to allocate space, to show progress, and as a cheap first check before hashing. A blob of the wrong size is rejected.

Descriptor. The small JSON object used whenever one piece of content points to another. Its required fields are mediaType, digest and size, and it can also carry annotations, urls, artifactType, platform (in indexes) and data (small content embedded inline). A descriptor is a typed, sized, verifiable pointer.

Image config. JSON that describes the image's runtime defaults and root filesystem: architecture, os, the config object (Env, Cmd, Entrypoint, WorkingDir, User, ExposedPorts, Labels, and so on), rootfs.diff_ids, and history.

Image manifest. JSON for one platform that points to one config and an ordered list of layers.

Image index. JSON that points to multiple manifests, each usually annotated with a platform. Docker's older name for this is manifest list.

Platform. An os/architecture pair, plus an optional variant (for example linux/arm64/v8 or linux/arm/v7) and, for Windows, os.version.

Artifact. Content stored with this same manifest and blob machinery that is not necessarily a runnable image. See section 7.

The manifest for container-lab:v1

Here is a simplified but structurally realistic manifest. The digests are truncated, and the sizes and base layer count are illustrative:

{
  "schemaVersion": 2,
  "mediaType": "application/vnd.oci.image.manifest.v1+json",
  "config": {
    "mediaType": "application/vnd.oci.image.config.v1+json",
    "digest": "sha256:7c1e0d…",
    "size": 6214
  },
  "layers": [
    { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:a2f4…", "size": 28230123 },
    { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:b81c…", "size": 3511467 },
    { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:c09d…", "size": 12814029 },
    { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:d5e7…", "size": 250 },
    { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:e3aa…", "size": 93 },
    { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:f41b…", "size": 206 },
    { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:0c7f…", "size": 1874212 },
    { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:19be…", "size": 812 }
  ],
  "annotations": {
    "org.opencontainers.image.created": "2026-09-25T10:00:00Z"
  }
}

Field by field:

  • schemaVersion: 2: required, and fixed at 2 for backward compatibility with Docker's schema 2. It does not track the spec version.
  • mediaType: declares that this document is an OCI image manifest. It is technically optional in the spec, but should always be set, because registries and clients use it to decide how to parse the document.
  • config: a descriptor for the image config blob. Its mediaType (application/vnd.oci.image.config.v1+json) is what makes this manifest a runnable container image. An artifact would use a different config media type or the empty descriptor (section 7).
    • digest: the SHA-256 of the config JSON bytes.
    • size: the config's length in bytes.
  • layers: an ordered array of layer descriptors. Index 0 is the base (bottom) layer. The runtime applies them in array order.
    • mediaType: how to decompress and interpret the layer (tar, tar+gzip, tar+zstd).
    • digest: the SHA-256 of the compressed bytes, the address to fetch from the registry.
    • size: the compressed size, which is what you download.
  • annotations: optional string key/value metadata. org.opencontainers.image.* keys are predefined, such as created, source, revision and licenses.

Reading the layers: the first three are the base image's (Debian rootfs and Python; the real count varies). The 93-byte one is plausibly the WORKDIR directory. Then come requirements.txt, the pip install and the application code.

The config

Also simplified and illustrative:

{
  "architecture": "arm64",
  "variant": "v8",
  "os": "linux",
  "config": {
    "Env": [
      "PATH=/usr/local/bin:/usr/local/sbin:/usr/sbin:/usr/bin:/sbin:/bin",
      "LANG=C.UTF-8",
      "PYTHON_VERSION=3.12.x",
      "PYTHONUNBUFFERED=1"
    ],
    "WorkingDir": "/app",
    "ExposedPorts": { "8000/tcp": {} },
    "Cmd": ["python", "app.py"]
  },
  "rootfs": {
    "type": "layers",
    "diff_ids": ["sha256:5d1f…", "sha256:6a90…", "…one per layer, same order…"]
  },
  "history": [
    { "created_by": "ENV PYTHONUNBUFFERED=1", "empty_layer": true },
    { "created_by": "WORKDIR /app" },
    { "created_by": "COPY requirements.txt . # buildkit" },
    { "created_by": "RUN /bin/sh -c pip install --no-cache-dir -r requirements.txt # buildkit" }
  ]
}
  • architecture/variant/os: the platform this image's binaries were built for.
  • config: the defaults the runtime uses to build the OCI runtime configuration. docker run flags such as -e and --workdir can override them.
  • rootfs.diff_ids: the uncompressed layer hashes, in the same order as the manifest's layers. After unpacking a layer, the runtime checks its diff ID against this list.
  • history: human-oriented provenance. It is the source of docker history's "CREATED BY" column. It is informational and not verified against the layers.

How the objects connect

§6

A tag names an index. The index holds one descriptor per platform manifest. Each manifest holds one config descriptor and an ordered list of layer descriptors. Every edge is a digest, so the whole graph can be verified from the top. Some builders also add attestation manifests to the index, which non-runtime consumers such as policy engines read.

Multi-platform images and Apple Silicon

An x86 server and an M-series Mac can both run docker pull python:3.12-slim because the tag resolves to an index. The client selects the manifest whose platform matches the host (linux/arm64/v8 on Apple Silicon, linux/amd64 on most servers) and ignores the rest.

The failure mode is common:

  1. You build container-lab:v1 on an Apple Silicon Mac. By default, the image is built for linux/arm64 only.
  2. You push it and deploy it to an amd64 cluster.
  3. The pull may fail with "no matching manifest for linux/amd64", or, if the manifest carries no platform information, the container starts and dies with exec format error: the kernel refused to execute an ARM binary.

The fix is to build for every target platform:

docker buildx build \
  --platform linux/amd64,linux/arm64 \
  -t username/container-lab:v1 \
  --push .

BuildKit builds one manifest per platform, using QEMU emulation, native nodes or cross-compilation, and ties them together with an index. Going the other way, Docker Desktop can run amd64-only images on Apple Silicon through emulation (Rosetta or QEMU) and prints a platform-mismatch warning. That works, but it is slower and not always faithful.

Try It Yourself

docker buildx imagetools inspect python:3.12-slim

This lists every platform manifest in the official image's index. For your own pushed images you may also see an unknown/unknown platform entry. That is not a broken image. It is an attestation manifest that BuildKit attached.

What you should remember

  • OCI has three specs: Image (the format), Distribution (the registry API) and Runtime (bundles and lifecycle).
  • A descriptor (mediaType + digest + size) is the universal typed, verifiable pointer.
  • A manifest is one platform: a config plus ordered layers. An index is many manifests, usually one per platform.
  • layers[0] is the bottom layer. The config's diff_ids list the same layers, uncompressed, in the same order.
  • Multi-platform tags point to an index. Architecture mismatches show up as "no matching manifest" or exec format error.

7. What "artifact" really means

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

"Artifact" is one of the most overloaded words in this field. It has three distinct meanings.

Meaning 1: build artifact (general software engineering)

Any file produced by a build process: a compiled binary, a .jar, a Python wheel, a minified JavaScript bundle, a coverage report. "Artifact" here just means an output worth keeping.

Meaning 2: CI/CD artifact

A build artifact that a pipeline stores or passes between stages, such as the "artifacts" section of a GitHub Actions workflow or a GitLab job. A container image is often the deployable artifact of a pipeline. In this sense "artifact" says nothing about format.

Meaning 3: OCI artifact

Content stored in an OCI registry using OCI manifests, descriptors and blobs where the content is not (necessarily) a runnable container image. Registries were designed around digests and blobs and can store arbitrary bytes. The ecosystem therefore reuses them to distribute:

  • Helm charts (helm push chart.tgz oci://registry.example.com/charts)
  • SBOMs: software bills of materials in SPDX or CycloneDX format
  • Signatures, from Sigstore Cosign or Notary Project Notation
  • Provenance and other attestations: in-toto statements, SLSA provenance
  • WebAssembly modules, policy bundles, ML model weights, configuration packages

How OCI 1.1 represents artifacts

An artifact is an image manifest that identifies what it contains with the artifactType field:

{
  "schemaVersion": 2,
  "mediaType": "application/vnd.oci.image.manifest.v1+json",
  "artifactType": "application/spdx+json",
  "config": {
    "mediaType": "application/vnd.oci.empty.v1+json",
    "digest": "sha256:44136fa355b3678a1146ad16f7e8649e94fb4fc21fe77e8310c060f61caaff8a",
    "size": 2
  },
  "layers": [
    { "mediaType": "application/spdx+json", "digest": "sha256:9f2b…", "size": 48213 }
  ],
  "subject": {
    "mediaType": "application/vnd.oci.image.index.v1+json",
    "digest": "sha256:<container-lab index digest>",
    "size": 1609
  }
}
  • artifactType says "this is an SPDX SBOM", not an image.
  • config uses the empty descriptor. Its content is the two bytes {}, and sha256:44136fa… is genuinely the SHA-256 of those two bytes. Artifacts that don't need a config use it because the manifest schema requires some config descriptor.
  • layers hold the payload. They don't have to be tar archives. Here the "layer" is the SBOM JSON itself.
  • subject points to the manifest or index this artifact is about. This is what links an SBOM, signature or attestation to our image.

The referrers API (GET /v2/<name>/referrers/<digest>) returns an index of every manifest whose subject is the given digest, optionally filtered by artifactType. For registries that don't implement the API, the spec defines a fallback: a tag named after the digest (sha256-<hex>) that points to an index of referrers.

§7

The arrows point from the artifact to the image, because the image never changes and can't list its own attachments. Adding a signature later must not change the image's digest. The registry indexes those reverse links and serves them through the referrers API.

Deep Dive Not every attestation uses subject. BuildKit stores its provenance and SBOM attestations as extra manifests inside the image index, with platform unknown/unknown and the annotation vnd.docker.reference.type: attestation-manifest. Cosign has historically stored signatures under a tag derived from the image digest, and newer versions can use referrers. The goal is the same in each case: attach metadata to an immutable digest without changing it.

Three terms, side by side

Build artifactOCI artifactContainer image
ScopeAny build outputContent in an OCI registry using OCI manifestsA specific kind of OCI content
FormatAnythingManifest + blobs, typed by artifactType or media typesManifest + image config + filesystem layers
Runnable?MaybeUsually not (SBOMs, charts, signatures)Yes, by an OCI runtime
Exampleapp.whlcontainer-lab's SBOMcontainer-lab:v1

Every container image uses the same OCI data model as artifacts. In everyday usage, though, "OCI artifact" refers to the non-image content that reuses that model.

What you should remember

  • "Artifact" can mean any build output, a CI/CD stored file, or OCI registry content that is not an image.
  • OCI artifacts are regular manifests with an artifactType, often the empty config descriptor, and arbitrary payload "layers".
  • subject plus the referrers API attach SBOMs, signatures and provenance to an image without changing its digest.
  • Registries have become general content-addressed stores, not just "Docker image hosts".

8. Registries: push and pull on the wire

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

The registry's data model

  • A registry is a service at a host, for example registry-1.docker.io, ghcr.io or 123456789.dkr.ecr.region.amazonaws.com.
  • A repository is a namespace within it, such as username/container-lab.
  • Within a repository:
    • blobs are addressed by digest
    • manifests are addressed by digest or by tag
    • tags are mutable pointers to manifests

Blobs are logically scoped to a repository for access control, but backends usually store each digest once and deduplicate across repositories.

docker push username/container-lab:v1

The name expands to docker.io/username/container-lab:v1. First, tag the local image: docker tag container-lab:v1 username/container-lab:v1.

§8

Step by step:

  1. Authentication. docker login stored your credentials, usually in an OS credential helper. The registry's 401 response names a token service, and the client exchanges credentials for a short-lived bearer token scoped to repository:username/container-lab:push,pull.
  2. Existence checks. For every blob the manifest references, the client asks HEAD /v2/<repo>/blobs/<digest>. The digests are known in advance, so the question is cheap.
  3. "Layer already exists" is Docker's message for a 200 on that HEAD. The registry already has those bytes for this repository, typically the unchanged base layers or layers from your previous push.
  4. "Mounted from library/python" means a cross-repository blob mount: the client asked the registry to link a blob it already holds in another repository, with no upload.
  5. Upload and verification. Missing blobs are uploaded (in chunks or monolithically) and finalized with PUT …?digest=sha256:X. The registry hashes what it received and rejects the upload on a mismatch.
  6. The manifest is uploaded last. When the manifest (or, for multi-platform images, each platform manifest by digest followed by the index) is PUT under the tag, everything it references already exists. A tag never points to incomplete content. The response's Docker-Content-Digest header is the digest you'd use to pin this exact push.

That sequence diagram is every message on the wire. Compressed to just the requests and their status codes, a push of an image with one new base layer, two new app layers and no prior manifest looks like this:

Interactive
docker push, on the wire
HEAD/v2/aryan/lab/blobs/sha256:0f2c9c…200
HEAD/v2/aryan/lab/blobs/sha256:a91e4d…404
POST/v2/aryan/lab/blobs/uploads/202
PATCH/v2/aryan/lab/blobs/uploads/3f7c…202
PUT/v2/aryan/lab/blobs/uploads/3f7c…?digest=sha256:a91e4d…201
HEAD/v2/aryan/lab/blobs/sha256:7bd21f…404
POST/v2/aryan/lab/blobs/uploads/202
PUT/v2/aryan/lab/blobs/uploads/9a2e…?digest=sha256:7bd21f…201
PUT/v2/aryan/lab/manifests/v1201
→tag v1 now points at this manifest's digest

Only the two new layers make the full upload trip; the base image layer was already on the registry from a previous push (200 on the HEAD, nothing else sent). The manifest PUT at the end is what actually moves the tag — until that lands, the tag still points at the old digest.

docker pull username/container-lab:v1

§8
  1. Resolve the tag. The client sends an Accept header listing every manifest and index media type it understands. The registry returns the document and its digest.
  2. Select the platform. If the document is an index, the client picks the manifest matching the host platform and fetches it by digest.
  3. Check local content. Layers already present are skipped. That is why docker pull prints "Already exists" for shared base layers.
  4. Fetch the missing blobs. Many registries answer blob requests with a redirect to object storage or a CDN. The client hashes the bytes as they stream in.
  5. Unpack. Each layer is decompressed and applied to a snapshot, and the result is checked against the config's diff_ids (section 9).
  6. Record the name. The local store maps username/container-lab:v1 to the top-level digest.

Common Misconception "The registry stores images." It stores blobs and manifests, and tags that point to manifests. An "image" is a graph that emerges from those pointers. Deleting a tag doesn't delete any blobs. Registry garbage collection removes blobs only when nothing references them.

What you should remember

  • A push uploads blobs first, checked with HEAD and verified by digest on upload, and the manifest last, so a tag is never dangling.
  • "Layer already exists" = HEAD returned 200. "Mounted from" = a cross-repository blob link with no upload.
  • A pull resolves a tag to a digest, selects a platform, downloads only missing blobs, verifies every one, and unpacks.
  • Auth is a bearer token scoped to a repository and an action (pull, push).

9. The local image store (and why macOS is different)

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

Two implementations inside Docker Engine

Docker Engine has two image-store implementations:

  • Classic graph drivers (overlay2 being the default). Images and layers are managed by dockerd itself. This is still in use on hosts that were upgraded from older Docker versions.
  • The containerd image store. Docker delegates image and snapshot management to containerd. It has been the default for fresh installs of Docker Engine 29.0+, and for new Docker Desktop installs for somewhat longer. Upgraded engines keep using graph drivers until you switch.

Run docker info and look at the "Storage Driver" section and the "driver-type" line to see which one you have.

What gets stored

Conceptually, whichever implementation you use, there are four kinds of state:

ComponentHoldsKeyed by
Content storeRaw blobs: indexes, manifests, configs, compressed layersDigest
Metadata storeImage records (name → target descriptor), labels, leases, garbage-collection referencesName / ID
SnapshotterUnpacked layers as snapshots. Committed snapshots are read-only image layers. Active snapshots are writable container layersChain ID / snapshot key
Container recordsContainer config, the writable snapshot reference, logs, stateContainer ID
§9

A pull writes compressed blobs into the content store and records the image name in the metadata database. Unpacking turns each layer into a committed snapshot. Creating a container adds a writable active snapshot on top. Garbage collection deletes content and snapshots that are no longer reachable from any image, container or lease.

Typical default paths on a Linux host. These are configurable, differ between versions, and should not be edited by hand:

  • Classic Docker: /var/lib/docker/image/overlay2/ (metadata) and /var/lib/docker/overlay2/ (unpacked layers and container layers)
  • containerd: /var/lib/containerd/io.containerd.content.v1.content/blobs/sha256/ (blobs) and /var/lib/containerd/io.containerd.snapshotter.v1.overlayfs/snapshots/ (snapshots)

Under the Hood The classic store kept only the unpacked layers and discarded the compressed blobs after a pull. The containerd store keeps both the compressed blobs and the unpacked snapshots, so each layer uses more disk space. In exchange, it can keep multi-platform indexes and attestations locally and re-push exactly the same bytes, which means exactly the same digests.

Try It Yourself (Linux host, containerd image store)

sudo ctr -n moby images ls
sudo ctr -n moby content ls | head
sudo ctr -n moby snapshots ls | head

ctr is containerd's low-level debugging client. Docker's objects live in the containerd namespace moby. Kubernetes' CRI uses k8s.io.

Why Docker Desktop on macOS is different

Linux containers need Linux kernel features: namespaces, cgroups, OverlayFS, and the Linux system call interface that the binaries inside the image expect. macOS runs the XNU kernel, which has none of them. An ELF binary from python:3.12-slim can't run on macOS directly any more than a macOS app can run on Linux.

Docker Desktop therefore runs a lightweight Linux virtual machine, using Apple's Virtualization framework in current versions. Inside that VM run the Linux kernel, dockerd, containerd, the image store and every container. The docker CLI on your Mac is only a client: it talks to the daemon through a socket that Docker Desktop forwards out of the VM (for example ~/.docker/run/docker.sock).

§9

The practical consequences:

  • There is no /var/lib/docker on your Mac. It is inside the VM's disk image.
  • uname -r in a container shows the VM's kernel.
  • Bind mounts cross a VM boundary through a file-sharing layer, which is why heavy file I/O on bind mounts can be slower than on Linux.
  • Published ports are forwarded from macOS into the VM before normal Docker networking applies.
  • Resource limits you set in Docker Desktop cap the VM. Container cgroups are nested inside that.

On Windows, Linux containers run in a WSL 2 VM in much the same way. Windows containers are a separate technology built on Windows kernel features.

What you should remember

  • Local storage has four parts: content store (blobs by digest), metadata (names), snapshotter (unpacked layers and writable layers) and container records.
  • Docker Engine 29+ fresh installs use the containerd image store. Upgraded hosts may still use overlay2 graph drivers.
  • The containerd store keeps compressed and unpacked copies: more disk, better fidelity.
  • On macOS and Windows, Linux containers run in a Linux VM. The host CLI is a remote client.

10. From image to container: docker run step by step

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

Concept

docker run  ≈  (docker pull, if missing)  +  docker create  +  docker start  (+ attach / --rm cleanup)

docker create produces a container record and its writable layer without starting a process. docker start asks the runtime to build the isolated environment and execute the command. Here is what happens for:

docker run -d --name lab -p 8080:8000 -v lab-data:/data --memory=512m --cpus=1 container-lab:v1
  1. Resolve the image. container-lab:v1 is normalized to docker.io/library/container-lab:v1. The image record gives a top-level digest, the index (if present) is resolved to the host platform's manifest, and the manifest to its config.
  2. Locate or download content. If the image isn't local, dockerd performs the pull from section 8.
  3. Ensure it is unpacked. Every layer must exist as a committed snapshot, identified by chain ID. This usually already happened during the pull or build.
  4. Prepare the filesystem view. The snapshotter creates an active snapshot whose parent is the image's top committed snapshot, and returns mount instructions: an overlay mount with the image snapshots as lowerdirs.
  5. Create the writable layer. That active snapshot's upperdir is the container's writable layer. It starts empty.
  6. Generate the OCI runtime configuration. dockerd and containerd merge the image config (Cmd, Env, WorkingDir, User) with the run flags (--memory, --cpus, -v, the default capability set, seccomp and AppArmor or SELinux profiles, hostname) into a config.json. That file plus the root filesystem forms an OCI bundle.
  7. Configure namespaces. containerd starts a shim, which invokes runc create. runc creates new PID, mount, network, UTS, IPC and cgroup namespaces as the config requests, or joins existing ones for options like --network container:x.
  8. Configure cgroups. runc creates the container's cgroup and writes memory.max, cpu.max, pids.max and so on, either directly or through systemd.
  9. Configure networking. Docker's networking component creates a veth pair, moves one end into the container's network namespace as eth0, assigns an IP address, and installs the NAT and forwarding rules for -p 8080:8000. The exact ordering relative to runc is an implementation detail.
  10. Mount the filesystem. Inside the new mount namespace, runc mounts the overlay root, then /proc, /sys, /dev and /dev/shm. Next come the lab-data volume at /data and Docker-managed /etc/hosts, /etc/hostname and /etc/resolv.conf. It then switches root to the new root with pivot_root, applies capabilities, seccomp, the LSM profile and the user, and waits.
  11. Start the process. runc start signals the waiting init process, which calls execve("python", ["python", "app.py"]) with the configured environment. Python is now PID 1 in its PID namespace. runc exits. The shim stays behind as the process's parent: it holds stdio, reaps the process, and reports its exit status.
§10

The left half (steps up to G) is image and storage work. The right half is kernel configuration. Everything before runc start is setup, and after the execve there is no container "program" running around your app. There is just your process, whose view and limits the kernel enforces.

The OCI bundle

This is an illustrative excerpt of what runc receives:

{
  "ociVersion": "1.2.0",
  "process": {
    "args": ["python", "app.py"],
    "env": ["PATH=/usr/local/bin:…", "PYTHONUNBUFFERED=1", "HOSTNAME=3f2c9a1b7d4e"],
    "cwd": "/app",
    "user": { "uid": 0, "gid": 0 },
    "capabilities": { "bounding": ["CAP_CHOWN", "CAP_NET_BIND_SERVICE", "…"] }
  },
  "root": { "path": "rootfs" },
  "hostname": "3f2c9a1b7d4e",
  "mounts": [
    { "destination": "/proc", "type": "proc", "source": "proc" },
    { "destination": "/data", "type": "bind", "source": "/var/lib/docker/volumes/lab-data/_data", "options": ["rbind"] }
  ],
  "linux": {
    "namespaces": [{ "type": "pid" }, { "type": "mount" }, { "type": "network", "path": "/var/run/docker/netns/…" }, { "type": "uts" }, { "type": "ipc" }, { "type": "cgroup" }],
    "resources": { "memory": { "limit": 536870912 }, "cpu": { "quota": 100000, "period": 100000 } },
    "seccomp": { "defaultAction": "SCMP_ACT_ERRNO", "syscalls": ["…"] }
  }
}

Notice where each piece came from: args, cwd and part of env come from the image config. resources and mounts come from the run flags. namespaces, capabilities and seccomp are runtime defaults. The image never contains any of the linux section.

What you should remember

  • docker run ≈ pull (if needed) + create + start.
  • Create = container record + writable snapshot + generated OCI spec. Start = namespaces, cgroups, mounts, network, security, then execve.
  • The OCI runtime spec (config.json) is where image defaults and your flags meet.
  • After start, runc exits. The shim remains as the process's parent.

11. OverlayFS and copy-on-write

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

Concept

How do eight read-only layer directories plus one empty writable directory appear to Python as a single /? On most Linux hosts the answer is OverlayFS, a union filesystem that has been in the mainline kernel since 3.18. containerd's default snapshotter and Docker's overlay2 driver both use it.

An overlay mount has four parts:

OptionMeaning
lowerdirOne or more read-only directories: the unpacked image layers. In the colon-separated list, the leftmost is the top
upperdirOne writable directory: the container's layer
workdirAn empty scratch directory on the same filesystem as upperdir, used internally to make copy-up and rename operations atomic
merged (mount point)The unified view the container sees as /

Conceptually, the runtime issues:

mount -t overlay overlay \
  -o lowerdir=/snap/8/fs:/snap/7/fs:…:/snap/1/fs,upperdir=/snap/9/fs,workdir=/snap/9/work \
  /run/…/rootfs

These paths are illustrative. Real paths depend on the snapshotter or driver.

§11

Lookups start at the top. For any path, OverlayFS checks upperdir first, then each lowerdir from top to bottom, and returns the first match. Directories that exist in several layers are merged: their listings are combined.

Five operations on our running container

1. Read a file from a lower layer.

docker exec lab python -c "import flask; print(flask.__file__)"

flask/__init__.py is found in the pip-install lowerdir and read directly from it, with no copying. Because the lower file is the same inode for every container using that snapshot, the page cache is shared too. Ten containers from container-lab:v1 don't hold ten copies of Flask in memory.

2. Create a new file.

docker exec lab sh -c 'echo scratch > /tmp/scratch.txt'

It is written straight into upperdir/tmp/scratch.txt. OverlayFS also creates upperdir/tmp/ as needed to hold it.

3. Modify a file that lives in a lower layer (copy-up).

docker exec lab sh -c 'echo "# edited" >> /app/app.py'

The lower layer is read-only, so OverlayFS performs a copy-up: it copies app.py, with its metadata, from the COPY . . lowerdir into upperdir/app/app.py, preparing it through workdir so the switch is atomic, and then applies the write to the copy. From then on, the upper copy hides the lower one. The image layer is untouched.

This is copy-on-write, with an important cost: copy-up copies the whole file, even for a one-byte append. Appending to a 2 GB file in a lower layer first copies 2 GB. That is one reason databases and logs belong on volumes, not in the container layer. (The metacopy mount option can defer data copying for metadata-only changes such as chmod. Whether it is enabled depends on the kernel and the runtime.)

4. Delete a file that lives in a lower layer (whiteout).

docker exec lab rm /app/requirements.txt

The lower file can't be removed, so OverlayFS creates a whiteout in upperdir: a character device with device number 0/0, named requirements.txt. The merged view treats it as "this path does not exist". The bytes remain in the image layer, and every other container still sees the file.

5. Delete and recreate a directory (opaque directory).

If a directory from a lower layer is removed and recreated, the new upper directory is marked opaque with an extended attribute (trusted.overlay.opaque="y", or the user.overlay.* namespace for unprivileged mounts). Lower-layer contents of that directory are then hidden entirely.

Step through the same four operations against the stack above — each one only ever adds to upperdir; the image layers underneath never move:

Interactiveoverlayfs — one container, four operations

merged view (what the process sees)

  • app.py (from lower)
  • requirements.txt (from lower)
  • logs/ (from lower)

upperdir (writable layer)

empty

lowerdir (image layers, read-only)

  • app.py
  • requirements.txt
  • logs/
    • boot.log
    • access.log

nothing — a read of an unmodified file never touches upperdir. Nothing in lowerdir is ever mutated; every write, delete and directory replacement is expressed as new content in upperdir, which is exactly what makes the same image layers safely shareable, read-only, across every container running from them.

Now ask Docker what changed:

docker diff lab
C /app
C /app/app.py
D /app/requirements.txt
C /tmp
A /tmp/scratch.txt

This output is illustrative, and you may see a few extra entries such as Python bytecode caches. A = added, C = changed, D = deleted. docker diff is effectively a readout of the upperdir. /data doesn't appear because it is a volume mount, not part of the overlay.

Whiteouts in two formats

This distinction is easy to miss:

WhereDeletion markerOpaque directory marker
OCI layer tarball (image spec, portable)An empty file named .wh.<name>, e.g. app/.wh.requirements.txtA file named .wh..wh..opq inside the directory
OverlayFS on disk (kernel format)A character device 0/0 with the original nameAn xattr trusted.overlay.opaque="y"

When a layer is unpacked onto an overlay snapshotter, .wh. entries are converted to the kernel format. When a build captures a filesystem diff into a layer, the conversion runs the other way. A Dockerfile RUN rm /some/file therefore produces a layer that contains a .wh.file entry and none of the file's bytes.

Try It Yourself (Linux host)

docker exec lab cat /proc/mounts | head -1
# overlay / overlay rw,relatime,lowerdir=…,upperdir=…,workdir=… 0 0

With the classic overlay2 driver, docker inspect lab --format '{{json .GraphDriver.Data}}' prints the LowerDir, UpperDir, MergedDir and WorkDir paths. With the containerd image store, that field may be empty or different. Look for the overlay mount on the host with findmnt -t overlay instead.

Common Misconception "Every container has a full copy of the image." Containers share the read-only lower snapshots. Each container's only private storage is its upperdir, which holds just the files it created, modified or deleted.

What you should remember

  • OverlayFS merges read-only lowerdirs (the image) and one writable upperdir (the container) into one view. workdir is internal scratch space.
  • Reads come from the topmost layer that has the path. Lower files and their page cache are shared.
  • Writes to lower files trigger a whole-file copy-up. Deletes create whiteouts. Recreated directories become opaque.
  • OCI layer tars use .wh. files, while OverlayFS uses 0/0 character devices and xattrs, and the runtime converts between them.

12. Linux namespaces: what a process can see

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

Concept

A namespace wraps a global system resource so that processes inside it see their own isolated instance of that resource. Every process on Linux belongs to exactly one namespace of each type. A "container" is a process whose namespaces differ from the host's. You can see a process's namespaces as symbolic links under /proc/<pid>/ns/:

# on a Linux host
PID=$(docker inspect --format '{{.State.Pid}}' lab)
sudo ls -l /proc/$PID/ns
sudo ls -l /proc/1/ns       # compare with the host's init
# illustrative
lrwxrwxrwx 1 root root 0 … net -> 'net:[4026532301]'
lrwxrwxrwx 1 root root 0 … pid -> 'pid:[4026532299]'
lrwxrwxrwx 1 root root 0 … uts -> 'uts:[4026532297]'
…

Two processes are in the same namespace exactly when those inode numbers match.

The namespaces, one by one

PID namespace (process IDs)

  • Purpose: give the container its own process-number space and its own "init".
  • Without isolation: the process sees every process on the host (ls /proc), including PID 1 (systemd).
  • Inside lab: it sees only its own processes. Python is PID 1.
  • Try it: docker exec lab ls /proc | grep -E '^[0-9]+$' shows a handful of numbers. docker top lab shows the host PIDs of the same processes.

NET namespace (network stack)

  • Purpose: a private set of interfaces, IP addresses, routes, firewall rules and port bindings.
  • Without isolation: all host interfaces are visible, and port 8000 would conflict with any host service.
  • Inside lab: only lo and eth0 (one end of a veth pair) exist, with an IP such as 172.17.0.2.
  • Try it: docker run --rm --network container:lab busybox ip addr. A second container joins lab's network namespace and sees the same interfaces. This is exactly how Kubernetes pods share networking.

MNT namespace (mount table)

  • Purpose: a private mount table, so the process gets its own / and other mounts.
  • Without isolation: the host's mount table and root filesystem.
  • Inside lab: the overlay root, /proc, /data and so on.
  • Try it: docker exec lab cat /proc/self/mounts.

UTS namespace (hostname, NIS domain name)

  • Purpose: an independent hostname.
  • Without isolation: the host's hostname.
  • Inside lab: the short container ID, for example 3f2c9a1b7d4e.
  • Try it: docker exec lab uname -n.

IPC namespace (System V IPC and POSIX message queues)

  • Purpose: isolate shared-memory segments, semaphores and message queues.
  • Without isolation: host-wide IPC objects are visible and attachable.
  • Inside lab: only its own IPC objects.
  • Try it: docker run --ipc=host … removes this isolation, which is occasionally needed and generally risky.

USER namespace (UID/GID mappings, capabilities)

  • Purpose: map UIDs inside the namespace to different UIDs outside, so "root" inside can be unprivileged outside.
  • Without isolation: UID 0 is host UID 0.
  • Inside lab, by default: not used by standard rootful Docker. uid: 0 in our app is host UID 0, reduced only by dropped capabilities, seccomp and the LSM profile.
  • Try it: docker exec lab cat /proc/self/uid_map shows 0 0 4294967295, the identity mapping. It changes under userns-remap or rootless mode.

CGROUP namespace (the cgroup tree view)

  • Purpose: virtualize the process's view of its cgroup path.
  • Without isolation: the full host path, e.g. 0::/system.slice/docker-3f2c….scope.
  • Inside lab: its cgroup appears as the root: 0::/.
  • Try it: docker exec lab cat /proc/self/cgroup. Docker uses a private cgroup namespace by default on cgroup v2 hosts.

TIME namespace (Linux 5.6+)

  • Purpose: offsets for the monotonic and boot-time clocks.
  • Without isolation: host clocks.
  • Inside lab: not used by Docker by default.
  • Try it: mostly relevant for checkpoint/restore.

User namespaces deserve emphasis. Rootless Docker, rootless Podman and Kubernetes pods with hostUsers: false all run with a user namespace, so UID 0 inside maps to an unprivileged UID range on the host. Standard rootful Docker does not unless you configure it.

Why PID 1 inside is not PID 1 outside

A new PID namespace is nested inside its parent. A process in it has one PID per level: its PID in its own namespace, and a different PID in every ancestor namespace. The kernel shows both:

PID=$(docker inspect --format '{{.State.Pid}}' lab)   # e.g. 48213 (illustrative)
grep NSpid /proc/$PID/status
# NSpid:  48213   1
Interactivetwo pids, one process

host process tree (default namespace)

  • 1/sbin/init
  • 742dockerd
  • 48198containerd-shim
  • 48213python app.py
    same task as container PID 1 →
pid namespace boundary

container process tree (its own pid namespace)

  • 1python app.py
    no signal handler installed

The host sees the Python process as PID 48213, a child of the shim. Inside its namespace the same process is PID 1, and a shell started with docker exec might be PID 7. The container can't see or signal anything to the left. The host can see and signal everything on the right. (The shim is shown as a child of systemd because shims daemonize and are reparented to init. Exact process trees vary.)

Being PID 1 has two consequences for your application:

  • Signals. The kernel does not apply default signal actions to a namespace's init process, even for signals sent from the host, except for SIGKILL/SIGSTOP sent from an ancestor namespace. If your app has no SIGTERM handler, docker stop appears to hang for 10 seconds and then kills it (section 16).
  • Zombie reaping. Orphaned processes are reparented to the namespace's PID 1, which must wait() for them. Apps that spawn children should run under a minimal init: docker run --init injects tini as PID 1.

Under the Hood Tools such as nsenter work because namespaces are kernel objects you can join. sudo nsenter -t $PID -n ip addr runs the host's ip binary inside lab's network namespace. This is how you debug a container whose image has no tools.

What you should remember

  • Namespaces virtualize visibility: PID, network, mounts, hostname, IPC, UIDs, cgroup path and clocks.
  • Two processes share a namespace exactly when their /proc/<pid>/ns/* inode numbers match. Pods and --network container:x rely on joining namespaces.
  • A containerized process has a different PID at each namespace level (NSpid).
  • PID 1 gets special signal and reaping semantics, so handle SIGTERM or use --init.
  • Standard rootful Docker does not use user namespaces. Root in the container is UID 0 on the host, with restrictions.

13. cgroups: what a process can use

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

The distinction

Namespaces → what a process can see. cgroups → how much a process can use, and how much it has used.

A PID namespace hides other processes but does nothing to stop a fork bomb. A network namespace gives the process its own interfaces but doesn't cap its bandwidth. Resource control is a separate kernel subsystem: control groups.

cgroup v2

Modern distributions use cgroup v2, a single unified hierarchy mounted at /sys/fs/cgroup. Every process belongs to exactly one cgroup, a directory in that tree, and each controller exposes interface files in that directory:

ControllerKey filesWhat it does
memorymemory.max, memory.high, memory.current, memory.events, memory.swap.maxHard and soft limits, usage accounting, OOM events
cpucpu.max, cpu.weight, cpu.statBandwidth quota per period, relative weight, throttling statistics
pidspids.max, pids.currentMaximum number of tasks (processes plus threads)
ioio.max, io.weight, io.statPer-device bandwidth and IOPS limits, proportional weight
cpusetcpuset.cpus, cpuset.memsPin to specific CPUs and memory nodes

From flag to kernel

For our run command, --memory=512m --cpus=1, plus --pids-limit=100 for illustration:

Interactive
from flag to enforcement

CLI flag

docker run --memory=512m --cpus=1.5 …

Engine API

"HostConfig":{"Memory":536870912,"NanoCpus":1500000000}

OCI runtime spec

linux.resources.memory.limit: 536870912 linux.resources.cpu.quota: 150000 / period: 100000

cgroup v2 file

/sys/fs/cgroup/…/memory.max → 536870912 /sys/fs/cgroup/…/cpu.max → 150000 100000

kernel enforcement

OOM killer past memory.max CFS throttles past cpu.max

Every stage after the first is just a translation of the same two numbers. Read the actual file on any Linux host running the container — cat /sys/fs/cgroup/…/memory.max — and it will show 536870912, in bytes, no daemon in the loop.

A human-friendly flag becomes a number in the Engine API, then a field in the OCI runtime spec, then a value runc writes into a cgroup interface file. From then on the kernel's scheduler and memory manager enforce it. No Docker component polices your process at run time. The path shown is typical of systemd-based hosts and is implementation-specific.

What each limit means in practice:

  • memory.max = 536870912. When the cgroup's usage approaches the limit, the kernel first reclaims memory, such as page cache, from this cgroup. If it can't reclaim enough, the cgroup's OOM killer kills a process in it. Docker then reports "OOMKilled": true and, for a killed main process, exit code 137 (128 + SIGKILL). Page cache counts toward the limit, which surprises people running file-heavy workloads.
  • cpu.max = "100000 100000". A quota of 100 ms of CPU time per 100 ms period, which equals one CPU's worth of time. That time may be spread across any cores, and exceeding it gets the cgroup throttled until the next period. It is not pinning to one core (that is --cpuset-cpus) and not a hidden-CPU illusion: os.cpu_count() in our Python app still reports the host's CPU count. Runtimes that size thread pools from the CPU count can over-subscribe and throttle.
  • pids.max = 100. The 101st task creation fails with EAGAIN. This is the actual defense against fork bombs.
  • cpu.weight (from --cpu-shares) is relative. It matters only when CPUs are contended.

Try It Yourself

docker exec lab cat /sys/fs/cgroup/memory.max      # 536870912
docker exec lab cat /sys/fs/cgroup/cpu.max         # 100000 100000
docker exec lab cat /sys/fs/cgroup/memory.current  # live usage in bytes
docker stats lab --no-stream

Inside the container, the cgroup namespace plus a cgroup mount make its own cgroup appear at /sys/fs/cgroup, so these paths work without knowing the host path.

Common Misconception "--memory reserves 512 MB for my container." It sets a ceiling. Nothing is reserved or pre-allocated, and other workloads can use that memory until your container needs it. Reservations and guarantees, as in Kubernetes requests, are a scheduling concept layered on top.

What you should remember

  • Namespaces = visibility. cgroups = resource limits and accounting. They are independent kernel features.
  • cgroup v2 is one hierarchy under /sys/fs/cgroup, and limits are plain files such as memory.max, cpu.max and pids.max.
  • The chain is: docker run flags → Engine API → OCI spec linux.resources → runc writes cgroup files → kernel enforces.
  • --cpus limits CPU time, not visible cores. --memory is a ceiling that triggers reclaim, then OOM kill.

14. Runtime architecture: Docker, containerd, shim, runc

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

The components

Architectures evolve, and not every installation has the same process topology. Kubernetes nodes, for example, don't need Docker at all. Here is the current, typical arrangement for Docker Engine on Linux:

§14

A request moves down the chain one level of abstraction at a time. The CLI speaks HTTP to dockerd, dockerd speaks gRPC to containerd, containerd starts a shim, and the shim executes runc, which programs the kernel and exits. The dotted edge is the relationship that persists while the container runs: the shim, not runc or containerd, is the container process's parent.

Docker CLI. A client that turns commands into Engine API calls. It holds no containers. On macOS it talks to a daemon inside a VM (section 9).

Docker Engine (dockerd). Docker's high-level features: the Engine API, the Docker object model (containers, networks, volumes), networking (bridges, NAT rules, embedded DNS), volume drivers, log drivers, BuildKit integration and Swarm. With the containerd image store it delegates image content and snapshots to containerd.

containerd. An industry-standard container runtime daemon (a CNCF graduated project) that manages the lifecycle of images and containers: pulling, content storage, snapshotters, container metadata and tasks (running processes). It is multi-tenant through namespaces (moby, k8s.io), and its built-in CRI plugin lets Kubernetes use it directly.

containerd-shim-runc-v2. A small per-container process, and the reason you can restart or upgrade containerd, or dockerd with live-restore, without killing your containers. It keeps the container's stdio pipes open, acts as the subreaper that collects the exit status, and exposes a ttrpc API to containerd.

runc. The reference implementation of the OCI Runtime Specification. Given a bundle (config.json + rootfs), it creates namespaces, configures cgroups, sets up mounts, drops privileges, applies seccomp and LSM profiles, and execves the process. It then exits.

Why runc doesn't stick around

runc's job is defined by the OCI Runtime Spec: create, start, kill, delete, state, operating on a bundle. The spec deliberately leaves out images, registries, networks, volumes and long-term supervision. Keeping the low-level runtime short-lived means:

  • No daemon has to stay alive per container. The small shim does the waiting.
  • The runtime can be swapped per container: crun (C), youki (Rust), gVisor's runsc, or Kata Containers. containerd only needs the matching shim.
  • Security-sensitive setup code, which runs with high privilege at the moment of container creation, is not left running.

A typical host process tree (illustrative; details vary):

systemd(1)
 ├─ containerd
 ├─ dockerd
 └─ containerd-shim-runc-v2 -namespace moby -id 3f2c9a1b7d4e… -address /run/containerd/containerd.sock
     └─ python app.py

Interview Insight "Did Kubernetes drop Docker?" In 1.24 Kubernetes removed dockershim, the adapter that let kubelet talk to dockerd. Images built with Docker are OCI images and run unchanged on containerd or CRI-O. What went away was one layer of indirection: kubelet → dockershim → dockerd → containerd became kubelet → containerd (CRI).

What you should remember

  • CLI → dockerd (Docker features) → containerd (images, snapshots, tasks) → shim (per-container parent) → runc (OCI runtime) → kernel.
  • runc implements the OCI Runtime Spec, then exits. The shim supervises the running process.
  • Shims let daemons restart without killing containers.
  • Kubernetes talks to containerd (or CRI-O) through CRI, with no Docker involved.

15. Container networking

1
Source
2
Build
3
Blobs
4
Registry
5
Content store
6
Snapshot
7
Runtime
8
Process

The building blocks

  • Network namespace: lab's private network stack (section 12).
  • veth pair: a virtual Ethernet cable with two ends. A packet sent into one end comes out the other. One end, renamed eth0, is moved into the container's namespace. The other (vethXXXX) stays on the host.
  • Linux bridge: a virtual layer-2 switch. Docker's default bridge network uses a bridge device called docker0 (e.g. 172.17.0.1/16). User-defined networks get their own bridges and an embedded DNS server that resolves container names.
  • Routing: inside the container, the default route points to the bridge's IP. On the host, IP forwarding must be enabled.
  • NAT: outbound traffic is masqueraded (source-NATed) to the host's IP. Inbound published ports are destination-NATed to the container's IP and port.
  • Port publishing: -p 8080:8000 programs those DNAT rules. Docker also runs a small userland proxy, docker-proxy, by default. It listens on the host port and handles cases the NAT rules miss, such as some loopback traffic. Docker programs the rules with iptables by default. Docker Engine 29 added experimental nftables support.

Tracing a request to container-lab

We use our own container (-p 8080:8000). docker run -p 8080:80 nginx works exactly the same way, just with port 80 inside.

§15

The packet arrives for host port 8080. A NAT rule rewrites its destination to the container's IP and port 8000 (or, on some paths, docker-proxy accepts the connection and opens a new one). Host routing then sends it to the docker0 bridge, across the veth pair, and into lab's namespace, where Python is listening. Replies follow the same path back, and connection tracking reverses the translation.

This is why the app must listen on 0.0.0.0, not 127.0.0.1. Inside the namespace, 127.0.0.1 is the container's own loopback, which traffic arriving on eth0 never reaches.

Outbound traffic, for example pip reaching PyPI during a RUN or our app calling an API, takes the reverse path. It leaves through eth0, crosses the veth pair to docker0, is routed out of the host's uplink, and is MASQUERADEd so it appears to come from the host's IP.

Try It Yourself (Linux host)

docker port lab                                  # 8000/tcp -> 0.0.0.0:8080
docker network inspect bridge --format '{{json .Containers}}'
ip link show master docker0                      # host-side veth ends
sudo iptables -t nat -L DOCKER -n                # DNAT rule (iptables backend)
sudo nsenter -t $(docker inspect -f '{{.State.Pid}}' lab) -n ip route

Rule layout and chain names differ between Docker versions and firewall backends. Read them to learn, don't depend on them.

On Docker Desktop, one more hop comes first: Docker Desktop forwards localhost:8080 on macOS or Windows into the Linux VM, and then the path above applies inside the VM.

What you should remember

  • Each container gets a network namespace. A veth pair connects it to a bridge in the host namespace.
  • Outbound traffic is masqueraded. Published ports are destination-NATed, with docker-proxy covering some edge paths.
  • Apps must bind 0.0.0.0 inside the container to be reachable through published ports.
  • Networking implementation details (iptables vs nftables, proxy behavior, Desktop forwarding) vary. The concepts don't.

16. The container lifecycle

§16

Docker's state names are created, running, paused, restarting, exited, removing and dead. "Stopped" in everyday speech means exited. A restart policy moves an exited container back to running automatically. A container you stopped explicitly with docker stop is not restarted by the policy. With always it comes back only when the daemon restarts, and with unless-stopped it doesn't come back at all. "Removed" is not a state Docker stores: the container record no longer exists.

What each command changes

CommandProcessesWritable layerAnonymous volumesNamed volumesLogsImage
docker createNone yetCreated (empty)CreatedCreated if missingEmptyRequired (pulled if absent), unchanged
docker startStarted (new PID)KeptKeptKeptAppendedUnchanged
docker stopSIGTERM (or the image's STOPSIGNAL), then SIGKILL after the timeout (10 s default)KeptKeptKeptKeptUnchanged
docker killSIGKILL (or --signal) immediatelyKeptKeptKeptKeptUnchanged
docker restartStopped, then a new process startedKeptKeptKeptAppendedUnchanged
docker rmMust be stopped (or use -f)DeletedDeleted only with -vKeptDeleted (for the default json-file and local drivers)Unchanged

Answering the common questions directly:

  • Does stopping delete the container? No. An exited container keeps its writable layer, config and logs. docker start resumes with the same filesystem. Memory state is lost, because processes don't survive a stop.
  • Does deleting the container delete the image? No. The image is shared, immutable content. Remove it with docker image rm, which is refused while any container still uses it.
  • What happens to the writable layer? It lives exactly as long as the container record and disappears with docker rm.
  • What happens to volumes? Named volumes survive docker rm. Anonymous volumes are removed only with docker rm -v or when the container was started with --rm.
  • What happens to logs? With the default json-file log driver, logs are stored with the container and deleted by docker rm. Drivers that ship logs elsewhere (journald, syslog, fluentd, cloud drivers) keep them in that system.
  • What happens to processes? stop asks, then forces. kill forces immediately, or sends whichever signal you choose. After either, the shim reports the exit status and the process is gone.

Common Misconception: "docker stop is slow" Try time docker stop lab with our app. It likely takes about 10 seconds, because our Python process is PID 1 and installs no SIGTERM handler. The kernel doesn't apply the default action to a PID namespace's init, so SIGTERM does nothing, Docker waits out the timeout, and then sends SIGKILL. To fix it, handle SIGTERM in the app, run with --init, or use a server that handles signals. Also always use the exec form (CMD ["python", "app.py"]), not the shell form, which puts /bin/sh at PID 1 as a signal-swallowing parent.

Try It Yourself

curl -s localhost:8080 >/dev/null; curl -s localhost:8080   # visits: 2
docker exec lab sh -c 'echo hi > /tmp/marker'
docker restart lab
docker exec lab cat /tmp/marker        # "hi": writable layer survived restart
docker rm -f lab
docker run -d --name lab -p 8080:8000 -v lab-data:/data container-lab:v1
docker exec lab cat /tmp/marker        # error: new writable layer
curl -s localhost:8080                 # visits: 3, because the volume survived

What you should remember

  • Stop/kill end processes. The container, its writable layer and its logs remain until docker rm.
  • docker rm deletes the writable layer and (with default drivers) the logs, but never the image or named volumes.
  • restart = same container and filesystem, new process.
  • PID 1 without a SIGTERM handler turns every docker stop into a 10-second wait followed by SIGKILL.

17. Image vs container vs volume

IMAGE      Immutable template.  Content-addressed layers + config.  Shared by every container made from it.
CONTAINER  Runtime instance.     Metadata + a private writable layer + (while running) namespaced, cgrouped processes.
VOLUME     Persistent data.      Managed by Docker, independent of any container's lifecycle, mounted into containers.

A bind mount is different from a volume. It mounts an arbitrary host path (for example -v "$(pwd)":/app or --mount type=bind,src=…,dst=…) into the container. Docker doesn't create or manage it: the host directory's lifetime, ownership and contents are yours. A tmpfs mount is memory-backed and disappears when the container stops.

§17

When lab is removed, everything in the lower box goes: its identity, its private writable layer, its logs and, optionally, its anonymous volumes. What it merely referenced stays: the shared image, the named volume, and any host directories.

Deep Dive A Docker-specific convenience: when an empty named volume is first mounted at a path that already contains files in the image, Docker copies the image's files into the volume. A bind mount never does this. It hides whatever the image had at that path. This behavior is Docker's, not an OCI requirement, and other runtimes may differ.

What you should remember

  • Image = immutable, shared template. Container = one instance plus private writable state. Volume = data with its own lifecycle.
  • Anything written outside a volume or bind mount lives in the writable layer and dies with the container.
  • Bind mounts are host paths you manage. Volumes are Docker-managed storage.

18. Names, tags, digests and IDs

People call a dozen different values "the image hash". Here they all are, for our image.

The reference grammar

[registry[:port]/][namespace/]repository[:tag][@algorithm:digest]

container-lab:v1                       → docker.io/library/container-lab:v1  (normalized)
username/container-lab:v1              → docker.io/username/container-lab:v1
ghcr.io/username/container-lab@sha256:…  (a digest reference: tag not needed)

Every identifier

IdentifierWhat is hashed / what it isMutable?Where you see it
RegistryHost name (docker.io, ghcr.io)n/aThe reference prefix
RepositoryPath in the registry (username/container-lab)n/aPush and pull targets
Image nameInformal: registry + repository (+ tag)n/aEverywhere
TagA pointer to a manifest or indexYes:v1, :latest
Index digestSHA-256 of the image index JSONNoWhat a multi-platform tag resolves to, RepoDigests
Manifest digestSHA-256 of one platform's manifest JSONNoIndex entries, single-platform RepoDigests
Config digestSHA-256 of the config JSONNoManifest config.digest, and the classic image ID
Layer digestSHA-256 of a compressed layer blobNoManifest layers[], push/pull progress lines
Diff IDSHA-256 of an uncompressed layer tarNoConfig rootfs.diff_ids, docker image inspect → RootFS.Layers
Chain IDHash of a stack of diff IDsNoSnapshot keys (internal)
Image IDImplementation-specific. Classic store: the config digest. containerd image store: the digest of the image's top-level index or manifestNodocker image ls IMAGE ID, docker image inspect .Id
Container ID64 random hex characters, not a content hashn/adocker ps, the default hostname

Why they differ even for "the same image"

  • Pull python:3.12-slim on an ARM Mac and an x86 server: the index digest is identical (both resolve the same index), but the manifest, config and layer digests differ, because they are different platform builds.
  • Push the same image to two registries: the manifest digest is identical if the bytes are pushed as-is. It changes if a tool rewrites or recompresses the content.
  • Compare docker image ls IMAGE IDs between a classic-store host and a containerd-store host for the same pulled image. They differ, because they are hashes of different documents (config vs index or manifest).
  • docker image inspect container-lab:v1 shows RootFS.Layers, which are diff IDs. They will not match the layer digests shown during docker push, because those are hashes of the compressed bytes.
§18

Each arrow is "contains or derives from". The only mutable link is the tag. The dotted edges show why the IMAGE ID column means different things on different Docker setups.

Interview Insight "What digest should I put in a Kubernetes manifest?" The one the registry reports for the tag you tested: the index digest for multi-platform images, the manifest digest otherwise. That is RepoDigests in docker image inspect, or the Docker-Content-Digest from the push. Never the image ID and never a layer digest.

What you should remember

  • Repository and tag are names. Digests are content identities. The container ID is random.
  • Index, manifest, config, layer and diff-ID digests hash different bytes, so they are always different values.
  • "Image ID" is implementation-specific: the config digest (classic) or the top-level index or manifest digest (containerd store).
  • Pin deployments with the registry digest (RepoDigests), not the image ID.

19. Lab: inspecting a real image

Everything so far can be observed on your own machine. The outputs below are illustrative: they show the shape and the important fields, not captures from a live system. Your digests, sizes, dates and layer counts will differ, and column layouts vary between Docker versions.

docker image ls --digests

docker image ls --digests container-lab
# illustrative
REPOSITORY      TAG   DIGEST        IMAGE ID       CREATED         SIZE
container-lab   v1    sha256:…      7c1e0d2f9a44   2 minutes ago   …
  • What it reveals: the local name-to-content mapping.
  • DIGEST is the registry digest (index or manifest). With the classic store, a never-pushed local build shows <none>, because the image hasn't been given a distribution digest yet. With the containerd store, one exists immediately, because the index or manifest blob exists locally.
  • IMAGE ID is implementation-specific (section 18).

docker image inspect

docker image inspect container-lab:v1 --format '{{json .RootFS.Layers}}'
docker image inspect container-lab:v1 --format '{{.Architecture}}/{{.Os}} cmd={{json .Config.Cmd}} wd={{.Config.WorkingDir}}'
docker image inspect container-lab:v1 --format '{{json .RepoDigests}}'
# illustrative
["sha256:5d1f…","sha256:6a90…","sha256:81c2…","sha256:9e04…","sha256:a7b3…","sha256:c1d8…","sha256:e25f…","sha256:f090…"]
arm64/linux cmd=["python","app.py"] wd=/app
["username/container-lab@sha256:…"]     # only after push, or after pulling from a registry
  • What it reveals: the image config as Docker sees it, plus local metadata.
  • RootFS.Layers are diff IDs: hashes of the uncompressed layers. Don't expect them to match push output.
  • Config.* holds the defaults every docker run starts from.
  • RepoDigests is the digest you should pin in deployments.

docker history

docker history container-lab:v1
# illustrative
IMAGE          CREATED         CREATED BY                                      SIZE      COMMENT
7c1e0d2f9a44   2 minutes ago   CMD ["python" "app.py"]                         0B        buildkit.dockerfile.v0
<missing>      2 minutes ago   EXPOSE map[8000/tcp:{}]                         0B        buildkit.dockerfile.v0
<missing>      2 minutes ago   COPY . . # buildkit                             4.1kB     buildkit.dockerfile.v0
<missing>      2 minutes ago   RUN /bin/sh -c pip install --no-cache-dir -r…   6.2MB     buildkit.dockerfile.v0
<missing>      2 minutes ago   COPY requirements.txt . # buildkit              13B       buildkit.dockerfile.v0
<missing>      2 minutes ago   WORKDIR /app                                    0B        buildkit.dockerfile.v0
<missing>      2 minutes ago   ENV PYTHONUNBUFFERED=1                          0B        buildkit.dockerfile.v0
<missing>      3 weeks ago     CMD ["python3"]                                 0B        buildkit.dockerfile.v0
<missing>      3 weeks ago     RUN /bin/sh -c set -eux; …                      …         buildkit.dockerfile.v0
…
  • What it reveals: the config's history array joined with layer sizes.
  • <missing> is normal. Only the top row corresponds to an image the engine knows by ID. BuildKit doesn't create intermediate images, and the base image's history rows came from another builder. <missing> does not mean a layer is missing.
  • 0B rows are metadata-only instructions (empty_layer). Sizes are uncompressed layer sizes.
  • requirements.txt is 13 bytes, which is exactly flask==3.0.3\n.
  • CREATED BY is informational text, not a verified record. Supply-chain verification uses attestations (section 21).
  • Use --no-trunc to see full commands.

docker manifest inspect and friends (registry side)

These read the registry's copy, so push first:

docker manifest inspect username/container-lab:v1
docker buildx imagetools inspect username/container-lab:v1
docker buildx imagetools inspect username/container-lab:v1 --raw
# illustrative (imagetools inspect)
Name:      docker.io/username/container-lab:v1
MediaType: application/vnd.oci.image.index.v1+json
Digest:    sha256:aaaa…

Manifests:
  Name:        docker.io/username/container-lab:v1@sha256:bbbb…
  MediaType:   application/vnd.oci.image.manifest.v1+json
  Platform:    linux/amd64

  Name:        docker.io/username/container-lab:v1@sha256:cccc…
  MediaType:   application/vnd.oci.image.manifest.v1+json
  Platform:    linux/arm64

  Name:        docker.io/username/container-lab:v1@sha256:dddd…
  MediaType:   application/vnd.oci.image.manifest.v1+json
  Platform:    unknown/unknown
  Annotations:
    vnd.docker.reference.digest: sha256:bbbb…
    vnd.docker.reference.type:   attestation-manifest
  • What it reveals: the index, its platform manifests, and attestation manifests.
  • --raw prints the exact JSON bytes the registry serves, the bytes whose SHA-256 is the digest.
  • docker manifest inspect shows the raw manifest or index with less decoration. --verbose adds resolved platform manifests.

Vendor-neutral tools do the same without a Docker daemon, for example crane manifest, crane config and crane digest (from go-containerregistry), skopeo inspect --raw docker://…, regctl manifest get, and oras discover (which walks referrers such as SBOMs and signatures).

Prove content addressing yourself

docker save writes an image to a tarball. Since Docker Engine 25 that tarball is an OCI image layout, with Docker's legacy manifest.json included for compatibility:

docker save container-lab:v1 -o lab.tar
mkdir lab-oci && tar -xf lab.tar -C lab-oci
ls lab-oci                       # blobs/  index.json  manifest.json  oci-layout  …
python3 -m json.tool lab-oci/index.json

cd lab-oci/blobs/sha256
for f in *; do
  printf '%s  %s\n' "$f" "$(shasum -a 256 "$f" | cut -c1-64)"
done

Every blob's filename equals its SHA-256. That is content-addressable storage, and you just verified it with a generic tool. To find the layer holding our code:

for f in *; do tar -tf "$f" 2>/dev/null | grep -qx 'app/app.py' && echo "app layer: $f"; done

Under the Hood The layer blobs inside a docker save archive may be stored uncompressed (the classic store re-exports from unpacked layers) or exactly as pulled (the containerd store). So the manifest inside the tarball may carry different layer digests than the registry's, even though the diff IDs, and therefore the filesystem, are identical. This is section 5's "same content, different compressed digest", observed directly.

Try It Yourself (Linux, containerd image store) sudo ctr -n moby content ls, sudo ctr -n moby images ls and sudo ctr -n moby snapshots ls show the same objects from containerd's side: blobs by digest, image records and snapshots by key.

What you should remember

  • image inspect shows the config and diff IDs. history shows build history with uncompressed sizes. manifest inspect and imagetools show the registry's documents.
  • <missing> in docker history is expected with BuildKit.
  • docker save gives you an OCI layout in which blob filenames are their SHA-256 digests.
  • Label anything you screenshot as version-specific. These outputs change between releases.

20. Build cache experiment

The two Dockerfiles

Dockerfile.bad:

# syntax=docker/dockerfile:1
FROM python:3.12-slim
ENV PYTHONUNBUFFERED=1
WORKDIR /app
COPY . .
RUN pip install --no-cache-dir -r requirements.txt
EXPOSE 8000
CMD ["python", "app.py"]

Our Dockerfile (the "good" one) copies requirements.txt first, installs dependencies, and then copies the rest.

The experiment

# 1. Warm both caches
docker build -f Dockerfile.bad -t container-lab:bad .
docker build -t container-lab:v1 .

# 2. Change application code only
echo '# tweak' >> app.py

# 3. Rebuild both and watch which steps run
docker build --progress=plain -f Dockerfile.bad -t container-lab:bad . 2>&1 | grep -E '^#[0-9]+ (\[|CACHED|DONE)'
docker build --progress=plain -t container-lab:v1 . 2>&1 | grep -E '^#[0-9]+ (\[|CACHED|DONE)'

Illustrative results:

# Dockerfile.bad
#5 [2/4] WORKDIR /app
#5 CACHED
#6 [3/4] COPY . .
#6 DONE 0.0s
#7 [4/4] RUN pip install --no-cache-dir -r requirements.txt
#7 DONE 7.8s                 ← dependencies reinstalled for a one-line code change

# Dockerfile (good)
#5 [2/5] WORKDIR /app
#5 CACHED
#6 [3/5] COPY requirements.txt .
#6 CACHED
#7 [4/5] RUN pip install --no-cache-dir -r requirements.txt
#7 CACHED                    ← reused
#8 [5/5] COPY . .
#8 DONE 0.0s

Why: the dependency graph

§20

Each step's cache key includes its parent's key, so invalidating one step invalidates everything below it. In the bad ordering, app.py is an input to the step above pip install, so every code edit reinstalls dependencies. In the good ordering, app.py only feeds the final COPY. Only a change to requirements.txt reaches pip install.

The rule: order instructions from least frequently changed to most frequently changed, and copy dependency manifests (requirements.txt, package.json + lockfile, go.mod/go.sum, pom.xml) before source code.

  • Cache mounts keep a package manager's download cache across builds without putting it in the image: RUN --mount=type=cache,target=/root/.cache/pip pip install -r requirements.txt. With this, even a dependency change avoids re-downloading unchanged wheels.
  • --no-cache ignores the cache entirely. --pull re-resolves the base image tag, so you pick up security rebuilds. --no-cache-filter <stage> busts a single stage.
  • CI runners start cold. Use --cache-to / --cache-from (for example type=registry) to share cache between runs.

Common Misconception "The cache noticed that a newer Flask exists." It didn't and can't. RUN cache keys are computed from the command text and the parent state, not from what the network would return today. Pin your dependency versions, and rebuild with --pull or --no-cache on a schedule if you want fresh upstream content.

What you should remember

  • A change to any input invalidates that step and every step after it.
  • Copy dependency manifests before source code so that code edits don't reinstall dependencies.
  • RUN is cached on its command text, not on what it downloads.
  • Cache mounts and remote cache export are the next level of build performance.

21. Security and the software supply chain

Hashes and OCI artifacts are the foundation of modern container supply-chain security, but each mechanism answers a different question.

QuestionMechanism
Did I get exactly the bytes that were named?Digest verification (automatic on every pull)
Will I get the same bytes tomorrow?Digest pinning (image@sha256:…)
Who produced these bytes, and were they approved?Signatures (e.g. Sigstore Cosign, Notary Project Notation)
How and from what source were they built?Provenance attestations (in-toto / SLSA provenance)
What software is inside?SBOM (SPDX, CycloneDX)
Is anything inside known to be vulnerable?Vulnerability scanning against CVE databases
Should this cluster run it?Admission policy that checks the above at deploy time

Integrity and pinning

Every pull verifies digests, so a registry, proxy or network can't substitute bytes without detection for the digest you asked for. The weak link is the step before that: resolving a tag to a digest.

FROM python:3.12-slim                         # whatever the tag points at when you build
FROM python:3.12-slim@sha256:<index digest>   # exactly this content, every time

image:latest is weaker for reproducibility because:

  • It is mutable: two builds or two nodes can silently get different content.
  • It is ambiguous: "latest" is just a tag name and says nothing about recency or stability.
  • A machine with a cached copy and a machine that pulls fresh can run different code under the same name.
  • Rollbacks and audits can't say exactly what ran.

A digest pin fixes what you run. It doesn't make it safe: a pinned image with a critical CVE stays vulnerable forever, and a pinned image from an attacker is still the attacker's image. Pin and update deliberately, for example with a dependency bot that proposes new digests. Pair pins with signature verification so you also know who produced them.

Signing, SBOMs and provenance as OCI artifacts

§21

The build produces an image and attestations, all addressed by digest. Signing binds an identity to that digest, and the signature is stored as another OCI object that refers to the image. At deploy time, a policy engine checks signatures and attestations for the exact digest being run. Because everything hangs off the digest, attaching metadata never changes the image, and verification never has to trust a tag.

Some specifics:

  • Signatures sign the manifest or index digest. Sigstore supports keyless signing: a short-lived certificate bound to a CI workload identity (via OIDC), with the event recorded in a public transparency log. Notary Project uses X.509 PKI.
  • SBOMs list the packages in the image (Debian packages, Python wheels). BuildKit can generate one at build time (--sbom=true), and scanners can generate them after the fact.
  • Provenance records the builder, source repository and commit, build parameters and materials. BuildKit can emit SLSA provenance (--provenance=mode=max). Attestations are claims by the builder, so their value depends on trusting the builder.
  • Scanning matches SBOM contents against vulnerability databases. Results are time-dependent: an unchanged image can go from clean to critical overnight when a new CVE is published. Rescan continuously, not just at build time.

Try It Yourself

docker buildx imagetools inspect username/container-lab:v1 --format '{{ json .Provenance }}'
docker buildx imagetools inspect username/container-lab:v1 --format '{{ json .SBOM }}'

Hygiene that the rest of this article explains

  • Secrets: never COPY then rm them. The bytes stay in a lower layer (section 4). Use RUN --mount=type=secret.
  • Root: our app runs as UID 0, which is host UID 0 without user namespaces (section 12). Add a USER instruction with an unprivileged UID for production.
  • Smaller bases (slim, distroless, minimal distributions) mean fewer packages to patch and a smaller attack surface for scanners to flag.
  • Keep the default seccomp, capability and LSM profiles. --privileged removes most of what makes a container a container.

What you should remember

  • Digests give integrity and reproducibility, not trustworthiness.
  • Signatures answer "who", provenance answers "how", SBOMs answer "what", and scanners answer "known vulnerable?"
  • All of them attach to the image's digest as OCI content, without changing the image.
  • image:latest is mutable and ambiguous. image@sha256:… is exact, but exact isn't the same as safe.

22. Containers vs virtual machines, precisely

DimensionContainers (standard runc-style)Virtual machines
KernelShared host kernel. The image carries userspace onlyEach VM boots its own guest kernel
Isolation boundaryKernel features: namespaces, cgroups, capabilities, seccomp, LSMsHardware virtualization enforced by the hypervisor (CPU virtualization extensions, EPT/NPT)
StartupProcess start plus rootfs setup, often under a second once the image is local. Real-world startup is usually dominated by image pull and application initializationKernel boot plus init. Seconds for traditional VMs. MicroVMs (e.g. Firecracker) can boot in the order of 100–200 ms
Memory overheadEssentially the process's own memory. Shared image files share page cacheGuest kernel plus guest OS services plus unused guest RAM, partly mitigated by ballooning and page sharing
PortabilitySame kernel family and CPU architecture (or emulation). Depends on kernel features and syscalls the app usesAny OS the hypervisor supports, same CPU architecture (or emulation)
Security boundaryWeaker by default: the whole kernel system-call interface is reachable (reduced by seccomp). Kernel bugs can mean escapesStronger: smaller hypervisor interface. Escapes are rarer but not impossible
OS requirementsLinux containers need a Linux kernel (hence Docker Desktop's VM on macOS). Windows containers need WindowsHypervisor support, and a guest OS image per VM
Operational modelImmutable images, declarative config, many small single-purpose units, orchestration (Kubernetes)Often long-lived, patched in place, configuration-managed. Also immutable images (AMIs) in modern setups

The distinction is blurring in useful ways. Sandboxed runtimes keep the container workflow (OCI images, Kubernetes) but put a VM (Kata Containers, Firecracker-based runtimes) or a user-space kernel (gVisor) behind each pod when you need a stronger boundary for untrusted code. Containers are not "always faster": a CPU-bound workload runs at about the same speed in either, and I/O through an overlay filesystem or a VM file-sharing layer can be slower than through a dedicated virtual disk.

What you should remember

  • The fundamental difference is one shared kernel vs one kernel per VM, and every other row follows from it.
  • Containers are lighter and faster to start, not inherently faster to run.
  • For hostile multi-tenant workloads, add a VM or sandbox boundary, which you can do without leaving the OCI ecosystem.

23. Containers and Kubernetes

Kubernetes doesn't run containers itself. It asks a container runtime to, through a standard interface.

§23

The kubelet translates a Pod into CRI gRPC calls. The CRI implementation (containerd's built-in CRI plugin, or CRI-O) creates a sandbox. With runc that is a tiny pause container whose job is to hold the pod's shared namespaces. A CNI plugin wires up the pod's network namespace. Each app container is then created with an OCI spec that joins those namespaces and gets its own cgroup limits, and is started by runc through a shim.

Everything in this article maps directly onto that flow:

Kubernetes conceptWhat it is underneathSection
image: username/container-lab@sha256:…An OCI reference, resolved and pulled by containerd5, 8, 18
PodA group of containers sharing network (and IPC, optionally PID) namespaces through a sandbox12
resources.limitscgroup memory.max / cpu.max for each container13
resources.requestsScheduling input plus cgroup weights, not a hard limit13
emptyDir, PVCMounts added to the OCI spec, living outside the writable layer11, 17
Container restartThe same pod sandbox, a new container process and a fresh writable layer16
securityContextUser, capabilities, seccomp, runAsNonRoot, user namespaces (hostUsers: false)12, 21
status.containerStatuses[].imageIDThe digest actually running18
crictl ps, crictl imagesInspect CRI-level containers and images on a node14

Once you know that a Pod is "a few processes sharing a network namespace, each with a cgroup and an overlay root from an OCI image", most Kubernetes behavior stops being magic. That includes OOMKilled, ImagePullBackOff, exec format error on mixed-architecture clusters, CrashLoopBackOff and slow terminations.

What you should remember

  • kubelet → CRI → containerd or CRI-O → shim → OCI runtime → Linux processes.
  • A Pod is a shared-namespace sandbox. Its containers join it.
  • Resource limits are cgroups, images are OCI content, and the pause container holds the namespaces.
  • Docker is not required on a Kubernetes node, but Docker-built OCI images run unchanged.

24. The master architecture diagram

§24

Read it top to bottom as the life of container-lab:v1:

  1. Build. BuildKit turns source code and a Dockerfile into content-addressed blobs: layers, a config and a manifest, optionally wrapped in an index. Each one points to the others by SHA-256.
  2. Distribute. A push uploads missing blobs and then the manifest, and a tag points at the top digest. SBOMs, signatures and provenance attach to that digest as OCI artifacts.
  3. Pull and store. A node, driven by docker or by kubelet through CRI, resolves the tag to a digest, fetches and verifies the missing blobs into the content store, and unpacks them into snapshots.
  4. Run. containerd prepares a writable snapshot and an OCI runtime spec. The shim invokes runc, which programs the kernel's namespaces, cgroups, OverlayFS and security features, then execves the application.
  5. What remains is an ordinary Linux process, parented by the shim, that sees a private world and can use a bounded share of the machine.

The diagram deliberately separates Docker from the parts that don't depend on it. Replace "docker CLI → dockerd" with "kubelet → CRI" and everything below containerd is identical.


The whole journey in one paragraph

docker build sends a Dockerfile to BuildKit, which compiles it into a graph, reuses cached steps, and exports layers (tar changesets), a config (runtime defaults and diff IDs) and a manifest (typed, sized pointers to both), all named by the SHA-256 of their bytes. A tag is a mutable name for the top digest. A registry stores blobs by digest and serves them over the OCI Distribution API, so push and pull skip whatever already exists and verify everything they transfer. Locally, containerd keeps the blobs in a content store and unpacks them into snapshots. docker run adds a writable snapshot, generates an OCI runtime spec, and hands it to runc through a shim. runc creates namespaces (what the process sees) and cgroups (what it can use), mounts an OverlayFS root, drops privileges, and execves your program. From that moment, the "container" is just your process, running directly on the host kernel.

Further reading