You have pulled images, listed them and deleted them, but you still do not know what they are on the inside. And that "inside" explains almost everything that will surprise you about Docker in the coming modules: why downloading the second image is faster than the first, why the sum of the sizes in docker images does not match the disk actually used, why the data you write inside a container disappears when you delete it, and why the latest tag is a trap. In this lesson you are going to dissect an image: its immutable nature, its layer model, the copy-on-write mechanism that separates image from container, the full naming scheme with registries and digests, the families of base images, and a hands-on exploration with docker image history and inspect on node:22-alpine, which will be the foundation of the Aurora Libros API.

Contents

  1. What exactly an image is
  2. The layer model
  3. The container's writable layer and copy-on-write
  4. The key consequence: data gets lost
  5. Full naming: registry, repository, tag and digest
  6. Why latest does not mean "the latest"
  7. Base images and official variants
  8. Hands-on exploration with node:22-alpine
  9. Proving that two images share layers

  1. What exactly an image is

A Docker image is two things packaged together:

  1. A filesystem: the complete set of directories and files the application will see when it runs. Binaries, libraries, configuration files, your code.
  2. Startup metadata: how it should be run. Which command to launch, with which environment variables, which user, in which working directory, which ports it declares.

And it has three properties that define its behavior:

  • It is immutable. Once built, it never changes. If you need a different version, you build another image.
  • It is read-only. No container can write to it.
  • It is cryptographically identifiable. Its content produces a SHA-256 hash that identifies it unambiguously.

Those three properties together are what make reproducibility possible: if two machines run the image with the same digest, they are running exactly the same bytes.

What an image does not contain, and it is worth being clear about from the start:

It does contain It does not contain
User-space binaries and libraries A kernel (it uses the host's)
Configuration files Your production database's data
Your application code Secrets that must vary by environment
Metadata: command, variables, user, ports Runtime state

  1. The layer model

Here is the central idea. An image is not a monolithic block: it is a stack of layers, each of which is a diff, a set of changes relative to the previous layer.

Picture how the Aurora Libros API image gets built:

  1. Layer 1: the base Alpine Linux filesystem (about 8 MB).
  2. Layer 2: on top, Node.js 22 is installed (new files are added).
  3. Layer 3: on top, the project's package.json is copied.
  4. Layer 4: on top, the npm dependencies are installed (node_modules).
  5. Layer 5: on top, the API's code is copied.
graph TB
    subgraph IMG["Image aurora-api:1.0.0 (read-only)"]
        L5["Layer 5 · API code · 120 KB"]
        L4["Layer 4 · node_modules · 48 MB"]
        L3["Layer 3 · package.json · 2 KB"]
        L2["Layer 2 · Node.js 22 · 42 MB"]
        L1["Layer 1 · Alpine Linux base · 8 MB"]
    end
    L1 --- L2 --- L3 --- L4 --- L5

When the container starts, all those layers are merged into a single unified view of the filesystem thanks to a union filesystem (on modern Linux, overlay2, the driver you saw in docker info). From inside the container you do not perceive layers: you see a normal / with /usr, /etc, /app and so on.

The merging rules are simple:

  • A file that exists in only one layer shows up.
  • If a file exists in several layers, the one in the topmost layer wins.
  • If a layer deletes a file from a lower one, it is marked with a special whiteout file and stops being visible (but it still takes up space in the lower layer; this consequence matters for image size and is exploited in lesson 05-04).

And now the property that makes everything efficient: layers are shared between images. If you have five different images built on node:22-alpine, the Alpine and Node layers are stored only once on disk and all five images reference them.

graph TB
    BASE["Layer: Alpine 3.20"] --> NODE["Layer: Node.js 22"]
    NODE --> API1["aurora-api:1.0.0<br/>+ 2 layers of its own"]
    NODE --> API2["aurora-api:1.1.0<br/>+ 2 layers of its own"]
    NODE --> WORKER["aurora-worker:1.0.0<br/>+ 3 layers of its own"]

Three very concrete benefits come out of that:

  • Disk savings: you do not multiply the base by each image.
  • Fast downloads: when you pull the second image, the daemon only fetches the layers it is missing. That is why you sometimes see Already exists instead of Pull complete.
  • Fast uploads: when you publish a new version of your API in which only the code changed, only that layer gets uploaded.

This last point will have enormous consequences when you write your first Dockerfile in module 2: the order of the instructions determines which layers get reused.

  1. The container's writable layer and copy-on-write

If the image is read-only, how can a container write files? The answer is the piece that joins the two concepts:

When you create a container, Docker stacks a thin read-write layer on top of the image's read-only layers. That layer belongs only to that container.

graph TB
    subgraph C1["Container A"]
        WA["Writable layer A<br/>(read/write)"]
    end
    subgraph C2["Container B"]
        WB["Writable layer B<br/>(read/write)"]
    end
    subgraph IMG["Image nginx:alpine (read-only, SHARED)"]
        I3["Layer 3"]
        I2["Layer 2"]
        I1["Layer 1"]
    end
    WA --> I3
    WB --> I3
    I3 --- I2 --- I1

Two containers from the same image share all of the image's layers and only own their topmost layer. That is why starting the tenth container from an image costs practically no disk.

The mechanism that makes this possible is called copy-on-write, and it works like this:

Operation inside the container What actually happens
Reading a file from the image It is read straight from the image's layer. Zero cost
Creating a new file It is created in the container's writable layer
Modifying a file from the image It is copied in full to the writable layer and modified there. The image is untouched
Deleting a file from the image It is marked as deleted in the writable layer (whiteout). It still exists in the image

Practical consequences you will notice:

  • Modifying a large file for the first time is slow, because it has to be copied in full first. Subsequent times are fast.
  • Write-intensive applications (databases, precisely) perform worse if they write to the container's layer. The solution is to use volumes, and that is why aurora-db will need one.

You can see the size of that writable layer:

docker ps -as
CONTAINER ID   IMAGE          NAMES             SIZE
3f2a9c1b7e4d   nginx:alpine   aurora-web-demo   1.09kB (virtual 52.5MB)

How to read it: 1.09kB is what this container has written to its own layer; virtual 52.5MB is the total including the image's layers, which are shared. If you start five nginx:alpine containers, disk usage does not grow by 5 × 52.5 MB, but by 52.5 MB plus five tiny writable layers.

  1. The key consequence: data gets lost

This is the most important practical implication of everything above, and it deserves its own block because it causes most beginners' scares:

The writable layer lives and dies with the container. When you delete a container, its writable layer is deleted with it. Everything the application had written there disappears forever.

See it for yourself, step by step:

docker run --name data-test alpine:3.20 \
  sh -c "echo 'Aurora Libros catalog' > /data.txt && cat /data.txt"
Aurora Libros catalog

The container wrote the file into its writable layer and displayed it. Now restart that same container with docker start -a (the -a is --attach, so you can see its output):

docker start -a data-test
Aurora Libros catalog

It still works: the container's writable layer persists for as long as the container exists, even when stopped. And now the revealing part:

docker rm data-test
docker run --name data-test alpine:3.20 cat /data.txt
cat: can't open '/data.txt': No such file or directory

What just happened: deleting the container removed its writable layer and /data.txt with it. The new container is born from the image, which never contained that file. The image is immutable: writing inside a container does not modify the image.

Translated to Aurora Libros: if you drop PostgreSQL into a container without further thought and delete that container, you lose the entire book catalog. In a development environment that is annoying; in production it is a serious incident.

The solution is called volumes: storage that lives outside the container's lifecycle. It is the full subject of lesson 03-06, and in lesson 01-06 you will make a very introductory first use of -v to serve an Aurora Libros page with Nginx. For now, hold on to the rule:

Any data that must outlive the container has to be in a volume.

  1. Full naming: registry, repository, tag and digest

When you type nginx:alpine, you are using a shorthand. An image's full name is:

[registry[:port]/][user-or-organization/]repository[:tag][@digest]

Examples, from most abbreviated to most explicit:

What you type How Docker resolves it
nginx docker.io/library/nginx:latest
nginx:alpine docker.io/library/nginx:alpine
bitnami/postgresql:16 docker.io/bitnami/postgresql:16
ghcr.io/aurora-libros/aurora-api:1.2.0 As-is: GitHub registry, organization, repository, tag
registry.auroralibros.local:5000/aurora-api:1.2.0 Private registry on port 5000

The auto-completion rules are:

  • If you do not state a registry, docker.io (Docker Hub) is assumed.
  • If you do not state a user or organization and the registry is Docker Hub, library/ is assumed, the namespace reserved for official images. That is why nginx is official and bitnami/postgresql is third-party.
  • If you do not state a tag, :latest is assumed.

The digest

Beyond name and tag, every image has a digest: the SHA-256 hash of its manifest.

docker images --digests nginx
REPOSITORY   TAG      DIGEST                                                                    IMAGE ID       SIZE
nginx        alpine   sha256:c15da6c91de8d2f436196f3a768483ad32c258ed4e1beb3d367a27ed67253e66   3f8a4339aadd   52.5MB

The difference between the three identifiers:

Identifier Example Nature
Tag nginx:alpine Mutable: it can be reassigned to another image at any time
Image ID 3f8a4339aadd Local hash of the configuration; immutable, but only meaningful on your machine
Digest sha256:c15da6c9... Immutable and global: it identifies exactly the same bytes on any registry and machine

And that is why you can pin an image with no ambiguity whatsoever:

docker pull nginx@sha256:c15da6c91de8d2f436196f3a768483ad32c258ed4e1beb3d367a27ed67253e66

This downloads that exact image, whatever happens to the tags. It is the recommended practice in production and in CI pipelines, where reproducibility matters more than convenience. We will come back to it in lessons 02-06 and 06-01.

  1. Why latest does not mean "the latest"

This is one of Docker's most expensive misunderstandings, so let's be blunt:

latest is not a function that returns the most recent version. It is simply the tag name Docker uses by default when you do not specify one.

Nothing forces latest to point at the newest version. It is a convention, and one that is frequently broken:

  • A maintainer can publish 2.0.0 without updating latest, which stays at 1.x.
  • Some projects use latest for the newest stable release and publish new major versions only under their number.
  • And the other way round: latest may point at an unstable development version.

The specific problems using it causes:

Problem Explanation
Not reproducible docker pull myapp:latest today and tomorrow can bring different images. Your CI can pass today and fail tomorrow without anyone touching the code
Silent updates A breaking major change can land in your deployment without you deciding it
Deceptive cache If you already have latest locally, Docker does not check the registry again unless you force the pull. You may be running a months-old version believing it is the newest
Impossible diagnostics "Which version is in production?" — "latest". That is not an answer

The correct practice:

# Bad: you do not know what you are running
docker pull node

# Better: major version and variant pinned
docker pull node:22-alpine

# Better still for production: a specific version
docker pull node:22.14.0-alpine3.21

# Maximum rigor: pinned by digest
docker pull node:22-alpine@sha256:9d2b8c0...

Rule of thumb: in development, node:22-alpine is a reasonable balance between stability and convenience. In production and CI, full version or digest. Throughout the course we will always use explicit tags: node:22-alpine, postgres:16-alpine, redis:7-alpine, nginx:alpine.

  1. Base images and official variants

A base image is the starting point you build on. These are the families you will come across:

Base image Approx. size Package manager Libc When to use it
scratch 0 bytes None None Static binaries (Go, Rust). Maximum security and minimum size
alpine ~8 MB apk musl The default when size matters
debian:bookworm-slim ~75 MB apt glibc Broad compatibility at a contained size
debian:bookworm ~120 MB apt glibc When you need system tooling
ubuntu:24.04 ~78 MB apt glibc Familiarity, Ubuntu packages
distroless images ~20 MB None glibc Hardened production, no shell

An important warning about Alpine: it uses musl as its C library instead of the usual glibc. 95% of the time it makes no difference, but it can cause problems with Python or Node packages that ship binaries compiled against glibc, and in specific cases there are performance differences in DNS name resolution. If something behaves oddly on Alpine and not on Debian, that is usually the cause.

Variants of the official language images

The official images publish several variants of each version, and choosing well matters. Take Node.js 22, which is what aurora-api will need:

Tag Base Approx. size Comment
node:22 Full Debian ~1.1 GB Ships compilers and tooling. Convenient, enormous
node:22-slim debian:slim ~230 MB Debian without extras. A good balance
node:22-alpine Alpine ~140 MB The smallest one with a package manager. The one we will use
node:22-bookworm Debian 12 ~1.1 GB Same as node:22 but pinning the Debian version

Databases follow the same pattern: postgres:16 (Debian) versus postgres:16-alpine, and redis:7 versus redis:7-alpine.

And a security recommendation that will make full sense in lesson 05-03: always prefer official or verified images. Anyone can publish an image called fast-postgresql on Docker Hub with whatever they like inside.

  1. Hands-on exploration with node:22-alpine

Let's dissect the image that will be the base of aurora-api.

Pull it

docker pull node:22-alpine
22-alpine: Pulling from library/node
f18232174bc9: Already exists
9c3b5e8d1a2f: Pull complete
7d4a2c9e0f31: Pull complete
b1e6f8a4c507: Pull complete
Digest: sha256:9d2b8c0f4e1a7b3d5c6e8f0a2b4d6e8f0a1c3e5d7f9b1d3f5a7c9e1b3d5f7a9c
Status: Downloaded newer image for node:22-alpine
docker.io/library/node:22-alpine

Look at the first layer line: Already exists. That layer (f18232174bc9) is the Alpine base system, and you already had it from the docker pull alpine:3.20 in the previous lesson. Docker did not download it again. You have just seen layer sharing at work.

List it

docker image ls node
REPOSITORY   TAG         IMAGE ID       CREATED       SIZE
node         22-alpine   c4b8e2f1a9d3   5 days ago    142MB

View its layer history

docker image history node:22-alpine
IMAGE          CREATED       CREATED BY                                      SIZE      COMMENT
c4b8e2f1a9d3   5 days ago    CMD ["node"]                                    0B
<missing>      5 days ago    ENTRYPOINT ["docker-entrypoint.sh"]             0B
<missing>      5 days ago    COPY docker-entrypoint.sh /usr/local/bin/ #…    388B
<missing>      5 days ago    RUN /bin/sh -c apk add --no-cache --virtual …   7.66MB
<missing>      5 days ago    ENV YARN_VERSION=1.22.22                        0B
<missing>      5 days ago    RUN /bin/sh -c addgroup -g 1000 node && addu…   125MB
<missing>      5 days ago    ENV NODE_VERSION=22.14.0                        0B
<missing>      3 weeks ago   /bin/sh -c #(nop) CMD ["/bin/sh"]               0B
<missing>      3 weeks ago   /bin/sh -c #(nop) ADD file:8f4a3d2b1c… in /     8.83MB

It reads from the bottom up, in build order. Here are the key points:

  • The last line (ADD file:... in /, 8.83 MB) is Alpine's base layer. It matches exactly the size of alpine:3.20 you saw earlier: it is literally the same layer.
  • RUN ... addgroup ... 125MB is the Node.js installation: by far the heaviest layer. That is where the bulk of the 142 MB lives.
  • The 0B layers (ENV, CMD, ENTRYPOINT) add no files: they only modify the startup metadata. They take up zero because they do not change the filesystem.
  • <missing> in the IMAGE column is not an error: it means that intermediate layer has no image ID of its own on your machine, because you downloaded the image instead of building it. It is entirely normal.

To see the output untruncated:

docker image history node:22-alpine --no-trunc --format "table {{.Size}}\t{{.CreatedBy}}"

Inspect its metadata

docker image inspect node:22-alpine

It returns a long JSON blob. Let's extract what matters:

docker image inspect node:22-alpine --format '{{json .Config}}'
{
  "Env": [
    "PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin",
    "NODE_VERSION=22.14.0",
    "YARN_VERSION=1.22.22"
  ],
  "Cmd": ["node"],
  "Entrypoint": ["docker-entrypoint.sh"],
  "WorkingDir": "",
  "Labels": null
}

These are the startup metadata from section 1:

  • Env: environment variables every container from this image will have.
  • Cmd: the default command. If you run docker run node:22-alpine with nothing else, it runs node and opens its interactive console for you.
  • Entrypoint: a script that runs first and receives Cmd as its argument.
  • WorkingDir: the initial working directory.

The distinction between Cmd and Entrypoint is subtle and important, and is covered thoroughly in lesson 02-04.

Other useful fields:

docker image inspect node:22-alpine --format 'Architecture: {{.Architecture}} / OS: {{.Os}}'
docker image inspect node:22-alpine --format 'Layers: {{len .RootFS.Layers}}'
docker image inspect node:22-alpine --format '{{range .RootFS.Layers}}{{println .}}{{end}}'
Architecture: amd64 / OS: linux
Layers: 4
sha256:f18232174bc91741fdf3da96d85011092101a032a93a388b79e99e69c2d5c870
sha256:9c3b5e8d1a2f...
sha256:7d4a2c9e0f31...
sha256:b1e6f8a4c507...

The last command lists the image's layer hashes. Keep hold of the first one in that list: you are going to compare it in the next section.

  1. Proving that two images share layers

Let's demonstrate the sharing beyond any doubt. First, make sure you have both images:

docker pull alpine:3.20
docker pull node:22-alpine

Now compare their layer lists:

docker image inspect alpine:3.20 --format '{{range .RootFS.Layers}}{{println .}}{{end}}'
sha256:f18232174bc91741fdf3da96d85011092101a032a93a388b79e99e69c2d5c870
docker image inspect node:22-alpine --format '{{range .RootFS.Layers}}{{println .}}{{end}}'
sha256:f18232174bc91741fdf3da96d85011092101a032a93a388b79e99e69c2d5c870
sha256:9c3b5e8d1a2f...
sha256:7d4a2c9e0f31...
sha256:b1e6f8a4c507...

The first layer is identical in both. Same hash, same bytes, stored only once in /var/lib/docker/overlay2. The Node image does not contain its own copy of Alpine: it references it.

You can automate the comparison:

comm -12 \
  <(docker image inspect alpine:3.20 --format '{{range .RootFS.Layers}}{{println .}}{{end}}' | sort) \
  <(docker image inspect node:22-alpine --format '{{range .RootFS.Layers}}{{println .}}{{end}}' | sort)

What it does: comm -12 compares two sorted lists and shows only the lines they have in common (it suppresses the ones exclusive to the first with -1 and to the second with -2). The <(...) constructs are bash process substitutions: they pass each command's output as though it were a file. The result is the list of shared layers.

And the confirmation in numbers:

docker image ls
docker system df
REPOSITORY   TAG         IMAGE ID       SIZE
node         22-alpine   c4b8e2f1a9d3   142MB
nginx        alpine      3f8a4339aadd   52.5MB
alpine       3.20        a8560b36e8b8   8.83MB
TYPE      TOTAL     ACTIVE    SIZE      RECLAIMABLE
Images    3         0         194.5MB   194.5MB (100%)

Add up the three images: 142 + 52.5 + 8.83 = 203.3 MB. But docker system df reports 194.5 MB. The difference of almost 9 MB is exactly Alpine's base layer, counted three times in docker image ls and only once on disk.

There you have the answer to one of the questions from the start of the lesson: the SIZE column of docker image ls shows each image's total size, without discounting what is shared. To know the real space used, use docker system df.

Common Mistakes and Tips

  • Believing an image can be modified. It cannot: it is immutable. What you write inside a container goes to its writable layer, not to the image. To "change" an image you build another one (module 2).
  • Losing data by not using volumes. This is the classic mistake. If your PostgreSQL container stores the Aurora Libros catalog in its writable layer, one docker rm wipes it all. Solution: volumes (lesson 03-06).
  • Using latest anywhere serious. It does not mean "the latest", it is not reproducible and it makes it impossible to answer "which version is deployed". Pin explicit tags, and digests in production.
  • Adding up the SIZE column of docker image ls. It gives a number larger than reality because of shared layers. Use docker system df.
  • Thinking that deleting a file in a layer shrinks the image. It does not: the lower layer still contains it, it is merely hidden. It is the reason a badly built image weighs more than it should (lesson 05-04).
  • Confusing Image ID and digest. The Image ID is local; the digest is global and it is the one you should use to pin an exact version.
  • Tip: docker image history before any optimization. It tells you at a glance which layer is fattening the image.
  • Tip: choose the right variant. node:22 weighs about eight times more than node:22-alpine. In a pipeline that builds twenty times a day, that difference shows up in time and in the bill.

Exercises

Exercise 1: layer arithmetic

Pull these three images and answer:

docker pull alpine:3.20
docker pull node:22-alpine
docker pull redis:7-alpine
  1. What do the sizes shown by docker image ls add up to?
  2. How much space do they really take on disk according to docker system df?
  3. Explain the difference by identifying which specific layer (by hash) is responsible.
  4. In the pull output, which ones showed Already exists lines, and why?

Exercise 2: the persistence lesson

Run this sequence and explain the outcome of each step:

docker run --name books alpine:3.20 sh -c "echo 'Rayuela, Julio Cortázar' > /catalog.txt && cat /catalog.txt"
docker start -a books
docker rm books
docker run --name books alpine:3.20 cat /catalog.txt

Then answer: (a) why does the second command find the file and the fourth one not?, (b) at exactly which moment was the data lost?, (c) what would have happened if instead of docker rm you had only run docker stop?

Exercise 3: identify an image unambiguously

For the node:22-alpine image on your machine, obtain: (a) its Image ID, (b) its digest, (c) how many layers it has, (d) its default command and its environment variables, and (e) the size of the largest layer and which instruction generated it. Then, write the docker pull command that would download exactly that same image two years from now, even if the 22-alpine tag has changed content.

Solutions

Solution to exercise 1

docker image ls --format "{{.Size}}\t{{.Repository}}:{{.Tag}}"
docker system df
  1. The sum of the SIZE column will be around 200 MB (roughly 8.8 + 142 + 41 depending on versions).
  2. docker system df will report a smaller figure, about 9 MB below the sum.
  3. The culprit is Alpine's base layer. Verify it like this:
for img in alpine:3.20 node:22-alpine redis:7-alpine; do
  echo "$img:"
  docker image inspect $img --format '{{index .RootFS.Layers 0}}'
done

This loop prints the first layer of each image (index .RootFS.Layers 0 takes element 0 of the array). All three show the same sha256:f1823217...: it is the same physical layer stored once and referenced three times. docker image ls counts it in each image; docker system df counts it once, hence the difference.

  1. Already exists appears in the second and third pull for that base layer. In the first one (alpine:3.20) it had to be downloaded; from then on, the daemon checks the hash, sees it already has it and skips it. That is the practical reason why successive downloads of images from the same family are much faster.

Solution to exercise 2

Step by step:

  1. docker run --name books ... creates a new container, writes /catalog.txt into its writable layer and displays it. Output: Rayuela, Julio Cortázar. The container finishes and is left as Exited (0).
  2. docker start -a books resumes the same container (it does not create another). It runs its command again, which rewrites and displays the file. The writable layer was still intact.
  3. docker rm books removes the container and its writable layer with it. /catalog.txt ceases to exist.
  4. docker run --name books alpine:3.20 cat /catalog.txt creates a new container from the image. Output:
cat: can't open '/catalog.txt': No such file or directory

Answers:

(a) The second command finds the file because it operates on the same container, whose writable layer persists for as long as the container exists. The fourth does not find it because it operates on a different container, born from the alpine:3.20 image, which never contained that file. Writing inside a container does not modify the image: the image is immutable.

(b) The data was lost exactly at the docker rm. That is the moment the writable layer is erased from disk.

(c) With docker stop the file would still be there. Stopping a container halts its process but preserves its writable layer; a later docker start -a books would have shown the content again. Only rm destroys the data. This distinction between "stopped" and "removed" is central to the lifecycle you will see in lesson 03-02.

Solution to exercise 3

# (a) Image ID
docker image inspect node:22-alpine --format '{{.Id}}'

# (b) Digest
docker image inspect node:22-alpine --format '{{index .RepoDigests 0}}'

# (c) Number of layers
docker image inspect node:22-alpine --format '{{len .RootFS.Layers}}'

# (d) Default command and environment variables
docker image inspect node:22-alpine --format 'Cmd: {{.Config.Cmd}}{{println}}Env: {{.Config.Env}}'

# (e) Largest layer and its instruction
docker image history node:22-alpine --no-trunc --format "{{.Size}}\t{{.CreatedBy}}" | sort -h -r | head -1

Comments:

  • .Id returns the full Image ID with the sha256: prefix; the first 12 characters are what you see in docker image ls.
  • .RepoDigests is an array because an image can be published in several repositories; index ... 0 takes the first. The format is node@sha256:....
  • len .RootFS.Layers counts the elements of the layer array: normally 4 for this image.
  • The largest layer will be the RUN that installs Node.js, around 125 MB. That single line accounts for most of the image's size.

The command that downloads exactly that image two years from now:

docker pull node@sha256:9d2b8c0f4e1a7b3d5c6e8f0a2b4d6e8f0a1c3e5d7f9b1d3f5a7c9e1b3d5f7a9c

Substituting the digest that part (b) returned to you. The key to the reasoning: the tag 22-alpine is a mutable pointer that the maintainer can reassign to a different image every time they publish a Node 22 patch. The digest, on the other hand, is the hash of the content: it identifies those specific bytes and nothing else. As long as the image still exists in the registry, that pull will always bring back the same thing. It is the standard practice for reproducibility in CI and production.

Conclusion

An image is a filesystem plus startup metadata, immutable, read-only and cryptographically identified. Internally it is a stack of layers, each one a diff of the previous, which merge into a single view and —this is the important part— are shared between images, which is where the incremental downloads and the disk savings come from, as you have verified with your own hands by comparing the hashes of alpine:3.20 and node:22-alpine.

When you create a container, Docker adds a writable layer of its own on top and applies copy-on-write: reading is free, modifying means copying. And from that comes the rule that prevents the most grief: that layer dies with the container, so any data that must survive has to go into a volume (lesson 03-06).

You also know now how to name images precisely —registry/user/repository:tag@digest—, why latest is a name and not a promise, and what families of base images exist, including node:22-alpine, which will be the foundation of the Aurora Libros API.

With image theory settled, nothing is stopping you from getting your hands dirty. In the next lesson, Creating Your First Docker Container, you will do a complete guided practice: a container in the foreground, an Nginx server in the background with its port published, an interactive shell inside Alpine, and your first Aurora Libros page served from a container.

Docker: From Beginner to Advanced

Module 1: Introduction to Docker

Module 2: Working with Docker Images

Module 3: Docker Containers

Module 4: Docker Compose

Module 5: Advanced Docker Concepts

Module 6: Docker in Production

Module 7: Docker Ecosystem and Tools

© Copyright 2026. All rights reserved