The auroralibros/aurora-api:1.2.0 image weighs 142 MB. That is not a disaster, but it is not optimized either: a C compiler travels inside it, along with the development dependencies and files nobody ever runs. In this lesson you put it on a diet measuring at every step, down to a specific figure, without losing any of the hardening from lesson 05-03.

Contents

  1. Why size matters and what must not be sacrificed
  2. Measure first: image ls and history
  3. dive: exploring layer by layer
  4. Choosing the base, with numbers
  5. Multi-stage builds: the concept
  6. Multi-stage applied to aurora-api
  7. Named stages and --target
  8. A common base image for several images
  9. Reducing layers and content
  10. Deleting a file does not reduce the size (demonstrated)
  11. Ordering for the cache
  12. Final verification and summary table

  1. Why size matters and what must not be sacrificed

Reason Concrete impact
Download time Every new node downloads the whole image the first time
Storage cost The registry, each node's cache, and every version you keep
Attack surface Fewer packages, fewer CVEs (lesson 05-03)
Deployment speed An urgent rollback is only as fast as the download
Autoscaling A traffic spike demands replicas now

On a platform with five nodes and fifteen deployments a month, shaving 40 MB off an image means roughly 3 GB less transfer per month and tens of seconds less on every cold start.

And what you do not sacrifice in exchange for weight: the HEALTHCHECK, the unprivileged user, the OCI labels for traceability, the ability to diagnose a production problem, and the reproducibility of the build. A 40 MB image nobody knows how to debug at three in the morning is a bad deal.

  1. Measure first: image ls and history

docker image ls auroralibros/aurora-api --format "{{.Tag}}\t{{.Size}}"
docker image history auroralibros/aurora-api:1.2.0 --format "{{.Size}}\t{{.CreatedBy}}" | head -8
1.2.0   142MB
21.4MB   RUN npm ci --omit=dev
3.1MB    COPY src/ ./src/
1.8MB    RUN apk add --no-cache tini
0B       ENV NODE_ENV=production
0B       USER node
0B       HEALTHCHECK &{["CMD-SHELL" "wget -qO- ..."]}
115MB    /bin/sh -c #(nop) ADD file:... in /

The diagnosis is immediate and it surprises almost everybody the first time: of the 142 MB, 115 MB are the base image. The dependencies are 21 MB and your own code, 3 MB. The main lever is not in your code: it is in which base you pick and what you add on top of it.

To compare precisely, use the real size (docker image inspect --format '{{.Size}}' | numfmt --to=iec) and not the one the registry shows, which is compressed and always looks smaller.

  1. dive: exploring layer by layer

docker history tells you how much each layer weighs; dive tells you what is inside it and how much is wasted.

docker run --rm -it \
  -v /var/run/docker.sock:/var/run/docker.sock \
  wagoodman/dive:latest auroralibros/aurora-api:1.2.0
│ Image Details │
Total Image size: 142 MB
Potential wasted space: 9.7 MB
Image efficiency score: 93 %

Count   Total Space  Path
    2         6.1 MB /app/node_modules/.cache
    2         2.4 MB /root/.npm/_cacache
    3         1.2 MB /app/.git

Three findings and three decisions: the npm cache should not be in the final image, node_modules/.cache is build junk, and .git should never have entered the build context in the first place (there is a line missing from .dockerignore).

The efficiency score measures how much content is written into one layer and then overwritten or deleted in another; below 95% there is work to do. To automate it in CI, add -e CI=true and --lowestEfficiency=0.95: dive exits with a non-zero code if the threshold is not met.

  1. Choosing the base, with numbers

Base Size libc Tooling When to choose it
node:22 ~1.1 GB glibc Everything: git, compilers, curl Only as a build stage
node:22-slim ~230 MB glibc Debian's minimum Native dependencies that insist on glibc
node:22-alpine ~135 MB musl BusyBox, apk The general case: the sweet spot
distroless/nodejs22 ~110 MB glibc None Hardened production
scratch 0 MB — None Static binaries only (Go, Rust)

Alpine's trade-off has a name: musl instead of glibc. The real consequences:

  • npm packages with binaries precompiled for glibc may not work and may have to be compiled on install (slower, and it requires build-base).
  • Some DNS-heavy or name-resolution-heavy workloads behave differently under musl.
  • Certain scientific and machine-learning libraries do not support musl.

aurora-api uses Express, pg and redis, none of which has problematic native dependencies: Alpine is the right choice. If one day a dependency caused trouble, node:22-slim costs 95 MB more and fixes it.

  1. Multi-stage builds: the concept

A multi-stage build uses several base images in the same Dockerfile. The intermediate stages compile and prepare; the final stage copies only the result. Everything else —compilers, caches, development dependencies, source code— stays out of the published image.

flowchart LR
  subgraph E1["Stage: dependencies (node:22-alpine)"]
    A["package*.json"] --> B["npm ci --omit=dev<br/>+ build-base to compile"]
  end
  subgraph E2["Stage: development"]
    C["full npm ci<br/>devDependencies, tests"]
  end
  subgraph E3["Final stage (node:22-alpine)"]
    D["COPY --from=dependencies node_modules"]
    E["COPY src/"]
  end
  B -.->|"COPY --from"| D
  E3 --> F["Published image:<br/>no compilers, no caches"]
  E2 -.->|"only with --target"| G["Development image"]

Two rules that sum up the mechanism:

  • Only the last stage (or the one named with --target) ends up in the image; the rest are discarded.
  • COPY --from=<stage> brings specific files over from an earlier stage, and nothing else travels: not its layers, not its history, not its secrets.

  1. Multi-stage applied to aurora-api

The starting point (a single stage, 142 MB) and the destination:

# syntax=docker/dockerfile:1
ARG NODE_VERSION=22

# ---------- Stage 1: production dependencies ----------
FROM node:${NODE_VERSION}-alpine AS dependencies
WORKDIR /app
# build-base lives only here: it compiles whatever is needed and does NOT travel to the final stage
RUN apk add --no-cache --virtual .build python3 make g++
COPY package.json package-lock.json ./
RUN npm ci --omit=dev && npm cache clean --force

# ---------- Stage 2: development and tests (not published) ----------
FROM node:${NODE_VERSION}-alpine AS development
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
CMD ["node", "--watch", "src/server.js"]

# ---------- Stage 3: final, minimal image ----------
FROM node:${NODE_VERSION}-alpine AS production
ARG VERSION=1.3.0
ARG REVISION=unknown
LABEL org.opencontainers.image.title="aurora-api" \
      org.opencontainers.image.version="${VERSION}" \
      org.opencontainers.image.revision="${REVISION}" \
      org.opencontainers.image.source="https://github.com/aurora-libros/aurora-api"
WORKDIR /app
ENV NODE_ENV=production PORT=3000
COPY --from=dependencies --chown=node:node /app/node_modules ./node_modules
COPY --chown=node:node package.json ./
COPY --chown=node:node src/ ./src/
USER node
EXPOSE 3000
HEALTHCHECK --interval=30s --timeout=3s --start-period=10s --retries=3 \
  CMD wget -qO- http://127.0.0.1:3000/health || exit 1
ENTRYPOINT ["node"]
CMD ["src/server.js"]
docker build -t auroralibros/aurora-api:1.3.0 --build-arg VERSION=1.3.0 ./api
docker image ls auroralibros/aurora-api --format "{{.Tag}}\t{{.Size}}"
1.3.0   121MB
1.2.0   142MB

142 MB → 121 MB: 15% less with a single structural change. Where do those 21 MB come from? From build-base (python3, make, g++), which used to be installed in the final image and now lives and dies inside the dependencies stage, and from the npm cache.

Notice that none of the hardening has been lost: USER node, the HEALTHCHECK, the OCI labels and the exec-form ENTRYPOINT/CMD are all still there. Optimizing is not trimming security.

  1. Named stages and --target

Named stages are not just organization: they are separately buildable artifacts.

docker build --target development -t aurora-api:dev ./api
docker build --target dependencies -t aurora-api:deps ./api
docker build -t auroralibros/aurora-api:1.3.0 ./api      # the last stage
docker image ls aurora-api --format "{{.Tag}}\t{{.Size}}"
dev     318MB
deps    198MB

The development image weighs 318 MB and that is exactly as it should be: it carries the devDependencies, the test framework and the debugging tools. It is precisely the target: development that compose.override.yaml has been using since lesson 04-07, and now you can see where it comes from.

A very useful pattern is a test stage that fails the build if the tests fail:

FROM development AS tests
RUN npm test
docker build --target tests ./api || echo "BUILD BROKEN: tests are red"

BuildKit does not run that stage when building the production one, because it is not part of its dependency graph (lesson 05-05). It only runs if you ask for it with --target.

  1. A common base image for several images

When you publish several Node services, it pays to have your own base with everything they share:

# aurora-base/Dockerfile -> auroralibros/aurora-base-node:22
FROM node:22-alpine
RUN apk add --no-cache tini tzdata && \
    addgroup -g 1001 aurora && adduser -u 1001 -G aurora -s /bin/sh -D aurora
ENV TZ=Europe/Madrid NODE_ENV=production
ENTRYPOINT ["/sbin/tini", "--"]
FROM auroralibros/aurora-base-node:22 AS production

Advantages: a single security update propagates to every service, the base layer is shared on disk and on download between them, and the shared decisions live in one place. Downside: you create an internal dependency that has to be maintained and versioned. It starts to pay off from three or four services onwards.

  1. Reducing layers and content

# ❌ Three layers, and the apk cache stays inside forever
RUN apk update
RUN apk add curl
RUN rm -rf /var/cache/apk/*

# ✅ One layer, no cache: what is never written takes no space
RUN apk add --no-cache curl
Technique What it does Typical saving
apk add --no-cache Avoids writing the package index 5-10 MB
apt-get ... && rm -rf /var/lib/apt/lists/* in the same RUN The same on Debian 30-50 MB
--no-install-recommends Does not install "suggested" packages 20-100 MB
npm ci --omit=dev No development dependencies 40-70% of node_modules
npm cache clean --force Removes ~/.npm/_cacache 10-30 MB
--virtual .build + apk del .build Temporary compilers in the same layer 100-200 MB
A well-tuned .dockerignore Keeps out what does not belong Variable, sometimes enormous

aurora-api's .dockerignore after what dive revealed:

.git
.gitignore
node_modules
npm-debug.log
Dockerfile*
compose*.yaml
.env
.env.*
!.env.example
coverage/
.nyc_output/
*.md
!README.md
.vscode/
.idea/
tests/
**/*.test.js
du -sh api/ && docker build -t aurora-api:ctx ./api 2>&1 | grep "transferring context"
94M     api/
=> transferring context: 812.43kB

From a 94 MB directory to 812 kB of context. That does not only make the build faster: it guarantees that .git and .env cannot sneak into any layer, which is a security matter as much as a size one.

  1. Deleting a file does not reduce the size (demonstrated)

FROM alpine:3
RUN dd if=/dev/zero of=/temp.bin bs=1M count=100    # layer A: +100 MB
RUN rm /temp.bin                                     # layer B: whiteout
docker build -q -t delete:bad /tmp/delete
docker image ls delete:bad --format "{{.Size}}"
docker run --rm delete:bad ls -la /temp.bin 2>&1
docker image history delete:bad --format "{{.Size}}\t{{.CreatedBy}}" | head -3
113MB
ls: /temp.bin: No such file or directory
0B       RUN rm /temp.bin
105MB    RUN dd if=/dev/zero of=/temp.bin bs=1M count=100
8MB      /bin/sh -c #(nop) ADD file:... in /

The file does not exist and the image still weighs 113 MB. The explanation is the one from lesson 05-02: layers are immutable, and rm in a later layer only creates a whiteout that hides the file; the 105 MB are still down there, downloaded on every pull and stored on every node.

The fix is to combine creation and deletion in the same instruction:

FROM alpine:3
RUN dd if=/dev/zero of=/temp.bin bs=1M count=100 && \
    echo "use the file..." && \
    rm /temp.bin
docker build -q -t delete:good /tmp/delete2 && docker image ls delete:good --format "{{.Size}}"
8.15MB

From 113 MB to 8.15 MB. The same logic, the same functional result, one && of difference. And it is exactly the reason for --virtual .build in the Dockerfile from section 6: install, compile and uninstall inside a single RUN.

  1. Ordering for the cache

A reminder from lesson 02-02, now combined with multi-stage: instructions are ordered from least to most likely to change.

COPY package.json package-lock.json ./   # rarely changes
RUN npm ci --omit=dev                    # expensive layer, reusable
COPY src/ ./src/                         # changes with every commit

With multi-stage the effect is multiplied: when you touch src/server.js, BuildKit reuses the entire dependencies stage —including the npm install— and only rebuilds the last two layers of the final stage.

touch api/src/server.js
time docker build -t auroralibros/aurora-api:1.3.0 ./api
 => CACHED [dependencies 4/4] RUN npm ci --omit=dev
 => [production 5/5] COPY --chown=node:node src/ ./src/     0.2s
real    0m3.184s

Three seconds against the forty-odd of a clean build. Optimizing size and optimizing build speed are, almost always, the same job.

  1. Final verification and summary table

An optimized image that does not work is worth nothing. The check is mandatory:

docker compose up -d --wait
docker compose exec aurora-api node --version
curl -s http://localhost:8080/api/health | jq -c
curl -s http://localhost:8080/api/books | jq -c '{source, total}'
curl -s http://localhost:8080/api/books/3 | jq -c '{title, author}'
docker inspect aurora-libros-aurora-api-1 --format 'user={{.Config.User}} health={{.State.Health.Status}}'
v22.11.0
{"status":"ok","db":"connected","cache":"connected","version":"1.3.0"}
{"source":"db","total":9}
{"title":"Cien años de soledad","author":"Gabriel García Márquez"}
user=node health=healthy

And the final step up, if your context allows it: switching the final base to distroless.

FROM gcr.io/distroless/nodejs22-debian12 AS production
COPY --from=dependencies --chown=1000:1000 /app/node_modules /app/node_modules
COPY --chown=1000:1000 src/ /app/src/
WORKDIR /app
USER 1000
CMD ["src/server.js"]        # distroless already has node as its ENTRYPOINT
docker build -t auroralibros/aurora-api:1.3.0-distroless ./api
docker image ls auroralibros/aurora-api --format "{{.Tag}}\t{{.Size}}"
1.3.0-distroless   102MB
1.3.0              121MB
1.2.0              142MB

142 MB → 102 MB: 28% less. The price, already familiar from lesson 05-03: with no sh there is no docker exec ... sh, and the wget-based HEALTHCHECK stops working, so the probe moves to Compose or to the orchestrator.

Technique Typical saving Cost
An -alpine base instead of the full one 800-900 MB musl instead of glibc
Multi-stage (compilers and devDependencies left out) 20-40% A slightly longer Dockerfile
--virtual + apk del in the same RUN 100-200 MB None
npm ci --omit=dev + cache clean 40-70% of node_modules None
A well-tuned .dockerignore Variable (here, 93 MB of context) It has to be maintained
Grouping RUNs with && Whatever the temporary files take Coarser cache granularity
A distroless base 15-25 MB more No shell: hard to debug
Squashing every layer into one Little Breaks the cache: almost never worth it

Common Mistakes and Tips

Optimizing without measuring. Without docker history and dive, you trim where it does not hurt and leave the 100 MB layer untouched.

Deleting files in a later RUN. The file disappears and the space does not. Creation and deletion belong in the same instruction.

Copying the whole project with COPY . . and no .dockerignore. In come .git, node_modules and, if you are lucky, some .env.

Installing build tools in the final image. Compilers in production are weight and attack surface. They belong in the dependencies stage.

Sacrificing the HEALTHCHECK or the USER to save a few megabytes. It is a terrible trade: security and operability are worth more than 3 MB.

Chasing the absolute minimum. A binary on scratch that you cannot debug costs more in its first incident than it saved in a year of downloads.

Tip: build measurement into your workflow. A CI step that compares the image's size with the previous version's and warns if it grows by more than 10% catches the bloated dependency on the day it lands, not six months later.

Exercises

Exercise 1. Analyze auroralibros/aurora-api:1.2.0 with history and with dive: identify the largest layer, work out what percentage of the total the base represents, and find at least one genuine piece of waste you can remove with .dockerignore.

Exercise 2. Convert aurora-api's Dockerfile to multi-stage with three stages (dependencies, development, production), measure before and after, and prove that the final image does not contain the compiler that is indeed in the dependencies stage.

Exercise 3. Prove that deleting a file in a later layer does not reduce the size, and fix it. Work out the exact saving and explain what would have happened if that file had been a credentials file.

Solutions

Solution 1.

docker image history auroralibros/aurora-api:1.2.0 --format "{{.Size}}\t{{.CreatedBy}}" \
  | sort -h -r | head -3
total=$(docker image inspect auroralibros/aurora-api:1.2.0 --format '{{.Size}}')
base=$(docker image inspect node:22-alpine --format '{{.Size}}')
echo "base: $((base * 100 / total)) % of the total"
115MB    /bin/sh -c #(nop) ADD file:... in /
21.4MB   RUN npm ci --omit=dev
3.1MB    COPY src/ ./src/
base: 81 % of the total

81% of the image is the base. That is the number that reorders your priorities: fighting over the 3 MB of your own code is irrelevant until you consciously decide which base you use. With dive the specific waste shows up:

Potential wasted space: 9.7 MB
    2         6.1 MB /app/node_modules/.cache
    2         2.4 MB /root/.npm/_cacache
    3         1.2 MB /app/.git

.git inside a production image is not just 1.2 MB of weight: it is the repository's full history traveling through the registry, along with any secret somebody committed and later deleted. It is fixed with one line in .dockerignore, and the check is straightforward:

docker run --rm auroralibros/aurora-api:1.3.0 ls -a /app | grep -c '^\.git$'   # -> 0

Solution 2.

docker build -t aurora-api:before -f api/Dockerfile.mono ./api
docker build -t aurora-api:after ./api
docker image ls aurora-api --format "{{.Tag}}\t{{.Size}}"
before   142MB
after    121MB
# Is the compiler present in each stage?
docker build -q --target dependencies -t aurora-api:deps ./api >/dev/null
docker run --rm aurora-api:deps sh -c 'which g++ make python3 | wc -l'
docker run --rm aurora-api:after sh -c 'which g++ make python3 2>/dev/null | wc -l'
docker run --rm aurora-api:after sh -c 'ls node_modules | wc -l'
docker run --rm aurora-api:deps sh -c 'ls node_modules | wc -l'
3
0
64
64

The dependencies stage has the three build binaries; the final one, none. And yet both have the same 64 packages in node_modules: COPY --from=dependencies has brought over exactly the result of the work and none of the tooling that produced it.

That is the central idea of multi-stage put as plainly as possible: what you need in order to build is not what you need in order to run. And the benefit is not only about size; it is also about security, because an attacker with command execution in the final image finds no compiler with which to prepare the next phase of their attack.

Solution 3.

mkdir -p /tmp/b1 /tmp/b2
printf 'FROM alpine:3\nRUN dd if=/dev/zero of=/t.bin bs=1M count=100\nRUN rm /t.bin\n' > /tmp/b1/Dockerfile
printf 'FROM alpine:3\nRUN dd if=/dev/zero of=/t.bin bs=1M count=100 && rm /t.bin\n' > /tmp/b2/Dockerfile
docker build -q -t bad /tmp/b1 >/dev/null && docker build -q -t good /tmp/b2 >/dev/null
docker image ls bad good --format "{{.Repository}}\t{{.Size}}"
docker run --rm bad ls /t.bin 2>&1
bad     113MB
good    8.15MB
ls: /t.bin: No such file or directory

The exact saving is 104.85 MB, 92.8% of the total, and the functional result is identical: in both images the file does not exist. The difference lies in whether those 100 MB ever became part of a consolidated layer. In bad, layer A wrote them and they became immutable; layer B only added a whiteout that hides them. In good, the file was born and died inside the same instruction, so by the time BuildKit consolidated the layer it was already gone.

The question about credentials takes the demonstration to its most uncomfortable conclusion:

printf 'FROM alpine:3\nRUN echo "aurora:S3cr3t0_2026" > /cred.txt && cat /cred.txt >/dev/null\nRUN rm /cred.txt\n' > /tmp/b1/Dockerfile
docker build -q -t cred /tmp/b1 >/dev/null
docker run --rm cred ls /cred.txt 2>&1
docker save cred | tar -xO --wildcards '*/layer.tar' 2>/dev/null | strings | grep -o 'S3cr3t0_[0-9]*'
ls: /cred.txt: No such file or directory
S3cr3t0_2026

The file does not exist and the password is recovered from the exported layer in a single command. If that image had been published on Docker Hub, anybody could extract it with docker pull and docker save, without ever running the container. That is why the same && that saves 105 MB is also a security control, and why the right solution for build-time secrets is not an && but RUN --mount=type=secret, which you will see in the next lesson.

Conclusion

aurora-api has gone from 142 MB to 102 MB, 28% less, without losing the unprivileged user, the HEALTHCHECK, the OCI labels or a single endpoint: /health, /books and /books/:id respond exactly as before and the container still reaches healthy. And, more important than the figure, you have the method: measure first. docker image ls for the total, docker image history to find the fat layer —which turned out to be the base, at 81% of the weight— and dive to see the real waste, which uncovered the npm cache and a .git that should never have entered the context.

You know how to choose a base deliberately rather than out of habit, aware of the musl-versus-glibc trade-off and of what each step between node:22, -slim, -alpine, distroless and scratch costs. You have mastered multi-stage builds: a stage that installs and compiles with build-base, a development stage with the devDependencies that feeds the target: development in your compose.override.yaml, a test stage that breaks the build if something fails, and a final one that receives only node_modules and src/ through COPY --from —demonstrated: three compilers in the intermediate stage, zero in the published one. And you know how to build each one separately with --target, as well as how to share a base image of your own across several services.

On the content side: --no-cache, --no-install-recommends, --virtual with its apk del, npm ci --omit=dev, npm cache clean and a .dockerignore that shrank the context from 94 MB to 812 kB. And the demonstration most worth remembering: deleting a file in a later layer does not reduce the size —113 MB versus 8.15 MB for one &&— and that same mechanism means a credential that was written and then deleted is still extractable from the published image.

Lesson 05-05 brings on stage the builder that has been doing all of this underneath: BuildKit, with its dependency graph and its parallelism between stages, and Buildx with its builders. You will see mounts in RUN —type=cache so you do not reinstall npm on every build, type=bind, type=tmpfs and above all type=secret, the correct answer to the credentials problem you have just demonstrated—, remote shared cache with --cache-from and --cache-to, multi-architecture builds for amd64 and arm64 with their manifest lists, and docker buildx bake to build every Aurora Libros image in one go.

Docker: From Beginner to Advanced

Module 1: Introduction to Docker

Module 2: Working with Docker Images

Module 3: Docker Containers

Module 4: Docker Compose

Module 5: Advanced Docker Concepts

Module 6: Docker in Production

Module 7: Docker Ecosystem and Tools

© Copyright 2026. All rights reserved