The auroralibros/aurora-api:1.2.0 image weighs 142 MB. That is not a disaster, but it is not optimized either: a C compiler travels inside it, along with the development dependencies and files nobody ever runs. In this lesson you put it on a diet measuring at every step, down to a specific figure, without losing any of the hardening from lesson 05-03.
Contents
- Why size matters and what must not be sacrificed
- Measure first:
image lsandhistory dive: exploring layer by layer- Choosing the base, with numbers
- Multi-stage builds: the concept
- Multi-stage applied to
aurora-api - Named stages and
--target - A common base image for several images
- Reducing layers and content
- Deleting a file does not reduce the size (demonstrated)
- Ordering for the cache
- Final verification and summary table
- Why size matters and what must not be sacrificed
| Reason | Concrete impact |
|---|---|
| Download time | Every new node downloads the whole image the first time |
| Storage cost | The registry, each node's cache, and every version you keep |
| Attack surface | Fewer packages, fewer CVEs (lesson 05-03) |
| Deployment speed | An urgent rollback is only as fast as the download |
| Autoscaling | A traffic spike demands replicas now |
On a platform with five nodes and fifteen deployments a month, shaving 40 MB off an image means roughly 3 GB less transfer per month and tens of seconds less on every cold start.
And what you do not sacrifice in exchange for weight: the HEALTHCHECK, the unprivileged user, the OCI labels for traceability, the ability to diagnose a production problem, and the reproducibility of the build. A 40 MB image nobody knows how to debug at three in the morning is a bad deal.
- Measure first:
image ls and history
image ls and historydocker image ls auroralibros/aurora-api --format "{{.Tag}}\t{{.Size}}"
docker image history auroralibros/aurora-api:1.2.0 --format "{{.Size}}\t{{.CreatedBy}}" | head -81.2.0 142MB
21.4MB RUN npm ci --omit=dev
3.1MB COPY src/ ./src/
1.8MB RUN apk add --no-cache tini
0B ENV NODE_ENV=production
0B USER node
0B HEALTHCHECK &{["CMD-SHELL" "wget -qO- ..."]}
115MB /bin/sh -c #(nop) ADD file:... in /The diagnosis is immediate and it surprises almost everybody the first time: of the 142 MB, 115 MB are the base image. The dependencies are 21 MB and your own code, 3 MB. The main lever is not in your code: it is in which base you pick and what you add on top of it.
To compare precisely, use the real size (docker image inspect --format '{{.Size}}' | numfmt --to=iec) and not the one the registry shows, which is compressed and always looks smaller.
dive: exploring layer by layer
dive: exploring layer by layerdocker history tells you how much each layer weighs; dive tells you what is inside it and how much is wasted.
docker run --rm -it \
-v /var/run/docker.sock:/var/run/docker.sock \
wagoodman/dive:latest auroralibros/aurora-api:1.2.0│ Image Details │
Total Image size: 142 MB
Potential wasted space: 9.7 MB
Image efficiency score: 93 %
Count Total Space Path
2 6.1 MB /app/node_modules/.cache
2 2.4 MB /root/.npm/_cacache
3 1.2 MB /app/.gitThree findings and three decisions: the npm cache should not be in the final image, node_modules/.cache is build junk, and .git should never have entered the build context in the first place (there is a line missing from .dockerignore).
The efficiency score measures how much content is written into one layer and then overwritten or deleted in another; below 95% there is work to do. To automate it in CI, add -e CI=true and --lowestEfficiency=0.95: dive exits with a non-zero code if the threshold is not met.
- Choosing the base, with numbers
| Base | Size | libc | Tooling | When to choose it |
|---|---|---|---|---|
node:22 |
~1.1 GB | glibc | Everything: git, compilers, curl |
Only as a build stage |
node:22-slim |
~230 MB | glibc | Debian's minimum | Native dependencies that insist on glibc |
node:22-alpine |
~135 MB | musl | BusyBox, apk |
The general case: the sweet spot |
distroless/nodejs22 |
~110 MB | glibc | None | Hardened production |
scratch |
0 MB | — | None | Static binaries only (Go, Rust) |
Alpine's trade-off has a name: musl instead of glibc. The real consequences:
- npm packages with binaries precompiled for glibc may not work and may have to be compiled on install (slower, and it requires
build-base). - Some DNS-heavy or name-resolution-heavy workloads behave differently under musl.
- Certain scientific and machine-learning libraries do not support musl.
aurora-api uses Express, pg and redis, none of which has problematic native dependencies: Alpine is the right choice. If one day a dependency caused trouble, node:22-slim costs 95 MB more and fixes it.
- Multi-stage builds: the concept
A multi-stage build uses several base images in the same Dockerfile. The intermediate stages compile and prepare; the final stage copies only the result. Everything else —compilers, caches, development dependencies, source code— stays out of the published image.
flowchart LR
subgraph E1["Stage: dependencies (node:22-alpine)"]
A["package*.json"] --> B["npm ci --omit=dev<br/>+ build-base to compile"]
end
subgraph E2["Stage: development"]
C["full npm ci<br/>devDependencies, tests"]
end
subgraph E3["Final stage (node:22-alpine)"]
D["COPY --from=dependencies node_modules"]
E["COPY src/"]
end
B -.->|"COPY --from"| D
E3 --> F["Published image:<br/>no compilers, no caches"]
E2 -.->|"only with --target"| G["Development image"]
Two rules that sum up the mechanism:
- Only the last stage (or the one named with
--target) ends up in the image; the rest are discarded. COPY --from=<stage>brings specific files over from an earlier stage, and nothing else travels: not its layers, not its history, not its secrets.
- Multi-stage applied to
aurora-api
aurora-apiThe starting point (a single stage, 142 MB) and the destination:
# syntax=docker/dockerfile:1
ARG NODE_VERSION=22
# ---------- Stage 1: production dependencies ----------
FROM node:${NODE_VERSION}-alpine AS dependencies
WORKDIR /app
# build-base lives only here: it compiles whatever is needed and does NOT travel to the final stage
RUN apk add --no-cache --virtual .build python3 make g++
COPY package.json package-lock.json ./
RUN npm ci --omit=dev && npm cache clean --force
# ---------- Stage 2: development and tests (not published) ----------
FROM node:${NODE_VERSION}-alpine AS development
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
CMD ["node", "--watch", "src/server.js"]
# ---------- Stage 3: final, minimal image ----------
FROM node:${NODE_VERSION}-alpine AS production
ARG VERSION=1.3.0
ARG REVISION=unknown
LABEL org.opencontainers.image.title="aurora-api" \
org.opencontainers.image.version="${VERSION}" \
org.opencontainers.image.revision="${REVISION}" \
org.opencontainers.image.source="https://github.com/aurora-libros/aurora-api"
WORKDIR /app
ENV NODE_ENV=production PORT=3000
COPY --from=dependencies --chown=node:node /app/node_modules ./node_modules
COPY --chown=node:node package.json ./
COPY --chown=node:node src/ ./src/
USER node
EXPOSE 3000
HEALTHCHECK --interval=30s --timeout=3s --start-period=10s --retries=3 \
CMD wget -qO- http://127.0.0.1:3000/health || exit 1
ENTRYPOINT ["node"]
CMD ["src/server.js"]docker build -t auroralibros/aurora-api:1.3.0 --build-arg VERSION=1.3.0 ./api
docker image ls auroralibros/aurora-api --format "{{.Tag}}\t{{.Size}}"142 MB → 121 MB: 15% less with a single structural change. Where do those 21 MB come from? From build-base (python3, make, g++), which used to be installed in the final image and now lives and dies inside the dependencies stage, and from the npm cache.
Notice that none of the hardening has been lost: USER node, the HEALTHCHECK, the OCI labels and the exec-form ENTRYPOINT/CMD are all still there. Optimizing is not trimming security.
- Named stages and
--target
--targetNamed stages are not just organization: they are separately buildable artifacts.
docker build --target development -t aurora-api:dev ./api
docker build --target dependencies -t aurora-api:deps ./api
docker build -t auroralibros/aurora-api:1.3.0 ./api # the last stage
docker image ls aurora-api --format "{{.Tag}}\t{{.Size}}"The development image weighs 318 MB and that is exactly as it should be: it carries the devDependencies, the test framework and the debugging tools. It is precisely the target: development that compose.override.yaml has been using since lesson 04-07, and now you can see where it comes from.
A very useful pattern is a test stage that fails the build if the tests fail:
BuildKit does not run that stage when building the production one, because it is not part of its dependency graph (lesson 05-05). It only runs if you ask for it with --target.
- A common base image for several images
When you publish several Node services, it pays to have your own base with everything they share:
# aurora-base/Dockerfile -> auroralibros/aurora-base-node:22
FROM node:22-alpine
RUN apk add --no-cache tini tzdata && \
addgroup -g 1001 aurora && adduser -u 1001 -G aurora -s /bin/sh -D aurora
ENV TZ=Europe/Madrid NODE_ENV=production
ENTRYPOINT ["/sbin/tini", "--"]Advantages: a single security update propagates to every service, the base layer is shared on disk and on download between them, and the shared decisions live in one place. Downside: you create an internal dependency that has to be maintained and versioned. It starts to pay off from three or four services onwards.
- Reducing layers and content
# ❌ Three layers, and the apk cache stays inside forever
RUN apk update
RUN apk add curl
RUN rm -rf /var/cache/apk/*
# ✅ One layer, no cache: what is never written takes no space
RUN apk add --no-cache curl| Technique | What it does | Typical saving |
|---|---|---|
apk add --no-cache |
Avoids writing the package index | 5-10 MB |
apt-get ... && rm -rf /var/lib/apt/lists/* in the same RUN |
The same on Debian | 30-50 MB |
--no-install-recommends |
Does not install "suggested" packages | 20-100 MB |
npm ci --omit=dev |
No development dependencies | 40-70% of node_modules |
npm cache clean --force |
Removes ~/.npm/_cacache |
10-30 MB |
--virtual .build + apk del .build |
Temporary compilers in the same layer | 100-200 MB |
A well-tuned .dockerignore |
Keeps out what does not belong | Variable, sometimes enormous |
aurora-api's .dockerignore after what dive revealed:
.git
.gitignore
node_modules
npm-debug.log
Dockerfile*
compose*.yaml
.env
.env.*
!.env.example
coverage/
.nyc_output/
*.md
!README.md
.vscode/
.idea/
tests/
**/*.test.jsFrom a 94 MB directory to 812 kB of context. That does not only make the build faster: it guarantees that .git and .env cannot sneak into any layer, which is a security matter as much as a size one.
- Deleting a file does not reduce the size (demonstrated)
FROM alpine:3
RUN dd if=/dev/zero of=/temp.bin bs=1M count=100 # layer A: +100 MB
RUN rm /temp.bin # layer B: whiteoutdocker build -q -t delete:bad /tmp/delete
docker image ls delete:bad --format "{{.Size}}"
docker run --rm delete:bad ls -la /temp.bin 2>&1
docker image history delete:bad --format "{{.Size}}\t{{.CreatedBy}}" | head -3113MB
ls: /temp.bin: No such file or directory
0B RUN rm /temp.bin
105MB RUN dd if=/dev/zero of=/temp.bin bs=1M count=100
8MB /bin/sh -c #(nop) ADD file:... in /The file does not exist and the image still weighs 113 MB. The explanation is the one from lesson 05-02: layers are immutable, and rm in a later layer only creates a whiteout that hides the file; the 105 MB are still down there, downloaded on every pull and stored on every node.
The fix is to combine creation and deletion in the same instruction:
FROM alpine:3
RUN dd if=/dev/zero of=/temp.bin bs=1M count=100 && \
echo "use the file..." && \
rm /temp.binFrom 113 MB to 8.15 MB. The same logic, the same functional result, one && of difference. And it is exactly the reason for --virtual .build in the Dockerfile from section 6: install, compile and uninstall inside a single RUN.
- Ordering for the cache
A reminder from lesson 02-02, now combined with multi-stage: instructions are ordered from least to most likely to change.
COPY package.json package-lock.json ./ # rarely changes
RUN npm ci --omit=dev # expensive layer, reusable
COPY src/ ./src/ # changes with every commitWith multi-stage the effect is multiplied: when you touch src/server.js, BuildKit reuses the entire dependencies stage —including the npm install— and only rebuilds the last two layers of the final stage.
=> CACHED [dependencies 4/4] RUN npm ci --omit=dev
=> [production 5/5] COPY --chown=node:node src/ ./src/ 0.2s
real 0m3.184sThree seconds against the forty-odd of a clean build. Optimizing size and optimizing build speed are, almost always, the same job.
- Final verification and summary table
An optimized image that does not work is worth nothing. The check is mandatory:
docker compose up -d --wait
docker compose exec aurora-api node --version
curl -s http://localhost:8080/api/health | jq -c
curl -s http://localhost:8080/api/books | jq -c '{source, total}'
curl -s http://localhost:8080/api/books/3 | jq -c '{title, author}'
docker inspect aurora-libros-aurora-api-1 --format 'user={{.Config.User}} health={{.State.Health.Status}}'v22.11.0
{"status":"ok","db":"connected","cache":"connected","version":"1.3.0"}
{"source":"db","total":9}
{"title":"Cien años de soledad","author":"Gabriel García Márquez"}
user=node health=healthyAnd the final step up, if your context allows it: switching the final base to distroless.
FROM gcr.io/distroless/nodejs22-debian12 AS production
COPY --from=dependencies --chown=1000:1000 /app/node_modules /app/node_modules
COPY --chown=1000:1000 src/ /app/src/
WORKDIR /app
USER 1000
CMD ["src/server.js"] # distroless already has node as its ENTRYPOINTdocker build -t auroralibros/aurora-api:1.3.0-distroless ./api
docker image ls auroralibros/aurora-api --format "{{.Tag}}\t{{.Size}}"142 MB → 102 MB: 28% less. The price, already familiar from lesson 05-03: with no sh there is no docker exec ... sh, and the wget-based HEALTHCHECK stops working, so the probe moves to Compose or to the orchestrator.
| Technique | Typical saving | Cost |
|---|---|---|
An -alpine base instead of the full one |
800-900 MB | musl instead of glibc |
| Multi-stage (compilers and devDependencies left out) | 20-40% | A slightly longer Dockerfile |
--virtual + apk del in the same RUN |
100-200 MB | None |
npm ci --omit=dev + cache clean |
40-70% of node_modules |
None |
A well-tuned .dockerignore |
Variable (here, 93 MB of context) | It has to be maintained |
Grouping RUNs with && |
Whatever the temporary files take | Coarser cache granularity |
| A distroless base | 15-25 MB more | No shell: hard to debug |
| Squashing every layer into one | Little | Breaks the cache: almost never worth it |
Common Mistakes and Tips
Optimizing without measuring. Without docker history and dive, you trim where it does not hurt and leave the 100 MB layer untouched.
Deleting files in a later RUN. The file disappears and the space does not. Creation and deletion belong in the same instruction.
Copying the whole project with COPY . . and no .dockerignore. In come .git, node_modules and, if you are lucky, some .env.
Installing build tools in the final image. Compilers in production are weight and attack surface. They belong in the dependencies stage.
Sacrificing the HEALTHCHECK or the USER to save a few megabytes. It is a terrible trade: security and operability are worth more than 3 MB.
Chasing the absolute minimum. A binary on scratch that you cannot debug costs more in its first incident than it saved in a year of downloads.
Tip: build measurement into your workflow. A CI step that compares the image's size with the previous version's and warns if it grows by more than 10% catches the bloated dependency on the day it lands, not six months later.
Exercises
Exercise 1. Analyze auroralibros/aurora-api:1.2.0 with history and with dive: identify the largest layer, work out what percentage of the total the base represents, and find at least one genuine piece of waste you can remove with .dockerignore.
Exercise 2. Convert aurora-api's Dockerfile to multi-stage with three stages (dependencies, development, production), measure before and after, and prove that the final image does not contain the compiler that is indeed in the dependencies stage.
Exercise 3. Prove that deleting a file in a later layer does not reduce the size, and fix it. Work out the exact saving and explain what would have happened if that file had been a credentials file.
Solutions
Solution 1.
docker image history auroralibros/aurora-api:1.2.0 --format "{{.Size}}\t{{.CreatedBy}}" \
| sort -h -r | head -3
total=$(docker image inspect auroralibros/aurora-api:1.2.0 --format '{{.Size}}')
base=$(docker image inspect node:22-alpine --format '{{.Size}}')
echo "base: $((base * 100 / total)) % of the total"115MB /bin/sh -c #(nop) ADD file:... in /
21.4MB RUN npm ci --omit=dev
3.1MB COPY src/ ./src/
base: 81 % of the total81% of the image is the base. That is the number that reorders your priorities: fighting over the 3 MB of your own code is irrelevant until you consciously decide which base you use. With dive the specific waste shows up:
Potential wasted space: 9.7 MB
2 6.1 MB /app/node_modules/.cache
2 2.4 MB /root/.npm/_cacache
3 1.2 MB /app/.git.git inside a production image is not just 1.2 MB of weight: it is the repository's full history traveling through the registry, along with any secret somebody committed and later deleted. It is fixed with one line in .dockerignore, and the check is straightforward:
Solution 2.
docker build -t aurora-api:before -f api/Dockerfile.mono ./api
docker build -t aurora-api:after ./api
docker image ls aurora-api --format "{{.Tag}}\t{{.Size}}"# Is the compiler present in each stage?
docker build -q --target dependencies -t aurora-api:deps ./api >/dev/null
docker run --rm aurora-api:deps sh -c 'which g++ make python3 | wc -l'
docker run --rm aurora-api:after sh -c 'which g++ make python3 2>/dev/null | wc -l'
docker run --rm aurora-api:after sh -c 'ls node_modules | wc -l'
docker run --rm aurora-api:deps sh -c 'ls node_modules | wc -l'The dependencies stage has the three build binaries; the final one, none. And yet both have the same 64 packages in node_modules: COPY --from=dependencies has brought over exactly the result of the work and none of the tooling that produced it.
That is the central idea of multi-stage put as plainly as possible: what you need in order to build is not what you need in order to run. And the benefit is not only about size; it is also about security, because an attacker with command execution in the final image finds no compiler with which to prepare the next phase of their attack.
Solution 3.
mkdir -p /tmp/b1 /tmp/b2
printf 'FROM alpine:3\nRUN dd if=/dev/zero of=/t.bin bs=1M count=100\nRUN rm /t.bin\n' > /tmp/b1/Dockerfile
printf 'FROM alpine:3\nRUN dd if=/dev/zero of=/t.bin bs=1M count=100 && rm /t.bin\n' > /tmp/b2/Dockerfile
docker build -q -t bad /tmp/b1 >/dev/null && docker build -q -t good /tmp/b2 >/dev/null
docker image ls bad good --format "{{.Repository}}\t{{.Size}}"
docker run --rm bad ls /t.bin 2>&1The exact saving is 104.85 MB, 92.8% of the total, and the functional result is identical: in both images the file does not exist. The difference lies in whether those 100 MB ever became part of a consolidated layer. In bad, layer A wrote them and they became immutable; layer B only added a whiteout that hides them. In good, the file was born and died inside the same instruction, so by the time BuildKit consolidated the layer it was already gone.
The question about credentials takes the demonstration to its most uncomfortable conclusion:
printf 'FROM alpine:3\nRUN echo "aurora:S3cr3t0_2026" > /cred.txt && cat /cred.txt >/dev/null\nRUN rm /cred.txt\n' > /tmp/b1/Dockerfile
docker build -q -t cred /tmp/b1 >/dev/null
docker run --rm cred ls /cred.txt 2>&1
docker save cred | tar -xO --wildcards '*/layer.tar' 2>/dev/null | strings | grep -o 'S3cr3t0_[0-9]*'The file does not exist and the password is recovered from the exported layer in a single command. If that image had been published on Docker Hub, anybody could extract it with docker pull and docker save, without ever running the container. That is why the same && that saves 105 MB is also a security control, and why the right solution for build-time secrets is not an && but RUN --mount=type=secret, which you will see in the next lesson.
Conclusion
aurora-api has gone from 142 MB to 102 MB, 28% less, without losing the unprivileged user, the HEALTHCHECK, the OCI labels or a single endpoint: /health, /books and /books/:id respond exactly as before and the container still reaches healthy. And, more important than the figure, you have the method: measure first. docker image ls for the total, docker image history to find the fat layer —which turned out to be the base, at 81% of the weight— and dive to see the real waste, which uncovered the npm cache and a .git that should never have entered the context.
You know how to choose a base deliberately rather than out of habit, aware of the musl-versus-glibc trade-off and of what each step between node:22, -slim, -alpine, distroless and scratch costs. You have mastered multi-stage builds: a stage that installs and compiles with build-base, a development stage with the devDependencies that feeds the target: development in your compose.override.yaml, a test stage that breaks the build if something fails, and a final one that receives only node_modules and src/ through COPY --from —demonstrated: three compilers in the intermediate stage, zero in the published one. And you know how to build each one separately with --target, as well as how to share a base image of your own across several services.
On the content side: --no-cache, --no-install-recommends, --virtual with its apk del, npm ci --omit=dev, npm cache clean and a .dockerignore that shrank the context from 94 MB to 812 kB. And the demonstration most worth remembering: deleting a file in a later layer does not reduce the size —113 MB versus 8.15 MB for one &&— and that same mechanism means a credential that was written and then deleted is still extractable from the published image.
Lesson 05-05 brings on stage the builder that has been doing all of this underneath: BuildKit, with its dependency graph and its parallelism between stages, and Buildx with its builders. You will see mounts in RUN —type=cache so you do not reinstall npm on every build, type=bind, type=tmpfs and above all type=secret, the correct answer to the credentials problem you have just demonstrated—, remote shared cache with --cache-from and --cache-to, multi-architecture builds for amd64 and arm64 with their manifest lists, and docker buildx bake to build every Aurora Libros image in one go.
Docker: From Beginner to Advanced
Module 1: Introduction to Docker
- What Is Docker?
- Installing Docker
- Docker Architecture
- Basic Docker Commands
- Understanding Docker Images
- Creating Your First Docker Container
- The Course Project: The Aurora Libros Platform
Module 2: Working with Docker Images
- Docker Hub and Repositories
- Building Docker Images
- Dockerfile Basics
- Advanced Dockerfile Instructions
- Managing Docker Images
- Tagging and Publishing Images
Module 3: Docker Containers
- Running Containers
- Container Lifecycle
- Managing Containers
- Inspecting and Debugging Containers
- Docker Networking
- Data Persistence with Volumes
- Resource Limits and Restart Policies
Module 4: Docker Compose
- Introduction to Docker Compose
- Defining Services in Docker Compose
- Docker Compose Commands
- Multi-Container Applications
- Environment Variables in Docker Compose
- Profiles, Overrides and Multiple Environments
- Local Development with Docker Compose
Module 5: Advanced Docker Concepts
- Docker Networking Deep Dive
- Docker Storage Options
- Docker Security Best Practices
- Optimizing Docker Images
- Advanced Builds with BuildKit and Buildx
- Logging and Monitoring in Docker
- The Runtime Inside: Namespaces, Cgroups and Layers
Module 6: Docker in Production
- Preparing an Image for Production
- CI/CD with Docker
- Orchestrating Containers with Docker Swarm
- Introduction to Kubernetes
- Deploying Docker Containers in Kubernetes
- Scaling and Load Balancing
- Deployment Strategies and Rollback
