When Module 4 ended we had catalog-service and orders-service running with npm run dev on a laptop, with MongoDB, PostgreSQL and RabbitMQ started by hand with docker run. That is fine for developing, but not for deploying: in production each service will run as several replicas, on machines nobody has prepared by hand, and it will have to start, die and start again hundreds of times without a human stepping in. The container is the unit that makes all of that possible: it packages the service with its exact Node.js and its dependencies into an immutable image that runs identically on Luis's laptop, on the CI runner and in the Kubernetes cluster. This lesson builds the production image of catalog-service, explains how it is tagged, run and shut down, and brings up TechCorp's complete system with Docker Compose for the first time, including the single E2E test that 04-05 left pending. Kubernetes (05-02) and the pipeline that will build these images on every commit (05-03) rest on what gets settled here.

Contents

  1. Why containers for microservices
  2. Concepts: image, layer, container, registry, volume and network
  3. The production Dockerfile of catalog-service, line by line
  4. .dockerignore, layer cache and instruction order
  5. Image tagging and the registry
  6. Basic Docker commands
  7. Signals, PID 1 and graceful shutdown inside the container
  8. Docker Compose: TechCorp's complete local environment
  9. Operating the environment: up, down, logs, ps, migrations and the E2E test
  10. Image size and basic security

  1. Why containers for microservices

A container is an isolated process that the Linux kernel runs with its own filesystem, its own view of the network and CPU and memory limits, starting from an image that contains everything that process needs (the Node binary, node_modules, code). There is no guest operating system and no hypervisor: the container shares the host's kernel, which is why it starts in milliseconds and takes up megabytes.

Aspect Virtual machine Container
What is virtualized The full hardware (CPU, disk, network) with its own kernel Only user space; shares the host's kernel
Typical size Gigabytes Tens or hundreds of megabytes
Startup Tens of seconds to minutes Milliseconds to seconds
Density per machine Units or tens Hundreds
Isolation Strong (own kernel) Good (namespaces and cgroups), not equivalent to a VM
Deployment unit A VM image (AMI, OVA), heavy to build A container image, built in seconds by CI

For TechCorp the fit is direct:

  • One image = one deployment unit. catalog-service is an image; deploying version 1.4.2 means running that image. There is no "install Node 20 on the server" and no "copy the folder and run npm install": all of that happened once, when the image was built.
  • Reproducibility. The very same image that passed the integration tests in CI is the one that runs in production, byte for byte. No more "works on my machine" and no more "production has a different version of pg".
  • Isolation. Six Node services on the same machine share no node_modules, no ports and no environment variables. Each one believes it has the machine to itself.
  • Fast startup, disposable. The graceful shutdown of 04-02 and the per-environment configuration of 04-03 were written with this in mind: a container is born, serves, receives SIGTERM and dies; another one takes its place.

  1. Concepts: image, layer, container, registry, volume and network

Concept What it is At TechCorp
Image Immutable, read-only template with the filesystem and metadata (start command, ports, user) ghcr.io/techcorp/catalog-service:1.4.2
Layer Every Dockerfile instruction that changes the filesystem produces a layer; the image is the stack of layers. Layers are cached and shared between images The six service images share the node:20-alpine base layer
Container A running instance of an image: the image's layers plus an ephemeral writable layer Each replica of catalog-service
Registry Image store you push to and pull from GitHub Container Registry (ghcr.io), next to the code and to @techcorp/common-http (04-01)
Volume Persistent storage outside the container's writable layer PostgreSQL, MongoDB and RabbitMQ data locally; TechCorp's services do not use volumes (they are stateless)
Network Virtual network in which containers resolve each other by name In Compose, orders-service calls http://catalog-service:3001, the same name the Kubernetes DNS will resolve (03-05)

That the services keep no state on disk is no accident: it is what allows killing and replicating them without a second thought. Everything that must survive lives in the database or in the broker.

  1. The production Dockerfile of catalog-service, line by line

The Dockerfile lives at the root of the techcorp/catalog-service repository and comes from the techcorp/node-service-template template (04-01, exercise 2: the Dockerfile is copied and adapted). It uses a multi-stage build: one stage installs dependencies and another, clean one keeps only what is needed to run.

# syntax=docker/dockerfile:1

# ---------- Stage 1: dependencies ----------
FROM node:20-alpine AS dependencies
WORKDIR /app
# Only the files that define the dependencies: if they don't change, this layer (and npm ci) is reused from the cache.
COPY package.json package-lock.json ./
# @techcorp/common-http lives in GitHub Packages (04-01): npm needs a token to download it.
# The .npmrc is mounted ONLY during this RUN as a build secret; it never ends up in any image layer.
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc \
    npm ci --omit=dev

# ---------- Stage 2: final image ----------
FROM node:20-alpine
ENV NODE_ENV=production
WORKDIR /app
# node_modules already installed and without dev dependencies; owned by the 'node' user that the base image provides.
COPY --from=dependencies --chown=node:node /app/node_modules ./node_modules
COPY --chown=node:node package.json ./
COPY --chown=node:node src ./src
COPY --chown=node:node scripts ./scripts
# From here on the process is not root.
USER node
EXPOSE 3001
# Optional: Docker (not Kubernetes) marks the container as unhealthy if /health/live doesn't answer 200.
HEALTHCHECK --interval=30s --timeout=3s --start-period=10s --retries=3 \
  CMD wget -qO- http://localhost:3001/health/live || exit 1
# Exec form (JSON): node is PID 1 and receives SIGTERM directly (section 7).
CMD ["node", "src/server.js"]

Instruction by instruction:

  • # syntax=docker/dockerfile:1: enables the modern BuildKit syntax (required for --mount=type=secret).
  • FROM node:20-alpine AS dependencies: official Node 20 base image on Alpine Linux (about 50 MB versus the 350 MB of node:20). AS dependencies names the stage so we can copy from it later.
  • WORKDIR /app: creates /app and makes it the working directory for the following instructions.
  • COPY package.json package-lock.json ./ before copying the code: this is the key to the cache (section 4).
  • RUN --mount=type=secret,id=npmrc ... npm ci --omit=dev: npm ci installs exactly what package-lock.json says (reproducible; it fails if the lock does not match package.json); --omit=dev leaves out nodemon, jest, supertest, @pact-foundation/pact, testcontainers. The GitHub Packages token is mounted as a file only during this command; it is passed at build time with --secret id=npmrc,src=$HOME/.npmrc.
  • FROM node:20-alpine (second time): the final image starts from scratch; nothing from the previous stage carries over except what is copied explicitly.
  • ENV NODE_ENV=production: Express disables debug messages and some libraries optimize; the config.js of 04-03 reads it as just another variable.
  • COPY --from=dependencies --chown=node:node /app/node_modules ./node_modules: brings only the already-installed node_modules. --chown=node:node assigns ownership to the node user (uid 1000) that the base image already defines; without it the files would belong to root and the non-root process might not be able to read them if permissions were restrictive.
  • Three separate COPY instructions for package.json, src and scripts: scripts/ comes in because seed.js runs from this same image (section 8). No tests, contracts or development configuration files are copied.
  • USER node: everything executed from here on (including CMD) runs unprivileged. It is the cheapest and most effective security measure in this lesson.
  • EXPOSE 3001: documents the port (it does not publish it; that is done with -p or in Compose/Kubernetes). It matches PORT=3001 from 04-02.
  • HEALTHCHECK: Docker runs wget against /health/live every 30 s, with a 10 s grace period at startup; three consecutive failures mark the container unhealthy. Alpine ships wget in BusyBox, so there is no need to install curl. Kubernetes ignores HEALTHCHECK and uses its own probes (05-02); in Compose it is useful for depends_on (section 8).
  • CMD ["node", "src/server.js"]: the start command, the same as the start script of 04-02, but without going through npm (section 7).

Why multi-stage if a JavaScript service compiles nothing? Because it separates two concerns: the first stage may hold tokens, the npm cache, build tools for native modules; the second holds only what runs. If one day a service is written in TypeScript, stage 1 becomes "install and compile" and stage 2 does not change.

  1. .dockerignore, layer cache and instruction order

COPY src ./src copies what is in the build context (the directory passed to docker build). Without a filter, a COPY . . would drag along the local node_modules (with dev dependencies and binaries for another operating system), .env with secrets, .git, pacts/, test coverage... The .dockerignore excludes them from the context:

node_modules
npm-debug.log
.env
.env.*
!.env.example
.git
.github
coverage
tests
pacts
*.md

About the cache: Docker builds the layers in order and, for each instruction, reuses the cached layer if the instruction and its inputs have not changed; as soon as one layer changes, all the following ones are rebuilt. Hence the order of the Dockerfile:

flowchart LR
    A[FROM node:20-alpine] --> B[COPY package*.json]
    B --> C[RUN npm ci]
    C --> D[COPY src, scripts]
    D --> E[CMD]
    style C fill:#dfe,stroke:#393
    style D fill:#fdd,stroke:#933

A change in src/routes/products.js invalidates only the red layer: the npm ci (the slow operation, 30-60 s with download) is reused from the cache. If we copied all the code first and installed afterwards, every commit would repeat the installation. Rule: what changes least goes at the top; what changes most goes at the bottom.

  1. Image tagging and the registry

An image is identified by registry/organization/name:tag. TechCorp uses two tags per build:

Tag Example Who uses it
Semantic version ghcr.io/techcorp/catalog-service:1.4.2 Kubernetes manifests, release notes, humans
Source commit ghcr.io/techcorp/catalog-service:sha-9f3c2ab Exact traceability: which code the image comes from; the pipeline of 05-03 creates it on every build

Both point to the same digest (sha256:...), which is the real, immutable identifier of the image. What is not used in production is latest:

  • latest does not mean "the most recent", it means "the last one somebody put that tag on". It is a mutable tag: today it points to 1.4.2 and tomorrow to 1.5.0 without any manifest changing.
  • With latest, two replicas of the same Deployment may run different versions depending on when they pulled, and a rollback is impossible because nobody knows what was there before.
  • The manifests of 05-02 always carry a specific tag; changing version means changing that line, and that ends up in git.

Manual build and publish (the pipeline of 05-03 automates exactly this):

# At the root of catalog-service. -t adds a tag; several can be given.
docker build \
  --secret id=npmrc,src=$HOME/.npmrc \
  -t ghcr.io/techcorp/catalog-service:1.4.2 \
  -t ghcr.io/techcorp/catalog-service:sha-9f3c2ab \
  .

# Authenticate against the registry with a GitHub token that has the write:packages permission
echo "$GITHUB_TOKEN" | docker login ghcr.io -u luis --password-stdin

docker push ghcr.io/techcorp/catalog-service:1.4.2
docker push ghcr.io/techcorp/catalog-service:sha-9f3c2ab

  1. Basic Docker commands

The ones used every day, with the freshly built image:

# Run in the background (-d), with a name, publishing the container's port 3001 on the host's 3001 (-p host:container)
# and passing the configuration through environment variables (04-03). --rm removes the container when it stops.
docker run -d --rm --name catalog -p 3001:3001 \
  -e MONGO_URL=mongodb://host.docker.internal:27017 -e MONGO_DB=catalog -e LOG_LEVEL=debug \
  ghcr.io/techcorp/catalog-service:1.4.2

docker ps                          # running containers: id, image, status (healthy/unhealthy), ports
docker logs -f catalog             # the process's stdout/stderr (pino writes JSON to stdout: that's why it works without log files)
docker exec -it catalog sh         # a shell inside the container (Alpine: sh, not bash) to inspect
docker exec catalog node scripts/seed.js   # run a one-off command with the container's own environment
docker stop catalog                # SIGTERM and, if it hasn't finished within 10 s, SIGKILL (section 7)
docker images                      # local images and their size

host.docker.internal is the name through which a container reaches the host (Docker Desktop; on Linux, --add-host=host.docker.internal:host-gateway). It only makes sense in this standalone experiment: in Compose and in Kubernetes the services find each other by name inside the same network.

  1. Signals, PID 1 and graceful shutdown inside the container

In 04-02 we wrote the graceful shutdown: on SIGTERM, /health/ready switches to 503, server.close() finishes in-flight requests, Mongo is closed and the process exits, with a 10 s limit. Inside a container there are three details that can undo that work:

  1. Who is PID 1. The process started by CMD is PID 1 of the container and is the only one that receives the signal from docker stop or from the kubelet. With the exec form CMD ["node", "src/server.js"], PID 1 is Node and our handler runs. With the shell form CMD node src/server.js, PID 1 is /bin/sh, which receives SIGTERM and does not forward it to Node: the service dies 10 s later by SIGKILL, with requests half done. And with CMD ["npm", "start"], PID 1 is npm, which does not forward signals reliably either. That is why the Dockerfile calls node directly.
  2. PID 1 has no default handlers. The kernel treats PID 1 specially: if it does not install a handler for a signal, it ignores it. Node does install ours (process.on('SIGTERM')), so we are covered; a service that forgot to do so would never die on docker stop, only on SIGKILL.
  3. Zombie processes. PID 1 must "reap" terminated child processes. Node spawns no children in our services, but if one day a service runs child_process, a minimal init is advisable: docker run --init (Docker injects tini) or, in Kubernetes, shareProcessNamespace or tini inside the image. It is a cheap line that avoids a problem that is hard to diagnose.

Timings: docker stop waits 10 s by default (-t changes it); Compose uses stop_grace_period; Kubernetes, terminationGracePeriodSeconds (30 s by default, 05-02). The internal 10 s limit of 04-02 is designed to fit inside any of them.

  1. Docker Compose: TechCorp's complete local environment

Docker Compose describes a set of containers, networks and volumes in a YAML file and brings them up with one command. It is the tool for the local environment and the E2E test (04-01, 04-05); it is not a production orchestrator. The file lives in the techcorp/platform repository, at local/compose.yaml, and assumes the service repositories are cloned as sibling directories (../../catalog-service, etc.).

# techcorp/platform/local/compose.yaml
name: techcorp

services:
  # ---------- Dependencies ----------
  postgres:
    image: postgres:16
    environment:
      POSTGRES_USER: svc_orders
      POSTGRES_PASSWORD: dev-orders            # local development only; in production, a Secret (05-02)
      POSTGRES_DB: orders
    volumes:
      - pg-data:/var/lib/postgresql/data       # data survives docker compose down (without -v)
    ports:
      - "5432:5432"                            # published only so psql/DBeaver can connect from the laptop
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U svc_orders -d orders"]
      interval: 5s
      timeout: 3s
      retries: 10

  mongo:
    image: mongo:7
    volumes:
      - mongo-data:/data/db
    ports:
      - "27017:27017"
    healthcheck:
      test: ["CMD", "mongosh", "--quiet", "--eval", "db.adminCommand('ping').ok"]
      interval: 5s
      timeout: 3s
      retries: 10

  rabbitmq:
    image: rabbitmq:3-management
    ports:
      - "5672:5672"                            # AMQP (the services)
      - "15672:15672"                          # web console http://localhost:15672 (guest/guest)
    volumes:
      - rabbitmq-data:/var/lib/rabbitmq
    healthcheck:
      test: ["CMD", "rabbitmq-diagnostics", "-q", "ping"]
      interval: 5s
      timeout: 5s
      retries: 12

  # ---------- One-off tasks ----------
  catalog-seed:                                # scripts/seed.js from 04-02, with the service's own image
    image: ghcr.io/techcorp/catalog-service:local
    build:
      context: ../../catalog-service
      secrets: [npmrc]
    command: ["node", "scripts/seed.js"]
    environment:
      MONGO_URL: mongodb://mongo:27017
      MONGO_DB: catalog
    depends_on:
      mongo: { condition: service_healthy }
    restart: "no"                              # finishes and is not restarted; it is idempotent (upsert)

  orders-migrations:                           # scripts/migrate.js from 04-04, before starting orders-service
    image: ghcr.io/techcorp/orders-service:local
    build:
      context: ../../orders-service
      secrets: [npmrc]
    command: ["node", "scripts/migrate.js"]
    environment:
      ORDERS_DB_URL: postgres://svc_orders:dev-orders@postgres:5432/orders
    depends_on:
      postgres: { condition: service_healthy }
    restart: "no"

  # ---------- Services ----------
  catalog-service:
    image: ghcr.io/techcorp/catalog-service:local
    build:
      context: ../../catalog-service
      secrets: [npmrc]
    environment:
      PORT: "3001"
      MONGO_URL: mongodb://mongo:27017
      MONGO_DB: catalog
      LOG_LEVEL: debug
    depends_on:
      mongo: { condition: service_healthy }
      catalog-seed: { condition: service_completed_successfully }
    stop_grace_period: 15s                     # room for the graceful shutdown (10 s internal + slack)

  orders-service:
    image: ghcr.io/techcorp/orders-service:local
    build:
      context: ../../orders-service
      secrets: [npmrc]
    environment:                               # exactly the "development" column of the 04-03 table, with Compose network names
      PORT: "3002"
      NODE_ENV: production
      LOG_LEVEL: debug
      ORDERS_DB_URL: postgres://svc_orders:dev-orders@postgres:5432/orders
      RABBITMQ_URL: amqp://rabbitmq:5672
      CATALOG_URL: http://catalog-service:3001
      CUSTOMERS_URL: http://customers-service:3004
      HTTP_TIMEOUT_MS: "2000"
      OUTBOX_INTERVAL_MS: "500"
      REMOTE_CATALOG: "true"
    depends_on:
      postgres: { condition: service_healthy }
      rabbitmq: { condition: service_healthy }
      orders-migrations: { condition: service_completed_successfully }
      catalog-service: { condition: service_started }
    stop_grace_period: 15s

  customers-service:                           # not extracted yet: the 04-04 stub (scripts/stubCustomers.js) answers c-1024
    image: ghcr.io/techcorp/orders-service:local
    command: ["node", "scripts/stubCustomers.js"]
    environment:
      PORT: "3004"

  gateway:                                     # the Express gateway of 03-04, with the /api/v1/* routes
    image: ghcr.io/techcorp/gateway:local
    build:
      context: ../../gateway
      secrets: [npmrc]
    ports:
      - "8080:8080"                            # the ONLY public door of the system
    environment:
      PORT: "8080"
      CATALOG_URL: http://catalog-service:3001
      ORDERS_URL: http://orders-service:3002
      CUSTOMERS_URL: http://customers-service:3004
      MONOLITH_URL: http://host.docker.internal:3000   # the monolith still runs on the laptop, if needed
    depends_on:
      - catalog-service
      - orders-service
      - customers-service

volumes:
  pg-data:
  mongo-data:
  rabbitmq-data:

secrets:
  npmrc:
    file: ${HOME}/.npmrc                       # GitHub Packages token for npm ci; it never gets into the image

Points worth understanding well:

  • Default network. Compose creates a techcorp_default network and connects every service to it; each one resolves the others by service name. That is why CATALOG_URL=http://catalog-service:3001 and RABBITMQ_URL=amqp://rabbitmq:5672 are identical to the ones we will use in Kubernetes: the configuration of 04-03 does not change between local and cluster.
  • ports only where needed. The gateway publishes 8080; the databases and RabbitMQ are published for development convenience. The *-service services do not publish ports: they are reached only through the gateway or from other containers, as in production.
  • healthcheck + depends_on: condition. Without a condition, depends_on only orders the startup, it does not wait for PostgreSQL to accept connections; orders-service would start, config.js would validate, the pool would fail and /health/ready would sit at 503 until it connected. With service_healthy Compose waits for the healthcheck; with service_completed_successfully it waits for the one-off task to finish with exit code 0. So the migrations are always applied before Orders starts, without wait scripts.
  • One-off tasks with the same image. catalog-seed and orders-migrations do not need another image: they use the service's and change command. That is the reason the Dockerfile copies scripts/. In Kubernetes they will be Jobs (05-02).
  • image + build. With both, docker compose build builds and tags the image as :local; docker compose up uses it. If tomorrow you want to try the image CI published, just change :local to :sha-9f3c2ab and drop the build.
  • customers-service as a stub. It reuses scripts/stubCustomers.js from 04-04 out of the Orders image. When the real service exists (extraction order of 02-02), it is replaced by its own build, with CUSTOMERS_DB_URL and the rest untouched.

  1. Operating the environment: up, down, logs, ps, migrations and the E2E test

cd platform/local

docker compose build                    # builds the :local images (uses the layer cache from section 4)
docker compose up -d                    # brings everything up in order: dependencies → seed/migrations → services → gateway
docker compose ps                       # status of each service: running (healthy), exited (0) for the one-off tasks
docker compose logs -f orders-service   # logs of one service; without a name, all of them, interleaved and prefixed
docker compose exec orders-service sh   # shell inside a service
docker compose run --rm orders-migrations     # re-run the migrations by hand (e.g. after adding 005-*.sql)
docker compose restart orders-service   # restart one (reloads variables if the YAML changed after an 'up')
docker compose down                     # stops and removes containers and network; volumes are kept
docker compose down -v                  # ...and removes the volumes too: clean database

Manual check of the order flow of the whole course, this time through the gateway and with every service in containers:

curl -s http://localhost:8080/api/v1/products?ids=p-501,p-777 | jq .data[].name
curl -s -X POST http://localhost:8080/api/v1/orders \
  -H 'Content-Type: application/json' -H 'Idempotency-Key: 7c1e0b3a-e2e-0001' \
  -d '{"customerId":"c-1024","lines":[{"productId":"p-501","quantity":1},{"productId":"p-777","quantity":2}]}'
# → 202 Accepted, {"id":"ord-...","status":"PENDING"} ; in the RabbitMQ console (15672) order.created shows up on techcorp.events

Since Inventory, Payments and Notifications are not extracted yet, locally their responses are simulated by publishing the events with scripts/publishEvent.js from 04-04, or by adding consumer stubs to Compose (exercise 2). The E2E test of 04-05 runs against this same environment from the platform repository:

docker compose up -d --wait             # --wait: doesn't return control until every healthcheck is green
GATEWAY_URL=http://localhost:8080 npm run test:e2e   # tests/e2e/createOrder.e2e.test.js: POST /api/v1/orders and polling until CONFIRMED (15 s)
docker compose down -v

That trio of commands is literally what the pipeline of 05-03 will run before promoting to staging: if a queue is bound wrongly or a variable is missing, it fails here and not in production.

  1. Image size and basic security

Practice Effect Status at TechCorp
node:20-alpine base (or node:20-slim if a native module does not compile against musl) Final image of ~130 MB instead of ~450 MB; smaller attack surface Alpine by default in the template
Multi-stage + npm ci --omit=dev No Jest, Pact or Testcontainers in production Yes
.dockerignore Small context, without .env or .git Yes
Non-root user (USER node) A flaw in the service does not grant root inside the container Yes; Kubernetes will enforce it with runAsNonRoot (05-02)
No secrets in the image No .env, no .npmrc, no ARG with passwords (ARGs remain in the image history) --mount=type=secret for the npm token
Pin the base by digest (node:20-alpine@sha256:...) Reproducible builds even if the tag moves Managed by the pipeline with Renovate/Dependabot (05-03)
Vulnerability scanning (Trivy, Grype) Detect CVEs in the base and in node_modules Pipeline stage; details in 07-04
docker history / dive See which layer is heavy and why Diagnostic tool

The rule in short: the image contains code and dependencies; nothing else. Configuration and secrets come from the environment (04-03), data lives in volumes or external services, and the process is not root.

Common Mistakes and Tips

  • COPY . . at the top of the Dockerfile. Every code change repeats npm ci. First package*.json, then install, then the code.
  • Forgetting .dockerignore. The laptop's node_modules (with macOS binaries) and the .env with the PostgreSQL password end up inside the image published to ghcr.io.
  • CMD npm start or the shell form. SIGTERM never reaches Node; every deployment cuts requests. Exec form and node directly.
  • latest anywhere other than the laptop. No traceability and no rollback.
  • depends_on without condition. It "works" on a fast laptop and fails on the CI runner, where PostgreSQL takes 8 s to accept connections. healthcheck on every dependency and service_healthy.
  • Publishing the internal services' ports "for testing". They are tested through the gateway (or with docker compose exec); publishing 3002 gets you used to bypassing the only door that will exist in production.
  • Tip: run docker compose config to see the final YAML with variables substituted; and docker compose up --build --wait as the single command to start the day.

Exercises

Exercise 1. Write the production Dockerfile of orders-service (04-04) starting from Catalog's. State which lines change and why, bearing in mind that Orders needs migrations/ and scripts/migrate.js inside the image, listens on 3002 and its HEALTHCHECK must not be used in Kubernetes.

Exercise 2. The Orders team wants the E2E test to pass locally without real Inventory or Payments. Add to compose.yaml a simulated-saga service that runs a Node script (scripts/simulatedSaga.js, already written: it consumes order.created from its own queue and publishes stock.reserved and payment.confirmed) using the Orders image. Which variables does it need, what does it depend on, and why must it not publish ports?

Exercise 3. A colleague runs docker stop catalog-service and notices that it takes exactly 10 s and that "shutting down" never appears in the logs. Their Dockerfile ends with CMD npm start. Explain the cause, the fix, and what other symptom they would see in Kubernetes with terminationGracePeriodSeconds: 30.

Solutions

Solution 1. What changes: COPY --chown=node:node migrations ./migrations added next to src and scripts (the Job of 05-02 and Compose's orders-migrations run node scripts/migrate.js from this image and read migrations/*.sql); EXPOSE 3002; the HEALTHCHECK points to http://localhost:3002/health/live. Since Kubernetes ignores HEALTHCHECK (it uses livenessProbe), it can be kept for Compose or removed; the important thing is not to confuse it with the probe. Everything else (two stages, npm ci --omit=dev with the npmrc secret, --chown=node:node, USER node, CMD ["node","src/server.js"]) is identical: it is what the 04-01 template gives you ready-made.

Solution 2.

  simulated-saga:
    image: ghcr.io/techcorp/orders-service:local
    command: ["node", "scripts/simulatedSaga.js"]
    environment:
      RABBITMQ_URL: amqp://rabbitmq:5672
      LOG_LEVEL: debug
    depends_on:
      rabbitmq: { condition: service_healthy }
      orders-service: { condition: service_started }   # so the topology (techcorp.events exchange) is already declared
    restart: unless-stopped

It only needs RABBITMQ_URL (it talks exclusively through events, like Inventory and Payments in 03-04: "not exposed"). It publishes no ports because it exposes no HTTP and because nothing outside the Compose network should talk to it; if one day something did, it would go through the gateway. It depends on a healthy RabbitMQ and on Orders having started (it declares the topology when it connects; alternatively the script can declare it itself with the library's messaging/topology, and then the second dependency is unnecessary).

Solution 3. With CMD npm start, PID 1 is npm, which starts node src/server.js as a child. docker stop sends SIGTERM to PID 1 (npm), which does not forward it to Node; the process.on('SIGTERM') handler of 04-02 never runs, so there is no "shutting down" line and no server.close(). After 10 s Docker sends SIGKILL and everything dies abruptly: in-flight requests get a connection reset. Fix: CMD ["node", "src/server.js"] (exec form, Node as PID 1). In Kubernetes the symptom would be that every rolling update (05-04) takes 30 s per pod instead of 1-2 s and that, during those 30 s, the pod keeps receiving traffic until the readinessProbe fails, because /health/ready never switched to 503; with the exec form, the service marks itself not ready the instant SIGTERM arrives.

Conclusion

We have turned the Node processes of the previous modules into deployable units: one image per service, built with a multi-stage Dockerfile on node:20-alpine (npm ci --omit=dev with the npm token as a build secret, COPY --chown=node:node, USER node, EXPOSE, optional HEALTHCHECK on /health/live, CMD ["node","src/server.js"] in exec form so SIGTERM reaches the graceful shutdown of 04-02), with a .dockerignore and an instruction order that exploits the layer cache, tagged with a semantic version and sha-<commit> in ghcr.io/techcorp/ and never with latest. And we have brought the complete system up for the first time with Docker Compose (platform/local/compose.yaml): PostgreSQL 16, MongoDB 7 and RabbitMQ with healthchecks, one-off tasks for the seed and the migrations, catalog-service, orders-service, the Customers stub and the gateway on 8080, with the same environment variables of 04-03 and the same DNS names of 03-05, and on top of it the course's single E2E test. Compose is perfect for a laptop, but it does not restart a crashed container on another machine, does not spread replicas and does not manage zero-downtime deployments: that is what Kubernetes is for, and in the next lesson it will run these very images with the ConfigMaps and Secrets that 04-03 left named.

Microservices Course

Module 1: Introduction to Microservices

Module 2: Microservice Design

Module 3: Communication between Microservices

Module 4: Implementing Microservices

Module 5: Deployment and Orchestration

Module 6: Monitoring and Maintenance

Module 7: Security in Microservices

Module 8: Case Studies and Practical Examples

© Copyright 2026. All rights reserved