When Module 4 ended we had catalog-service and orders-service running with npm run dev on a laptop, with MongoDB, PostgreSQL and RabbitMQ started by hand with docker run. That is fine for developing, but not for deploying: in production each service will run as several replicas, on machines nobody has prepared by hand, and it will have to start, die and start again hundreds of times without a human stepping in. The container is the unit that makes all of that possible: it packages the service with its exact Node.js and its dependencies into an immutable image that runs identically on Luis's laptop, on the CI runner and in the Kubernetes cluster. This lesson builds the production image of catalog-service, explains how it is tagged, run and shut down, and brings up TechCorp's complete system with Docker Compose for the first time, including the single E2E test that 04-05 left pending. Kubernetes (05-02) and the pipeline that will build these images on every commit (05-03) rest on what gets settled here.
Contents
- Why containers for microservices
- Concepts: image, layer, container, registry, volume and network
- The production
Dockerfileofcatalog-service, line by line .dockerignore, layer cache and instruction order- Image tagging and the registry
- Basic Docker commands
- Signals, PID 1 and graceful shutdown inside the container
- Docker Compose: TechCorp's complete local environment
- Operating the environment:
up,down,logs,ps, migrations and the E2E test - Image size and basic security
- Why containers for microservices
A container is an isolated process that the Linux kernel runs with its own filesystem, its own view of the network and CPU and memory limits, starting from an image that contains everything that process needs (the Node binary, node_modules, code). There is no guest operating system and no hypervisor: the container shares the host's kernel, which is why it starts in milliseconds and takes up megabytes.
| Aspect | Virtual machine | Container |
|---|---|---|
| What is virtualized | The full hardware (CPU, disk, network) with its own kernel | Only user space; shares the host's kernel |
| Typical size | Gigabytes | Tens or hundreds of megabytes |
| Startup | Tens of seconds to minutes | Milliseconds to seconds |
| Density per machine | Units or tens | Hundreds |
| Isolation | Strong (own kernel) | Good (namespaces and cgroups), not equivalent to a VM |
| Deployment unit | A VM image (AMI, OVA), heavy to build | A container image, built in seconds by CI |
For TechCorp the fit is direct:
- One image = one deployment unit.
catalog-serviceis an image; deploying version 1.4.2 means running that image. There is no "install Node 20 on the server" and no "copy the folder and runnpm install": all of that happened once, when the image was built. - Reproducibility. The very same image that passed the integration tests in CI is the one that runs in production, byte for byte. No more "works on my machine" and no more "production has a different version of
pg". - Isolation. Six Node services on the same machine share no
node_modules, no ports and no environment variables. Each one believes it has the machine to itself. - Fast startup, disposable. The graceful shutdown of 04-02 and the per-environment configuration of 04-03 were written with this in mind: a container is born, serves, receives SIGTERM and dies; another one takes its place.
- Concepts: image, layer, container, registry, volume and network
| Concept | What it is | At TechCorp |
|---|---|---|
| Image | Immutable, read-only template with the filesystem and metadata (start command, ports, user) | ghcr.io/techcorp/catalog-service:1.4.2 |
| Layer | Every Dockerfile instruction that changes the filesystem produces a layer; the image is the stack of layers. Layers are cached and shared between images |
The six service images share the node:20-alpine base layer |
| Container | A running instance of an image: the image's layers plus an ephemeral writable layer | Each replica of catalog-service |
| Registry | Image store you push to and pull from |
GitHub Container Registry (ghcr.io), next to the code and to @techcorp/common-http (04-01) |
| Volume | Persistent storage outside the container's writable layer | PostgreSQL, MongoDB and RabbitMQ data locally; TechCorp's services do not use volumes (they are stateless) |
| Network | Virtual network in which containers resolve each other by name | In Compose, orders-service calls http://catalog-service:3001, the same name the Kubernetes DNS will resolve (03-05) |
That the services keep no state on disk is no accident: it is what allows killing and replicating them without a second thought. Everything that must survive lives in the database or in the broker.
- The production
Dockerfile of catalog-service, line by line
Dockerfile of catalog-service, line by lineThe Dockerfile lives at the root of the techcorp/catalog-service repository and comes from the techcorp/node-service-template template (04-01, exercise 2: the Dockerfile is copied and adapted). It uses a multi-stage build: one stage installs dependencies and another, clean one keeps only what is needed to run.
# syntax=docker/dockerfile:1
# ---------- Stage 1: dependencies ----------
FROM node:20-alpine AS dependencies
WORKDIR /app
# Only the files that define the dependencies: if they don't change, this layer (and npm ci) is reused from the cache.
COPY package.json package-lock.json ./
# @techcorp/common-http lives in GitHub Packages (04-01): npm needs a token to download it.
# The .npmrc is mounted ONLY during this RUN as a build secret; it never ends up in any image layer.
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc \
npm ci --omit=dev
# ---------- Stage 2: final image ----------
FROM node:20-alpine
ENV NODE_ENV=production
WORKDIR /app
# node_modules already installed and without dev dependencies; owned by the 'node' user that the base image provides.
COPY --from=dependencies --chown=node:node /app/node_modules ./node_modules
COPY --chown=node:node package.json ./
COPY --chown=node:node src ./src
COPY --chown=node:node scripts ./scripts
# From here on the process is not root.
USER node
EXPOSE 3001
# Optional: Docker (not Kubernetes) marks the container as unhealthy if /health/live doesn't answer 200.
HEALTHCHECK --interval=30s --timeout=3s --start-period=10s --retries=3 \
CMD wget -qO- http://localhost:3001/health/live || exit 1
# Exec form (JSON): node is PID 1 and receives SIGTERM directly (section 7).
CMD ["node", "src/server.js"]Instruction by instruction:
# syntax=docker/dockerfile:1: enables the modern BuildKit syntax (required for--mount=type=secret).FROM node:20-alpine AS dependencies: official Node 20 base image on Alpine Linux (about 50 MB versus the 350 MB ofnode:20).AS dependenciesnames the stage so we can copy from it later.WORKDIR /app: creates/appand makes it the working directory for the following instructions.COPY package.json package-lock.json ./before copying the code: this is the key to the cache (section 4).RUN --mount=type=secret,id=npmrc ... npm ci --omit=dev:npm ciinstalls exactly whatpackage-lock.jsonsays (reproducible; it fails if the lock does not matchpackage.json);--omit=devleaves outnodemon,jest,supertest,@pact-foundation/pact,testcontainers. The GitHub Packages token is mounted as a file only during this command; it is passed at build time with--secret id=npmrc,src=$HOME/.npmrc.FROM node:20-alpine(second time): the final image starts from scratch; nothing from the previous stage carries over except what is copied explicitly.ENV NODE_ENV=production: Express disables debug messages and some libraries optimize; theconfig.jsof 04-03 reads it as just another variable.COPY --from=dependencies --chown=node:node /app/node_modules ./node_modules: brings only the already-installednode_modules.--chown=node:nodeassigns ownership to thenodeuser (uid 1000) that the base image already defines; without it the files would belong to root and the non-root process might not be able to read them if permissions were restrictive.- Three separate
COPYinstructions forpackage.json,srcandscripts:scripts/comes in becauseseed.jsruns from this same image (section 8). No tests, contracts or development configuration files are copied. USER node: everything executed from here on (includingCMD) runs unprivileged. It is the cheapest and most effective security measure in this lesson.EXPOSE 3001: documents the port (it does not publish it; that is done with-por in Compose/Kubernetes). It matchesPORT=3001from 04-02.HEALTHCHECK: Docker runswgetagainst/health/liveevery 30 s, with a 10 s grace period at startup; three consecutive failures mark the containerunhealthy. Alpine shipswgetin BusyBox, so there is no need to installcurl. Kubernetes ignoresHEALTHCHECKand uses its own probes (05-02); in Compose it is useful fordepends_on(section 8).CMD ["node", "src/server.js"]: the start command, the same as thestartscript of 04-02, but without going throughnpm(section 7).
Why multi-stage if a JavaScript service compiles nothing? Because it separates two concerns: the first stage may hold tokens, the npm cache, build tools for native modules; the second holds only what runs. If one day a service is written in TypeScript, stage 1 becomes "install and compile" and stage 2 does not change.
.dockerignore, layer cache and instruction order
.dockerignore, layer cache and instruction orderCOPY src ./src copies what is in the build context (the directory passed to docker build). Without a filter, a COPY . . would drag along the local node_modules (with dev dependencies and binaries for another operating system), .env with secrets, .git, pacts/, test coverage... The .dockerignore excludes them from the context:
About the cache: Docker builds the layers in order and, for each instruction, reuses the cached layer if the instruction and its inputs have not changed; as soon as one layer changes, all the following ones are rebuilt. Hence the order of the Dockerfile:
flowchart LR
A[FROM node:20-alpine] --> B[COPY package*.json]
B --> C[RUN npm ci]
C --> D[COPY src, scripts]
D --> E[CMD]
style C fill:#dfe,stroke:#393
style D fill:#fdd,stroke:#933
A change in src/routes/products.js invalidates only the red layer: the npm ci (the slow operation, 30-60 s with download) is reused from the cache. If we copied all the code first and installed afterwards, every commit would repeat the installation. Rule: what changes least goes at the top; what changes most goes at the bottom.
- Image tagging and the registry
An image is identified by registry/organization/name:tag. TechCorp uses two tags per build:
| Tag | Example | Who uses it |
|---|---|---|
| Semantic version | ghcr.io/techcorp/catalog-service:1.4.2 |
Kubernetes manifests, release notes, humans |
| Source commit | ghcr.io/techcorp/catalog-service:sha-9f3c2ab |
Exact traceability: which code the image comes from; the pipeline of 05-03 creates it on every build |
Both point to the same digest (sha256:...), which is the real, immutable identifier of the image. What is not used in production is latest:
latestdoes not mean "the most recent", it means "the last one somebody put that tag on". It is a mutable tag: today it points to 1.4.2 and tomorrow to 1.5.0 without any manifest changing.- With
latest, two replicas of the sameDeploymentmay run different versions depending on when they pulled, and arollbackis impossible because nobody knows what was there before. - The manifests of 05-02 always carry a specific tag; changing version means changing that line, and that ends up in git.
Manual build and publish (the pipeline of 05-03 automates exactly this):
# At the root of catalog-service. -t adds a tag; several can be given.
docker build \
--secret id=npmrc,src=$HOME/.npmrc \
-t ghcr.io/techcorp/catalog-service:1.4.2 \
-t ghcr.io/techcorp/catalog-service:sha-9f3c2ab \
.
# Authenticate against the registry with a GitHub token that has the write:packages permission
echo "$GITHUB_TOKEN" | docker login ghcr.io -u luis --password-stdin
docker push ghcr.io/techcorp/catalog-service:1.4.2
docker push ghcr.io/techcorp/catalog-service:sha-9f3c2ab
- Basic Docker commands
The ones used every day, with the freshly built image:
# Run in the background (-d), with a name, publishing the container's port 3001 on the host's 3001 (-p host:container)
# and passing the configuration through environment variables (04-03). --rm removes the container when it stops.
docker run -d --rm --name catalog -p 3001:3001 \
-e MONGO_URL=mongodb://host.docker.internal:27017 -e MONGO_DB=catalog -e LOG_LEVEL=debug \
ghcr.io/techcorp/catalog-service:1.4.2
docker ps # running containers: id, image, status (healthy/unhealthy), ports
docker logs -f catalog # the process's stdout/stderr (pino writes JSON to stdout: that's why it works without log files)
docker exec -it catalog sh # a shell inside the container (Alpine: sh, not bash) to inspect
docker exec catalog node scripts/seed.js # run a one-off command with the container's own environment
docker stop catalog # SIGTERM and, if it hasn't finished within 10 s, SIGKILL (section 7)
docker images # local images and their sizehost.docker.internal is the name through which a container reaches the host (Docker Desktop; on Linux, --add-host=host.docker.internal:host-gateway). It only makes sense in this standalone experiment: in Compose and in Kubernetes the services find each other by name inside the same network.
- Signals, PID 1 and graceful shutdown inside the container
In 04-02 we wrote the graceful shutdown: on SIGTERM, /health/ready switches to 503, server.close() finishes in-flight requests, Mongo is closed and the process exits, with a 10 s limit. Inside a container there are three details that can undo that work:
- Who is PID 1. The process started by
CMDis PID 1 of the container and is the only one that receives the signal fromdocker stopor from the kubelet. With the exec formCMD ["node", "src/server.js"], PID 1 is Node and our handler runs. With the shell formCMD node src/server.js, PID 1 is/bin/sh, which receives SIGTERM and does not forward it to Node: the service dies 10 s later by SIGKILL, with requests half done. And withCMD ["npm", "start"], PID 1 isnpm, which does not forward signals reliably either. That is why theDockerfilecallsnodedirectly. - PID 1 has no default handlers. The kernel treats PID 1 specially: if it does not install a handler for a signal, it ignores it. Node does install ours (
process.on('SIGTERM')), so we are covered; a service that forgot to do so would never die ondocker stop, only on SIGKILL. - Zombie processes. PID 1 must "reap" terminated child processes. Node spawns no children in our services, but if one day a service runs
child_process, a minimalinitis advisable:docker run --init(Docker injectstini) or, in Kubernetes,shareProcessNamespaceortiniinside the image. It is a cheap line that avoids a problem that is hard to diagnose.
Timings: docker stop waits 10 s by default (-t changes it); Compose uses stop_grace_period; Kubernetes, terminationGracePeriodSeconds (30 s by default, 05-02). The internal 10 s limit of 04-02 is designed to fit inside any of them.
- Docker Compose: TechCorp's complete local environment
Docker Compose describes a set of containers, networks and volumes in a YAML file and brings them up with one command. It is the tool for the local environment and the E2E test (04-01, 04-05); it is not a production orchestrator. The file lives in the techcorp/platform repository, at local/compose.yaml, and assumes the service repositories are cloned as sibling directories (../../catalog-service, etc.).
# techcorp/platform/local/compose.yaml
name: techcorp
services:
# ---------- Dependencies ----------
postgres:
image: postgres:16
environment:
POSTGRES_USER: svc_orders
POSTGRES_PASSWORD: dev-orders # local development only; in production, a Secret (05-02)
POSTGRES_DB: orders
volumes:
- pg-data:/var/lib/postgresql/data # data survives docker compose down (without -v)
ports:
- "5432:5432" # published only so psql/DBeaver can connect from the laptop
healthcheck:
test: ["CMD-SHELL", "pg_isready -U svc_orders -d orders"]
interval: 5s
timeout: 3s
retries: 10
mongo:
image: mongo:7
volumes:
- mongo-data:/data/db
ports:
- "27017:27017"
healthcheck:
test: ["CMD", "mongosh", "--quiet", "--eval", "db.adminCommand('ping').ok"]
interval: 5s
timeout: 3s
retries: 10
rabbitmq:
image: rabbitmq:3-management
ports:
- "5672:5672" # AMQP (the services)
- "15672:15672" # web console http://localhost:15672 (guest/guest)
volumes:
- rabbitmq-data:/var/lib/rabbitmq
healthcheck:
test: ["CMD", "rabbitmq-diagnostics", "-q", "ping"]
interval: 5s
timeout: 5s
retries: 12
# ---------- One-off tasks ----------
catalog-seed: # scripts/seed.js from 04-02, with the service's own image
image: ghcr.io/techcorp/catalog-service:local
build:
context: ../../catalog-service
secrets: [npmrc]
command: ["node", "scripts/seed.js"]
environment:
MONGO_URL: mongodb://mongo:27017
MONGO_DB: catalog
depends_on:
mongo: { condition: service_healthy }
restart: "no" # finishes and is not restarted; it is idempotent (upsert)
orders-migrations: # scripts/migrate.js from 04-04, before starting orders-service
image: ghcr.io/techcorp/orders-service:local
build:
context: ../../orders-service
secrets: [npmrc]
command: ["node", "scripts/migrate.js"]
environment:
ORDERS_DB_URL: postgres://svc_orders:dev-orders@postgres:5432/orders
depends_on:
postgres: { condition: service_healthy }
restart: "no"
# ---------- Services ----------
catalog-service:
image: ghcr.io/techcorp/catalog-service:local
build:
context: ../../catalog-service
secrets: [npmrc]
environment:
PORT: "3001"
MONGO_URL: mongodb://mongo:27017
MONGO_DB: catalog
LOG_LEVEL: debug
depends_on:
mongo: { condition: service_healthy }
catalog-seed: { condition: service_completed_successfully }
stop_grace_period: 15s # room for the graceful shutdown (10 s internal + slack)
orders-service:
image: ghcr.io/techcorp/orders-service:local
build:
context: ../../orders-service
secrets: [npmrc]
environment: # exactly the "development" column of the 04-03 table, with Compose network names
PORT: "3002"
NODE_ENV: production
LOG_LEVEL: debug
ORDERS_DB_URL: postgres://svc_orders:dev-orders@postgres:5432/orders
RABBITMQ_URL: amqp://rabbitmq:5672
CATALOG_URL: http://catalog-service:3001
CUSTOMERS_URL: http://customers-service:3004
HTTP_TIMEOUT_MS: "2000"
OUTBOX_INTERVAL_MS: "500"
REMOTE_CATALOG: "true"
depends_on:
postgres: { condition: service_healthy }
rabbitmq: { condition: service_healthy }
orders-migrations: { condition: service_completed_successfully }
catalog-service: { condition: service_started }
stop_grace_period: 15s
customers-service: # not extracted yet: the 04-04 stub (scripts/stubCustomers.js) answers c-1024
image: ghcr.io/techcorp/orders-service:local
command: ["node", "scripts/stubCustomers.js"]
environment:
PORT: "3004"
gateway: # the Express gateway of 03-04, with the /api/v1/* routes
image: ghcr.io/techcorp/gateway:local
build:
context: ../../gateway
secrets: [npmrc]
ports:
- "8080:8080" # the ONLY public door of the system
environment:
PORT: "8080"
CATALOG_URL: http://catalog-service:3001
ORDERS_URL: http://orders-service:3002
CUSTOMERS_URL: http://customers-service:3004
MONOLITH_URL: http://host.docker.internal:3000 # the monolith still runs on the laptop, if needed
depends_on:
- catalog-service
- orders-service
- customers-service
volumes:
pg-data:
mongo-data:
rabbitmq-data:
secrets:
npmrc:
file: ${HOME}/.npmrc # GitHub Packages token for npm ci; it never gets into the imagePoints worth understanding well:
- Default network. Compose creates a
techcorp_defaultnetwork and connects every service to it; each one resolves the others by service name. That is whyCATALOG_URL=http://catalog-service:3001andRABBITMQ_URL=amqp://rabbitmq:5672are identical to the ones we will use in Kubernetes: the configuration of 04-03 does not change between local and cluster. portsonly where needed. The gateway publishes 8080; the databases and RabbitMQ are published for development convenience. The*-serviceservices do not publish ports: they are reached only through the gateway or from other containers, as in production.healthcheck+depends_on: condition. Without a condition,depends_ononly orders the startup, it does not wait for PostgreSQL to accept connections;orders-servicewould start,config.jswould validate, the pool would fail and/health/readywould sit at 503 until it connected. Withservice_healthyCompose waits for the healthcheck; withservice_completed_successfullyit waits for the one-off task to finish with exit code 0. So the migrations are always applied before Orders starts, without wait scripts.- One-off tasks with the same image.
catalog-seedandorders-migrationsdo not need another image: they use the service's and changecommand. That is the reason theDockerfilecopiesscripts/. In Kubernetes they will beJobs (05-02). image+build. With both,docker compose buildbuilds and tags the image as:local;docker compose upuses it. If tomorrow you want to try the image CI published, just change:localto:sha-9f3c2aband drop thebuild.customers-serviceas a stub. It reusesscripts/stubCustomers.jsfrom 04-04 out of the Orders image. When the real service exists (extraction order of 02-02), it is replaced by its ownbuild, withCUSTOMERS_DB_URLand the rest untouched.
- Operating the environment:
up, down, logs, ps, migrations and the E2E test
up, down, logs, ps, migrations and the E2E testcd platform/local
docker compose build # builds the :local images (uses the layer cache from section 4)
docker compose up -d # brings everything up in order: dependencies → seed/migrations → services → gateway
docker compose ps # status of each service: running (healthy), exited (0) for the one-off tasks
docker compose logs -f orders-service # logs of one service; without a name, all of them, interleaved and prefixed
docker compose exec orders-service sh # shell inside a service
docker compose run --rm orders-migrations # re-run the migrations by hand (e.g. after adding 005-*.sql)
docker compose restart orders-service # restart one (reloads variables if the YAML changed after an 'up')
docker compose down # stops and removes containers and network; volumes are kept
docker compose down -v # ...and removes the volumes too: clean databaseManual check of the order flow of the whole course, this time through the gateway and with every service in containers:
curl -s http://localhost:8080/api/v1/products?ids=p-501,p-777 | jq .data[].name
curl -s -X POST http://localhost:8080/api/v1/orders \
-H 'Content-Type: application/json' -H 'Idempotency-Key: 7c1e0b3a-e2e-0001' \
-d '{"customerId":"c-1024","lines":[{"productId":"p-501","quantity":1},{"productId":"p-777","quantity":2}]}'
# → 202 Accepted, {"id":"ord-...","status":"PENDING"} ; in the RabbitMQ console (15672) order.created shows up on techcorp.eventsSince Inventory, Payments and Notifications are not extracted yet, locally their responses are simulated by publishing the events with scripts/publishEvent.js from 04-04, or by adding consumer stubs to Compose (exercise 2). The E2E test of 04-05 runs against this same environment from the platform repository:
docker compose up -d --wait # --wait: doesn't return control until every healthcheck is green
GATEWAY_URL=http://localhost:8080 npm run test:e2e # tests/e2e/createOrder.e2e.test.js: POST /api/v1/orders and polling until CONFIRMED (15 s)
docker compose down -vThat trio of commands is literally what the pipeline of 05-03 will run before promoting to staging: if a queue is bound wrongly or a variable is missing, it fails here and not in production.
- Image size and basic security
| Practice | Effect | Status at TechCorp |
|---|---|---|
node:20-alpine base (or node:20-slim if a native module does not compile against musl) |
Final image of ~130 MB instead of ~450 MB; smaller attack surface | Alpine by default in the template |
Multi-stage + npm ci --omit=dev |
No Jest, Pact or Testcontainers in production | Yes |
.dockerignore |
Small context, without .env or .git |
Yes |
Non-root user (USER node) |
A flaw in the service does not grant root inside the container | Yes; Kubernetes will enforce it with runAsNonRoot (05-02) |
| No secrets in the image | No .env, no .npmrc, no ARG with passwords (ARGs remain in the image history) |
--mount=type=secret for the npm token |
Pin the base by digest (node:20-alpine@sha256:...) |
Reproducible builds even if the tag moves | Managed by the pipeline with Renovate/Dependabot (05-03) |
| Vulnerability scanning (Trivy, Grype) | Detect CVEs in the base and in node_modules |
Pipeline stage; details in 07-04 |
docker history / dive |
See which layer is heavy and why | Diagnostic tool |
The rule in short: the image contains code and dependencies; nothing else. Configuration and secrets come from the environment (04-03), data lives in volumes or external services, and the process is not root.
Common Mistakes and Tips
COPY . .at the top of theDockerfile. Every code change repeatsnpm ci. Firstpackage*.json, then install, then the code.- Forgetting
.dockerignore. The laptop'snode_modules(with macOS binaries) and the.envwith the PostgreSQL password end up inside the image published toghcr.io. CMD npm startor the shell form. SIGTERM never reaches Node; every deployment cuts requests. Exec form andnodedirectly.latestanywhere other than the laptop. No traceability and no rollback.depends_onwithoutcondition. It "works" on a fast laptop and fails on the CI runner, where PostgreSQL takes 8 s to accept connections.healthcheckon every dependency andservice_healthy.- Publishing the internal services' ports "for testing". They are tested through the gateway (or with
docker compose exec); publishing3002gets you used to bypassing the only door that will exist in production. - Tip: run
docker compose configto see the final YAML with variables substituted; anddocker compose up --build --waitas the single command to start the day.
Exercises
Exercise 1. Write the production Dockerfile of orders-service (04-04) starting from Catalog's. State which lines change and why, bearing in mind that Orders needs migrations/ and scripts/migrate.js inside the image, listens on 3002 and its HEALTHCHECK must not be used in Kubernetes.
Exercise 2. The Orders team wants the E2E test to pass locally without real Inventory or Payments. Add to compose.yaml a simulated-saga service that runs a Node script (scripts/simulatedSaga.js, already written: it consumes order.created from its own queue and publishes stock.reserved and payment.confirmed) using the Orders image. Which variables does it need, what does it depend on, and why must it not publish ports?
Exercise 3. A colleague runs docker stop catalog-service and notices that it takes exactly 10 s and that "shutting down" never appears in the logs. Their Dockerfile ends with CMD npm start. Explain the cause, the fix, and what other symptom they would see in Kubernetes with terminationGracePeriodSeconds: 30.
Solutions
Solution 1. What changes: COPY --chown=node:node migrations ./migrations added next to src and scripts (the Job of 05-02 and Compose's orders-migrations run node scripts/migrate.js from this image and read migrations/*.sql); EXPOSE 3002; the HEALTHCHECK points to http://localhost:3002/health/live. Since Kubernetes ignores HEALTHCHECK (it uses livenessProbe), it can be kept for Compose or removed; the important thing is not to confuse it with the probe. Everything else (two stages, npm ci --omit=dev with the npmrc secret, --chown=node:node, USER node, CMD ["node","src/server.js"]) is identical: it is what the 04-01 template gives you ready-made.
Solution 2.
simulated-saga:
image: ghcr.io/techcorp/orders-service:local
command: ["node", "scripts/simulatedSaga.js"]
environment:
RABBITMQ_URL: amqp://rabbitmq:5672
LOG_LEVEL: debug
depends_on:
rabbitmq: { condition: service_healthy }
orders-service: { condition: service_started } # so the topology (techcorp.events exchange) is already declared
restart: unless-stoppedIt only needs RABBITMQ_URL (it talks exclusively through events, like Inventory and Payments in 03-04: "not exposed"). It publishes no ports because it exposes no HTTP and because nothing outside the Compose network should talk to it; if one day something did, it would go through the gateway. It depends on a healthy RabbitMQ and on Orders having started (it declares the topology when it connects; alternatively the script can declare it itself with the library's messaging/topology, and then the second dependency is unnecessary).
Solution 3. With CMD npm start, PID 1 is npm, which starts node src/server.js as a child. docker stop sends SIGTERM to PID 1 (npm), which does not forward it to Node; the process.on('SIGTERM') handler of 04-02 never runs, so there is no "shutting down" line and no server.close(). After 10 s Docker sends SIGKILL and everything dies abruptly: in-flight requests get a connection reset. Fix: CMD ["node", "src/server.js"] (exec form, Node as PID 1). In Kubernetes the symptom would be that every rolling update (05-04) takes 30 s per pod instead of 1-2 s and that, during those 30 s, the pod keeps receiving traffic until the readinessProbe fails, because /health/ready never switched to 503; with the exec form, the service marks itself not ready the instant SIGTERM arrives.
Conclusion
We have turned the Node processes of the previous modules into deployable units: one image per service, built with a multi-stage Dockerfile on node:20-alpine (npm ci --omit=dev with the npm token as a build secret, COPY --chown=node:node, USER node, EXPOSE, optional HEALTHCHECK on /health/live, CMD ["node","src/server.js"] in exec form so SIGTERM reaches the graceful shutdown of 04-02), with a .dockerignore and an instruction order that exploits the layer cache, tagged with a semantic version and sha-<commit> in ghcr.io/techcorp/ and never with latest. And we have brought the complete system up for the first time with Docker Compose (platform/local/compose.yaml): PostgreSQL 16, MongoDB 7 and RabbitMQ with healthchecks, one-off tasks for the seed and the migrations, catalog-service, orders-service, the Customers stub and the gateway on 8080, with the same environment variables of 04-03 and the same DNS names of 03-05, and on top of it the course's single E2E test. Compose is perfect for a laptop, but it does not restart a crashed container on another machine, does not spread replicas and does not manage zero-downtime deployments: that is what Kubernetes is for, and in the next lesson it will run these very images with the ConfigMaps and Secrets that 04-03 left named.
Microservices Course
Module 1: Introduction to Microservices
- Basic Concepts of Microservices
- Advantages and Disadvantages of Microservices
- Comparison with the Monolithic Architecture
- When to Adopt Microservices: Decision Criteria
- The Course Case Study: TechCorp's Online Store
Module 2: Microservice Design
- Microservice Design Principles
- Decomposing Monolithic Applications
- Defining Bounded Contexts
- Data Management: One Database per Service
- Distributed Consistency: Sagas, CQRS and Event Sourcing
Module 3: Communication between Microservices
- RESTful APIs
- Asynchronous Messaging
- Communication Protocols: gRPC, GraphQL
- API Gateway and Backend for Frontend
- Service Discovery and Load Balancing
- API Contracts and Versioning
Module 4: Implementing Microservices
- Choosing Technologies and Tools
- Building a Simple Microservice
- Configuration Management
- Hands-On Integration: Consuming APIs and Publishing Events
- Testing Microservices: Unit, Integration and Contract Tests
Module 5: Deployment and Orchestration
- Containers and Docker
- Orchestration with Kubernetes
- CI/CD for Microservices
- Deployment Strategies: Rolling, Blue-Green and Canary
- Service Mesh: Istio and Linkerd
Module 6: Monitoring and Maintenance
- Monitoring and Logging
- Distributed Tracing with OpenTelemetry
- Error Handling and Recovery
- Scalability and Performance
- SLOs, Alerts and Incident Management
Module 7: Security in Microservices
- Authentication and Authorization
- Communication Security
- Security Practices
- Container and Kubernetes Security
