The previous lesson ended with an uncomfortable question: aurora-db and aurora-cache have been running happily for minutes, but aurora-api died after two seconds with exit code 1. And if you try docker run -d ubuntu:24.04, you will see something even more baffling: the container stops all by itself, with no error, no log and nobody touching it. This is not a Docker bug; it is the direct consequence of the most important rule in the whole module: a container lives exactly as long as its PID 1 process lives. Not one second longer.
In this lesson you are going to walk through the seven states a container can be in and the transitions between them, you are going to understand what really happens when you type docker stop —the SIGTERM, grace period, SIGKILL sequence—, you are going to time with a stopwatch why the shell form of CMD always costs ten seconds, and you are going to learn to read exit codes: that 1, and also 125, 137 and 143, which stop being cryptic numbers and become diagnoses. And you will teach server.js to shut down properly, closing the PostgreSQL pool and the Redis client before it goes.
Contents
- The states of a container
- PID 1 and the fundamental rule
- Why
ubuntudies andnginxdoes not start,stopandrestartpauseandunpause: freezing without killingkilland sending specific signals- Graceful stop versus forced stop
- The shell form of
CMDand its ten seconds, measured - Exit codes and their diagnostic table
- Graceful shutdown for
aurora-api docker waitanddocker rename
- The states of a container
A container is not simply "on" or "off". Docker handles seven states:
stateDiagram-v2
[*] --> created: docker create<br/>(or the 1st half of docker run)
created --> running: docker start
running --> paused: docker pause
paused --> running: docker unpause
running --> exited: PID 1 ends<br/>docker stop / docker kill
running --> restarting: restart policy<br/>(lesson 03-07)
restarting --> running: successful retry
restarting --> exited: retries exhausted
exited --> running: docker start
exited --> removing: docker rm
paused --> exited: docker stop
removing --> [*]: container removed
running --> dead: daemon or file system<br/>failure
dead --> removing: docker rm -f
And the table you need to keep at hand:
| State | How you get there | Is there a process? | Does it use RAM/CPU? | What is preserved |
|---|---|---|---|---|
created |
docker create |
No | No | Empty writable layer and configuration |
running |
docker start, docker run, unpause |
Yes | Yes | Everything |
paused |
docker pause |
Yes, frozen | RAM yes, CPU no | Everything, including the process's memory |
restarting |
--restart policy after a failure |
Transiently no | Little | Everything |
exited |
PID 1 ends, docker stop, docker kill |
No | No | Writable layer, logs and configuration |
dead |
The daemon could not remove it properly | No | No | Leftovers; it can only be deleted |
removing |
docker rm in progress |
No | No | Nothing, it is transient |
The two states people confuse most are exited and removing, and the difference is enormous:
- An
exitedcontainer still exists. It occupies its name, keeps its writable layer with every file it wrote, holds on to its logs and its configuration, and you can start it again withdocker start. It is a sleeping container, not a deleted one. docker rmis what really removes it, and with it the writable layer and the logs are gone for good.
Check any container's state precisely:
docker inspect --format '{{.State.Status}}' aurora-db
docker inspect --format 'State: {{.State.Status}} | PID: {{.State.Pid}} | Started: {{.State.StartedAt}}' aurora-dbThat PID: 24817 is the process identifier on your machine: containers are not virtual machines, they are host processes isolated with namespaces, as you saw in lesson 01-01.
- PID 1 and the fundamental rule
Inside its process namespace, the container's main command sees itself as PID 1. Check it:
PostgreSQL is PID 1 inside aurora-db. And from that comes the rule:
The container exists as long as its PID 1 exists. When that process ends —well or badly, with an error or without one—, the container moves to
exited. It does not matter if there were a hundred child processes inside: they all die with it.
Being PID 1 has two implications you feel day to day:
| Implication | Practical consequence |
|---|---|
| PID 1 has no default signal handlers for SIGTERM in some languages | A process that ignores SIGTERM hangs until the SIGKILL arrives |
| PID 1 is responsible for reaping zombie processes | A PID 1 that does not do so accumulates zombies; solved with --init |
The --init option injects a minimal init (tini) as PID 1, which forwards signals and reaps zombies:
Now nginx is PID 7 and docker-init is PID 1. For aurora-api you do not need it: node is a proper PID 1 as long as you use the exec form, which you have been doing since lesson 02-04.
- Why
ubuntu dies and nginx does not
ubuntu dies and nginx does notThis is the demonstration that clears up 80% of "my container stops by itself" in one go:
docker run -d --name test-ubuntu ubuntu:24.04
docker run -d --name test-nginx nginx:alpine
sleep 2
docker ps -a --filter name=test- --format "table {{.Names}}\t{{.Status}}\t{{.Command}}"NAMES STATUS COMMAND
test-nginx Up 2 seconds "/docker-entrypoint.…"
test-ubuntu Exited (0) 2 seconds ago "/bin/bash"The difference is in the COMMAND column:
- The
ubuntu:24.04image hasCMD ["/bin/bash"]. Abashwith no terminal and no input has nothing to read: it reaches the end of its standard input and ends successfully, with code 0. The container did not fail; it did exactly what you asked, which was to run a shell that had nothing to do. - The
nginx:alpineimage runsnginx -g "daemon off;". Thatdaemon offis the key: it forces Nginx to stay in the foreground instead of turning into a background daemon. The process never ends, so neither does the container.
Make the Ubuntu one live by giving it something to do:
docker rm test-ubuntu
docker run -d --name test-ubuntu ubuntu:24.04 sleep 60
docker ps --filter name=test-ubuntu --format "{{.Names}}: {{.Status}}"Or give it a terminal, which is what you were doing in lesson 01-06 with -it:
docker rm -f test-ubuntu
docker run -d -it --name test-ubuntu ubuntu:24.04
docker ps --filter name=test-ubuntu --format "{{.Names}}: {{.Status}}"With -it, bash has an open terminal waiting for input and does not end. And from there comes the classic mistake:
A container is not a machine you switch on: it is a process you run. If you want it to last, give it a process that lasts. And the opposite mistake —"I set
daemon onin Nginx so it works like on my server"— turns the container into an instant suicide: Nginx goes to the background, the original process ends and Docker considers the container finished.
It is also the explanation of what happened to aurora-api: its PID 1, node, called process.exit(1) when it could not connect to Redis. The process ended, so the container ended.
start, stop and restart
start, stop and restart| Command | What it does | Same container? | Same PID? |
|---|---|---|---|
docker stop |
SIGTERM, wait, SIGKILL | — | — |
docker start |
Launches the original command again | Yes, same ID and same writable layer | No, new PID |
docker restart |
stop + start in a single step |
Yes | No, new PID |
What matters is understanding what survives a restart and what does not, because it is a constant source of surprises:
docker exec aurora-cache redis-cli SET book:favorite "Rayuela"
docker exec aurora-cache sh -c 'echo "temporary note" > /tmp/note.txt'
docker restart aurora-cache
sleep 2
docker exec aurora-cache cat /tmp/note.txt
docker exec aurora-cache redis-cli GET book:favoriteTwo opposite results in the same command:
/tmp/note.txtis still there: it is written in the container's writable layer, which survives stops and starts. It only disappears withdocker rm.- The Redis key is gone: it lived in the process's memory, and the process is new. Redis with no persistence configured loses everything when it restarts.
That distinction between "what is on disk inside the container" and "what is in the process's memory" is essential, and there is still a third category missing —"what is outside the container and survives even docker rm"—, which is the subject of lesson 03-06.
You can also operate on several containers at once:
A warning about docker start: it does not accept configuration changes. docker start -p 5433:5432 aurora-db does not exist. It starts the container again exactly as it was created, with its original ports, variables and mounts. That is the consequence of the create/start separation from the previous lesson.
And a useful option of start:
-a (--attach) starts the container and hooks your terminal to its output, as if you had launched it in the foreground. It is handy for watching the startup of a service you suspect is failing.
pause and unpause: freezing without killing
pause and unpause: freezing without killingdocker pause freezes all the container's processes using the cgroups freezer. The process does not even notice: it receives no signal, it simply stops getting CPU time. And from outside:
That command hangs indefinitely: the container cannot answer because it is not executing. Press Ctrl+C and unfreeze it:
pause |
stop |
|
|---|---|---|
| The process | Frozen, still exists | Terminated |
| The process's memory | Preserved intact | Lost |
| Signals sent | None | SIGTERM and, if necessary, SIGKILL |
| Open network connections | Kept, but unanswered (they eventually time out) | Closed |
| On resuming | Continues exactly where it was | Starts from scratch |
| RAM usage | Still occupied | Released |
Real uses of pause: freeing up CPU momentarily without losing the state of a long-running process, freezing a container while you take a consistent backup of its files, or temporarily halting one application while you diagnose another. It does not save memory: the RAM stays occupied.
kill and sending specific signals
kill and sending specific signalsdocker kill aurora-cache
docker ps -a --filter name=aurora-cache --format "{{.Names}}: {{.Status}}"
docker start aurora-cachedocker kill sends SIGKILL immediately, with no grace period. The process cannot catch it, ignore it or negotiate: the kernel terminates it on the spot. That is why the exit code is 137, which we will talk about in section 9.
But docker kill is good for much more than killing, because it sends any signal:
| Signal | Typical use in containers |
|---|---|
SIGTERM (15) |
A polite request to terminate. It is what docker stop sends |
SIGKILL (9) |
Immediate, uncatchable termination. What docker kill sends by default |
SIGHUP (1) |
Reload configuration without restarting (Nginx, HAProxy) |
SIGQUIT (3) |
Graceful shutdown in Nginx; thread dump in the JVM |
SIGUSR1/SIGUSR2 (10/12) |
Custom signals: rotating logs in Nginx, dumping the heap in Node |
SIGINT (2) |
What Ctrl+C sends |
A real example: reloading Nginx's configuration without dropping a single connection.
docker run -d --name reload-demo -p 8080:80 nginx:alpine
docker kill -s SIGHUP reload-demo
docker ps --filter name=reload-demo --format "{{.Names}}: {{.Status}}"Still Up: the signal did not kill it, Nginx read it as "reread your configuration". You will use it in lesson 03-05 when you set up aurora-web as a reverse proxy.
- Graceful stop versus forced stop
This is the central section of the lesson. docker stop does not kill the container outright: it runs a three-step sequence.
sequenceDiagram
participant U as You
participant D as dockerd
participant P as Container PID 1
U->>D: docker stop aurora-api
D->>P: 1. SIGTERM (or the image's STOPSIGNAL)
Note over P: Grace period: 10 s by default
alt The process ends in time
P-->>D: Closes connections, releases resources and exits
D-->>U: Container exited with its code
else The process does not respond
Note over D,P: The 10 s run out
D->>P: 2. SIGKILL (uncatchable)
P-->>D: Abrupt termination, code 137
D-->>U: Container exited (137)
end
The three steps are:
- Docker sends the image's STOPSIGNAL (
SIGTERMby default) to PID 1. Remember from lesson 02-04 that Nginx declaresSTOPSIGNAL SIGQUIT. - It waits the grace period: 10 seconds by default.
- If the process is still alive, it sends SIGKILL.
The grace period is adjusted with --time (or -t):
docker stop --time 30 aurora-db # waits up to 30 seconds
docker stop --time 0 aurora-cache # immediate SIGKILL, equivalent to docker killWhen is it worth raising it? When the process needs time to close properly: a database flushing its buffer to disk, a worker finishing the task it had in hand, an API waiting for in-flight requests to complete. For PostgreSQL, 30 seconds is a reasonable figure; with 10 you could force a recovery on the next startup.
docker stop |
docker kill |
|
|---|---|---|
| Initial signal | The image's STOPSIGNAL (SIGTERM) |
SIGKILL (or the one from -s) |
| Can it be caught? | Yes | No |
| Grace period | 10 s by default, adjustable with -t |
None |
| Buffered data | The process can flush it | Lost |
| Typical exit code | 0 or 143 | 137 |
| When to use it | Always, by default | Only if the process does not respond |
And the summary worth memorizing: stop asks, kill forces. A docker kill on PostgreSQL is the equivalent of unplugging the server.
- The shell form of
CMD and its ten seconds, measured
CMD and its ten seconds, measuredHere comes the bill for a decision that looked cosmetic in lesson 02-03. We are going to time it with a stopwatch.
Create a demonstration image that uses the shell form:
# ~/aurora-libros/api/Dockerfile.shell — ONLY for this demonstration
FROM auroralibros/aurora-api:1.1.0
ENTRYPOINT []
CMD echo "[startup] starting aurora-api" && node server.jsNow start both versions. Since the API still cannot find its dependencies, we will use Nginx so the comparison is clean and reproducible without depending on anything:
# Exec form: nginx is PID 1 and receives the signal directly
docker run -d --name stop-exec nginx:alpine
# Shell form: sh is PID 1 and nginx is its child
docker run -d --name stop-shell nginx:alpine \
sh -c 'echo "[startup] starting nginx" && nginx -g "daemon off;"'
docker exec stop-exec ps -o pid,comm | head -3
docker exec stop-shell ps -o pid,comm | head -3There is the difference, visible in a single line: in the first one nginx is PID 1; in the second, PID 1 is sh and nginx is one of its children. Time the stops:
Eighty times slower. And the exit codes tell the rest of the story:
Let's reconstruct what happened in the second case:
docker stopsends SIGQUIT (Nginx's STOPSIGNAL) to PID 1, which issh.shhas no handler for that signal and knows nothing about forwarding it to its children. It ignores it or terminates on its own, butnginxnever finds out.- Docker waits its full ten seconds.
- SIGKILL. The whole container dies at once, with connections cut mid-request. Code 137.
An honest nuance almost nobody mentions: some shells optimize the simplest case. If the CMD were exactly CMD node server.js, many shells (including dash and BusyBox's ash) replace their own process with node via exec, and the problem does not appear. But all it takes is two commands chained with &&, a redirection or a pipe —as in the example, and as in 90% of real-world CMDs— for the shell to have to stay as PID 1 and for the failure to show up. It is not worth gambling your shutdown behavior on a shell optimization:
Always use the exec form:
CMD ["node", "server.js"]. If you need shell logic, write it in adocker-entrypoint.shthat ends withexec "$@", as you did in lesson 02-04.
Clean up the demonstration:
- Exit codes and their diagnostic table
When a container ends, it leaves a number behind. That number is the first clue in any investigation.
docker ps -a --filter name=aurora --format "table {{.Names}}\t{{.Status}}"
docker inspect --format '{{.State.ExitCode}}' aurora-apiThe table that solves most cases:
| Code | Meaning | Usual cause in practice |
|---|---|---|
| 0 | Successful termination | The process did its job and exited. Also bash with no input |
| 1 | Generic application error | Uncaught exception, process.exit(1), invalid configuration |
| 125 | Failure of docker run itself |
Misspelled option, non-existent --env-file, duplicate name. The container was not even created |
| 126 | The command exists but could not be executed | Missing execute bit (chmod +x), or an attempt to execute a directory |
| 127 | Command not found | Typo in the CMD, or a binary missing from a minimal image |
| 137 | 128 + 9 → SIGKILL | docker kill, docker stop timing out, or the OOM killer (lesson 03-07) |
| 139 | 128 + 11 → SIGSEGV | Segmentation fault: a bug in native code or a binary incompatible with the architecture |
| 143 | 128 + 15 → SIGTERM | The process ended on SIGTERM with no handler of its own. This is a normal stop |
The rule that explains half the table: if the code is greater than 128, the process died from a signal, and the signal number is code − 128. 137 − 128 = 9 (SIGKILL); 143 − 128 = 15 (SIGTERM); 139 − 128 = 11 (SIGSEGV).
And the most important distinction for debugging, which separates two worlds:
| Range | Who failed | Where to look |
|---|---|---|
| 125, 126, 127 | Docker or the command's startup | In your command line and in the Dockerfile |
| 1, 2, … , 124 | Your application | In docker logs |
| >128 | An external or kernel signal | In docker inspect (OOMKilled, Error) |
Let's trigger them to see them for real:
docker run --name exit-0 alpine:3.20 true
docker run --name exit-1 alpine:3.20 sh -c 'exit 1'
docker run --name exit-127 alpine:3.20 command-that-does-not-exist
docker run --made-up-option alpine:3.20
docker run --name exit-126 alpine:3.20 /etcdocker: Error response from daemon: failed to create task for container:
failed to create shim task: OCI runtime create failed: exec: "command-that-does-not-exist": executable file not found in $PATH
unknown flag: --made-up-option
docker: Error response from daemon: failed to create task for container:
failed to create shim task: OCI runtime create failed: exec: "/etc": permission denieddocker ps -a --filter name=exit- --format "table {{.Names}}\t{{.Status}}"
echo "Exit status of the invalid docker run: $?"NAMES STATUS
exit-126 Created
exit-127 Created
exit-1 Exited (1) 8 seconds ago
exit-0 Exited (0) 12 seconds ago
125Notice the revealing detail: exit-127 and exit-126 stayed in the Created state, not Exited. The container did get created but never got to start, because the executable did not exist or could not be executed. exit-1, in contrast, did start and its application decided to exit with an error. That is the difference between "it could not begin" and "it began and failed", and it saves you looking in the wrong place.
Applied to Aurora Libros, the Exited (1) from the previous lesson is now diagnosed without ambiguity: the application started and failed on its own, it was not Docker and not a signal. The logs confirm the reason (getaddrinfo ENOTFOUND aurora-cache) and the process.exit(1) in server.js explains the exact number.
- Graceful shutdown for
aurora-api
aurora-apiThe time has come to fix two things at once in server.js: that the process should not die because an optional dependency is unavailable, and that when it is asked to terminate it should do so closing what it has open.
Replace the final block of ~/aurora-libros/api/server.js (the one that starts at // --- Startup ---) with this:
// --- Startup ---
let server;
async function start() {
// The cache is an OPTIONAL dependency: if Redis does not respond, the API
// must keep serving the catalog from PostgreSQL, not die at startup.
cache.connect().catch((err) =>
console.error('[cache] not available at startup:', err.message)
);
server = app.listen(PORT, '0.0.0.0', () => {
console.log(`[aurora-api] listening on port ${PORT}`);
console.log(`[aurora-api] database: ${DB_HOST}:${DB_PORT}/${DB_NAME}`);
console.log(`[aurora-api] cache: ${REDIS_HOST}:${REDIS_PORT}`);
});
}
// --- Graceful shutdown ---
let shuttingDown = false;
async function shutdown(signal) {
if (shuttingDown) return; // A second Ctrl+C must not re-enter here
shuttingDown = true;
console.log(`[aurora-api] received ${signal}, shutting down gracefully...`);
// Safety net: leave under our own steam BEFORE the SIGKILL arrives.
// 8 seconds < the 10 of docker stop's grace period.
const forceExit = setTimeout(() => {
console.error('[aurora-api] shutdown stalled, forcing exit');
process.exit(1);
}, 8000);
forceExit.unref();
server.close(async () => {
try {
await pool.end();
console.log('[aurora-api] PostgreSQL pool closed');
if (cache.isOpen) {
await cache.quit();
console.log('[aurora-api] Redis client closed');
}
} catch (err) {
console.error('[aurora-api] error during shutdown:', err.message);
}
console.log('[aurora-api] shut down cleanly');
process.exit(0);
});
}
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
start().catch((err) => {
console.error('[aurora-api] failed to start:', err.message);
process.exit(1);
});Let's go over the decisions one by one, because each one answers something you have learned in this lesson:
| Line | Why it is there |
|---|---|
cache.connect().catch(...) without await |
Startup does not block on an optional dependency. The container stays running and you can study it, instead of dying with code 1 |
server = app.listen(...) saved in a variable |
Without the reference to the server you cannot call .close() |
if (shuttingDown) return; |
Stops two signals in a row from launching two simultaneous shutdowns |
setTimeout(..., 8000) |
8 < 10: if something stalls, we exit ourselves before the SIGKILL and with a code of our own, not with a 137 |
.unref() |
Stops that timer from keeping the process alive if the shutdown goes well |
server.close(callback) |
Stops accepting new connections and waits for the in-flight ones to finish |
await pool.end() |
Closes the PostgreSQL connections. Without this, the database keeps them open until they time out |
if (cache.isOpen) |
quit() on a client that never connected would throw an exception |
process.exit(0) |
An explicit, successful exit: code 0, not 143 |
SIGINT as well as SIGTERM |
So that Ctrl+C in the foreground behaves just like docker stop |
Build the new version. Since you are changing behavior without breaking the API, bump the MINOR number according to what you decided in lesson 02-06:
docker build -t auroralibros/aurora-api:1.2.0 \
-t auroralibros/aurora-api:1.2 \
--build-arg VERSION=1.2.0 \
~/aurora-libros/apiAnd now the timed check. Compare the old version with the new one:
# Version 1.1.0: it died at startup. With 1.2.0 the container stays alive.
docker run -d --name api-shutdown --env-file ~/aurora-libros/aurora.env \
auroralibros/aurora-api:1.2.0
sleep 3
docker ps --filter name=api-shutdown --format "{{.Names}}: {{.Status}}"The container survives. It no longer dies for not finding Redis; it simply logs it and carries on. Now stop it, measuring the time:
time docker stop api-shutdown
docker logs --tail 6 api-shutdown
docker inspect --format 'Exit code: {{.State.ExitCode}}' api-shutdownapi-shutdown
real 0m0.187s
[cache] not available at startup: getaddrinfo ENOTFOUND aurora-cache
[aurora-api] listening on port 3000
[aurora-api] database: aurora-db:5432/aurora_books
[aurora-api] received SIGTERM, shutting down gracefully...
[aurora-api] PostgreSQL pool closed
[aurora-api] shut down cleanly
Exit code: 0The three things we were after, confirmed in a single output:
- 0.187 seconds, compared with the 10.271 of the shell form. The signal reached
nodebecause it is PID 1 and the exec form lets it through. - The logs tell the story of the shutdown: the PostgreSQL pool was closed explicitly before exiting. The Redis client does not appear because it never got opened (
cache.isOpenwas false), exactly as the code anticipated. - Exit code 0, not 143. The difference between "the process was terminated by a signal" and "the process decided to terminate properly because it was asked nicely".
In an orchestrator, this difference is what separates a deployment with no errors from one with requests cut mid-flight and orphaned connections in the database.
docker wait and docker rename
docker wait and docker renameTwo small commands that solve specific problems.
docker wait
It blocks until the container ends and prints its exit code:
docker run -d --name slow-task alpine:3.20 sh -c 'sleep 5; exit 3'
docker wait slow-task
docker rm slow-taskIt sits there for five seconds and then prints 3. It is the piece that lets you chain things in a script: "wait for the database migration to finish and continue only if it went well".
docker run -d --name migration alpine:3.20 sh -c 'echo "migrating..."; sleep 3'
if [ "$(docker wait migration)" -eq 0 ]; then
echo "Migration succeeded, starting the API"
else
echo "Migration failed, aborting the deployment"
fi
docker rm migrationdocker rename
It changes a container's name on the fly, without stopping it or recreating it:
docker rename aurora-cache aurora-cache-old
docker ps --format "{{.Names}}: {{.Status}}" --filter name=aurora
docker rename aurora-cache-old aurora-cacheIt is more useful than it looks: in a zero-downtime deployment, you rename the old container to aurora-api-old, start the new one with the good name and delete the old one once you confirm everything is fine. And a warning that connects with lesson 03-05: if another container was resolving it over DNS by its previous name, it will stop finding it the moment you rename it.
Common Mistakes and Tips
- Believing an
Exitedcontainer is deleted. It still exists, it occupies its name and it keeps its writable layer and logs. You delete it withdocker rm, not withdocker stop. - Putting a service into daemon mode inside the container.
nginxwithoutdaemon off,httpd -k start,postgreswithpg_ctl start: the main process ends immediately and the container stops. Services in containers run in the foreground, always. - Using
docker killout of habit. A SIGKILL on PostgreSQL forces a recovery on startup and can lose whatever was buffered.docker stopfirst;killonly if the other one does not respond. - Not catching SIGTERM in the application. The process dies outright with its connections open. With a handler, you close the pool and exit with code 0.
- Setting the safety timer to 10 seconds or more. It must be less than the grace period; otherwise it never fires, because the SIGKILL arrives first.
- Always reading 137 as "I killed it". The OOM killer also produces it when the container goes over its memory. You tell them apart by looking at
.State.OOMKilledindocker inspect(lesson 03-07). - Confusing 125 with 1. 125 is Docker's: the container does not even exist, so there are no logs to look at. Review the command line.
- Tip: raise the
--timeon databases.docker stop --time 30 aurora-dbgives PostgreSQL room to do its checkpoint and close cleanly. - Tip: use
docker start -awhen a container starts and dies right away: you see the startup live without having to chase it withdocker logs.
Exercises
Exercise 1: walk through the seven states
Using alpine:3.20 and a container called full-lifecycle that runs sleep 600, take the container through this sequence and note after each step the result of docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}': create without starting → start → pause → resume → stop gracefully → start again → kill → remove. Answer: what is the exit code after docker stop and why? And after docker kill? In which state is the process's memory still occupied?
Exercise 2: demonstrate data loss according to the type of stop
With a redis:7-alpine container called cache-test, write the key catalog:version with value 1 and also a file /tmp/marker.txt with the same content. Then:
- Run
docker restartand check what survives of the two things. - Run
docker stopfollowed bydocker startand check the same. - Run
docker rm -fand create a container again with the same name and image. Check the same.
Explain the results in terms of process memory, writable layer and container.
Exercise 3: measure the cost of not handling signals
Prepare three containers from node:22-alpine that run these three programs and time docker stop on each one, also noting the exit code:
- A:
node -e "setInterval(()=>{},1000)"— a process that does nothing and catches no signals. - B:
sh -c 'echo start && node -e "setInterval(()=>{},1000)"'— the same one, but wrapped in a shell with&&. - C:
node -e "process.on('SIGTERM',()=>{console.log('bye');process.exit(0)});setInterval(()=>{},1000)"— with a SIGTERM handler.
Explain the three times and the three codes, and say which of the three corresponds to auroralibros/aurora-api:1.1.0 and which to 1.2.0.
Solutions
Solution to exercise 1
docker create --name full-lifecycle alpine:3.20 sleep 600
docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycledocker start full-lifecycle && docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
docker pause full-lifecycle && docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
docker unpause full-lifecycle && docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycletime docker stop full-lifecycle
docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
docker start full-lifecycle
docker kill full-lifecycle
docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
docker rm full-lifecyclefull-lifecycle
real 0m0.156s
exited / exit code 143
full-lifecycle
full-lifecycle
exited / exit code 137The three answers:
- After
docker stop: code 143, that is, 128 + 15 = SIGTERM.sleepinstalls no signal handler, so SIGTERM's default action applies: terminate immediately. That is whyrealis barely 0.156 s and no SIGKILL was needed. It is a clean stop from Docker's point of view, even though the process ran no shutdown logic of its own: exactly the case ofaurora-api:1.1.0. - After
docker kill: code 137 (128 + 9 = SIGKILL), instantly and with no grace period. Compare the two numbers: 143 means "it was asked to terminate and it terminated"; 137 means "it was terminated without being asked". A 137 accompanied by a ten-second wait would additionally point to a process that ignored the request. - The memory is still occupied in
paused. The processes are frozen by the cgroups freezer, but they still exist with all their memory reserved. Inexitedthere is no process and the memory is released; what persists is the writable layer on disk.
Solution to exercise 2
docker run -d --name cache-test redis:7-alpine
docker exec cache-test redis-cli SET catalog:version 1
docker exec cache-test sh -c 'echo 1 > /tmp/marker.txt'1. After docker restart:
docker restart cache-test && sleep 2
docker exec cache-test redis-cli GET catalog:version
docker exec cache-test cat /tmp/marker.txt2. After docker stop + docker start:
docker stop cache-test && docker start cache-test && sleep 2
docker exec cache-test redis-cli GET catalog:version
docker exec cache-test cat /tmp/marker.txt3. After docker rm -f and recreating:
docker rm -f cache-test
docker run -d --name cache-test redis:7-alpine && sleep 2
docker exec cache-test redis-cli GET catalog:version
docker exec cache-test cat /tmp/marker.txt
docker rm -f cache-testThe explanation on three levels:
| Level | What survives restart/stop+start |
What survives rm |
|---|---|---|
| Process memory (the Redis key) | No. The process is new, its memory starts empty | No |
Writable layer (/tmp/marker.txt) |
Yes. It is disk, and the container is the same one | No. It is destroyed with the container |
| Volume (lesson 03-06) | Yes | Yes, and that is the key |
Cases 1 and 2 are equivalent: restart is literally stop + start. What is interesting is that no level of this table survives a docker rm, and that means that as things stand today, if you delete aurora-db, you lose the Aurora Libros catalog. That is exactly the demonstration that opens lesson 03-06.
Solution to exercise 3
docker run -d --name signal-a node:22-alpine node -e "setInterval(()=>{},1000)"
docker run -d --name signal-b node:22-alpine sh -c 'echo start && node -e "setInterval(()=>{},1000)"'
docker run -d --name signal-c node:22-alpine node -e "process.on('SIGTERM',()=>{console.log('bye');process.exit(0)});setInterval(()=>{},1000)"
for c in signal-a signal-b signal-c; do
echo "--- $c ---"
{ time docker stop "$c" ; } 2>&1 | grep real
docker inspect --format 'exit code {{.State.ExitCode}}' "$c"
done--- signal-a ---
real 0m0.142s
exit code 143
--- signal-b ---
real 0m10.238s
exit code 137
--- signal-c ---
real 0m0.118s
exit code 0| Case | Time | Code | Why |
|---|---|---|---|
| A | 0.14 s | 143 | node is PID 1 and receives SIGTERM. With no handler of its own, the default action applies: terminate. Fast, but abrupt: it closed nothing |
| B | 10.24 s | 137 | PID 1 is sh, which does not forward the signal. Ten seconds of waiting and a SIGKILL. The worst of all worlds |
| C | 0.12 s | 0 | node catches SIGTERM, runs its shutdown logic and exits of its own accord with code 0 |
How it maps to Aurora Libros:
auroralibros/aurora-api:1.1.0is case A. Exec form (ENTRYPOINT ["node"]), so the signal arrives, but with no handler: it stopped in 0.3 seconds with code 143 and left the PostgreSQL pool open. Those were the "0.3 seconds" you were celebrating in lesson 02-04, correct but incomplete.auroralibros/aurora-api:1.2.0is case C. The same speed, but closing the pool and the Redis client and exiting with code 0.- Case B must never exist in your images. It is what you would get with
CMD node server.js && echo doneor with an entrypoint that lacksexec "$@".
Conclusion
You now know why some containers live and others die instantly, and there is no mystery to it: a container lasts exactly as long as its PID 1 process lasts. ubuntu:24.04 runs a bash with no input that ends cleanly with code 0; nginx:alpine runs nginx -g "daemon off;", which never ends. And aurora-api was dying because its own code called process.exit(1). You know the seven states and their transitions, and the crucial difference between exited —asleep, with its writable layer, its logs and its name intact— and truly deleted with docker rm.
You have mastered the lifecycle verbs: start, stop, restart, pause/unpause with its cgroups freezer that freezes without sending a single signal and without releasing memory, and kill as a sender of arbitrary signals, including that SIGHUP that reloads Nginx without dropping a single connection. And you understand what really happens after a docker stop: SIGTERM, ten seconds of grace adjustable with --time, and SIGKILL. You have measured it: 0.128 seconds with the exec form versus 10.271 seconds and code 137 with a shell in between that does not forward signals.
Exit codes have stopped being noise. You know that above 128 there is a signal hiding (137 = SIGKILL, 143 = SIGTERM, 139 = SIGSEGV), that 125, 126 and 127 point at Docker or at the command's startup and not at your application, and that a container in the Created state after a docker run means the executable did not even exist. And you have taken the theory into the project: auroralibros/aurora-api:1.2.0 catches SIGTERM and SIGINT, closes the PostgreSQL pool and the Redis client, protects itself with an 8-second timer deliberately shorter than the 10 of the grace period, and exits with code 0 in 0.187 seconds. On top of that, it no longer kills itself because Redis is not there: it stays alive and logs it, which is what a serious service should do.
With aurora-db and aurora-cache running and aurora-api finally able to stay on its feet, you are starting to have something resembling a fleet. And a fleet needs to be watched and kept in order. In the next lesson, Managing Containers, you will squeeze docker ps for everything it has: every column explained, the real size of the writable layer with -s, the filters by state, name, image, label and health, and the --format templates you will use to build a small control panel for the four Aurora Libros services. You will learn to compose commands with -q to operate on dozens of containers at once without wrecking anything, to copy files between the host and a container with docker cp, to see with docker diff what a container has changed with respect to its image, and why docker commit must never be used to build real images.
Docker: From Beginner to Advanced
Module 1: Introduction to Docker
- What Is Docker?
- Installing Docker
- Docker Architecture
- Basic Docker Commands
- Understanding Docker Images
- Creating Your First Docker Container
- The Course Project: The Aurora Libros Platform
Module 2: Working with Docker Images
- Docker Hub and Repositories
- Building Docker Images
- Dockerfile Basics
- Advanced Dockerfile Instructions
- Managing Docker Images
- Tagging and Publishing Images
Module 3: Docker Containers
- Running Containers
- Container Lifecycle
- Managing Containers
- Inspecting and Debugging Containers
- Docker Networking
- Data Persistence with Volumes
- Resource Limits and Restart Policies
Module 4: Docker Compose
- Introduction to Docker Compose
- Defining Services in Docker Compose
- Docker Compose Commands
- Multi-Container Applications
- Environment Variables in Docker Compose
- Profiles, Overrides and Multiple Environments
- Local Development with Docker Compose
Module 5: Advanced Docker Concepts
- Docker Networking Deep Dive
- Docker Storage Options
- Docker Security Best Practices
- Optimizing Docker Images
- Advanced Builds with BuildKit and Buildx
- Logging and Monitoring in Docker
- The Runtime Inside: Namespaces, Cgroups and Layers
Module 6: Docker in Production
- Preparing an Image for Production
- CI/CD with Docker
- Orchestrating Containers with Docker Swarm
- Introduction to Kubernetes
- Deploying Docker Containers in Kubernetes
- Scaling and Load Balancing
- Deployment Strategies and Rollback
