The previous lesson ended with an uncomfortable question: aurora-db and aurora-cache have been running happily for minutes, but aurora-api died after two seconds with exit code 1. And if you try docker run -d ubuntu:24.04, you will see something even more baffling: the container stops all by itself, with no error, no log and nobody touching it. This is not a Docker bug; it is the direct consequence of the most important rule in the whole module: a container lives exactly as long as its PID 1 process lives. Not one second longer.

In this lesson you are going to walk through the seven states a container can be in and the transitions between them, you are going to understand what really happens when you type docker stop —the SIGTERM, grace period, SIGKILL sequence—, you are going to time with a stopwatch why the shell form of CMD always costs ten seconds, and you are going to learn to read exit codes: that 1, and also 125, 137 and 143, which stop being cryptic numbers and become diagnoses. And you will teach server.js to shut down properly, closing the PostgreSQL pool and the Redis client before it goes.

Contents

  1. The states of a container
  2. PID 1 and the fundamental rule
  3. Why ubuntu dies and nginx does not
  4. start, stop and restart
  5. pause and unpause: freezing without killing
  6. kill and sending specific signals
  7. Graceful stop versus forced stop
  8. The shell form of CMD and its ten seconds, measured
  9. Exit codes and their diagnostic table
  10. Graceful shutdown for aurora-api
  11. docker wait and docker rename

  1. The states of a container

A container is not simply "on" or "off". Docker handles seven states:

stateDiagram-v2
    [*] --> created: docker create<br/>(or the 1st half of docker run)
    created --> running: docker start
    running --> paused: docker pause
    paused --> running: docker unpause
    running --> exited: PID 1 ends<br/>docker stop / docker kill
    running --> restarting: restart policy<br/>(lesson 03-07)
    restarting --> running: successful retry
    restarting --> exited: retries exhausted
    exited --> running: docker start
    exited --> removing: docker rm
    paused --> exited: docker stop
    removing --> [*]: container removed
    running --> dead: daemon or file system<br/>failure
    dead --> removing: docker rm -f

And the table you need to keep at hand:

State How you get there Is there a process? Does it use RAM/CPU? What is preserved
created docker create No No Empty writable layer and configuration
running docker start, docker run, unpause Yes Yes Everything
paused docker pause Yes, frozen RAM yes, CPU no Everything, including the process's memory
restarting --restart policy after a failure Transiently no Little Everything
exited PID 1 ends, docker stop, docker kill No No Writable layer, logs and configuration
dead The daemon could not remove it properly No No Leftovers; it can only be deleted
removing docker rm in progress No No Nothing, it is transient

The two states people confuse most are exited and removing, and the difference is enormous:

  • An exited container still exists. It occupies its name, keeps its writable layer with every file it wrote, holds on to its logs and its configuration, and you can start it again with docker start. It is a sleeping container, not a deleted one.
  • docker rm is what really removes it, and with it the writable layer and the logs are gone for good.

Check any container's state precisely:

docker inspect --format '{{.State.Status}}' aurora-db
docker inspect --format 'State: {{.State.Status}} | PID: {{.State.Pid}} | Started: {{.State.StartedAt}}' aurora-db
running
State: running | PID: 24817 | Started: 2026-08-04T19:28:41.113927Z

That PID: 24817 is the process identifier on your machine: containers are not virtual machines, they are host processes isolated with namespaces, as you saw in lesson 01-01.

  1. PID 1 and the fundamental rule

Inside its process namespace, the container's main command sees itself as PID 1. Check it:

docker exec aurora-db ps -o pid,comm
PID   COMMAND
    1 postgres
   67 postgres
   68 postgres
  ...

PostgreSQL is PID 1 inside aurora-db. And from that comes the rule:

The container exists as long as its PID 1 exists. When that process ends —well or badly, with an error or without one—, the container moves to exited. It does not matter if there were a hundred child processes inside: they all die with it.

Being PID 1 has two implications you feel day to day:

Implication Practical consequence
PID 1 has no default signal handlers for SIGTERM in some languages A process that ignores SIGTERM hangs until the SIGKILL arrives
PID 1 is responsible for reaping zombie processes A PID 1 that does not do so accumulates zombies; solved with --init

The --init option injects a minimal init (tini) as PID 1, which forwards signals and reaps zombies:

docker run -d --init --name with-init nginx:alpine
docker exec with-init ps -o pid,comm | head -3
PID   COMMAND
    1 /sbin/docker-init
    7 nginx

Now nginx is PID 7 and docker-init is PID 1. For aurora-api you do not need it: node is a proper PID 1 as long as you use the exec form, which you have been doing since lesson 02-04.

docker rm -f with-init

  1. Why ubuntu dies and nginx does not

This is the demonstration that clears up 80% of "my container stops by itself" in one go:

docker run -d --name test-ubuntu ubuntu:24.04
docker run -d --name test-nginx nginx:alpine
sleep 2
docker ps -a --filter name=test- --format "table {{.Names}}\t{{.Status}}\t{{.Command}}"
NAMES          STATUS                     COMMAND
test-nginx     Up 2 seconds               "/docker-entrypoint.…"
test-ubuntu    Exited (0) 2 seconds ago   "/bin/bash"

The difference is in the COMMAND column:

  • The ubuntu:24.04 image has CMD ["/bin/bash"]. A bash with no terminal and no input has nothing to read: it reaches the end of its standard input and ends successfully, with code 0. The container did not fail; it did exactly what you asked, which was to run a shell that had nothing to do.
  • The nginx:alpine image runs nginx -g "daemon off;". That daemon off is the key: it forces Nginx to stay in the foreground instead of turning into a background daemon. The process never ends, so neither does the container.

Make the Ubuntu one live by giving it something to do:

docker rm test-ubuntu
docker run -d --name test-ubuntu ubuntu:24.04 sleep 60
docker ps --filter name=test-ubuntu --format "{{.Names}}: {{.Status}}"
test-ubuntu: Up 3 seconds

Or give it a terminal, which is what you were doing in lesson 01-06 with -it:

docker rm -f test-ubuntu
docker run -d -it --name test-ubuntu ubuntu:24.04
docker ps --filter name=test-ubuntu --format "{{.Names}}: {{.Status}}"
test-ubuntu: Up 2 seconds

With -it, bash has an open terminal waiting for input and does not end. And from there comes the classic mistake:

A container is not a machine you switch on: it is a process you run. If you want it to last, give it a process that lasts. And the opposite mistake —"I set daemon on in Nginx so it works like on my server"— turns the container into an instant suicide: Nginx goes to the background, the original process ends and Docker considers the container finished.

It is also the explanation of what happened to aurora-api: its PID 1, node, called process.exit(1) when it could not connect to Redis. The process ended, so the container ended.

docker rm -f test-ubuntu test-nginx

  1. start, stop and restart

docker stop aurora-cache
docker start aurora-cache
docker restart aurora-cache
aurora-cache
aurora-cache
aurora-cache
Command What it does Same container? Same PID?
docker stop SIGTERM, wait, SIGKILL — —
docker start Launches the original command again Yes, same ID and same writable layer No, new PID
docker restart stop + start in a single step Yes No, new PID

What matters is understanding what survives a restart and what does not, because it is a constant source of surprises:

docker exec aurora-cache redis-cli SET book:favorite "Rayuela"
docker exec aurora-cache sh -c 'echo "temporary note" > /tmp/note.txt'
docker restart aurora-cache
sleep 2
docker exec aurora-cache cat /tmp/note.txt
docker exec aurora-cache redis-cli GET book:favorite
OK
aurora-cache
temporary note
(nil)

Two opposite results in the same command:

  • /tmp/note.txt is still there: it is written in the container's writable layer, which survives stops and starts. It only disappears with docker rm.
  • The Redis key is gone: it lived in the process's memory, and the process is new. Redis with no persistence configured loses everything when it restarts.

That distinction between "what is on disk inside the container" and "what is in the process's memory" is essential, and there is still a third category missing —"what is outside the container and survives even docker rm"—, which is the subject of lesson 03-06.

You can also operate on several containers at once:

docker stop aurora-db aurora-cache
docker start aurora-db aurora-cache

A warning about docker start: it does not accept configuration changes. docker start -p 5433:5432 aurora-db does not exist. It starts the container again exactly as it was created, with its original ports, variables and mounts. That is the consequence of the create/start separation from the previous lesson.

And a useful option of start:

docker start -a aurora-cache

-a (--attach) starts the container and hooks your terminal to its output, as if you had launched it in the foreground. It is handy for watching the startup of a service you suspect is failing.

  1. pause and unpause: freezing without killing

docker pause aurora-cache
docker ps --filter name=aurora-cache --format "{{.Names}}: {{.Status}}"
aurora-cache
aurora-cache: Up 8 minutes (Paused)

docker pause freezes all the container's processes using the cgroups freezer. The process does not even notice: it receives no signal, it simply stops getting CPU time. And from outside:

docker exec aurora-cache redis-cli PING

That command hangs indefinitely: the container cannot answer because it is not executing. Press Ctrl+C and unfreeze it:

docker unpause aurora-cache
docker exec aurora-cache redis-cli PING
aurora-cache
PONG
pause stop
The process Frozen, still exists Terminated
The process's memory Preserved intact Lost
Signals sent None SIGTERM and, if necessary, SIGKILL
Open network connections Kept, but unanswered (they eventually time out) Closed
On resuming Continues exactly where it was Starts from scratch
RAM usage Still occupied Released

Real uses of pause: freeing up CPU momentarily without losing the state of a long-running process, freezing a container while you take a consistent backup of its files, or temporarily halting one application while you diagnose another. It does not save memory: the RAM stays occupied.

  1. kill and sending specific signals

docker kill aurora-cache
docker ps -a --filter name=aurora-cache --format "{{.Names}}: {{.Status}}"
docker start aurora-cache
aurora-cache
aurora-cache: Exited (137) 3 seconds ago
aurora-cache

docker kill sends SIGKILL immediately, with no grace period. The process cannot catch it, ignore it or negotiate: the kernel terminates it on the spot. That is why the exit code is 137, which we will talk about in section 9.

But docker kill is good for much more than killing, because it sends any signal:

docker kill -s SIGHUP <container>
Signal Typical use in containers
SIGTERM (15) A polite request to terminate. It is what docker stop sends
SIGKILL (9) Immediate, uncatchable termination. What docker kill sends by default
SIGHUP (1) Reload configuration without restarting (Nginx, HAProxy)
SIGQUIT (3) Graceful shutdown in Nginx; thread dump in the JVM
SIGUSR1/SIGUSR2 (10/12) Custom signals: rotating logs in Nginx, dumping the heap in Node
SIGINT (2) What Ctrl+C sends

A real example: reloading Nginx's configuration without dropping a single connection.

docker run -d --name reload-demo -p 8080:80 nginx:alpine
docker kill -s SIGHUP reload-demo
docker ps --filter name=reload-demo --format "{{.Names}}: {{.Status}}"
reload-demo: Up 20 seconds

Still Up: the signal did not kill it, Nginx read it as "reread your configuration". You will use it in lesson 03-05 when you set up aurora-web as a reverse proxy.

docker rm -f reload-demo

  1. Graceful stop versus forced stop

This is the central section of the lesson. docker stop does not kill the container outright: it runs a three-step sequence.

sequenceDiagram
    participant U as You
    participant D as dockerd
    participant P as Container PID 1
    U->>D: docker stop aurora-api
    D->>P: 1. SIGTERM (or the image's STOPSIGNAL)
    Note over P: Grace period: 10 s by default
    alt The process ends in time
        P-->>D: Closes connections, releases resources and exits
        D-->>U: Container exited with its code
    else The process does not respond
        Note over D,P: The 10 s run out
        D->>P: 2. SIGKILL (uncatchable)
        P-->>D: Abrupt termination, code 137
        D-->>U: Container exited (137)
    end

The three steps are:

  1. Docker sends the image's STOPSIGNAL (SIGTERM by default) to PID 1. Remember from lesson 02-04 that Nginx declares STOPSIGNAL SIGQUIT.
  2. It waits the grace period: 10 seconds by default.
  3. If the process is still alive, it sends SIGKILL.

The grace period is adjusted with --time (or -t):

docker stop --time 30 aurora-db     # waits up to 30 seconds
docker stop --time 0 aurora-cache   # immediate SIGKILL, equivalent to docker kill

When is it worth raising it? When the process needs time to close properly: a database flushing its buffer to disk, a worker finishing the task it had in hand, an API waiting for in-flight requests to complete. For PostgreSQL, 30 seconds is a reasonable figure; with 10 you could force a recovery on the next startup.

docker stop docker kill
Initial signal The image's STOPSIGNAL (SIGTERM) SIGKILL (or the one from -s)
Can it be caught? Yes No
Grace period 10 s by default, adjustable with -t None
Buffered data The process can flush it Lost
Typical exit code 0 or 143 137
When to use it Always, by default Only if the process does not respond

And the summary worth memorizing: stop asks, kill forces. A docker kill on PostgreSQL is the equivalent of unplugging the server.

  1. The shell form of CMD and its ten seconds, measured

Here comes the bill for a decision that looked cosmetic in lesson 02-03. We are going to time it with a stopwatch.

Create a demonstration image that uses the shell form:

# ~/aurora-libros/api/Dockerfile.shell — ONLY for this demonstration
FROM auroralibros/aurora-api:1.1.0
ENTRYPOINT []
CMD echo "[startup] starting aurora-api" && node server.js
docker build -f ~/aurora-libros/api/Dockerfile.shell -t aurora-api:shell-form ~/aurora-libros/api

Now start both versions. Since the API still cannot find its dependencies, we will use Nginx so the comparison is clean and reproducible without depending on anything:

# Exec form: nginx is PID 1 and receives the signal directly
docker run -d --name stop-exec nginx:alpine

# Shell form: sh is PID 1 and nginx is its child
docker run -d --name stop-shell nginx:alpine \
  sh -c 'echo "[startup] starting nginx" && nginx -g "daemon off;"'

docker exec stop-exec ps -o pid,comm | head -3
docker exec stop-shell ps -o pid,comm | head -3
PID   COMMAND
    1 nginx
   30 nginx

PID   COMMAND
    1 sh
    7 nginx
   14 nginx

There is the difference, visible in a single line: in the first one nginx is PID 1; in the second, PID 1 is sh and nginx is one of its children. Time the stops:

time docker stop stop-exec
time docker stop stop-shell
stop-exec
real    0m0.128s

stop-shell
real    0m10.271s

Eighty times slower. And the exit codes tell the rest of the story:

docker inspect --format '{{.Name}} → exit code {{.State.ExitCode}}' stop-exec stop-shell
/stop-exec → exit code 0
/stop-shell → exit code 137

Let's reconstruct what happened in the second case:

  1. docker stop sends SIGQUIT (Nginx's STOPSIGNAL) to PID 1, which is sh.
  2. sh has no handler for that signal and knows nothing about forwarding it to its children. It ignores it or terminates on its own, but nginx never finds out.
  3. Docker waits its full ten seconds.
  4. SIGKILL. The whole container dies at once, with connections cut mid-request. Code 137.

An honest nuance almost nobody mentions: some shells optimize the simplest case. If the CMD were exactly CMD node server.js, many shells (including dash and BusyBox's ash) replace their own process with node via exec, and the problem does not appear. But all it takes is two commands chained with &&, a redirection or a pipe —as in the example, and as in 90% of real-world CMDs— for the shell to have to stay as PID 1 and for the failure to show up. It is not worth gambling your shutdown behavior on a shell optimization:

Always use the exec form: CMD ["node", "server.js"]. If you need shell logic, write it in a docker-entrypoint.sh that ends with exec "$@", as you did in lesson 02-04.

Clean up the demonstration:

docker rm stop-exec stop-shell
docker image rm aurora-api:shell-form

  1. Exit codes and their diagnostic table

When a container ends, it leaves a number behind. That number is the first clue in any investigation.

docker ps -a --filter name=aurora --format "table {{.Names}}\t{{.Status}}"
docker inspect --format '{{.State.ExitCode}}' aurora-api
NAMES          STATUS
aurora-cache   Up 4 minutes
aurora-db      Up 12 minutes
1

The table that solves most cases:

Code Meaning Usual cause in practice
0 Successful termination The process did its job and exited. Also bash with no input
1 Generic application error Uncaught exception, process.exit(1), invalid configuration
125 Failure of docker run itself Misspelled option, non-existent --env-file, duplicate name. The container was not even created
126 The command exists but could not be executed Missing execute bit (chmod +x), or an attempt to execute a directory
127 Command not found Typo in the CMD, or a binary missing from a minimal image
137 128 + 9 → SIGKILL docker kill, docker stop timing out, or the OOM killer (lesson 03-07)
139 128 + 11 → SIGSEGV Segmentation fault: a bug in native code or a binary incompatible with the architecture
143 128 + 15 → SIGTERM The process ended on SIGTERM with no handler of its own. This is a normal stop

The rule that explains half the table: if the code is greater than 128, the process died from a signal, and the signal number is code − 128. 137 − 128 = 9 (SIGKILL); 143 − 128 = 15 (SIGTERM); 139 − 128 = 11 (SIGSEGV).

And the most important distinction for debugging, which separates two worlds:

Range Who failed Where to look
125, 126, 127 Docker or the command's startup In your command line and in the Dockerfile
1, 2, … , 124 Your application In docker logs
>128 An external or kernel signal In docker inspect (OOMKilled, Error)

Let's trigger them to see them for real:

docker run --name exit-0 alpine:3.20 true
docker run --name exit-1 alpine:3.20 sh -c 'exit 1'
docker run --name exit-127 alpine:3.20 command-that-does-not-exist
docker run --made-up-option alpine:3.20
docker run --name exit-126 alpine:3.20 /etc
docker: Error response from daemon: failed to create task for container:
failed to create shim task: OCI runtime create failed: exec: "command-that-does-not-exist": executable file not found in $PATH

unknown flag: --made-up-option

docker: Error response from daemon: failed to create task for container:
failed to create shim task: OCI runtime create failed: exec: "/etc": permission denied
docker ps -a --filter name=exit- --format "table {{.Names}}\t{{.Status}}"
echo "Exit status of the invalid docker run: $?"
NAMES      STATUS
exit-126   Created
exit-127   Created
exit-1     Exited (1) 8 seconds ago
exit-0     Exited (0) 12 seconds ago
125

Notice the revealing detail: exit-127 and exit-126 stayed in the Created state, not Exited. The container did get created but never got to start, because the executable did not exist or could not be executed. exit-1, in contrast, did start and its application decided to exit with an error. That is the difference between "it could not begin" and "it began and failed", and it saves you looking in the wrong place.

docker rm exit-0 exit-1 exit-126 exit-127

Applied to Aurora Libros, the Exited (1) from the previous lesson is now diagnosed without ambiguity: the application started and failed on its own, it was not Docker and not a signal. The logs confirm the reason (getaddrinfo ENOTFOUND aurora-cache) and the process.exit(1) in server.js explains the exact number.

  1. Graceful shutdown for aurora-api

The time has come to fix two things at once in server.js: that the process should not die because an optional dependency is unavailable, and that when it is asked to terminate it should do so closing what it has open.

Replace the final block of ~/aurora-libros/api/server.js (the one that starts at // --- Startup ---) with this:

// --- Startup ---
let server;

async function start() {
  // The cache is an OPTIONAL dependency: if Redis does not respond, the API
  // must keep serving the catalog from PostgreSQL, not die at startup.
  cache.connect().catch((err) =>
    console.error('[cache] not available at startup:', err.message)
  );

  server = app.listen(PORT, '0.0.0.0', () => {
    console.log(`[aurora-api] listening on port ${PORT}`);
    console.log(`[aurora-api] database: ${DB_HOST}:${DB_PORT}/${DB_NAME}`);
    console.log(`[aurora-api] cache: ${REDIS_HOST}:${REDIS_PORT}`);
  });
}

// --- Graceful shutdown ---
let shuttingDown = false;

async function shutdown(signal) {
  if (shuttingDown) return;      // A second Ctrl+C must not re-enter here
  shuttingDown = true;
  console.log(`[aurora-api] received ${signal}, shutting down gracefully...`);

  // Safety net: leave under our own steam BEFORE the SIGKILL arrives.
  // 8 seconds < the 10 of docker stop's grace period.
  const forceExit = setTimeout(() => {
    console.error('[aurora-api] shutdown stalled, forcing exit');
    process.exit(1);
  }, 8000);
  forceExit.unref();

  server.close(async () => {
    try {
      await pool.end();
      console.log('[aurora-api] PostgreSQL pool closed');
      if (cache.isOpen) {
        await cache.quit();
        console.log('[aurora-api] Redis client closed');
      }
    } catch (err) {
      console.error('[aurora-api] error during shutdown:', err.message);
    }
    console.log('[aurora-api] shut down cleanly');
    process.exit(0);
  });
}

process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));

start().catch((err) => {
  console.error('[aurora-api] failed to start:', err.message);
  process.exit(1);
});

Let's go over the decisions one by one, because each one answers something you have learned in this lesson:

Line Why it is there
cache.connect().catch(...) without await Startup does not block on an optional dependency. The container stays running and you can study it, instead of dying with code 1
server = app.listen(...) saved in a variable Without the reference to the server you cannot call .close()
if (shuttingDown) return; Stops two signals in a row from launching two simultaneous shutdowns
setTimeout(..., 8000) 8 < 10: if something stalls, we exit ourselves before the SIGKILL and with a code of our own, not with a 137
.unref() Stops that timer from keeping the process alive if the shutdown goes well
server.close(callback) Stops accepting new connections and waits for the in-flight ones to finish
await pool.end() Closes the PostgreSQL connections. Without this, the database keeps them open until they time out
if (cache.isOpen) quit() on a client that never connected would throw an exception
process.exit(0) An explicit, successful exit: code 0, not 143
SIGINT as well as SIGTERM So that Ctrl+C in the foreground behaves just like docker stop

Build the new version. Since you are changing behavior without breaking the API, bump the MINOR number according to what you decided in lesson 02-06:

docker build -t auroralibros/aurora-api:1.2.0 \
             -t auroralibros/aurora-api:1.2 \
             --build-arg VERSION=1.2.0 \
             ~/aurora-libros/api

And now the timed check. Compare the old version with the new one:

# Version 1.1.0: it died at startup. With 1.2.0 the container stays alive.
docker run -d --name api-shutdown --env-file ~/aurora-libros/aurora.env \
  auroralibros/aurora-api:1.2.0
sleep 3
docker ps --filter name=api-shutdown --format "{{.Names}}: {{.Status}}"
api-shutdown: Up 3 seconds (health: starting)

The container survives. It no longer dies for not finding Redis; it simply logs it and carries on. Now stop it, measuring the time:

time docker stop api-shutdown
docker logs --tail 6 api-shutdown
docker inspect --format 'Exit code: {{.State.ExitCode}}' api-shutdown
api-shutdown
real    0m0.187s

[cache] not available at startup: getaddrinfo ENOTFOUND aurora-cache
[aurora-api] listening on port 3000
[aurora-api] database: aurora-db:5432/aurora_books
[aurora-api] received SIGTERM, shutting down gracefully...
[aurora-api] PostgreSQL pool closed
[aurora-api] shut down cleanly
Exit code: 0

The three things we were after, confirmed in a single output:

  • 0.187 seconds, compared with the 10.271 of the shell form. The signal reached node because it is PID 1 and the exec form lets it through.
  • The logs tell the story of the shutdown: the PostgreSQL pool was closed explicitly before exiting. The Redis client does not appear because it never got opened (cache.isOpen was false), exactly as the code anticipated.
  • Exit code 0, not 143. The difference between "the process was terminated by a signal" and "the process decided to terminate properly because it was asked nicely".

In an orchestrator, this difference is what separates a deployment with no errors from one with requests cut mid-flight and orphaned connections in the database.

docker rm api-shutdown

  1. docker wait and docker rename

Two small commands that solve specific problems.

docker wait

It blocks until the container ends and prints its exit code:

docker run -d --name slow-task alpine:3.20 sh -c 'sleep 5; exit 3'
docker wait slow-task
docker rm slow-task
7d3f9a1b2c48
3

It sits there for five seconds and then prints 3. It is the piece that lets you chain things in a script: "wait for the database migration to finish and continue only if it went well".

docker run -d --name migration alpine:3.20 sh -c 'echo "migrating..."; sleep 3'
if [ "$(docker wait migration)" -eq 0 ]; then
  echo "Migration succeeded, starting the API"
else
  echo "Migration failed, aborting the deployment"
fi
docker rm migration
Migration succeeded, starting the API

docker rename

It changes a container's name on the fly, without stopping it or recreating it:

docker rename aurora-cache aurora-cache-old
docker ps --format "{{.Names}}: {{.Status}}" --filter name=aurora
docker rename aurora-cache-old aurora-cache
aurora-cache-old: Up 25 minutes
aurora-db: Up 33 minutes

It is more useful than it looks: in a zero-downtime deployment, you rename the old container to aurora-api-old, start the new one with the good name and delete the old one once you confirm everything is fine. And a warning that connects with lesson 03-05: if another container was resolving it over DNS by its previous name, it will stop finding it the moment you rename it.

Common Mistakes and Tips

  • Believing an Exited container is deleted. It still exists, it occupies its name and it keeps its writable layer and logs. You delete it with docker rm, not with docker stop.
  • Putting a service into daemon mode inside the container. nginx without daemon off, httpd -k start, postgres with pg_ctl start: the main process ends immediately and the container stops. Services in containers run in the foreground, always.
  • Using docker kill out of habit. A SIGKILL on PostgreSQL forces a recovery on startup and can lose whatever was buffered. docker stop first; kill only if the other one does not respond.
  • Not catching SIGTERM in the application. The process dies outright with its connections open. With a handler, you close the pool and exit with code 0.
  • Setting the safety timer to 10 seconds or more. It must be less than the grace period; otherwise it never fires, because the SIGKILL arrives first.
  • Always reading 137 as "I killed it". The OOM killer also produces it when the container goes over its memory. You tell them apart by looking at .State.OOMKilled in docker inspect (lesson 03-07).
  • Confusing 125 with 1. 125 is Docker's: the container does not even exist, so there are no logs to look at. Review the command line.
  • Tip: raise the --time on databases. docker stop --time 30 aurora-db gives PostgreSQL room to do its checkpoint and close cleanly.
  • Tip: use docker start -a when a container starts and dies right away: you see the startup live without having to chase it with docker logs.

Exercises

Exercise 1: walk through the seven states

Using alpine:3.20 and a container called full-lifecycle that runs sleep 600, take the container through this sequence and note after each step the result of docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}': create without starting → start → pause → resume → stop gracefully → start again → kill → remove. Answer: what is the exit code after docker stop and why? And after docker kill? In which state is the process's memory still occupied?

Exercise 2: demonstrate data loss according to the type of stop

With a redis:7-alpine container called cache-test, write the key catalog:version with value 1 and also a file /tmp/marker.txt with the same content. Then:

  1. Run docker restart and check what survives of the two things.
  2. Run docker stop followed by docker start and check the same.
  3. Run docker rm -f and create a container again with the same name and image. Check the same.

Explain the results in terms of process memory, writable layer and container.

Exercise 3: measure the cost of not handling signals

Prepare three containers from node:22-alpine that run these three programs and time docker stop on each one, also noting the exit code:

  • A: node -e "setInterval(()=>{},1000)" — a process that does nothing and catches no signals.
  • B: sh -c 'echo start && node -e "setInterval(()=>{},1000)"' — the same one, but wrapped in a shell with &&.
  • C: node -e "process.on('SIGTERM',()=>{console.log('bye');process.exit(0)});setInterval(()=>{},1000)" — with a SIGTERM handler.

Explain the three times and the three codes, and say which of the three corresponds to auroralibros/aurora-api:1.1.0 and which to 1.2.0.

Solutions

Solution to exercise 1

docker create --name full-lifecycle alpine:3.20 sleep 600
docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
created / exit code 0
docker start full-lifecycle && docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
docker pause full-lifecycle && docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
docker unpause full-lifecycle && docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
running / exit code 0
paused / exit code 0
running / exit code 0
time docker stop full-lifecycle
docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
docker start full-lifecycle
docker kill full-lifecycle
docker inspect --format '{{.State.Status}} / exit code {{.State.ExitCode}}' full-lifecycle
docker rm full-lifecycle
full-lifecycle
real    0m0.156s
exited / exit code 143

full-lifecycle
full-lifecycle
exited / exit code 137

The three answers:

  • After docker stop: code 143, that is, 128 + 15 = SIGTERM. sleep installs no signal handler, so SIGTERM's default action applies: terminate immediately. That is why real is barely 0.156 s and no SIGKILL was needed. It is a clean stop from Docker's point of view, even though the process ran no shutdown logic of its own: exactly the case of aurora-api:1.1.0.
  • After docker kill: code 137 (128 + 9 = SIGKILL), instantly and with no grace period. Compare the two numbers: 143 means "it was asked to terminate and it terminated"; 137 means "it was terminated without being asked". A 137 accompanied by a ten-second wait would additionally point to a process that ignored the request.
  • The memory is still occupied in paused. The processes are frozen by the cgroups freezer, but they still exist with all their memory reserved. In exited there is no process and the memory is released; what persists is the writable layer on disk.

Solution to exercise 2

docker run -d --name cache-test redis:7-alpine
docker exec cache-test redis-cli SET catalog:version 1
docker exec cache-test sh -c 'echo 1 > /tmp/marker.txt'

1. After docker restart:

docker restart cache-test && sleep 2
docker exec cache-test redis-cli GET catalog:version
docker exec cache-test cat /tmp/marker.txt
(nil)
1

2. After docker stop + docker start:

docker stop cache-test && docker start cache-test && sleep 2
docker exec cache-test redis-cli GET catalog:version
docker exec cache-test cat /tmp/marker.txt
(nil)
1

3. After docker rm -f and recreating:

docker rm -f cache-test
docker run -d --name cache-test redis:7-alpine && sleep 2
docker exec cache-test redis-cli GET catalog:version
docker exec cache-test cat /tmp/marker.txt
docker rm -f cache-test
(nil)
cat: can't open '/tmp/marker.txt': No such file or directory

The explanation on three levels:

Level What survives restart/stop+start What survives rm
Process memory (the Redis key) No. The process is new, its memory starts empty No
Writable layer (/tmp/marker.txt) Yes. It is disk, and the container is the same one No. It is destroyed with the container
Volume (lesson 03-06) Yes Yes, and that is the key

Cases 1 and 2 are equivalent: restart is literally stop + start. What is interesting is that no level of this table survives a docker rm, and that means that as things stand today, if you delete aurora-db, you lose the Aurora Libros catalog. That is exactly the demonstration that opens lesson 03-06.

Solution to exercise 3

docker run -d --name signal-a node:22-alpine node -e "setInterval(()=>{},1000)"
docker run -d --name signal-b node:22-alpine sh -c 'echo start && node -e "setInterval(()=>{},1000)"'
docker run -d --name signal-c node:22-alpine node -e "process.on('SIGTERM',()=>{console.log('bye');process.exit(0)});setInterval(()=>{},1000)"

for c in signal-a signal-b signal-c; do
  echo "--- $c ---"
  { time docker stop "$c" ; } 2>&1 | grep real
  docker inspect --format 'exit code {{.State.ExitCode}}' "$c"
done
--- signal-a ---
real    0m0.142s
exit code 143
--- signal-b ---
real    0m10.238s
exit code 137
--- signal-c ---
real    0m0.118s
exit code 0
Case Time Code Why
A 0.14 s 143 node is PID 1 and receives SIGTERM. With no handler of its own, the default action applies: terminate. Fast, but abrupt: it closed nothing
B 10.24 s 137 PID 1 is sh, which does not forward the signal. Ten seconds of waiting and a SIGKILL. The worst of all worlds
C 0.12 s 0 node catches SIGTERM, runs its shutdown logic and exits of its own accord with code 0

How it maps to Aurora Libros:

  • auroralibros/aurora-api:1.1.0 is case A. Exec form (ENTRYPOINT ["node"]), so the signal arrives, but with no handler: it stopped in 0.3 seconds with code 143 and left the PostgreSQL pool open. Those were the "0.3 seconds" you were celebrating in lesson 02-04, correct but incomplete.
  • auroralibros/aurora-api:1.2.0 is case C. The same speed, but closing the pool and the Redis client and exiting with code 0.
  • Case B must never exist in your images. It is what you would get with CMD node server.js && echo done or with an entrypoint that lacks exec "$@".
docker rm signal-a signal-b signal-c

Conclusion

You now know why some containers live and others die instantly, and there is no mystery to it: a container lasts exactly as long as its PID 1 process lasts. ubuntu:24.04 runs a bash with no input that ends cleanly with code 0; nginx:alpine runs nginx -g "daemon off;", which never ends. And aurora-api was dying because its own code called process.exit(1). You know the seven states and their transitions, and the crucial difference between exited —asleep, with its writable layer, its logs and its name intact— and truly deleted with docker rm.

You have mastered the lifecycle verbs: start, stop, restart, pause/unpause with its cgroups freezer that freezes without sending a single signal and without releasing memory, and kill as a sender of arbitrary signals, including that SIGHUP that reloads Nginx without dropping a single connection. And you understand what really happens after a docker stop: SIGTERM, ten seconds of grace adjustable with --time, and SIGKILL. You have measured it: 0.128 seconds with the exec form versus 10.271 seconds and code 137 with a shell in between that does not forward signals.

Exit codes have stopped being noise. You know that above 128 there is a signal hiding (137 = SIGKILL, 143 = SIGTERM, 139 = SIGSEGV), that 125, 126 and 127 point at Docker or at the command's startup and not at your application, and that a container in the Created state after a docker run means the executable did not even exist. And you have taken the theory into the project: auroralibros/aurora-api:1.2.0 catches SIGTERM and SIGINT, closes the PostgreSQL pool and the Redis client, protects itself with an 8-second timer deliberately shorter than the 10 of the grace period, and exits with code 0 in 0.187 seconds. On top of that, it no longer kills itself because Redis is not there: it stays alive and logs it, which is what a serious service should do.

With aurora-db and aurora-cache running and aurora-api finally able to stay on its feet, you are starting to have something resembling a fleet. And a fleet needs to be watched and kept in order. In the next lesson, Managing Containers, you will squeeze docker ps for everything it has: every column explained, the real size of the writable layer with -s, the filters by state, name, image, label and health, and the --format templates you will use to build a small control panel for the four Aurora Libros services. You will learn to compose commands with -q to operate on dozens of containers at once without wrecking anything, to copy files between the host and a container with docker cp, to see with docker diff what a container has changed with respect to its image, and why docker commit must never be used to build real images.

Docker: From Beginner to Advanced

Module 1: Introduction to Docker

Module 2: Working with Docker Images

Module 3: Docker Containers

Module 4: Docker Compose

Module 5: Advanced Docker Concepts

Module 6: Docker in Production

Module 7: Docker Ecosystem and Tools

© Copyright 2026. All rights reserved