In the previous lesson you started threads with start(), waited for them with join() and cancelled them with interrupt(). Along the way two getState() values appeared —NEW and TERMINATED— with no explanation of what they are or how many more there are.
This lesson explains them all. A Java thread is always in exactly one of six states, and knowing which one it is in and why is the difference between diagnosing a production problem in five minutes or in two days. When an application "hangs", the question is not "why does it not work?" but "what state are its threads in and what are they waiting for?". That question is answered with a thread dump, and reading one is a skill that on its own is worth the price of this lesson.
You will also see the lowest-level coordination mechanism in Java —wait, notify and notifyAll—, which is what sits underneath all the convenient utilities of 08-05 and 08-06; why it must always be wrapped in a while loop and why an if there is a genuine bug, not a style preference; and why three Thread methods that looked extremely useful —stop, suspend and resume— have been marked as dangerous for over twenty years.
A note on level.
wait/notifyis the only part of this module you will probably never write in production:BlockingQueueandCountDownLatchdo it better and without the bugs. It is studied because it explains how those tools work and because it turns up constantly when reading other people's code and thread dumps.
Contents
- The six
Thread.Statevalues - The complete lifecycle diagram
NEWandTERMINATED: the extremesRUNNABLE: the nuance that confuses everybodyBLOCKEDversusWAITING: the distinction that makes a dump readableTIMED_WAITING: waiting with a deadlinewait,notifyandnotifyAll- The mandatory
whileloop notifyversusnotifyAllyieldandonSpinWait- The withdrawn methods:
stop,suspend,resume - Diagnosis in practice: the thread dump
- Reading a deadlock detected by the JVM
- Inspection from inside the program
- BiblioTech: tracking the state of its threads
- Common Mistakes and Tips
- Exercises
- The six
Thread.State values
Thread.State valuesThread.State is an enum —of the 04-07 kind— with six constants. A thread is always in one and only one of them.
| State | Meaning | How you get there | How you leave |
|---|---|---|---|
NEW |
Thread object created, not yet started |
new Thread(...) |
start() |
RUNNABLE |
Running, or ready to run and waiting for a core | start(); returning from a wait |
Finishing run(); blocking; waiting |
BLOCKED |
Waiting to acquire an object's monitor | Entering a busy synchronized |
The monitor becomes free and is acquired |
WAITING |
Waiting indefinitely for a signal from another thread | wait(), join(), LockSupport.park() |
notify/notifyAll, end of the awaited thread, interruption |
TIMED_WAITING |
Same as WAITING but with a deadline |
sleep(ms), wait(ms), join(ms), parkNanos |
Signal, deadline expired or interruption |
TERMINATED |
run() has finished (well or with an exception) |
Return or uncaught exception from run() |
Never: it is terminal |
Two facts worth fixing from the start:
TERMINATEDis irreversible. A finished thread cannot be restarted.start()on it throwsIllegalThreadStateException. That is the deep reason pools exist: if the thread object cannot be reused, you have to reuse the thread — that is, keep it alive and feed it new tasks.Thread.Stateis a JVM view, not an operating-system one. The state you see is what the JVM knows; the system scheduler has its own, finer-grained idea.
- The complete lifecycle diagram
stateDiagram-v2
[*] --> NEW : new Thread(task)
NEW --> RUNNABLE : start()
RUNNABLE --> BLOCKED : enters a busy synchronized
BLOCKED --> RUNNABLE : obtains the monitor
RUNNABLE --> WAITING : wait()
RUNNABLE --> WAITING : join() with no deadline
RUNNABLE --> WAITING : LockSupport.park()
WAITING --> RUNNABLE : notify or notifyAll
WAITING --> RUNNABLE : the awaited thread finishes
WAITING --> RUNNABLE : interrupt()
RUNNABLE --> TIMED_WAITING : sleep(ms)
RUNNABLE --> TIMED_WAITING : wait(ms)
RUNNABLE --> TIMED_WAITING : join(ms)
TIMED_WAITING --> RUNNABLE : deadline expired
TIMED_WAITING --> RUNNABLE : notify or interrupt
RUNNABLE --> TERMINATED : run() returns
RUNNABLE --> TERMINATED : uncaught exception
TERMINATED --> [*]
Two structural observations about this diagram:
First: everything goes through RUNNABLE. There is no direct transition between BLOCKED, WAITING and TIMED_WAITING. A thread that is waiting for a signal and then needs a monitor passes through RUNNABLE in between. This matters when reading dumps: a thread coming out of wait() must reacquire the monitor before continuing, and at that instant it can end up BLOCKED.
Second: there is only one exit from TERMINATED, and it leads out. There is no arrow back. That is the point that makes "reusing a thread" impossible by design.
This small program walks through nearly all the states:
import java.util.concurrent.TimeUnit;
public class StateWalkthrough {
private static final Object lock = new Object();
public static void main(String[] args) throws InterruptedException {
Thread watched = new Thread(() -> {
try {
TimeUnit.MILLISECONDS.sleep(300); // TIMED_WAITING
synchronized (lock) { // BLOCKED (if it is busy)
lock.wait(500); // TIMED_WAITING
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}, "watched-thread");
System.out.println("Before start() : " + watched.getState()); // NEW
watched.start();
System.out.println("Right after start() : " + watched.getState()); // RUNNABLE
TimeUnit.MILLISECONDS.sleep(100);
System.out.println("Sleeping : " + watched.getState()); // TIMED_WAITING
// We take the lock so the watched thread ends up BLOCKED when asking for it.
synchronized (lock) {
TimeUnit.MILLISECONDS.sleep(400);
System.out.println("Lock held by main : " + watched.getState()); // BLOCKED
}
TimeUnit.MILLISECONDS.sleep(100);
System.out.println("Inside wait(500) : " + watched.getState()); // TIMED_WAITING
watched.join();
System.out.println("After join() : " + watched.getState()); // TERMINATED
}
}Output:
Before start() : NEW
Right after start() : RUNNABLE
Sleeping : TIMED_WAITING
Lock held by main : BLOCKED
Inside wait(500) : TIMED_WAITING
After join() : TERMINATEDThis program is inherently fragile and it is honest to say so: it depends on the watched thread's
sleeps falling at exactly the right moments. On a loaded machine it might print something else. It is useful for seeing the states, not as a diagnostic technique: the dumps of section 12 are for that.
NEW and TERMINATED: the extremes
NEW and TERMINATED: the extremesNEW is the state of a freshly constructed Thread object. The operating-system thread does not exist yet: new Thread(...) only allocates a Java object. The native thread is created in start().
From this comes the diagnosis of the 08-02 mistake: if a thread shows as NEW when you expected it to be working, it means nobody called start() —or that run() was called by mistake—.
TERMINATED is the state after run() ends, by either of its two routes:
public class TwoWaysToFinish {
public static void main(String[] args) throws InterruptedException {
Thread good = new Thread(() -> System.out.println("[good] work done"), "good");
Thread bad = new Thread(() -> { throw new RuntimeException("failure"); }, "bad");
good.start(); good.join();
bad.start(); bad.join();
System.out.println("good -> " + good.getState()); // TERMINATED
System.out.println("bad -> " + bad.getState()); // TERMINATED
}
}Both end up TERMINATED. The state does not tell success from failure, and that is exactly the trap of 08-02: join() returns normally whether the thread did its work or died from an exception. To tell them apart you need either an uncaught exception handler or a Future (08-05), which does store the exception and rethrows it in get().
RUNNABLE: the nuance that confuses everybody
RUNNABLE: the nuance that confuses everybodyMany textbooks draw a RUNNING state separate from RUNNABLE. Java does not have one. Thread.State.RUNNABLE groups together two very different situations:
- Running right now on a core.
- Ready to run, in the scheduler's queue, waiting to be given a core.
The JVM does not distinguish them because it cannot: who runs at each instant is decided by the operating-system scheduler, without informing the JVM. With eight RUNNABLE threads on a four-core machine, at any instant at most four are genuinely running and at least four are waiting their turn, and they all show as RUNNABLE.
There is a third situation, even more counter-intuitive:
A thread blocked on classic I/O shows as
RUNNABLE.
A thread stopped in InputStream.read() waiting for the disk has not executed an instruction for a while, but getState() says RUNNABLE. The reason is that the JVM has no idea that thread is parked inside a system call: from its point of view, the thread is executing native code.
// This thread can spend 30 seconds here and its state will be RUNNABLE
// the whole time, even though it does not consume a single CPU cycle.
Thread reader = new Thread(() -> {
try (BufferedReader br = Files.newBufferedReader(Path.of("data/huge.csv"))) {
String line;
while ((line = br.readLine()) != null) { process(line); }
} catch (IOException e) { /* ... */ }
}, "bibliotech-reader");Practical consequence for diagnosis: seeing lots of RUNNABLE threads does not mean the CPU is saturated. You have to look at the stack trace:
- If the top is
java.net.SocketInputStream.socketRead0orFileInputStream.readBytes, that thread is waiting on I/O, not computing. - If the top is your own code in a loop, then it really is consuming CPU.
This distinction is what separates "I need more cores" from "I need more threads because I am waiting for the disk" —the CPU-bound versus I/O-bound table of 08-01, now seen from the other side—.
BLOCKED versus WAITING: the distinction that makes a dump readable
BLOCKED versus WAITING: the distinction that makes a dump readableThis is the most valuable distinction in the lesson for real work.
BLOCKED |
WAITING |
|
|---|---|---|
| What it waits for | A monitor (the key to a synchronized) |
A signal from another thread |
| How you get there | Trying to enter a busy synchronized |
wait(), join(), LockSupport.park() |
| How you leave | The monitor's owner releases it | Another thread calls notify/notifyAll, the awaited thread finishes, or an interruption arrives |
| Does it depend on another thread? | Yes, the one holding the monitor | Yes, the one that must signal |
| In a dump | waiting to lock <0x...> |
in Object.wait() / parking to wait for <0x...> |
| What it usually indicates | Contention: many threads fighting over a lock | Coordination: empty queue, pending task, join |
| If it lasts a long time | Suspect a deadlock or a long critical section | Normal in an idle pool; suspicious if nobody is going to signal |
The short way to remember it:
BLOCKED= "they will not let me in".WAITING= "I am waiting to be told".
Demonstration of both states at once:
import java.util.concurrent.TimeUnit;
public class BlockedVsWaiting {
private static final Object lock = new Object();
public static void main(String[] args) throws InterruptedException {
// Thread 1: takes the lock and stays inside doing wait().
// It releases the monitor on entering wait(), but stays WAITING.
Thread waiter = new Thread(() -> {
synchronized (lock) {
try {
lock.wait(); // WAITING (and RELEASES the monitor)
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
}, "waiting-thread");
// Thread 2: takes the lock and holds it while sleeping.
// sleep does NOT release the monitor (08-02, section 8).
Thread hogger = new Thread(() -> {
synchronized (lock) {
try {
TimeUnit.SECONDS.sleep(5); // TIMED_WAITING, holding the monitor
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
}, "hogging-thread");
// Thread 3: wants the lock and cannot get it -> BLOCKED.
Thread blocked = new Thread(() -> {
synchronized (lock) {
System.out.println("[blocked] I have finally got in");
}
}, "blocked-thread");
waiter.start();
TimeUnit.MILLISECONDS.sleep(200); // let it reach wait()
hogger.start();
TimeUnit.MILLISECONDS.sleep(200); // let it take the lock
blocked.start();
TimeUnit.MILLISECONDS.sleep(200); // let it hit the lock
System.out.println("waiter : " + waiter.getState());
System.out.println("hogger : " + hogger.getState());
System.out.println("blocked : " + blocked.getState());
// Cleanup
hogger.interrupt();
synchronized (lock) { lock.notifyAll(); }
}
}Output:
Three states at once, each for a different reason. Note the crucial detail about waiter: it entered the synchronized, that is, it had the monitor, and on calling wait() it released it. That is why hogger could get in afterwards. This is the essential difference between wait() and sleep():
Thread.sleep(ms) |
object.wait() |
|
|---|---|---|
| Does it release the monitor? | No | Yes |
| Where it can be called | Anywhere | Only inside a synchronized on that object |
| How it wakes up | Deadline expired or interruption | notify/notifyAll, deadline, interruption or spurious |
| Resulting state | TIMED_WAITING |
WAITING or TIMED_WAITING |
| What it is for | Pausing | Coordinating |
TIMED_WAITING: waiting with a deadline
TIMED_WAITING: waiting with a deadlineTIMED_WAITING is the deadline version of WAITING. The methods that lead to it:
Thread.sleep(1000); // TIMED_WAITING
TimeUnit.SECONDS.sleep(1); // same
object.wait(1000); // TIMED_WAITING, you release the monitor
thread.join(1000); // TIMED_WAITING
LockSupport.parkNanos(1_000_000_000L); // TIMED_WAITING
lock.tryLock(1, TimeUnit.SECONDS); // TIMED_WAITING (08-04)
queue.poll(1, TimeUnit.SECONDS); // TIMED_WAITING (08-06)In a dump, TIMED_WAITING (sleeping) and TIMED_WAITING (on object monitor) are distinguished from each other, which is useful: the first is a thread that chose to pause, the second is a thread waiting for a signal with a deadline.
The diagnostic detail that saves hours: a thread permanently in TIMED_WAITING (sleeping) inside a loop is almost always polling, and almost always an improvable design:
// ANTIPATTERN: polling. The thread wakes up 10 times a second
// just to check whether there is anything, and usually there is not.
while (reservationQueue.isEmpty()) {
Thread.sleep(100);
}
Reservation r = reservationQueue.poll();
// CORRECT: block until there is something, with no pointless wake-ups (08-06).
Reservation r = reservationQueue.take(); // BlockingQueue
wait, notify and notifyAll
wait, notify and notifyAllThese three methods are not on Thread: they are on Object, the class from 03-09. Any Java object can serve as a coordination point, because every object has an associated monitor.
| Method | What it does |
|---|---|
object.wait() |
Releases object's monitor and leaves the thread WAITING until it receives a signal |
object.wait(ms) |
Same, but it also wakes up when the deadline expires (TIMED_WAITING) |
object.notify() |
Wakes one thread (arbitrarily chosen) of those waiting on object |
object.notifyAll() |
Wakes all the threads waiting on object |
Rule number one, which the compiler does not check:
All three methods must be called holding that object's monitor, that is, inside a
synchronizedblock or method on that same object. OtherwiseIllegalMonitorStateExceptionis thrown at runtime.
Object lock = new Object();
// WRONG: without the monitor -> IllegalMonitorStateException
lock.wait();
// RIGHT
synchronized (lock) {
lock.wait();
}Why that rule exists —and it is a classic interview question—: without it there would be an unfixable race condition. Imagine you could call wait() without the monitor:
- The consumer checks
queue.isEmpty()→true. - Right here the producer adds an element and calls
notify(). - The consumer calls
wait()… and misses the signal, because it arrived before it started waiting.
Result: the consumer waits forever with a full queue. Requiring the monitor makes "checking the condition" and "starting to wait" a single indivisible operation from the producer's point of view, because the producer needs the same monitor to signal.
Minimal coordination example:
import java.util.ArrayDeque;
import java.util.Deque;
public class BasicCoordination {
// The locking AND signalling object: the queue itself.
private static final Deque<String> reservations = new ArrayDeque<>();
public static void main(String[] args) throws InterruptedException {
Thread consumer = new Thread(() -> {
try {
synchronized (reservations) {
// A while LOOP, not an if. Section 8.
while (reservations.isEmpty()) {
System.out.println("[consumer] queue empty, waiting");
reservations.wait(); // releases the monitor and waits
}
System.out.println("[consumer] serving: " + reservations.poll());
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}, "bibliotech-consumer");
Thread producer = new Thread(() -> {
try {
Thread.sleep(500);
synchronized (reservations) {
reservations.add("Reservation by Marta Ruiz: 978-0000000001");
System.out.println("[producer] reservation added, signalling");
reservations.notifyAll(); // wakes those who are waiting
} // the monitor is released HERE
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}, "bibliotech-producer");
consumer.start();
producer.start();
consumer.join();
producer.join();
}
}Output:
[consumer] queue empty, waiting
[producer] reservation added, signalling
[consumer] serving: Reservation by Marta Ruiz: 978-0000000001A detail that is easy to miss: notifyAll() does not release the monitor. The signalling thread stays inside the synchronized until the closing brace. The consumer wakes up, but cannot continue until the monitor is free; meanwhile it goes from WAITING to BLOCKED. It is the transition from section 2 —everything goes through RUNNABLE— seen in action, and the reason it is best to signal at the end of the block, not at the beginning.
A note on ordering.
synchronizedis studied in depth in 08-04: what exactly a monitor is, reentrancy, private locks and granularity. Here the operational rule is enough: to callwait,notifyornotifyAllon an object, you must be inside asynchronizedon that same object.
- The mandatory
while loop
while loopThis is the most broken rule and the one that produces the hardest bugs to reproduce:
wait()ALWAYS goes inside awhileloop that checks the condition. Never inside anif.
// WRONG - a real bug, not a matter of style
synchronized (reservations) {
if (reservations.isEmpty()) {
reservations.wait();
}
Reservation r = reservations.poll(); // may be null
serve(r); // NullPointerException
}
// RIGHT
synchronized (reservations) {
while (reservations.isEmpty()) {
reservations.wait();
}
Reservation r = reservations.poll(); // guaranteed non-null
serve(r);
}There are three independent reasons, and each one would be enough on its own:
Reason 1: spurious wake-ups. The Java specification explicitly allows wait() to return without anyone having called notify. It is not an implementation flaw: it is a deliberate concession to how the underlying operating systems' synchronisation primitives work (pthread_cond_wait carries the same warning). They are rare, but they happen, and they happen more on loaded machines —that is, in production and not on your laptop—.
Reason 2: notifyAll wakes everybody, but only one can win. With three consumers waiting and a single element in the queue, notifyAll() wakes all three. The first to reacquire the monitor takes the element; the other two check again and —with while— go back to waiting. With if, they carry on and find the queue empty.
Reason 3: the condition may have changed between the signal and the wake-up. Time passes between notify() and the moment the woken thread reacquires the monitor, and in that time another thread may have come in and consumed the element. The signal says "something changed", not "your condition now holds".
The third reason is the most important and the one that makes the while a logical necessity rather than a precaution: notify conveys no information about the state, it only says "look again". The only source of truth is the condition, and that is why it has to be re-evaluated.
Demonstration of the failure with if:
import java.util.ArrayDeque;
import java.util.Deque;
import java.util.concurrent.TimeUnit;
public class TheIfBug {
private static final Deque<Integer> queue = new ArrayDeque<>();
static Runnable consumerWithIf = () -> {
try {
synchronized (queue) {
if (queue.isEmpty()) { // BUG: it should be while
queue.wait();
}
Integer v = queue.poll();
System.out.println("[" + Thread.currentThread().getName()
+ "] consumes: " + v
+ (v == null ? " <-- FAILURE: empty queue" : ""));
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
};
public static void main(String[] args) throws InterruptedException {
// Three consumers waiting...
for (int i = 1; i <= 3; i++) {
new Thread(consumerWithIf, "consumer-" + i).start();
}
TimeUnit.MILLISECONDS.sleep(300);
// ...and ONE single element with notifyAll.
synchronized (queue) {
queue.add(42);
queue.notifyAll(); // wakes ALL THREE
}
TimeUnit.MILLISECONDS.sleep(500);
}
}Output:
[consumer-1] consumes: 42
[consumer-2] consumes: null <-- FAILURE: empty queue
[consumer-3] consumes: null <-- FAILURE: empty queueTwo of the three consumers get null and in a real program that would be a NullPointerException on the next line. Change if for while and run again: consumers 2 and 3 go back to waiting silently, which is exactly right.
notify versus notifyAll
notify versus notifyAllnotify() |
notifyAll() |
|
|---|---|---|
| How many it wakes | One, chosen arbitrarily | All those waiting on that object |
| Cost | Lower | Higher: they all wake up and compete for the monitor |
| Risk | Lost signal: it may wake the wrong thread | None to correctness; only inefficiency |
| When it is safe | All the waiters are interchangeable and wait on the same condition | Always |
The recommendation is blunt: use notifyAll() unless you can prove it is safe to use notify().
The danger of notify() is the lost wake-up. If threads wait on the same object for different conditions —some because the queue is empty, others because it is full—, notify() may wake precisely one that cannot make progress. That thread checks its condition, it still does not hold, it goes back to wait()… and the signal has been consumed. The thread that could have made progress was never told, and the system grinds to a halt with no error and no exception.
// DANGEROUS: two different conditions waiting on the SAME object
synchronized (buffer) {
while (buffer.isFull()) buffer.wait(); // producers
// ...
buffer.notify(); // may wake ANOTHER producer
}
synchronized (buffer) {
while (buffer.isEmpty()) buffer.wait(); // consumers
// ...
buffer.notify(); // may wake ANOTHER consumer
}In that design, notify() can end in a total system standstill. With notifyAll() the problem disappears: everybody checks, those who can make progress do, those who cannot go back to waiting. You pay a little cost for a correctness guarantee; it is almost always the best possible trade.
The elegant solution —one wait set per condition— does exist: they are the
Conditionobjects ofReentrantLock, which allowsignal()on exactly the right group of threads. They are covered in 08-04.
yield and onSpinWait
yield and onSpinWaitTwo minor methods, in a note.
Thread.yield() is a hint to the scheduler: "if there is another ready thread, let it through". It guarantees nothing; the scheduler can ignore it entirely, and the thread can be chosen again immediately. It does not change the state (still RUNNABLE) and it releases no monitor. Do not use it for correctness; its legitimate place is stress tests, where forcing more context switches helps latent race conditions to show up.
// Legitimate use: in a test, to provoke different interleavings
// and increase the chance of a race bug manifesting.
for (int i = 0; i < 1000; i++) {
counter.increment();
if (i % 10 == 0) Thread.yield();
}Thread.onSpinWait() (Java 9) is a hint to the CPU inside a very short busy wait: it emits an instruction like x86's PAUSE, which reduces power consumption and improves performance on leaving the loop. It only makes sense in nanosecond-scale waiting loops and in very low-level code:
// A VERY short busy wait, in high-performance code.
// In ordinary application code this is an antipattern:
// it burns a whole core. Use wait/notify or a BlockingQueue.
while (!ready) {
Thread.onSpinWait();
}Neither of the two solves a coordination problem. If you think you need yield() for your program to work, what you need is synchronisation.
- The withdrawn methods:
stop, suspend, resume
stop, suspend, resumeThread has three methods that did exactly what one would want —stop a thread, pause it and resume it— and that have been marked deprecated since Java 1.2 (1998). In Java 20 stop() started throwing UnsupportedOperationException unconditionally, and suspend/resume were marked for removal. It is worth understanding why, because the reason teaches something real.
Thread.stop(): it killed the thread by throwing a ThreadDeath at it wherever it happened to be.
The problem: it releases all its monitors instantly, in the middle of whatever it was doing.
// Suppose the thread is HERE when stop() reaches it:
synchronized (catalog) {
catalog.removeIsbn(isbn); // already executed
// <-- stop() lands right here
catalog.removeFromIndex(isbn); // NEVER executed
}Result: the ISBN was removed from the set but not from the index. The catalogue is left inconsistent, the monitor is released and the other threads walk right in to read a corrupt structure. And there is no exception or trace to indicate it: the failure shows up later, somewhere else, with no apparent connection.
It is exactly the problem the cooperative cancellation of 08-02 solves: interrupt() does not interrupt wherever it fancies, but at the points the thread itself has designated, where the state is consistent.
Thread.suspend() and resume(): they paused and resumed a thread.
The problem: suspend() does not release the monitors. A thread suspended inside a synchronized holds the lock indefinitely. If the thread that was going to call resume() needs that same lock to get there, perfect deadlock, with no error and no trace.
// The deadlock recipe:
workerThread.suspend(); // suspended while holding 'catalog's monitor
synchronized (catalog) { // main ends up BLOCKED here, forever
workerThread.resume(); // this line is never reached
}What is used instead:
| Withdrawn method | Replacement |
|---|---|
stop() |
interrupt() + cooperative protocol (08-02) |
suspend() / resume() |
volatile flag + wait/notify, or Semaphore/CyclicBarrier (08-05) |
destroy() |
It was never implemented; removed |
Pause and resume done properly:
public class PausableTask implements Runnable {
private final Object monitor = new Object();
private boolean paused = false; // guarded by 'monitor'
private volatile boolean stopped = false;
public void pause() { synchronized (monitor) { paused = true; } }
public void resume() {
synchronized (monitor) {
paused = false;
monitor.notifyAll(); // wakes the worker
}
}
public void stop() { stopped = true; }
@Override
public void run() {
try {
while (!stopped && !Thread.currentThread().isInterrupted()) {
// PAUSE POINT chosen by the thread itself: here the
// state is consistent and no business lock is being
// held. Nothing like suspend().
synchronized (monitor) {
while (paused && !stopped) {
monitor.wait(); // releases 'monitor', goes WAITING
}
}
processOneBatch();
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
private void processOneBatch() { /* work */ }
}The essential difference from suspend(): the thread chooses where it pauses. At that point the business state is consistent and no catalogue lock is held. It is the same principle as cooperative interruption.
- Diagnosis in practice: the thread dump
We now reach the part you will genuinely use. A thread dump is an instantaneous snapshot of all the JVM's threads: their name, their state, their stack trace and the locks they hold or are waiting for. It is the number-one tool for diagnosing a hung or slow application.
12.1 How to obtain one
Option A: Ctrl+\ in the terminal (Linux/macOS; on Windows, Ctrl+Break). The JVM prints the dump on System.err. It needs no tools and does not kill the process.
Option B: jstack, included in the JDK. It is the professional option:
# 1. Locate the Java process
jps -l
# 48213 com.nexussoftware.bibliotech.BiblioTechApp
# 48291 jdk.jcmd/sun.tools.jps.Jps
# 2. Dump its threads
jstack 48213
# 3. Save it to analyse at leisure
jstack 48213 > /tmp/dump-1.txt
# 4. Take THREE dumps five seconds apart:
# comparing the three tells a stuck thread from one that is progressing
for i in 1 2 3; do jstack 48213 > /tmp/dump-$i.txt; sleep 5; doneOption C: jcmd, more modern and with more information:
The most valuable operational tip in the whole section: always take THREE dumps a few seconds apart. A single dump tells you where each thread is at that instant, and a thread may just be passing through. Three dumps where the same thread appears on the same line mean it is stuck, not busy. It is the difference between a diagnosis and a guess.
12.2 Anatomy of an entry
"bibliotech-importer" #21 prio=5 os_prio=0 cpu=1243.55ms elapsed=18.32s tid=0x00007f3c1c0e1800 nid=0x5f2a waiting for monitor entry [0x00007f3c0a1fe000]
java.lang.Thread.State: BLOCKED (on object monitor)
at com.nexussoftware.bibliotech.service.Catalog.add(Catalog.java:88)
- waiting to lock <0x000000071ab34d18> (a com.nexussoftware.bibliotech.service.Catalog)
at com.nexussoftware.bibliotech.persistence.ImportTask.run(ImportTask.java:64)
at java.base/java.lang.Thread.run(Thread.java:1583)Field by field:
| Field | Meaning |
|---|---|
"bibliotech-importer" |
The thread name. This is where the 08-02 tip pays off |
#21 |
The thread's identifier in the JVM |
prio=5 |
Java priority (remember 08-02: it barely means anything) |
cpu=1243.55ms |
Total CPU consumed. A "stuck" thread with rising CPU is in a loop, not blocked |
elapsed=18.32s |
The thread's lifetime |
nid=0x5f2a |
Native identifier, in hexadecimal. Cross-reference with top -H to see which thread is burning CPU |
waiting for monitor entry |
Textual summary of the state |
java.lang.Thread.State: BLOCKED |
The state from section 1 |
- waiting to lock <0x...> |
Which lock it is waiting for, with its identity and its class |
at ... |
The stack trace: exactly where it is |
The three lines to look at first, in this order: the state, the - waiting to lock / - locked, and the first at line of your own code (ignoring the java.base ones).
Common shapes:
--- A thread genuinely WORKING (consuming CPU) ---
"bibliotech-compute" #24 ... cpu=8420.11ms ... runnable
java.lang.Thread.State: RUNNABLE
at com.nexussoftware.bibliotech.service.FineCalculator.recalculate(FineCalculator.java:52)
--- A thread waiting on I/O: RUNNABLE but NOT consuming CPU ---
"bibliotech-reader" #25 ... cpu=12.30ms ... runnable
java.lang.Thread.State: RUNNABLE
at java.base/java.io.FileInputStream.readBytes(Native Method)
at java.base/java.io.BufferedInputStream.read(BufferedInputStream.java:...)
--- An IDLE pool thread: normal, not a problem ---
"bibliotech-notices-3" #31 ... cpu=340.02ms ... waiting on condition
java.lang.Thread.State: WAITING (parking)
- parking to wait for <0x000000071b0c4470> (a java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject)
at java.base/java.util.concurrent.LinkedBlockingQueue.take(LinkedBlockingQueue.java:...)
at java.base/java.util.concurrent.ThreadPoolExecutor.getTask(ThreadPoolExecutor.java:...)
--- A SLEEPING thread ---
"bibliotech-monitor" #33 ... sleeping
java.lang.Thread.State: TIMED_WAITING (sleeping)
at java.base/java.lang.Thread.sleep(Native Method)How to read these four patterns:
RUNNABLEwith high and risingcpubetween dumps: working or in an infinite loop. Compare the three dumps: if the trace does not change and the CPU goes up, it is a loop with no way out.RUNNABLEwith lowcpuandreadBytesat the top: waiting on I/O. It is not a CPU problem; it is the I/O-bound task case of 08-01.WAITING (parking)in aThreadPoolExecutor'sgetTask: a pool thread waiting for work. This is normal, do not chase it. You will see them by the dozen in 08-05.BLOCKED (on object monitor)in several threads on the same<0x...>: contention. If on top of that it persists across the three dumps, find who owns that lock —it will appear with- locked <that same 0x...>— and why it is taking so long.
- Reading a deadlock detected by the JVM
The best thing about dumps is that the JVM detects monitor deadlocks for you and explains them at the end of the dump. This is what it looks like, with BiblioTech code:
public class BiblioTechDeadlock {
private static final Object CATALOG = new Object();
private static final Object REGISTRY = new Object();
public static void main(String[] args) {
// Thread A: catalog -> registry
new Thread(() -> {
synchronized (CATALOG) {
pause(200);
synchronized (REGISTRY) {
System.out.println("A: never gets here");
}
}
}, "bibliotech-loans").start();
// Thread B: registry -> catalog (REVERSE ORDER: the cause)
new Thread(() -> {
synchronized (REGISTRY) {
pause(200);
synchronized (CATALOG) {
System.out.println("B: never gets here");
}
}
}, "bibliotech-returns").start();
}
static void pause(long ms) {
try { Thread.sleep(ms); } catch (InterruptedException e) { Thread.currentThread().interrupt(); }
}
}This program hangs: it prints nothing, throws no exception, consumes no CPU and never finishes. It is the worst kind of failure: silent and total. Run jstack on it and the end of the dump says:
Found one Java-level deadlock:
=============================
"bibliotech-loans":
waiting to lock monitor 0x00007f3c14006e00 (object 0x000000071ab34d18, a java.lang.Object),
which is held by "bibliotech-returns"
"bibliotech-returns":
waiting to lock monitor 0x00007f3c14009480 (object 0x000000071ab34d28, a java.lang.Object),
which is held by "bibliotech-loans"
Java stack information for the threads listed above:
===================================================
"bibliotech-loans":
at BiblioTechDeadlock.lambda$main$0(BiblioTechDeadlock.java:14)
- waiting to lock <0x000000071ab34d18> (a java.lang.Object)
- locked <0x000000071ab34d28> (a java.lang.Object)
at java.base/java.lang.Thread.run(Thread.java:1583)
"bibliotech-returns":
at BiblioTechDeadlock.lambda$main$1(BiblioTechDeadlock.java:25)
- waiting to lock <0x000000071ab34d28> (a java.lang.Object)
- locked <0x000000071ab34d18> (a java.lang.Object)
at java.base/java.lang.Thread.run(Thread.java:1583)
Found 1 deadlock.How to read it, in three steps:
Found one Java-level deadlockis the keyword to search for. If it appears, there is nothing to discuss: there is a deadlock.- The section above gives the cycle:
loanswaits for a lock held byreturns, andreturnswaits for a lock held byloans. A two-way cycle; they can be three-way or longer. - The stack traces give the exact lines —
BiblioTechDeadlock.java:14and:25— and, above all, the combination- locked <A>+- waiting to lock <B>on one thread and- locked <B>+- waiting to lock <A>on the other. That is the signature of a deadlock caused by reverse acquisition order, and the solution —imposing a global ordering— is 08-04.
Two important limitations of the automatic detection:
- It detects deadlocks on
synchronizedmonitors and onReentrantLock(throughThreadMXBean), but it does not detect logical mutual waits: two threads each waiting for a signal from the other withwait()do not produce a declared "deadlock", even though the effect is the same. There you will only see two threads inWAITINGforever. - It does not detect livelock (08-04), in which the threads do move but make no progress.
- Inspection from inside the program
Sometimes it is useful to inspect the threads from inside the application, typically for a diagnostic endpoint or to dump the state before aborting.
Thread.getAllStackTraces() returns a map of all the live threads with their traces:
import java.util.Map;
public class InternalDump {
/** Dumps every live thread and its trace. Useful in a global handler. */
public static void dumpThreads() {
Map<Thread, StackTraceElement[]> all = Thread.getAllStackTraces();
System.out.println("=== INTERNAL DUMP: " + all.size() + " threads ===");
all.forEach((thread, trace) -> {
System.out.printf("%n\"%s\" id=%d state=%s daemon=%s%n",
thread.getName(), thread.threadId(), thread.getState(), thread.isDaemon());
// Only the first 5 lines: the rest rarely adds anything
int n = Math.min(5, trace.length);
for (int i = 0; i < n; i++) {
System.out.println(" at " + trace[i]);
}
if (trace.length > n) {
System.out.println(" ... " + (trace.length - n) + " more");
}
});
}
}ThreadMXBean (from java.lang.management) gives far more information, including programmatic deadlock detection:
import java.lang.management.ManagementFactory;
import java.lang.management.ThreadInfo;
import java.lang.management.ThreadMXBean;
public class DeadlockDetector {
private static final ThreadMXBean MX = ManagementFactory.getThreadMXBean();
/** Returns a deadlock report, or null if there is none. */
public static String detect() {
long[] ids = MX.findDeadlockedThreads(); // null if there is none
if (ids == null) {
return null;
}
StringBuilder sb = new StringBuilder("DEADLOCK DETECTED:\n");
for (ThreadInfo info : MX.getThreadInfo(ids, true, true)) {
sb.append(" Thread '").append(info.getThreadName())
.append("' (").append(info.getThreadState()).append(")\n")
.append(" waits for : ").append(info.getLockName()).append('\n')
.append(" held by : ").append(info.getLockOwnerName()).append('\n');
}
return sb.toString();
}
/**
* Watchdog that checks every 10 seconds whether there is a deadlock.
* A daemon: it is purely informational and must not prevent shutdown (08-02).
*/
public static void startWatchdog() {
Thread watchdog = new Thread(() -> {
while (!Thread.currentThread().isInterrupted()) {
String report = detect();
if (report != null) {
System.err.println(report);
}
try {
Thread.sleep(10_000);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
}, "bibliotech-deadlock-watchdog");
watchdog.setDaemon(true);
watchdog.start();
}
}A watchdog like that, logging to the 06-07 logger, turns "the application has hung and we do not know why" into a log entry saying exactly which two threads and which two locks. It is worth the thirty lines it costs.
Graphical tools.
jconsole(included in the JDK) and VisualVM (a separate download) show, in real time, the number of threads, their state, an activity graph and a "Detect Deadlock" button. For a first exploration they are far more convenient thanjstack; for serious analysis or on a system with no graphical interface,jstackis still the tool. Java Flight Recorder + JDK Mission Control is the professional tier, with continuous low-overhead sampling.
- BiblioTech: tracking the state of its threads
A practical application: a monitor that shows the state of BiblioTech's threads while they work.
package com.nexussoftware.bibliotech.infrastructure;
import java.util.List;
import java.util.concurrent.TimeUnit;
/**
* BiblioTech thread monitor.
*
* It is a DAEMON thread: purely informational, it must not prevent the
* application from shutting down (08-02, section 7).
*/
public class ThreadMonitor implements Runnable {
private final List<Thread> watched;
private final long periodMs;
public ThreadMonitor(List<Thread> watched, long periodMs) {
this.watched = List.copyOf(watched); // immutable copy: safe
this.periodMs = periodMs;
}
@Override
public void run() {
try {
while (!Thread.currentThread().isInterrupted()) {
boolean anyAlive = false;
StringBuilder sb = new StringBuilder("[monitor] ");
for (Thread t : watched) {
Thread.State s = t.getState();
sb.append(String.format("%-24s %-14s | ", t.getName(), s));
if (s != Thread.State.TERMINATED && s != Thread.State.NEW) {
anyAlive = true;
}
}
System.out.println(sb);
if (!anyAlive) {
System.out.println("[monitor] every thread has finished");
return;
}
TimeUnit.MILLISECONDS.sleep(periodMs);
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
/** Starts the monitor as a daemon and returns its thread. */
public static Thread launch(List<Thread> watched, long periodMs) {
Thread t = new Thread(new ThreadMonitor(watched, periodMs),
"bibliotech-thread-monitor");
t.setDaemon(true);
t.start();
return t;
}
}Usage, on the 08-02 import and a notices task:
package com.nexussoftware.bibliotech.presentation;
import java.nio.file.Path;
import java.util.List;
import java.util.concurrent.TimeUnit;
public class StateDemo {
public static void main(String[] args) throws InterruptedException {
Catalog catalog = new Catalog();
Thread importer = new Thread(
new ImportTask(Path.of("data/inventory.csv"), catalog),
"bibliotech-importer");
Thread notices = new Thread(() -> {
for (int i = 1; i <= 20; i++) {
try {
// Simulated I/O: in a dump, this thread would appear
// in TIMED_WAITING (sleeping).
TimeUnit.MILLISECONDS.sleep(150);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
return;
}
}
}, "bibliotech-notices");
// The monitor starts FIRST: that way it captures the NEW state.
ThreadMonitor.launch(List.of(importer, notices), 400);
TimeUnit.MILLISECONDS.sleep(500);
importer.start();
notices.start();
importer.join();
notices.join();
TimeUnit.MILLISECONDS.sleep(500); // margin for the last line
}
}Output:
[monitor] bibliotech-importer NEW | bibliotech-notices NEW |
[monitor] bibliotech-importer RUNNABLE | bibliotech-notices TIMED_WAITING |
[monitor] bibliotech-importer TIMED_WAITING | bibliotech-notices TIMED_WAITING |
[monitor] bibliotech-importer RUNNABLE | bibliotech-notices TIMED_WAITING |
[monitor] bibliotech-importer TERMINATED | bibliotech-notices TIMED_WAITING |
[monitor] bibliotech-importer TERMINATED | bibliotech-notices TERMINATED |
[monitor] every thread has finishedWhat this output teaches:
- The importer alternates between
RUNNABLE(reading and validating lines) andTIMED_WAITING(the simulated pauses). It is the profile of a mixed task. - The notices thread is almost always in
TIMED_WAITING: it is the profile of an I/O-bound task. This is exactly the empirical observation that justifies using many more threads than cores for the notices (08-01, section 6) and that will be exploited in 08-05. - The monitor detects the end of both and finishes on its own. Since it is a daemon, even if it did not, it would not prevent the JVM from shutting down.
Common Mistakes and Tips
Mistake 1: believing RUNNABLE means "consuming CPU". A thread blocked in read() shows as RUNNABLE and consumes nothing. Look at the stack trace and the dump's cpu= field, not just the state.
Mistake 2: using if instead of while around wait(). A bug, not a style. Spurious wake-ups, notifyAll with several waiters, and conditions that change between the signal and the wake-up: three independent reasons.
Mistake 3: calling wait/notify without the monitor. IllegalMonitorStateException at runtime. The compiler does not help here.
Mistake 4: using notify() when threads are waiting for different conditions. Lost signal and a silent standstill. Use notifyAll(), or 08-04's Condition if performance justifies it.
Mistake 5: calling notifyAll() at the start of the synchronized block. It does not release the monitor: the woken threads go to BLOCKED and wait just the same. Signal as close to the end of the block as possible.
Mistake 6: using stop, suspend or resume. Corrupt state and guaranteed deadlocks. stop() does not even work any more in Java 20+.
Mistake 7: taking a single thread dump and drawing conclusions. A thread may just be passing through that line. Take three, five seconds apart, and compare.
Mistake 8: chasing WAITING threads in a pool. A ThreadPoolExecutor's idle threads are in WAITING (parking) inside getTask(). That is their normal state when there is no work.
Mistake 9: polling with sleep in a loop. It burns CPU for nothing, and adds latency equal to the polling period. Block on the condition: wait, a BlockingQueue or a CountDownLatch.
Tip 1: learn to read a dump before you need to. Provoke a deadlock on purpose with the code from section 13 and run jstack on it. When it happens in production at three in the morning, you will be glad you have seen it before.
Tip 2: cross-reference the dump's nid with top -H -p <pid>. top -H shows the threads with their CPU usage in decimal; convert it to hexadecimal and look for it as nid= in the dump. That is how you identify exactly which Java thread is burning a core.
Tip 3: install a deadlock detector with ThreadMXBean. Thirty lines and a daemon thread turn a mysterious hang into a concrete log entry.
Tip 4: do not use wait/notify in new code. Use it to understand and to read other people's code. To write, use a BlockingQueue (08-06), a CountDownLatch or a Semaphore (08-05): they do the same thing, without the three classic mistakes of this section.
Exercises
Exercise 1: State observer
Write StateObserver which launches a thread called thread-under-observation whose work deliberately passes through at least four different states: TIMED_WAITING (sleeping), BLOCKED (on a lock main is holding), WAITING (on a wait() with no deadline) and TERMINATED. A second thread, a daemon called observer, must sample the first one's state every 50 ms and print only the state changes (not repeating the same state on consecutive lines), together with the time elapsed since start-up measured with System.nanoTime().
Exercise 2: Reservation queue with wait/notifyAll
Implement CoordinatedReservationQueue, a bounded queue of BiblioTech reservations using only synchronized, wait and notifyAll (nothing from java.util.concurrent). It must have:
- A fixed maximum capacity (5, for example).
put(Reservation r): if it is full, wait until there is room.take(): if it is empty, wait until there is something.- The correct
whileloop in both methods, andnotifyAllafter every modification. - Interruption support: propagate
InterruptedException.
Write a main with two producers and three consumers demonstrating that the queue never exceeds its capacity and that nobody gets null.
Exercise 3: Provoking and diagnosing a deadlock
Write DiagnosedDeadlock which:
- Deliberately provokes a deadlock between two threads called
bibliotech-loansandbibliotech-returns, which take theCATALOGandREGISTRYlocks in opposite order. - Launches a third daemon thread,
bibliotech-detector, which usesThreadMXBean.findDeadlockedThreads()to detect it, print a report with the thread names, the locks involved and who owns each one, and then end the program withSystem.exit(1). - Documents in a comment how you would obtain the same diagnosis with
jstackand which line you would look for.
Then write a second version, NoDeadlock, that solves the problem by imposing a global acquisition order on the locks, and check that the detector no longer finds anything.
Solutions
Solution to Exercise 1
import java.util.concurrent.TimeUnit;
public class StateObserver {
private static final Object lock = new Object();
private static final Object signal = new Object();
public static void main(String[] args) throws InterruptedException {
final long start = System.nanoTime();
Thread watched = new Thread(() -> {
try {
// 1) TIMED_WAITING
TimeUnit.MILLISECONDS.sleep(400);
// 2) BLOCKED: main is holding 'lock'
synchronized (lock) {
// (we barely do anything inside)
}
// 3) WAITING: indefinite wait until main signals
synchronized (signal) {
signal.wait();
}
// 4) TERMINATED on leaving run()
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}, "thread-under-observation");
Thread observer = new Thread(() -> {
Thread.State previous = null;
try {
while (true) {
Thread.State current = watched.getState();
// We only print the CHANGES: otherwise the output
// would be unreadable, hundreds of repeated lines.
if (current != previous) {
long ms = (System.nanoTime() - start) / 1_000_000;
System.out.printf("%5d ms %-16s%n", ms, current);
previous = current;
}
if (current == Thread.State.TERMINATED) {
return;
}
TimeUnit.MILLISECONDS.sleep(50);
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}, "observer");
// Daemon: it is informational, it must not prevent the JVM shutting down.
observer.setDaemon(true);
observer.start();
// NEW state visible for 200 ms before starting.
TimeUnit.MILLISECONDS.sleep(200);
watched.start();
// At 600 ms we take the lock for 500 ms: the watched thread,
// which arrives at about 400 ms, will end up BLOCKED.
TimeUnit.MILLISECONDS.sleep(400);
synchronized (lock) {
TimeUnit.MILLISECONDS.sleep(500);
}
// We let it enter wait() and then we wake it up.
TimeUnit.MILLISECONDS.sleep(400);
synchronized (signal) {
signal.notifyAll();
}
watched.join();
TimeUnit.MILLISECONDS.sleep(200); // margin for the last line
}
}Output:
0 ms NEW
201 ms RUNNABLE
252 ms TIMED_WAITING
604 ms BLOCKED
1104 ms RUNNABLE
1155 ms WAITING
1508 ms TERMINATEDThe six states in order, with their timings. Note the move from BLOCKED to WAITING: it goes through RUNNABLE at 1104 ms, exactly as the diagram in section 2 said. There are no direct transitions between waiting states.
Solution to Exercise 2
package com.nexussoftware.bibliotech.service;
import java.util.ArrayDeque;
import java.util.Deque;
import java.util.concurrent.TimeUnit;
/**
* Bounded reservation queue implemented with synchronized/wait/notifyAll.
*
* TEACHING NOTE: in production you would use an ArrayBlockingQueue (08-06),
* which does exactly this better and with no room for error. This class exists
* so you can see the mechanism from the inside.
*/
public class CoordinatedReservationQueue {
private final Deque<String> queue = new ArrayDeque<>();
private final int capacity;
public CoordinatedReservationQueue(int capacity) {
if (capacity <= 0) {
throw new IllegalArgumentException("capacity must be positive");
}
this.capacity = capacity;
}
/**
* Adds a reservation. If the queue is full, it waits.
* It propagates InterruptedException: the caller decides (08-02).
*/
public void put(String reservation) throws InterruptedException {
synchronized (queue) {
// WHILE, not if: on waking up we must RE-CHECK, because
// another producer may have filled the gap before us.
while (queue.size() == capacity) {
queue.wait();
}
queue.addLast(reservation);
// notifyAll and not notify: threads wait on this queue for TWO
// different conditions (full and empty). With notify we could
// wake a producer when the one that can make progress is
// a consumer, and lose the signal (section 9).
queue.notifyAll();
}
}
/** Removes a reservation. If the queue is empty, it waits. Never returns null. */
public String take() throws InterruptedException {
synchronized (queue) {
while (queue.isEmpty()) {
queue.wait();
}
String r = queue.pollFirst();
queue.notifyAll();
return r;
}
}
public int size() {
synchronized (queue) {
return queue.size();
}
}
// --- Demonstration ---
public static void main(String[] args) throws InterruptedException {
final int CAPACITY = 5;
final int PER_PRODUCER = 10;
CoordinatedReservationQueue queue = new CoordinatedReservationQueue(CAPACITY);
Runnable producer = () -> {
String me = Thread.currentThread().getName();
try {
for (int i = 1; i <= PER_PRODUCER; i++) {
queue.put(me + "-reservation-" + i);
System.out.printf(" [+] %-14s puts %2d (size=%d)%n",
me, i, queue.size());
TimeUnit.MILLISECONDS.sleep(30);
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
};
Runnable consumer = () -> {
String me = Thread.currentThread().getName();
try {
while (!Thread.currentThread().isInterrupted()) {
String r = queue.take();
// Without the 'while' in take(), this check would fail.
if (r == null) {
throw new IllegalStateException("null: coordination bug");
}
System.out.printf(" [-] %-14s takes %-28s (size=%d)%n",
me, r, queue.size());
TimeUnit.MILLISECONDS.sleep(70);
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
};
Thread p1 = new Thread(producer, "producer-1");
Thread p2 = new Thread(producer, "producer-2");
Thread c1 = new Thread(consumer, "consumer-1");
Thread c2 = new Thread(consumer, "consumer-2");
Thread c3 = new Thread(consumer, "consumer-3");
c1.start(); c2.start(); c3.start();
p1.start(); p2.start();
p1.join();
p2.join();
// We let the consumers drain it and then stop them with interrupt.
TimeUnit.MILLISECONDS.sleep(1000);
c1.interrupt(); c2.interrupt(); c3.interrupt();
c1.join(); c2.join(); c3.join();
System.out.println("Reservations pending at the end: " + queue.size());
}
}Output (fragment):
[+] producer-1 puts 1 (size=1)
[+] producer-2 puts 1 (size=1)
[-] consumer-1 takes producer-1-reservation-1 (size=1)
[-] consumer-2 takes producer-2-reservation-1 (size=0)
[+] producer-1 puts 2 (size=1)
...
Reservations pending at the end: 0Three teaching points:
- The size never exceeds 5. The
while (queue.size() == capacity)condition inputguarantees it. If the producers were much faster, they would block instead of letting the queue grow without limit: that is backpressure, and it is a hugely valuable property that is revisited withBlockingQueuein 08-06. - No consumer gets
null. Thewhileintakeguarantees it. Change thatwhilefor anifand run: with three consumers andnotifyAll, theIllegalStateExceptionfires within seconds. - The printed size may not match the operation.
System.out.printfruns outside the synchronized block, so between putting and reading the size another thread may have acted. It is not a bug in the queue: it is the demonstration that a read outside the critical section is only a snapshot of an instant already past. It is exactly the "check-then-act" problem of 08-06.
Solution to Exercise 3
import java.lang.management.ManagementFactory;
import java.lang.management.ThreadInfo;
import java.lang.management.ThreadMXBean;
import java.util.concurrent.TimeUnit;
/**
* Provokes a deadlock and diagnoses it from inside the program.
*
* EQUIVALENT EXTERNAL DIAGNOSIS:
* 1) jps -l -> locate the pid
* 2) jstack <pid> > dump.txt -> dump the threads
* 3) look in dump.txt for the line "Found one Java-level deadlock:"
* Below it appear the two threads, the lock each one waits for
* and who owns it, plus the traces with "- locked" and
* "- waiting to lock" for the same pair of addresses in reverse order.
*/
public class DiagnosedDeadlock {
// Two locks representing two domain resources.
private static final Object CATALOG = new Object();
private static final Object REGISTRY = new Object();
public static void main(String[] args) {
startDetector();
// Thread A: CATALOG -> REGISTRY
new Thread(() -> {
synchronized (CATALOG) {
pause(300); // gives the other thread time
synchronized (REGISTRY) { // it ends up BLOCKED here
System.out.println("A completed (never happens)");
}
}
}, "bibliotech-loans").start();
// Thread B: REGISTRY -> CATALOG <-- REVERSE ORDER: the root cause
new Thread(() -> {
synchronized (REGISTRY) {
pause(300);
synchronized (CATALOG) { // it ends up BLOCKED here
System.out.println("B completed (never happens)");
}
}
}, "bibliotech-returns").start();
}
/** Daemon watchdog that detects deadlocks and aborts with a report. */
private static void startDetector() {
Thread detector = new Thread(() -> {
ThreadMXBean mx = ManagementFactory.getThreadMXBean();
while (!Thread.currentThread().isInterrupted()) {
long[] ids = mx.findDeadlockedThreads(); // null if there is none
if (ids != null) {
System.err.println();
System.err.println("========================================");
System.err.println(" DEADLOCK DETECTED (" + ids.length + " threads)");
System.err.println("========================================");
// true, true: include the monitors and synchronisers held
for (ThreadInfo info : mx.getThreadInfo(ids, true, true)) {
System.err.printf("%nThread : %s (%s)%n",
info.getThreadName(), info.getThreadState());
System.err.printf("Waits : %s%n", info.getLockName());
System.err.printf("Held by: %s%n", info.getLockOwnerName());
System.err.println("Trace (first 2 lines):");
StackTraceElement[] trace = info.getStackTrace();
for (int i = 0; i < Math.min(2, trace.length); i++) {
System.err.println(" at " + trace[i]);
}
}
System.err.println();
System.err.println("CAUSE: the two threads acquire the same two");
System.err.println("locks in OPPOSITE ORDER. Solution in 08-04:");
System.err.println("impose a global acquisition order.");
System.exit(1);
}
pause(500);
}
}, "bibliotech-detector");
detector.setDaemon(true); // it must not prevent shutdown
detector.start();
}
private static void pause(long ms) {
try {
TimeUnit.MILLISECONDS.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
}Output:
========================================
DEADLOCK DETECTED (2 threads)
========================================
Thread : bibliotech-loans (BLOCKED)
Waits : java.lang.Object@5b6f7412
Held by: bibliotech-returns
Trace (first 2 lines):
at DiagnosedDeadlock.lambda$main$0(DiagnosedDeadlock.java:31)
at java.base/java.lang.Thread.run(Thread.java:1583)
Thread : bibliotech-returns (BLOCKED)
Waits : java.lang.Object@2d1ef81a
Held by: bibliotech-loans
Trace (first 2 lines):
at DiagnosedDeadlock.lambda$main$1(DiagnosedDeadlock.java:42)
at java.base/java.lang.Thread.run(Thread.java:1583)
CAUSE: the two threads acquire the same two locks in OPPOSITE ORDER.
Solution in 08-04: impose a global acquisition order.And the corrected version, with a global acquisition order:
/**
* Same functionality, with no deadlock possible.
*
* RULE: ALL the threads acquire the locks in the SAME order,
* always CATALOG before REGISTRY. With a total ordering over the
* locks no waiting cycle can form, and with no cycle there is no
* deadlock. It is the solution developed in 08-04.
*/
public class NoDeadlock {
private static final Object CATALOG = new Object();
private static final Object REGISTRY = new Object();
public static void main(String[] args) throws InterruptedException {
Runnable lend = () -> {
synchronized (CATALOG) { // 1st CATALOG
pause(300);
synchronized (REGISTRY) { // 2nd REGISTRY
System.out.println("[" + Thread.currentThread().getName()
+ "] loan registered");
}
}
};
Runnable returnItem = () -> {
synchronized (CATALOG) { // 1st CATALOG (it was REGISTRY before)
pause(300);
synchronized (REGISTRY) { // 2nd REGISTRY
System.out.println("[" + Thread.currentThread().getName()
+ "] return registered");
}
}
};
Thread a = new Thread(lend, "bibliotech-loans");
Thread b = new Thread(returnItem, "bibliotech-returns");
a.start(); b.start();
a.join(); b.join();
System.out.println("Both threads have finished: there is NO deadlock");
}
private static void pause(long ms) {
try { TimeUnit.MILLISECONDS.sleep(ms); }
catch (InterruptedException e) { Thread.currentThread().interrupt(); }
}
}Output:
[bibliotech-loans] loan registered
[bibliotech-returns] return registered
Both threads have finished: there is NO deadlockThe change is a single line —swapping the order of the two synchronized blocks in returnItem— and it completely eliminates a class of failure that can take down a production application without leaving a trace. It is the best cost/benefit ratio in the whole module, and in 08-04 it becomes an explicit design rule.
Conclusion
You no longer see a thread as a black box that starts and finishes: you know exactly what it can be doing and what moves it from one place to another.
The six Thread.State values and their transitions: NEW before start(), RUNNABLE when it is running or ready to run, BLOCKED waiting for a monitor, WAITING waiting for a signal, TIMED_WAITING waiting with a deadline, and TERMINATED as a terminal, irreversible state —which is the deep reason a Thread cannot be reused and pools exist—. With the structural observation that organises the diagram: every transition goes through RUNNABLE; a thread waking from wait() has to reacquire the monitor, and at that moment it may end up BLOCKED.
You know the two nuances that separate someone who can read a production system from someone who cannot. The first: RUNNABLE does not mean "consuming CPU", because it groups running, waiting for a turn and —this always surprises— being blocked on classic I/O; you have to look at the stack trace and the cpu= field, not the state. The second: BLOCKED is "they will not let me in" and WAITING is "I am waiting to be told", contention versus coordination, and that distinction turns an unreadable dump into a diagnosis.
You have mastered the low-level coordination mechanism: wait, notify and notifyAll live on Object, they are always called holding the monitor —or IllegalMonitorStateException—, and wait() releases the monitor whereas sleep() does not, which is the difference that defines what each is for. And you have the most broken rule of all: wait() always goes inside a while, never inside an if, for three independent reasons —spurious wake-ups, notifyAll with several waiters and a single element, and conditions that change between the signal and the wake-up—, with the idea that sums them up: notify does not say "your condition holds", it says "look again". And the derived recommendation: notifyAll by default, because notify with threads waiting for different conditions produces lost signals and silent standstills.
You know why stop, suspend and resume have been withdrawn since 1998: the first releases all its monitors in the middle of an operation and leaves structures corrupt without giving any error; the other two hold the monitors while the thread is paused, which produces perfect deadlocks. The alternative is the same idea in both cases: the thread chooses where it stops, which is the cooperative cancellation of 08-02 and the explicit pause point with wait/notify.
And, above all, you know how to diagnose. Obtain a dump with Ctrl+\, jstack or jcmd; read an entry field by field —name, state, cpu=, - locked and - waiting to lock, trace—; recognise the four usual profiles —working, waiting on I/O, idle pool thread, blocked—; and find the Found one Java-level deadlock the JVM writes for you, with the waiting cycle and the exact lines of code. With the operational rule that prevents false conclusions: three dumps a few seconds apart, because a thread may be passing through but it cannot be passing through the same line three times. And with Thread.getAllStackTraces and ThreadMXBean.findDeadlockedThreads() to diagnose from the inside, plus jconsole, VisualVM and JFR to do it with a graphical interface.
BiblioTech can now be observed. ThreadMonitor shows the state of the import and notices threads while they work, and that output teaches something that until now was theory: the importer alternates RUNNABLE and TIMED_WAITING —a mixed task—, while the notices thread is almost always in TIMED_WAITING —an I/O-bound task—. It is the empirical confirmation of the 08-01 table and the justification for everything that follows.
Because now you know how to observe the problem, and the time has come to solve it. You have seen a real deadlock, you have seen three consumers get null because of a misplaced if, and you still have the counter that lost half a million increments outstanding from 08-01.
In the next lesson, Synchronization —the central lesson of the module—, all of that is solved. You will see why counter++ is actually three operations and which exact interleaving loses the increments; the Java memory model, with its per-core caches, instruction reordering and the happens-before relationship that governs everything; what volatile guarantees —visibility— and what it does not —atomicity—, with the two examples that prove it; synchronized in all its forms, the intrinsic monitor, the private lock and the right granularity; the reproducible deadlock and its two solutions —global ordering and tryLock with a deadline—; ReentrantLock with its Conditions, ReadWriteLock for structures that are read a lot and written rarely, and the best strategy of all: do not share. By the end, BiblioTech's Catalog and LoanRegistry will be thread-safe, and case C of 08-01 —two employees working at the same time— will stop being a threat.
Java Programming Course
Module 1: Introduction to Java
- Introduction to Java
- Setting Up the Development Environment
- Basic Syntax and Structure
- Variables and Data Types
- Operators
- Console Input and Output
- Your First Complete Program: BiblioTech
Module 2: Control Flow
- Conditional Statements
- Loops
- Switch Statements
- Break and Continue
- Debugging and Execution Traces
- Project: The BiblioTech Interactive Menu
Module 3: Object-Oriented Programming
- Introduction to OOP
- Classes and Objects
- Methods
- Constructors
- Inheritance
- Polymorphism
- Encapsulation
- Abstraction
- The Object Class: equals, hashCode and toString
Module 4: Advanced Object-Oriented Programming
- Interfaces
- Abstract Classes
- Inner Classes
- Anonymous Classes
- Lambda Expressions
- Functional Interfaces and Method References
- Enums and Records
Module 5: Data Structures and Collections
- Arrays
- The Collections Framework
- ArrayList
- LinkedList
- HashMap
- HashSet
- Queue and Deque
- Stack
- Sorting and Searching Collections
Module 6: Exception Handling
- Introduction to Exceptions
- The Try-Catch Block
- Throw and Throws
- Custom Exceptions
- The Finally Block
- Try-with-resources and AutoCloseable
- Error Handling Strategies and Logging
Module 7: File Input/Output
- Reading Files
- Writing Files
- File Streams
- BufferedReader and BufferedWriter
- Serialization
- The NIO.2 API: Path and Files
- Interchange Formats: CSV and Properties
Module 8: Multithreading and Concurrency
- Introduction to Multithreading
- Creating Threads
- Thread Lifecycle
- Synchronization
- Concurrency Utilities
- Concurrent Collections and Atomic Variables
- Asynchronous Tasks with CompletableFuture
Module 9: Networking
- Introduction to Networking
- Sockets
- ServerSocket
- DatagramSocket and DatagramPacket
- URL and HttpURLConnection
- The Modern HTTP Client
Module 10: Advanced Topics
- Generics
- Annotations
- Reflection
- Java 8 Features: Streams and Optional
- Dates and Times with java.time
- Java 9 and Beyond
- Memory, Garbage Collection and Performance
Module 11: Java Frameworks and Libraries
- Introduction to Java Frameworks
- Spring Framework
- Hibernate
- JUnit
- Maven
- Advanced Testing with Mockito
- Essential Ecosystem Libraries
