An order service stops responding. Not slowly: requests that were being served finish, and then nothing new completes. CPU is near zero. Memory is flat. No exceptions in the logs. The health check endpoint times out, so Kubernetes restarts the pod, and it works fine for another forty minutes.
A thread dump taken before the restart shows sixteen threads in the same pool, all parked in CompletableFuture.join, all waiting for tasks that are sitting in that pool’s queue.
There is nobody left to run them. The threads that would have executed those tasks are the threads waiting for them.
The shape of the deadlock
ExecutorService pool = Executors.newFixedThreadPool(16);
public OrderSummary buildSummary(Long orderId) {
return CompletableFuture.supplyAsync(() -> {
Order order = loadOrder(orderId);
// Child work submitted to the SAME pool, then waited on.
CompletableFuture<List<Item>> items =
CompletableFuture.supplyAsync(() -> loadItems(orderId), pool);
CompletableFuture<Customer> customer =
CompletableFuture.supplyAsync(() -> loadCustomer(order.customerId()), pool);
// This thread now blocks, holding a pool slot, waiting for work
// that needs a pool slot.
return new OrderSummary(order, items.join(), customer.join());
}, pool).join();
}
With light traffic this works. A parent occupies one thread, two children take two more, everything completes, threads are returned.
With sixteen concurrent requests it stops forever.
flowchart TB
subgraph pool["Pool: 16 threads, all occupied"]
T1["Thread 1: parent, join()"]
T2["Thread 2: parent, join()"]
T16["Thread 16: parent, join()"]
end
subgraph queue["Queue: children waiting"]
C1["loadItems"]
C2["loadCustomer"]
C3["... 30 more"]
end
T1 -.->|"waiting for"| C1
T2 -.->|"waiting for"| C2
queue -.->|"needs a free thread"| pool
style pool stroke:#ef4444,stroke-width:3px,color:#fff
Every thread is busy. Utilisation reads 100 percent. Throughput is zero. Those two facts together are the signature, and they are why a dashboard showing thread pool utilisation looks like a capacity problem when it is actually a structural one.
Nothing detects it. join() has no timeout. The pool has no notion of dependency. The JVM’s deadlock detector only finds monitor cycles, not this. It sits there until something external kills the process.
Why a bigger pool is not a fix
The instinct is to raise the size, and it does make the symptom rarer, which is the worst possible property for a fix because it converts a reproducible failure into an intermittent one.
The condition for deadlock is simply enough concurrent parents to fill the pool. Sixteen threads needs sixteen parents. Two hundred threads needs two hundred parents. Under a traffic spike, or a slow downstream that makes each parent hold its thread longer, you get there.
The arithmetic is the same one that sizes any pool: concurrency equals arrival rate times holding time, and blocking on a child increases the holding time enormously. That is the queueing relationship that decides pool size generally, and here it is being applied to a pool whose own work is the thing inflating the holding time.
Raising the number buys you a larger traffic spike before it happens. It does not remove the cycle.
The common pool makes it a shared problem
The version of this I find genuinely nasty is that you can hit it without configuring any pool at all.
// No executor argument. Runs on ForkJoinPool.commonPool().
CompletableFuture.supplyAsync(() -> callService())
.thenApply(this::transform)
.join();
ForkJoinPool.commonPool() defaults to one fewer thread than your core count. On a 4 core container that is three threads, shared across every CompletableFuture without an explicit executor, every parallel stream, and any library that made the same default choice.
Three threads. Shared with code you did not write. A parallel stream in a dependency and a blocking supplyAsync in your handler are now competing for the same three slots, and neither author knew about the other.
The fix at this level is a rule rather than a tuning exercise: always pass an explicit executor. Every supplyAsync, thenApplyAsync, thenComposeAsync has an overload that takes one, and using it makes the dependency visible and bounded.
Composing instead of blocking
The structural answer is to stop blocking inside tasks at all.
public CompletableFuture<OrderSummary> buildSummary(Long orderId) {
return CompletableFuture
.supplyAsync(() -> loadOrder(orderId), ioPool)
.thenCompose(order -> {
CompletableFuture<List<Item>> items =
CompletableFuture.supplyAsync(() -> loadItems(orderId), ioPool);
CompletableFuture<Customer> customer =
CompletableFuture.supplyAsync(() -> loadCustomer(order.customerId()), ioPool);
// No join. The continuation runs when both complete,
// and no thread is held waiting in the meantime.
return items.thenCombine(customer,
(i, c) -> new OrderSummary(order, i, c));
});
}
Nothing blocks. When loadItems finishes it schedules the continuation, and until then no thread is occupied on its behalf. The deadlock is not mitigated, it is impossible, because there is no thread waiting on a thread.
The cost is that the method now returns a future and the caller has to deal with it, which propagates up until something at the boundary joins. That boundary is the right place for it: a servlet container thread or a controller returning CompletableFuture to the framework, where blocking is somebody else’s pool and the framework is designed for it.
Isolating the blocking that cannot be removed
Some calls are synchronous and are not going to change. A JDBC driver blocks. A legacy SDK blocks. That is fine, provided the blocking is confined.
// Separate pools. A parent in coordinationPool waiting on work submitted
// to blockingIoPool can never be waiting for its own thread.
ExecutorService coordinationPool = Executors.newFixedThreadPool(8);
ExecutorService blockingIoPool = new ThreadPoolExecutor(
32, 32, 0L, TimeUnit.MILLISECONDS,
new LinkedBlockingQueue<>(500), // bounded, deliberately
new ThreadPoolExecutor.AbortPolicy());
The bounded queue matters as much as the separation. An unbounded queue in front of the blocking pool turns saturation into unbounded latency instead of a fast rejection, which is the failure mode that makes overload unrecoverable.
Timeouts on every wait convert a permanent hang into an error you can see:
future.orTimeout(2, TimeUnit.SECONDS)
.exceptionally(ex -> fallbackSummary());
orTimeout arrived in Java 9 and is worth applying by default. A hang that becomes a timeout is a monitored failure rather than a silent one.
Virtual threads change the arithmetic here
Project Loom addresses this directly. A virtual thread that blocks unmounts from its carrier thread, so the carrier is free to run something else. The parent waiting on a child no longer occupies a platform thread, and the cycle cannot form.
That genuinely removes this failure for blocking I/O, and it does not remove it for CPU bound work or for blocking inside a synchronized block in older builds, which pins the carrier. It changes the calculus for blocking I/O specifically, and it is not a general fix for concurrency mistakes.
How to find it before it finds you
Take a thread dump when a service stops responding, before restarting it. jstack output showing many threads in WAITING on CompletableFuture internals, from the same pool, is diagnostic in seconds. The restart destroys the evidence, so the instinct to restart first is the thing that makes this take three incidents to identify.
Grep for supplyAsync, thenApplyAsync and runAsync without an executor argument. Each one is using the common pool and sharing fate with every library in the process.
Grep for .join() and .get() inside lambdas passed to an executor. That combination is the pattern, and it is easy to search for.
Export queue depth alongside active thread count. A pool at full utilisation with a growing queue and no completions is not busy. It is stuck, and those two metrics side by side make the difference obvious.
// SPONSORSHIP
If this research saved you time or improved your architecture, consider sponsoring my work on GitHub. All sponsorships go directly toward infrastructure and further technical research.
[ Become a Sponsor ]