A Java service is configured with -Xmx2g. The container limit is 2.5GB, which somebody chose as a comfortable margin. The heap usage graph oscillates between 800MB and 1.6GB, never approaching the limit.
The pod is killed roughly twice a day with exit code 137.
There is no OutOfMemoryError in the logs. No stack trace, no shutdown hook output, no final log line. The process simply stops mid-sentence, and Kubernetes reports OOMKilled.
The heap was never the problem. It was doing exactly what it was told.
Two different things are called OOM
The word gets used for two failures that are unrelated and behave completely differently.
A Java OutOfMemoryError means the JVM tried to allocate inside the heap, could not, and threw. It is an exception. Handlers run, logs are written, a heap dump can be produced, and there is a stack trace pointing at the allocation.
A kernel OOM kill means the whole process exceeded its cgroup memory limit and was sent SIGKILL. There is no exception because there is no Java involved. SIGKILL cannot be caught, blocked or handled, so nothing runs afterwards.
The absence of a stack trace is the diagnostic. If the logs simply stop and the exit code is 137, the JVM did not decide anything. The kernel did.
Everything the JVM allocates that is not the heap
-Xmx bounds one region. The container limit covers the process.
flowchart TB
LIMIT["Container limit: 2.5GB"] --> HEAP["Java heap: -Xmx2g"]
LIMIT --> META["Metaspace<br/>class metadata, ~200MB"]
LIMIT --> CODE["JIT code cache<br/>up to 240MB reserved"]
LIMIT --> STACK["Thread stacks<br/>~1MB each"]
LIMIT --> DIRECT["Direct byte buffers<br/>Netty, NIO"]
LIMIT --> GCM["GC structures<br/>~1% of heap for G1"]
LIMIT --> NATIVE["Native allocations<br/>inside dependencies"]
style HEAP stroke:#4ade80,stroke-width:2px,color:#fff
style DIRECT stroke:#ef4444,stroke-width:3px,color:#fff
style NATIVE stroke:#ef4444,stroke-width:3px,color:#fff
Metaspace holds class metadata and grows with the number of loaded classes. A Spring application with a large dependency tree loads tens of thousands of classes, and a couple of hundred megabytes is unremarkable. It is unbounded by default, which means a classloader leak can consume the container without touching the heap at all.
Thread stacks are around 1MB each by default. Three hundred threads is 300MB of address space, and a service with a large servlet pool plus various client library threads gets there easily.
The JIT code cache reserves up to 240MB by default. It only commits what it uses, and on a large application it uses a lot.
Direct byte buffers are the ones that catch people. Netty, NIO and most modern HTTP clients allocate off-heap buffers because it avoids a copy when writing to a socket. They are freed when the referencing Java object is collected, which means their lifetime is tied to garbage collection timing rather than to when you finished with them. A burst of traffic can allocate a lot of direct memory that is not reclaimed until a collection happens, and if -XX:MaxDirectMemorySize is unset it defaults to roughly the max heap, which means a 2GB heap permits another 2GB of direct buffers.
Native allocations inside dependencies are invisible to everything Java side. Compression libraries, cryptographic providers, database drivers with native components, and anything using JNI allocate through malloc and appear only as resident memory.
The rough guidance I use is that a JVM needs 25 to 50 percent above -Xmx in container memory, and the correct number is measured rather than assumed.
Measuring instead of guessing
Native Memory Tracking is built in and gives the breakdown directly.
-XX:NativeMemoryTracking=summary
jcmd <pid> VM.native_memory summary
The output splits reserved and committed memory by category: Java heap, class, thread, code, GC, compiler, internal, symbol. The committed column is what actually counts against the container.
The important caveat is that this accounts for memory the JVM allocated. A native library calling malloc directly does not appear, and if NMT totals are well below RSS, that gap is where the problem lives. Running with jemalloc and its profiling enabled is the next step there, and it is a real amount of effort, which is why it is worth ruling out everything else first.
For direct buffers specifically, the JMX bean is easier:
BufferPoolMXBean direct = ManagementFactory
.getPlatformMXBeans(BufferPoolMXBean.class).stream()
.filter(b -> b.getName().equals("direct"))
.findFirst().orElseThrow();
// Worth exporting as a metric. It is invisible on every heap graph
// and it is the most common cause of a container OOM in a JVM service.
log.info("direct buffers: count={} memoryUsed={}MB",
direct.getCount(), direct.getMemoryUsed() / (1024 * 1024));
Configuring it so the JVM knows about the limit
Modern JVMs are container aware and read the cgroup limit, which is the behaviour you want.
-XX:MaxRAMPercentage=60.0
-XX:MaxDirectMemorySize=256m
-XX:MaxMetaspaceSize=512m
-XX:NativeMemoryTracking=summary
MaxRAMPercentage sizes the heap as a fraction of the container limit rather than as an absolute value, so changing the pod limit resizes the heap automatically instead of silently leaving the two inconsistent. Sixty percent is a starting point for a service with meaningful off-heap usage, and 75 is reasonable for one without.
Bounding direct memory and metaspace explicitly converts a container kill into a Java exception. That is a large improvement in diagnosability: an OutOfMemoryError: Direct buffer memory tells you exactly what happened, while exit code 137 tells you nothing. I would set both even if the values are generous, purely for the error message.
Setting the request equal to the limit is worth doing too, because a pod whose request is lower is in a burstable QoS class and gets evicted before guaranteed pods when the node comes under pressure. That produces kills that have nothing to do with your service’s own usage.
The relationship to RSS
None of this is JVM specific at the bottom. The kernel is counting resident pages, which is the same accounting that decides why RSS matters more than what your allocator reported. The JVM reserving 240MB of code cache address space costs nothing until pages are touched; the kill happens when they are.
What the JVM adds is a layer of indirection that makes the number people watch the wrong one. The heap graph is prominent, well instrumented, and covers maybe 70 percent of the process. The remaining 30 percent has no dashboard by default and is where the failure comes from.
What I would check today
Compare container RSS against heap usage on any JVM service. If the gap is more than about 40 percent of the heap size, something off-heap is worth identifying before it becomes an incident.
Export direct buffer usage as a metric, because it is the most common culprit and the least visible.
Bound metaspace and direct memory explicitly so failures arrive as exceptions rather than as a process that vanishes.
And treat exit code 137 as a distinct failure class from OutOfMemoryError. They have different causes, different evidence, and different fixes, and conflating them sends the investigation into the heap dump for a problem that was never on the heap.
// SPONSORSHIP
If this research saved you time or improved your architecture, consider sponsoring my work on GitHub. All sponsorships go directly toward infrastructure and further technical research.
[ Become a Sponsor ]