JVM Internals: Memory, GC & Class Loading
A guide to the Java Virtual Machine β runtime memory areas, garbage collection algorithms and collectors, class loading, and monitoring tools.
1. JVM Architecture Overview
The HotSpot JVM splits work across runtime data areas and an execution engine. Shared across all threads: the Heap (objects and arrays, GC-managed) and the Method Area (class metadata / bytecode β Metaspace since Java 8). Private per thread: the JVM Stack (frames, locals, operand stack), the PC Register (address of the current bytecode instruction), and the Native Method Stack (JNI / C frames). The execution engine runs bytecode via the interpreter and promotes hot methods with the JIT (C1/C2) into the Code Cache; the GC reclaims unreachable heap objects.
Click a block in the diagram for responsibilities, tuning flags, and common failure modes.
ClassLoader Subsystem
Overview: Responsible for dynamically loading, linking, and initializing Java class files (.class) at runtime.
- Key Responsibilities:
- Loading: Reads binary bytecode data and creates a Class object.
- Linking: Verifies bytecode safety, prepares static fields with default values, and resolves symbolic references.
- Initialization: Executes class initializers (<clinit>) and assigns static variables to actual values.
- JVM configuration flags: -verbose:class (prints loaded classes) | -Xbootclasspath (modifies bootstrap loader path)
- Failure modes:
- java.lang.ClassNotFoundException
- java.lang.NoClassDefFoundError
π‘ Click on any component (ClassLoader, Heap, JIT Compiler, VM Stack, etc.) in the diagram to inspect its execution parameters.
2. On-Heap vs. Off-Heap Memory Layout
The memory footprint of a Java Virtual Machine (JVM) process is split into two primary areas at the operating system level: On-Heap Memory and Off-Heap (Native) Memory.
JVM Memory Architecture
Below is the complete memory layout of a JVM process, illustrating how physical RAM is partitioned between the garbage-collected Heap and the Native (Off-Heap) areas.
Eden Space (Young Generation)
Overview: The entry point for newly created objects allocated by the application threads.
- Tuning flags: -Xms / -Xmx (Heap size bounds) | -XX:NewRatio (Young to Old ratio) | -XX:SurvivorRatio (Eden to Survivor size ratio)
- GC Management: Extremely fast Minor GCs. Garbage collection runs frequently, evacuating surviving objects to Survivor spaces and reclaiming dead objects instantly (pointer-bump reset).
- OOM Risks & Failure modes:
- Causes frequent Minor GCs if sized too small, leading to high CPU latency.
- Contributes to OutOfMemoryError: Java heap space if Old Gen is also full.
- Under the Hood Details:
- Objects are allocated using Thread Local Allocation Buffers (TLABs) to avoid thread synchronization overhead.
- A pointer-bumping allocator is used: allocating an object simply moves the allocation pointer forward, which is extremely cheap.
- Sizing Eden correctly is critical: too small causes excessive GC frequency; too large increases GC pause durations.
π‘ Click on any partition (Eden, Survivor, Old Space, Metaspace, Stacks, Direct Memory, etc.) in the diagram to inspect its parameters.
Detailed Side-by-Side Comparison
| Feature | On-Heap Memory | Off-Heap (Native) Memory |
|---|---|---|
| Location | Inside the Java Heap (managed by the JVM). | Outside the Java Heap, in the OS virtual memory. |
| Garbage Collection | Fully managed by JVM Garbage Collectors (G1, ZGC, Parallel, etc.). | Not managed by GC. Must be released manually or via cleanups/cleaner wrappers. |
| Allocation Cost | Very low (pointer bumping or free-list allocation). | High (requires native OS memory allocation system calls like malloc). |
| Access Speed | Extremely fast (direct reference access). | Slightly slower due to JNI boundary or serialization overhead if copying to Heap. |
| Zero-Copy I/O | Impossible directly. Must be copied to off-heap before OS can access. | Supported. OS kernel can perform DMA (Direct Memory Access) directly. |
| Typical Use Cases | Standard Java objects, variables, collections, application state. | Large data caches (to avoid GC overhead), Netty network buffers, JIT compilation. |
| Common Failure Mode | java.lang.OutOfMemoryError: Java heap space | java.lang.OutOfMemoryError: Metaspace or process killed by OS (OOM Killer). |
| Key Tuning Flags | -Xms, -Xmx, -Xmn, -XX:NewRatio | -XX:MaxMetaspaceSize, -XX:MaxDirectMemorySize, -Xss |
The JVM Process Memory Equation (RAM Sizing)
When configuring a container (such as a Docker container or Kubernetes Pod memory limit), you must budget for the entire OS-level Resident Set Size (RSS) of the JVM process, not just the Java Heap. Sizing a container to match only -Xmx will lead to the process being terminated by the OS Out-of-Memory (OOM) Killer.
The total memory consumed by a JVM process at the OS level is calculated using the following equation:
Breakdown of the Equation:β
- Heap (
-Xms/-Xmx):- What it stores: All Java object instances and arrays.
- Sizing impact: This is usually the largest component. If it reaches
-Xmxand cannot reclaim space, it throwsjava.lang.OutOfMemoryError: Java heap space.
- Metaspace (
-XX:MaxMetaspaceSize):- What it stores: Class definitions, method bytecode, annotations, and the constant pool.
- Sizing impact: Scaled dynamically based on the number of loaded classes. Unbounded by default; always set a limit to detect classloader leaks.
- Thread Stacks (
-Xss* Thread Count):- What it stores: Local variables, active method frames, and execution state for every running thread.
- Sizing impact: Default is typically
1MBper thread. If you have 500 active threads, they will consume500MBof native RAM just for stack allocations.
- Code Cache (
-XX:ReservedCodeCacheSize):- What it stores: JIT-compiled native machine code (compiled from bytecode by C1/C2 compilers for hot methods).
- Sizing impact: Typically defaults to
240MB. If full, JIT compilation stops, and performance degrades severely.
- Direct Memory (
-XX:MaxDirectMemorySize):- What it stores: Native/off-heap buffers allocated via
ByteBuffer.allocateDirect()(heavily used by network libraries like Netty and frameworks like gRPC/Kafka for zero-copy I/O). - Sizing impact: Defaults to matching the maximum heap size (
-Xmx) if not explicitly capped.
- What it stores: Native/off-heap buffers allocated via
- JVM Overhead (GC, JIT, C++ Heap, Symbol Tables, etc.):
- What it stores: Memory used by the JVM's internal C++ engines. This includes:
- Garbage Collector metadata: Card tables, mark bitmaps, and remembered sets (RSets) used to track references. G1 and ZGC can require an overhead of 10% to 20% of the heap size just for internal metadata.
- JIT Compiler queues: Native memory used by compiler threads during optimization.
- Symbol Tables: Interned strings and method/field symbols.
- JNI/Native Allocations: Memory allocated by custom C/C++ libraries loaded into the JVM.
- OS Page Cache / JVM Overhead: Miscellaneous OS overhead, process control blocks, etc.
- What it stores: Memory used by the JVM's internal C++ engines. This includes:
β οΈ Kubernetes/Docker Container Sizing Heuristicβ
Always leave a non-heap headroom buffer of at least 25% to 30% of the total container memory.
If you set your Kubernetes resource limit to 4GB, your -Xmx should not exceed 3GB.
On-Heap Memory Components
Heap (Shared, GC-managed)β
πΆ Beginner Concept: The "Warehouse and the Desk"β
- The Heap (The Warehouse): This is a massive, shared storage facility where every object you create (
new User(),new ArrayList()) permanently lives. It is huge, fully shared by all threads, but requires a Garbage Collector janitor to clean up abandoned items. - The Stack (The Desk): Every thread gets its own tiny, private working desk. You cannot put a giant
ArrayListon the desk. You can only put tiny primitives (int,boolean) and Remote Controls (Pointers/References) on the desk. When a method finishes, the entire desk is instantly wiped clean.
The largest memory area. Stores all object instances and arrays. Divided into generations for GC efficiency:
Heap Memory (GC-Managed Shared Area)
Capacity Ratio: 100% of Heap Allocation
Overview: The primary memory workspace of the JVM where all objects, class instances, and array payloads are allocated.
- GC Role & Management: Scanned continuously by minor, major, and mixed GC threads to recover unreachable blocks.
- JVM Sizing Flags:
- -Xms (Initial Heap size)
- -Xmx (Maximum Heap size)
- -XX:MinHeapFreeRatio
- -XX:MaxHeapFreeRatio
π‘ Click on any segment (Heap, Young Gen, Eden, S0, S1, or Old Gen) in the diagram above to inspect JVM memory sizing parameters.
- Eden: New objects are allocated here.
- Survivors: Objects that survive a minor GC move between S0 and S1.
- Old Generation: Long-lived objects promoted from young gen after surviving multiple GC cycles (default threshold: 15).
Off-Heap (Native) Memory Components
Metaspace (Shared)β
Stores class metadata, static variables, constant pool, and compiled code.
- JDK 7 and earlier: PermGen (permanent generation) β fixed size, prone to
OutOfMemoryError: PermGen space - JDK 8+: Metaspace β stored in native memory (not heap), grows dynamically
# Configuration flags
-XX:MetaspaceSize=256m -XX:MaxMetaspaceSize=512m
VM Stack (Per-Thread)β
Each thread has its own stack. Each method call creates a stack frame containing:
- Local variable array β method parameters and local variables
- Operand stack β intermediate computation values
- Frame data β constant pool reference, return address
π§ Senior Deep Dive: Escape Analysis & Scalar Replacementβ
Seniors know a critical JVM optimization: Objects do NOT always go to the Heap. Since Java 1.6, the JIT Compiler runs Escape Analysis. If the compiler proves that an object created inside a method never "escapes" that method (it isn't returned, nor passed to another thread), it performs Scalar Replacement, allocating its fields directly on the stack or CPU registers.
For complete code examples, optimizations (like Lock Elimination and Lock Coarsening), and mechanics, see the Escape Analysis Deep Dive in Stack vs. Heap Memory.
Errors:
StackOverflowErrorβ too many nested calls (e.g., infinite recursion)OutOfMemoryErrorβ cannot allocate new thread stacks
# Configuration flag (per thread stack size)
-Xss1m
Program Counter (Per-Thread)β
A small memory area holding the address of the current bytecode instruction being executed. Undefined for native methods.
Native Method Stack (Per-Thread)β
Similar to the VM stack but for native (JNI) methods. HotSpot JVM combines native method stack and VM stack.
Direct Memory (Buffer Pool)β
Allows Java applications to bypass the garbage collector and allocate memory directly from the operating system's native memory pool.
- Allocation: Allocated via
ByteBuffer.allocateDirect(capacity)or via third-party libraries usingsun.misc.Unsafe. - Performance Advantage (Zero-Copy): When writing/reading data from disk or sockets, standard Java heap buffers (
byte[]) require the JVM to copy the data into an intermediate native buffer first because the GC could move the heap array in physical memory at any point. Direct memory buffers are pinned (they do not move), allowing the OS kernel to read/write directly from them via Direct Memory Access (DMA), avoiding copy cycles. - Deallocation: Freed using internal Cleaners (which use phantom references) when the heap wrapper is garbage collected, or manually using
unsafe.freeMemory().
# Maximum limit of direct memory allocation
-XX:MaxDirectMemorySize=2g
Code Cacheβ
Used by the JIT compiler to store compiled native machine code. If the Code Cache becomes full, the JIT compiler is disabled, and code runs in interpreted mode, leading to massive performance degradation.
# Code Cache sizing
-XX:InitialCodeCacheSize=24m
-XX:ReservedCodeCacheSize=240m
GC Internal Structuresβ
Modern GC engines like G1 or ZGC require native memory outside the heap to keep track of their own metadata:
- Card Tables & Remembered Sets (RSets): G1 uses RSets to keep track of cross-region references (e.g., an object in region A referencing an object in region B) so it can perform GC on individual regions without scanning the whole heap.
- Mark Bitmaps: Bit arrays used to trace live objects during concurrent marking.
- Load Barriers: ZGC uses native memory to track page state and colored pointers metadata.
3. Object Lifecycle
Object Creation
When the JVM encounters a new instruction:
- Class loading check β Is the class loaded? If not, trigger class loading.
- Memory allocation β Allocate space in Eden. Two strategies:
- Bump-the-pointer β if heap is compacted, just move the pointer forward
- Free list β if heap is fragmented, find a suitable gap
- Initialize to zero β Set all fields to default values (0, null, false)
- Set object header β Store class pointer, hash code, GC age, lock info
- Execute
<init>β Run the constructor
Object Memory Layout
Mark Word (Object Header)
Size: 8 Bytes (64-bit JVM) | 4 Bytes (32-bit JVM)
Overview: Contains essential runtime identity and synchronization metadata for the object instance.
- Internal Structure:
- Identity Hashcode: Lazy computed hash value of the object (never changed after compute).
- GC Age: Tracks how many Minor GC cycles this object survived (4 bits, limits age limit to 15).
- Biased Lock Flag & Lock State Bits: Indicates if lock is Unlocked (001), Biased (101), Lightweight (00), Heavyweight (10), or Marked for GC (11).
- JVM Optimization Features:
- Biased Locking (-XX:+UseBiasedLocking) - legacy optimization (deprecated in Java 15).
- Lock Inflation - automatically escalates stack-based lightweight locks to native heavyweight monitor mutexes under heavy thread contention.
π‘ Click on any partition (Mark Word, Class Pointer, Instance Data, or Padding) in the diagram to inspect its binary footprint.
Compressed OOPs & Object Alignment (Memory Optimization)
On 64-bit JVMs, object references (known as Ordinary Object Pointers / OOPs) occupy 8 bytes (64 bits) of memory. This pointer widening increases heap consumption by 30% to 40% compared to 32-bit JVMs. To mitigate this, the JVM uses an optimization called Compressed OOPs (-XX:+UseCompressedOops).
The 8-Byte Object Alignment Trickβ
In HotSpot JVM, all objects allocated on the heap are aligned to 8-byte boundaries. This means an object's memory size is always a multiple of 8, and the JVM adds 1 to 7 bytes of Padding at the end of the object layout to satisfy this constraint.
Because every object address is a multiple of 8, the lower 3 bits of any object memory address are always 000:
- Address
8is00001000 - Address
16is00010000 - Address
24is00011000
The JVM exploits this by shifting the 32-bit pointer left by 3 bits when loading it from CPU registers, and shifting it right by 3 bits when storing it back to the heap.
This bit-shifting trick allows a 32-bit pointer (which can only address 4GB of memory space) to reference up to 32 GB of heap space:
β οΈ The 32GB Heap Threshold Trap (Interview Critical)β
When the heap size configured (-Xmx) exceeds 32GB (roughly 32GB to 35GB depending on the OS and JVM vendor), the JVM disables Compressed OOPs and reverts to raw 64-bit pointers.
- The trap: When Compressed OOPs are disabled, all pointers instantly widen from 4 bytes to 8 bytes.
- The impact: A heap configured for 33GB can hold fewer actual objects than a heap configured for 31GB because the wider 64-bit references consume more memory, leading to higher GC pressure.
- Senior Heuristic: Never set your heap size just over the threshold (e.g. 33β36GB). If you need more than 31GB of heap, jump straight to 40GB+ to compensate for pointer widening.
Object Access
Two approaches:
- Direct pointer (HotSpot): Reference points directly to the object. Faster access.
- Handle pool: Reference points to a handle containing pointers to both instance data and class data. More resilient during GC (only handle pointer changes).
4. Garbage Collection
While the JVM automates memory allocation, garbage collection (GC) is responsible for sweeping and defragmenting the heap. Understanding how the GC manages this space is vital for writing high-performance backend systems.
How GC Identifies Garbage
β Reference Counting (Not Used by JVM)β
In reference counting, each object maintains a counter tracks how many active references point to it. The object is reclaimed when the count hits zero.
- The Circular Reference Flaw: If Object A references Object B, and Object B references Object A, but neither is referenced by any other active part of the application, their counters remain at
1. The memory is permanently leaked.
β Reachability Analysis (Used by JVM)β
To resolve circular references, Java uses Reachability Analysis based on tracing reference graphs from a set of starting nodes called GC Roots:
- GC Roots are references that are guaranteed to be active, including:
- Local variables and input parameters inside active thread stacks.
- Static variables declared in classes loaded in the Method Area.
- Active threads.
- JNI (Java Native Interface) global/local references.
- Active JVM internal system classes.
- The Algorithm: The GC starts at the roots and traverses all references. Any object that can be reached from a GC Root is marked "alive" (reachable). Any object that is unreachable (even if they form a closed loop of references among themselves) is identified as garbage and reclaimed.
The Dock Analogy: Imagine GC Roots as secure docks on a riverbank. Boats that are tied directly to the docks, or tied to other boats that eventually lead back to a dock, are safe. A cluster of boats floating freely in the middle of the river, even if tied tightly to one another, will be swept away by the current (garbage collected) because they have no line connecting them to a dock.
Core GC Algorithms
1. Mark-Sweepβ
- Mark: Traces reference paths from GC Roots and marks all reachable objects as alive.
- Sweep: Scans the heap and releases memory blocks of unmarked (dead) objects.
- Drawback: Leaves Memory Fragmentation (holes of free space scattered between live objects). If the JVM needs to allocate a large contiguous array, it may fail and throw an
OutOfMemoryErroreven if total free space is sufficient.
2. Mark-Compact (Mark-Sweep-Compact)β
- Mark & Sweep: Same as above.
- Compact: Relocates (slides) all surviving objects to one end of the memory block, creating a single, contiguous block of free memory.
- Drawback: Relocating objects requires pausing threads to update reference memory addresses, introducing higher CPU overhead.
3. Copying (Mark-Copy)β
- Mechanism: Memory is split into active and inactive zones. The GC traces and copies surviving objects from the active zone to the inactive zone, then completely wipes the active zone in one bulk delete.
- Drawback: Consumes twice the memory space, but it is extremely fast and leaves no fragmentation.
Generational Collection
The JVM applies these algorithms based on the Weak Generational Hypothesis:
- Most objects die very young (e.g., temporary variables in a method loop).
- Objects that survive initial GC cycles tend to remain active for a very long time (e.g., Spring singleton beans, connection pools, caches).
To exploit this, the JVM splits the heap into two main generations:
| Generation | Memory Space | Main Algorithm | Trigger | Collection Name |
|---|---|---|---|---|
| Young Generation | Eden + Survivor 0 (S0) + Survivor 1 (S1) | Copying (Mark-Copy) | Eden Space is full | Minor GC (Young GC) |
| Old Generation | Tenured Space | Mark-Compact | Old Generation is full | Major GC (Old GC) |
| Entire Heap | Young + Old + Metaspace | Combined | Various triggers | Full GC (Stop-The-World) |
Minor GC Lifecycle Flowβ
- Allocation: New objects are created in Eden.
- First GC: Eden fills up, triggering a Minor GC. The GC copies surviving objects from Eden into
Survivor 0 (S0 / From), then wipes Eden entirely. - Subsequent GCs: In the next Minor GC, survivors from Eden and
S0are copied intoSurvivor 1 (S1 / To). Eden andS0are wiped. The survivor spaces swap roles. - Aging & Promotion: Each copy increment's an object's age. When an object's age exceeds the threshold (configured by
-XX:MaxTenuringThreshold, default is15), it is promoted (copied) into the Old Generation.
Interactive object lifecycle (Eden β S0/S1 β Old β reclaimed), algorithms, and STW evolution: Garbage Collection Deep Dive.
Stop-The-World (STW) Pauses & Container Risks
During certain garbage collection phases, the JVM must freeze all application execution threads so the GC can safely manipulate memory reference pointers without race conditions. This freeze is called a Stop-The-World (STW) pause.
- Minor GC: STW pauses are usually negligible (a few milliseconds) because the Young Gen is small and the copy operation is fast.
- Major GC / Full GC: These run on the Old Generation (or the entire heap + Metaspace). Because the Old Gen is much larger, tracing and compacting it can take hundreds of milliseconds to several seconds.
- The Container Danger: If a containerized Java service experiences a long STW pause (e.g., a 2-second Full GC), it will stop responding to TCP traffic. Kubernetes liveness/readiness probes will fail, leading Kubernetes to assume the container has hung and trigger an unexpected restart (
OOMKilledor simple container reboot).
This danger led to the development of modern low-pause collectors like G1GC and ZGC/Shenandoah (which perform marking and compaction concurrently with application threads, keeping STW pauses under 1 millisecond).
Memory Leaks: The GC Cannot Save You
A common misconception is that Java's automatic garbage collection prevents memory leaks. This is false.
A memory leak in Java occurs when objects are logically abandoned by your application but remain physically reachable from a GC Root.
- Common Causes: A static
Mapwhere entries are added but never removed, forgotten event listener registrations, or uncleanedThreadLocalvariables in a thread-pool reuse environment. - The Consequence: Because these object graphs are reachable from GC Roots, the GC is forced to keep them in the Old Generation. Eventually, the Old Gen fills up, triggering continuous Full GCs that fail to reclaim memory, culminating in a
java.lang.OutOfMemoryError: Java heap space.
5. Garbage Collectors
Serial Collector (-XX:+UseSerialGC)
Single-threaded, stop-the-world. Suitable for small heaps and single-CPU machines.
Parallel Collector (-XX:+UseParallelGC)
Multi-threaded young + old gen collection. Throughput-oriented β minimizes total GC time at the cost of longer individual pauses. Default in JDK 8.
CMS (Concurrent Mark Sweep) (-XX:+UseConcMarkSweepGC)
Low-latency collector for old generation. Most work is done concurrently with application threads:
- Initial Mark (STW) β mark GC roots
- Concurrent Mark β traverse object graph concurrently
- Remark (STW) β fix changes during concurrent mark
- Concurrent Sweep β free dead objects concurrently
Downsides: CPU-intensive, produces fragmentation (no compaction), "concurrent mode failure" if old gen fills during collection. Deprecated since JDK 9, removed in JDK 14.
G1 (Garbage First) (-XX:+UseG1GC)
Region-based collector. Divides the heap into equal-sized regions (~2048). Each region can be Eden, Survivor, Old, or Humongous (for large objects).
Eden Region (Young Generation)
Overview: Logical young generation region. Sized dynamically based on allocation throughput.
- Tuning flags: -XX:G1NewSizePercent (Initial size, default 5%) | -XX:G1MaxNewSizePercent (Max size, default 60%)
- GC Collection Behavior: Minor GCs evacuate surviving objects to Survivor regions, emptying the Eden regions completely to become Free.
- Under the Hood Mechanics:
- Objects are allocated here via TLABs to avoid synchronization locks.
- Sizing is adjusted dynamically by G1 between collections to meet user pause time targets (-XX:MaxGCPauseMillis).
π‘ Click on any region box (Eden, Survivor, Old, Humongous, Free) on the grid to inspect its logical partition rules.
Key features:
- Predictable pause times:
-XX:MaxGCPauseMillis=200(target, not guarantee) - Mixed collections: Can collect young + some old regions selectively
- Compacting: Copies live objects between regions β no fragmentation
- Default in JDK 9+
ZGC (-XX:+UseZGC)
Ultra-low-latency collector (sub-millisecond pauses) using colored pointers and load barriers.
- Pauses are < 1ms regardless of heap size.
- Supports multi-terabyte heaps (from 16MB to 16TB).
- Concurrent relocation (moves objects in memory concurrently while application threads are running, resolving fragmentation without STW pauses).
- Production-ready since JDK 15.
π§ Senior Deep Dive: Generational ZGC (Java 21+ / JEP 439)β
Historically, ZGC was a single-generation collector, meaning it concurrently scanned the entire heap during every GC cycle. Under high allocation rate workloads, this design led to allocation stalls (where application threads ran out of memory before the concurrent collector finished scanning, freezing the application).
To solve this, Java 21 introduced Generational ZGC (-XX:+UseZGC -XX:+ZGenerational), which leverages the weak generational hypothesis (most objects die young) by splitting the heap into two logical generations:
- Young Generation: Collected frequently in a very fast, low-overhead cycle.
- Old Generation: Collected less frequently.
Key Benefits over Non-Generational ZGC:β
- Higher Throughput: Collecting only young objects requires scanning a fraction of the heap, releasing CPU cycles back to application threads.
- Preventing Allocation Stalls: Rapid reclamation of short-lived objects makes allocation stalls extremely rare under heavy load.
- Sub-millisecond Latency: Retains the core concurrent guarantees of ZGC, keeping pause times under 1 millisecond (typically under 100 microseconds).
Collector Selection Guide
| Collector | Pause Target | Heap Size | Use Case |
|---|---|---|---|
| Serial | N/A | Small (< 100 MB) | Embedded, single-core |
| Parallel | High throughput | Medium | Batch processing |
| G1 | < 200ms | Medium-Large | General purpose (default) |
| ZGC | < 1ms | Any (up to TB) | Latency-critical apps |
| Shenandoah | < 10ms | Large | Low-latency alternative |
6. Class Loading
Class Loading Process
1. Loading Phase
Phase Group: Loading
Overview: Finds the binary representation of a class (compiled .class byte stream) by its fully qualified name and loads it into memory.
- Internal Steps & Actions:
- Generates a binary stream from filesystem, JAR file, network, or dynamic proxies.
- Parses the stream into class metadata structures in the Metaspace (method area).
- Instantiates a corresponding java.lang.Class object in the Heap to represent the type.
- Phase Output / Outcome: Class metadata created in Metaspace; Class instance reference created on Heap.
π‘ Click on any step (Loading, Verification, Preparation, Resolution, or Initialization) in the pipeline to inspect compiler execution.
Class Loaders
Java uses a hierarchical delegation model (parent delegation):
Application ClassLoader (System ClassLoader)
Parent Loader: Platform ClassLoader
Overview: Loads standard application-level classes and dependency jars.
- Source Directories / Targets:
- Class paths specified by environment variable CLASSPATH
- Class paths specified by startup arguments -classpath or -cp
- Main application JAR / classes
- Delegation & Scope Rules: Delegates upward to Platform. If Platform and Bootstrap both fail, Application searches classpath.
π‘ Click on any ClassLoader in the tree to inspect its library classpath and upward parent delegation path.
Parent Delegation Model
When a class needs to be loaded:
- Check if already loaded
- Delegate to parent class loader first
- If parent can't load it, try loading it yourself
protected Class<?> loadClass(String name, boolean resolve) {
// 1. Already loaded?
Class<?> c = findLoadedClass(name);
if (c == null) {
try {
// 2. Delegate to parent
c = parent.loadClass(name, false);
} catch (ClassNotFoundException e) {
// 3. Parent failed β load it ourselves
c = findClass(name);
}
}
return c;
}
Why parent delegation?
- Security: Prevents malicious code from replacing core classes (e.g., custom
java.lang.String) - Consistency: Ensures core classes are loaded by the same loader
Thread Context ClassLoader (TCCL)
While the Parent Delegation model is excellent for security and consistency, it has a fundamental design flaw: Core classes loaded by parent loaders cannot load classes that only exist in child loaders.
The Service Provider Interface (SPI) Conundrumβ
Consider the Java Database Connectivity (JDBC) API:
- The JDBC framework class
java.sql.DriverManageris part of the core Java API and is loaded by the Bootstrap ClassLoader. - When
DriverManagertries to establish a connection, it uses Java's SPI (ServiceLoader) to find and load concrete database driver implementations (likecom.mysql.cj.jdbc.Driver) present on your application's classpath. - However, the classpath is loaded by the Application ClassLoader. Since the Bootstrap ClassLoader is a parent loader, it cannot see classes loaded by its child (the Application ClassLoader). Parent delegation only goes up, not down.
The Parent Delegation Conundrum (Visibility Boundary)
Loader Scope: ClassLoader Delegation Architecture Violation
Visibility Issue: Parent classloaders CANNOT see or load classes from child classloaders (e.g. Bootstrap ClassLoader cannot see Application ClassLoader JARs).
Workaround / Resolution: This breaks standard Service Provider Interfaces (SPI) like JDBC, JNDI, and JAXB where core APIs are in the Java standard library but implementations are user dependencies.
// Standard class loading fails: // Parents only search upward, never downward! BootstrapLoader -> (no visibility) -> ApplicationLoader classpath
π‘ Click on any component or connection path (Direct Load vs TCCL Workaround) in the diagram above to inspect how JVM SPI loaders bypass visibility boundaries.
Breaking the Hierarchyβ
To solve this chicken-and-egg problem, Java introduced the Thread Context ClassLoader (TCCL). Each thread holds a reference to a ClassLoader (Thread.currentThread().getContextClassLoader()), which defaults to the Application ClassLoader.
Core classes in the parent ClassLoader can "break" the hierarchy by fetching the context loader from the current running thread and using it to load the child classes:
// How DriverManager breaks parent delegation (simplified)
ClassLoader cl = Thread.currentThread().getContextClassLoader();
ServiceLoader<Driver> loadedDrivers = ServiceLoader.load(Driver.class, cl);
β οΈ Senior Context: ClassLoader Memory Leaks in Containersβ
In application servers (like Tomcat) or plug-in systems where applications are deployed/undeployed dynamically, TCCL can cause severe memory leaks:
- When a web application is deployed, Tomcat creates a custom
WebappClassLoaderand sets it as the TCCL for the request thread. - If the application starts a thread pool or registers a ThreadLocal that isn't cleaned up, the thread retains a strong reference to the
WebappClassLoadervia its context class loader. - When the web application is undeployed, the GC cannot reclaim the classloader or any of the classes it loaded because the thread context pointer is still active. This leads to
OutOfMemoryError: Metaspace. - Mitigation: Always restore the original context classloader in a
finallyblock or clean up custom threads upon application shutdown.
7. Class File Structure
Every .class file follows a strict binary format:
ClassFile {
u4 magic; // 0xCAFEBABE
u2 minor_version;
u2 major_version; // Java 17 = 61
u2 constant_pool_count;
cp_info constant_pool[]; // literals, type refs, method refs
u2 access_flags; // public, final, abstract, etc.
u2 this_class;
u2 super_class;
u2 interfaces_count;
u2 interfaces[];
u2 fields_count;
field_info fields[];
u2 methods_count;
method_info methods[];
u2 attributes_count;
attribute_info attributes[];
}
Use javap -verbose MyClass.class to inspect the structure.
8. Important JVM Parameters
Heap Sizing
# Initial and maximum heap size
-Xms512m # initial heap (set equal to -Xmx to avoid resizing)
-Xmx2g # maximum heap
# Young generation size
-Xmn512m # young gen size
-XX:NewRatio=2 # old:young ratio (default 2 β old is 2x young)
# Metaspace
-XX:MetaspaceSize=256m
-XX:MaxMetaspaceSize=512m
GC Configuration
# Select collector
-XX:+UseG1GC # G1 (default JDK 9+)
-XX:+UseZGC # ZGC
-XX:+UseParallelGC # Parallel (default JDK 8)
# G1 tuning
-XX:MaxGCPauseMillis=200 # target pause time
-XX:G1HeapRegionSize=4m # region size (1-32 MB, power of 2)
# GC logging (JDK 9+)
-Xlog:gc*:file=gc.log:time,uptime,level,tags
Thread Stack
-Xss512k # thread stack size (default ~1MB)
Troubleshooting
# Heap dump on OOM
-XX:+HeapDumpOnOutOfMemoryError
-XX:HeapDumpPath=/path/to/dump.hprof
# Print GC details
-verbose:gc
9. JIT Compilation (HotSpot C1 / C2)
The JVM doesn't just interpret bytecode β it dynamically compiles hot code to native machine code. Understanding the tiers is critical for diagnosing startup slowdowns and latency spikes.
Compilation Tiers
| Tier | Compiler | Description |
|---|---|---|
| 0 | Interpreter | Execute bytecode directly (cold start) |
| 1β3 | C1 (Client) | Quick compilation with basic optimization |
| 4 | C2 (Server) | Aggressive optimization: inlining, loop unrolling, escape analysis |
"Compilation storm": At startup, many methods reach the hot threshold simultaneously β C2 compiler overwhelmed β CPU spike, latency increase. Common in Kubernetes when pods receive traffic immediately.
Mitigation: GraalVM Native Image (AOT) for instant startup; JVM Tiered Compilation (-XX:+TieredCompilation) for warmup.
Deoptimization
JIT makes optimistic assumptions β e.g., that a virtual method is called with only one concrete type (monomorphic call). When assumptions break:
// JIT inlines Dog.speak() for all calls β optimized for monomorphic dispatch
void speak(Animal a) { a.speak(); }
// First Cat appears β JIT's inline prediction invalid β deoptimize β interpreter
speak(new Cat());
Cold code paths with rare types cause unexpected production latency spikes even after warm-up.
-XX:+PrintCompilation # See which methods JIT compiles
-XX:CompileThreshold=10000 # Invocations before C2 trigger (default)
10. G1 GC β Internal Mechanics
Humongous Objects
Objects larger than 50% of a region size are allocated directly in humongous regions (multiple contiguous Old gen regions). These are only collected during a full GC unless explicitly triggered.
# Fix: increase region size to reduce humongous allocations
-XX:G1HeapRegionSize=32m
Remembered Sets (RSet) and SATB
Remembered Sets: Each G1 region tracks external references into it. Required so G1 can collect a single region without scanning the entire heap.
SATB (Snapshot-At-The-Beginning): G1's write barrier during concurrent marking. When a reference is overwritten, G1 records the old value in an SATB log buffer. This ensures that objects alive at mark-start remain live even if pointers are nulled during marking.
obj.field = newRef;
// SATB write barrier fires here β logs old obj.field reference
Without SATB, a concurrent mutator could hide a live object from the marking thread, causing premature collection.
Mixed Collections
After a full concurrent mark cycle, G1 picks the highest-garbage-density Old regions and collects them alongside Young gen:
-XX:G1MixedGCLiveThresholdPercent=85 # Only collect Old regions < 85% live data
-XX:G1HeapWastePercent=5 # Stop mixed GC if < 5% heap is reclaimable
# Diagnosing pauses:
-Xlog:gc*:file=gc.log:time,uptime,level,tags
# Look for: "Pause Full" β means G1 fell back to stop-the-world (bad!)
11. JDK Monitoring & Troubleshooting Tools
Command-Line Tools
| Tool | Purpose | Example |
|---|---|---|
jps | List running JVM processes | jps -lv |
jstat | GC and memory statistics | jstat -gcutil <pid> 1000 |
jinfo | View/modify JVM flags | jinfo -flags <pid> |
jmap | Heap dump and histogram | jmap -dump:format=b,file=heap.hprof <pid> |
jstack | Thread dump (diagnose deadlocks) | jstack <pid> |
jcmd | All-in-one diagnostic tool | jcmd <pid> GC.heap_info |
Graphical Tools
- JVisualVM β bundled with JDK (up to JDK 8), monitors heap, threads, CPU
- JConsole β JMX-based monitoring console
- Eclipse MAT β heap dump analysis, find memory leaks
- Arthas β powerful runtime diagnostic tool (bytecode-level debugging)
Common Troubleshooting Scenarios
OutOfMemoryError: Java heap space
- Generate heap dump:
-XX:+HeapDumpOnOutOfMemoryError - Analyze with Eclipse MAT β find objects consuming most memory
- Check for memory leaks (growing collections, unclosed resources)
High CPU usage
top -H -p <pid>β find the CPU-intensive thread (note the TID)jstack <pid>β find the thread by TID (convert to hex)- Analyze the stack trace
Deadlock detection
jstack <pid>β JVM automatically detects and reports deadlocks- Look for "Found one Java-level deadlock" in the output
Frequent Full GC
jstat -gcutil <pid> 1000β monitor GC frequency and duration- Check if old gen is filling up (memory leak?) or if young gen is too small (premature promotion)
- Consider switching to G1 or ZGC for better pause behavior
12. Java Agents & Instrumentation (Telemetry Hooks)
For Senior and Lead developers working on APM (Application Performance Monitoring) tools or custom frameworks, understanding Java Agents is essential.
What is a Java Agent?
A Java Agent is a pluggable JVM-level tool that uses the Java Instrumentation API (java.lang.instrument) to intercept and modify the bytecode of classes loaded into the JVM.
Execution Mechanisms
A Java Agent can be loaded in two ways:
1. Static Loading (premain)β
The agent is specified at JVM startup using the -javaagent flag. The JVM runs the agent's premain method before the application's main method starts.
// Command: java -javaagent:myagent.jar -jar myapp.jar
public static void premain(String agentArgs, Instrumentation inst) {
inst.addTransformer(new MyClassFileTransformer());
}
2. Dynamic Attachment (agentmain)β
The agent is dynamically loaded into a running JVM using the VirtualMachine API (from the tools.jar Attach API) after the application has already started.
public static void agentmain(String agentArgs, Instrumentation inst) {
inst.addTransformer(new MyClassFileTransformer(), true);
// Force retransformation of already-loaded classes
inst.retransformClasses(TargetClass.class);
}
Bytecode Modification
Inside the ClassFileTransformer, you inspect the class bytes, modify them (usually using libraries like ByteBuddy, ASM, or Javassist), and return the modified byte array:
public class MyClassFileTransformer implements ClassFileTransformer {
@Override
public byte[] transform(ClassLoader loader, String className, Class<?> classBeingRedefined,
ProtectionDomain protectionDomain, byte[] classfileBuffer) {
if ("com/example/service/BillingService".equals(className)) {
// Intercept billing methods, inject entry/exit logs or latency trackers
return injectLatencyProfilingBytes(classfileBuffer);
}
return null; // Return null to indicate no changes
}
}
Real-world Use Cases:
- APM Tooling (Datadog, Dynatrace, New Relic): Auto-instruments database drivers, HTTP controllers, and outbound clients to record transaction traces and execution metrics without changing application code.
- Dynamic Profiling (async-profiler, Arthas): Inspects class bytecode and system metrics dynamically in production.
- Frameworks & Testing (Lombok, Mockito): Lombok uses compile-time annotation processing, but Mockito uses runtime bytecode generation (ByteBuddy) to mock interfaces and classes.
13. Reference Types & GC
Java provides four reference types that influence garbage collection behavior:
| Reference Type | Class | GC Behavior | Use Case |
|---|---|---|---|
| Strong | (default) | Never collected while reachable | Normal references |
| Soft | SoftReference<T> | Collected when JVM is low on memory | Memory-sensitive caches |
| Weak | WeakReference<T> | Collected at next GC | WeakHashMap, canonicalizing maps |
| Phantom | PhantomReference<T> | Enqueued after finalization | Resource cleanup tracking |
// Soft reference: cache that yields to memory pressure
SoftReference<byte[]> cache = new SoftReference<>(new byte[1024 * 1024]);
byte[] data = cache.get(); // may be null if GC reclaimed it
// Weak reference: doesn't prevent GC
WeakReference<ExpensiveObject> ref = new WeakReference<>(new ExpensiveObject());
ExpensiveObject obj = ref.get(); // null after GC
ReferenceQueues for Cleanups
To cleanly handle post-mortem resources, you can register soft, weak, or phantom references with a ReferenceQueue.
When the garbage collector decides to reclaim the referent (the object referenced), it automatically clears the reference (sets it to null) and appends the reference container itself (the SoftReference or WeakReference instance) to the registered ReferenceQueue.
The application can poll or block on this queue in a background thread to safely release associated native resources (like database connections, file handles, or off-heap memory) without using slow, deprecated finalize() methods.
ReferenceQueue<ExpensiveObject> queue = new ReferenceQueue<>();
WeakReference<ExpensiveObject> ref = new WeakReference<>(new ExpensiveObject(), queue);
// ... later, after ExpensiveObject has been garbage-collected ...
Reference<? extends ExpensiveObject> clearedRef = queue.poll();
if (clearedRef != null) {
// Perform resource cleanup associated with this reference
}
π» Phantom References Require ReferenceQueueβ
Unlike Soft and Weak references, a PhantomReference's get() method always returns null. This prevents the application from accidentally resurrecting the object during garbage collection.
A PhantomReference is completely useless without a ReferenceQueue. It is used purely as a notification mechanism to know exactly when an object has been fully finalized and its memory reclaimed by the GC.
βοΈ Production Example: DirectByteBuffer & Cleanerβ
The most notable use of PhantomReference and ReferenceQueue is Java's off-heap memory management:
- When you allocate off-heap memory using
ByteBuffer.allocateDirect(10 * 1024), the JVM creates aDirectByteBufferobject on the heap. - This heap object references a native memory address allocated outside the JVM heap.
- To prevent memory leaks,
DirectByteBufferregisters a phantom reference with aCleaner(which uses aReferenceQueueinternally). - When the heap-based
DirectByteBufferis garbage-collected, the phantom reference is enqueued in theReferenceQueue. - A system-level daemon thread polls this queue and frees the associated off-heap native memory using
unsafe.freeMemory().
13. Common OOM Scenarios & Solutions
| Error | Cause | Solution |
|---|---|---|
OutOfMemoryError: Java heap space | Heap exhausted | Increase -Xmx, fix memory leaks |
OutOfMemoryError: Metaspace | Too many classes loaded | Increase -XX:MaxMetaspaceSize, fix classloader leaks |
OutOfMemoryError: GC overhead limit | GC consuming over 98% CPU for under 2% heap recovery | Fix memory leaks, increase heap |
StackOverflowError | Deep/infinite recursion | Fix recursion, increase -Xss |
OutOfMemoryError: unable to create new native thread | Too many threads | Use thread pools, reduce stack size |
Advanced Editorial Pass: JVM Internals for Operational Excellence
Senior-Level Focus
- GC tuning is workload-specific and must be tied to SLO outcomes.
- Heap, metaspace, and thread configuration are architecture choices, not defaults.
- Classloading and JIT behavior can materially impact startup and latency profiles.
Failure Modes in Production
- Over-tuned JVM flags copied between services with different traffic patterns.
- Memory leaks masked by oversized heaps until incident windows.
- Misinterpreting GC logs without correlating application-level latency.
Practical Heuristics
- Treat JVM tuning as iterative experimentation with measurable hypotheses.
- Baseline key metrics before any flag change.
- Keep service-specific runbooks for memory, GC, and thread incidents.
Compare Next
- Java Concurrency: Threads, Locks & Concurrent Utilities
- Java Fundamentals: Core Language Concepts
- Java Interview Questions & Answers
Interview Questions
Q: How do you choose between G1 and ZGC for a backend service?
A: G1 is a strong default for balanced throughput and latency; ZGC is preferred for strict low-latency requirements with larger heaps.
Q: What metrics indicate GC tuning is required?
A: Rising tail latency, frequent long pauses, promotion failures, and high GC CPU share under normal load.
Q: Why is allocation rate often more important than heap size?
A: High allocation churn drives GC pressure even on large heaps, so reducing object churn often beats increasing memory.
Q: How do classloader leaks usually appear in production?
A: Metaspace growth over time after redeploy/plugin cycles and inability to reclaim old class metadata.
Q: What is a practical JVM tuning workflow for senior engineers?
A: Baseline, form a hypothesis, apply one controlled change, validate with load and latency data, then iterate.
Q: Why are full GC events high priority incidents?
A: They are stop-the-world and can trigger latency spikes, timeouts, and cascading failures.
Q: How do you explain JIT warmup impact during autoscaling?
A: New pods initially run colder code paths, so p95/p99 latency can temporarily degrade until optimization stabilizes.
