ZGC & Generational ZGC: Sub-Millisecond Pause Architecture
The Z Garbage Collector (ZGC) is HotSpot's scalable low-latency garbage collector. Designed to manage heaps ranging from 8 megabytes to 16 terabytes, ZGC performs all heavy mark, evacuate, and reference relocation phases concurrently while application worker threads run without interruption.
In Java 21 LTS, Generational ZGC (JEP 439) introduced generational separation, delivering consistent p99.99 pause times under 1 millisecond even under extreme allocation pressure.
1. The Low-Latency Paradigm: ZGC vs. G1
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β STOP-THE-WORLD (STW) PAUSE COMPARISON β
β β
β Traditional G1 GC: β
β β’ Young Evacuation: 20ms - 200ms (Pauses proportional to live object countβ
β β’ Mixed Evacuation: 50ms - 300ms STW pause β
β β
β Generational ZGC (JDK 21+): β
β β’ Pause Mark Start: < 0.1ms (STW scans GC Roots only) β
β β’ Concurrent Mark: 0ms STW (Mutators run at full speed) β
β β’ Pause Mark End: < 0.1ms (STW synchronizes thread buffers) β
β β’ Concurrent Relocate: 0ms STW (Objects moved concurrently!) β
β β’ Max p99.99 Pause: < 1.0ms guaranteed regardless of heap size! β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
2. Colored Pointers & Virtual Memory Multi-Mapping
On 64-bit systems, a 48-bit address space can address 256 Terabytes of RAM. ZGC exploits the unused high-order bits of standard 64-bit reference pointers to store metadata flags directly on the pointer (Colored Pointers):
+-------------------+-------------+-----------------------------------------------+
| 16 Unused Bits | 4 Color Bits| 44-bit Object Offset (Addresses up to 16 TB) |
+-------------------+-------------+-----------------------------------------------+
| | | |
| | | +-- Finalizable (Weak/Phantom reference tracking)
| | +---- Remapped (Address points to new relocated location)
| +------ Marked1 (Active live marking phase)
+-------- Marked0 (Alternate live marking phase)
Virtual Memory Multi-Mapping (mmap)
To allow the CPU hardware to dereference a colored pointer without paying software bitmasking penalties on every memory read:
- The HotSpot VM invokes Linux kernel
mmap()syscalls to map three separate virtual memory address ranges (Marked0,Marked1, andRemapped) to the exact same underlying physical memory page:
Virtual Address 1 (Marked0): 0x0001_0000_1234_5000 βββ
Virtual Address 2 (Marked1): 0x0002_0000_1234_5000 βββΌβββΊ Physical RAM Page (0x45000)
Virtual Address 3 (Remapped): 0x0004_0000_1234_5000 βββ
When an application thread dereferences an object pointer, the CPU memory management unit (MMU) resolves the address to the identical physical page regardless of which color bit is set!
3. The Self-Healing Load Barrier
While G1 relies on Write Barriers to track modified fields, ZGC utilizes Load Barriers executed whenever a thread reads an object reference from the heap (obj.field):
[Application Thread Reads Reference: Object target = source.field]
β
βΌ
[ZGC Load Barrier (~2 Nanoseconds)]
β
ββββββββββββββββββββ΄βββββββββββββββββββ
β Is pointer color "Remapped"? β
ββββββββββββββββββββ¬βββββββββββββββββββ
YES β NO (Object was relocated by GC)
β βββ Read Forwarding Table (O(1))
β βββ Update reference to new address
β βββ "Heal" pointer in memory!
βΌ
[Return valid object reference immediately]
- Self-Healing Guarantee: Only the very first thread to read a relocated object executes the load barrier slow path. It immediately patches ("heals") the field in memory so all subsequent reads take the fast 2-nanosecond path.
4. Generational ZGC Architecture (Java 21 LTS - JEP 439)
Prior to Java 21, single-generation ZGC collected the entire heap uniformly. Under heavy allocation bursts, the concurrent marking phase could not keep up with allocation rates, triggering Allocation Stalls (threads temporarily blocked waiting for free memory).
Generational ZGC solved this by separating the heap into Young and Old generations:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β GENERATIONAL ZGC MEMORY LAYOUT β
β β
β ββββββββββββββββββββββββββββββββββββββββ βββββββββββββββββββββββββββββββββ β
β β Young Generation (Frequent Cycles) β β Old Generation (Rare Cycles) β β
β β β’ Collects request DTOs & builders β β β’ Holds long-lived pools β β
β β β’ Low CPU usage; instant recovery β β β’ Runs concurrent background β β
β ββββββββββββββββββββββββββββββββββββββββ βββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Key Innovations in Generational ZGC:
- Multi-Generation Load Barriers: Differentiates between young and old pointers using optimized color bit sequences.
- Store Barriers for Remembered Sets: Employs lightweight store barriers to track Old-to-Young references without full card table overhead.
- Automatic Worker Scaling: Dynamically allocates more GC worker threads to Young collections during sudden allocation spikes.
5. Enabling & Sizing Generational ZGC in Production
To enable Generational ZGC in Java 21 LTS:
java -XX:+UseZGC -XX:+ZGenerational \
-Xms16g -Xmx16g \
-XX:SoftMaxHeapSize=12g \
-Xlog:gc*,gc+phases=debug:file=gc.log:time,uptime,pid:filecount=5,filesize=100M \
-jar core-banking-engine.jar
-XX:SoftMaxHeapSize: Instructs ZGC to keep heap usage within this soft ceiling under normal operations, uncommitting unused memory back to the Linux OS.- Tuning Simplicity: Unlike G1, which requires balancing dozens of flags (
MaxGCPauseMillis,IHOP,SurvivorRatio,G1ReservePercent), Generational ZGC is self-tuning: you only need to configure-Xmx!
6. Principal Architect Review Checklist
- JDK 21 LTS Prerequisite: Are latency-sensitive microservices running on Java 21 or later with
-XX:+UseZGC -XX:+ZGenerationalenabled? - Memory Allocation Headroom: Is
-Xmxsized with at least 25%β35% headroom above steady-state live data to allow concurrent evacuation threads sufficient space to operate without stalls? - Linux Virtual Memory Limits: Is
sysctl vm.max_map_countconfigured to at least1048576on Linux hosts to support ZGC's virtual memory multi-mapping (mmap) calls? - Throughput vs. Latency Evaluation: Has throughput been verified against G1? (Generational ZGC achieves 95%β99% of G1 throughput while dropping p99.99 latency from 150ms down to <1ms).
