Skip to main content

LMAX Disruptor Architecture: Ultra-Low-Latency RingBuffer

Developed by the LMAX Exchange for financial trading, the Disruptor is an ultra-high-performance inter-thread messaging library. It achieves millions of operations per second with sub-microsecond p99 latency by eliminating lock contention, eliminating runtime garbage collection, and exploiting hardware cache line locality.


1. The Bottlenecks of Traditional Queues

Standard concurrent queues like java.util.concurrent.ArrayBlockingQueue suffer from three structural limitations under high throughput:

  1. Lock Contention: Relies on ReentrantLock instances (takeLock and putLock). Under heavy contention, CPU cores spend more time arbitrating OS mutexes than processing data.
  2. Cache Line False Sharing: The queue's head, tail, and size pointers reside close together in memory, causing continuous cache line bouncing between producer and consumer cores.
  3. Garbage Collection Pressure: If messages are wrapped in dynamic task objects, the JVM heap suffers continuous Eden allocation churn, triggering STW garbage collection pauses.

2. Core Pillars of the Disruptor Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ PRODUCER THREAD β”‚
β”‚ Claims Sequence: next() β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό (Ring Buffer Claim)
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ CIRCULAR RING BUFFER β”‚
β”‚ [Slot 0] [Slot 1] [Slot 2] [Slot 3] ... β”‚
β”‚ (Pre-allocated mutable event objects) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό (Lock-Free Sequence Barrier)
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ CONSUMER THREAD β”‚
β”‚ Polls SequenceBarrier: waitFor() β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

1. Pre-Allocated Circular Ring Buffer (Zero GC)

  • The RingBuffer is an array of pre-instantiated event objects created at system initialization.
  • Producers populate existing objects in-place rather than allocating new Event() on every message.
  • Impact: Generates zero heap garbage, completely removing Minor GC pauses from the critical processing path.

2. Power-of-Two Bitwise Indexing

  • The buffer capacity must strictly be a power of two (2n2^n, e.g. 65,536).
  • The array index is calculated using a single-cycle bitwise AND mask instead of an expensive integer modulo (%) division: ArrayΒ Index=SequenceΒ NumberΒ &Β (BufferΒ Sizeβˆ’1)\text{Array Index} = \text{Sequence Number} \ \& \ (\text{Buffer Size} - 1)

3. The Single-Writer Principle

  • If an application isolates event publishing to a single dedicated thread, zero CAS instructions or locks are required.
  • The sequence counter increments via plain writes, eliminating memory bus lock contention entirely.

3. Disruptor Wait Strategies

The consumer thread coordinates with the producer via a SequenceBarrier using configurable wait strategies:

Wait StrategyCPU UtilizationLatency JitterOptimal Use Case
BusySpinWaitStrategy100% of 1 CPU Core< 50nsFinancial matching engines, ultra-low-latency gateways where cores are pinned.
YieldingWaitStrategyHigh (Spins 100x then Thread.yield())~100ns – 250nsLow-latency applications sharing CPU cores with other threads.
SleepingWaitStrategyLow (Progressive spin β†’\rightarrow yield β†’\rightarrow park)~2Β΅s – 5Β΅sAsynchronous loggers (Apache Log4j 2), telemetry ingestion.
BlockingWaitStrategyMinimal (Uses Lock and Condition)~10Β΅s – 50Β΅sResource-constrained systems where CPU efficiency is prioritized over latency.

4. Performance Benchmark: Disruptor vs. ArrayBlockingQueue

Metricjava.util.concurrent.ArrayBlockingQueueLMAX Disruptor (Single Producer)
Underlying MechanismTwo ReentrantLock instances (takeLock, putLock)Lock-free sequence barriers on circular array
Garbage CreationLow (if reusing DTOs)Zero (pre-allocated ring buffer slots)
Cache Line ContentionHigh (head and tail pointers contend on shared lines)Zero (head and tail separated via cache line padding)
Max Throughput~150,000 to 450,000 ops/second6,000,000 to 25,000,000 ops/second
p99.9 Latency2,500 – 15,000 microseconds (lock contention stalls)< 50 nanoseconds to 1 microsecond

5. Principal Architect Review Checklist

  • Power-of-Two Buffer Sizing: Is the RingBuffer size strictly verified to be a power of two (2n2^n) to prevent bitwise wrapping corruption?
  • Event Handler Exceptions: Does every EventHandler implement an explicit ExceptionHandler to prevent unhandled exceptions from terminating consumer threads?
  • CPU Pinning for BusySpin: If BusySpinWaitStrategy is selected, are consumer threads pinned to dedicated OS CPU cores using thread affinity (e.g. Java-Thread-Affinity)?
  • Batching Optimization: Do consumer event handlers process batches of sequences up to the available sequence barrier before updating their sequence counter?

πŸ“–
Track Page Progress0 / 635 Read
Knowledge Base Completion0%