Skip to main content

Thread Pools, Netty, Tomcat & HikariCP

Who this guide is for
Core Prerequisite

Before learning how threads and connections are pooled, make sure you understand the fundamental difference between logical multi-tasking and physical simultaneous execution. Check out the dedicated guide: Concurrency vs. Parallelism.


1. Thread Pools β€” ThreadPoolExecutor

What is a Thread Pool?

A thread pool is a managed collection of pre-created threads that are reused to execute tasks. Instead of creating a new OS thread for every task (expensive: ~1MB stack + kernel call), the pool maintains a fixed number of threads that pick tasks from a queue.

🧡 Lifecycle Mode: Thread-Per-Task

Thread Creation/Destruction LoopTask 1πŸ› οΈ Spawn OS Threadβš™οΈ Execute Task❌ Destroy ThreadOverhead: Clogged by kernel allocation boundaries.

Without a Pool: Spawning Threads Per Task

Overhead Cost: ⚠️ High Overhead (~1ms creation, ~1ms destruction per task)

Overview: The JVM must request physical thread resources from the underlying operating system kernel for every single task, allocation of 1MB stack memory space, and context switching.

  • Execution Details:
    • Task 1 -> Create Thread-1 -> Run Task -> Destroy Thread-1
    • High OS kernel context-switching penalty under burst load.
    • Risk of OutOfMemoryError (OOM) if thread count grows uncontrolled.

πŸ’‘ Switch between "No Pool" and "With Pool" tabs to compare CPU/OS allocation boundaries.

How ThreadPoolExecutor Works Internally

java.util.concurrent.ThreadPoolExecutor is the engine behind all Java thread pools. Understanding its internals prevents catastrophic production failures.

Constructor Parameters​

ThreadPoolExecutor executor = new ThreadPoolExecutor(
4, // corePoolSize
8, // maximumPoolSize
60, TimeUnit.SECONDS, // keepAliveTime
new ArrayBlockingQueue<>(100), // workQueue
new ThreadFactory() { ... }, // threadFactory
new CallerRunsPolicy() // rejectionHandler
);
ParameterWhat It ControlsWhy It Matters
corePoolSizeThreads that stay alive even when idleToo low β†’ tasks queue; too high β†’ wasted RAM
maximumPoolSizeAbsolute thread ceiling under burst loadSafety valve β€” prevents unbounded thread creation
keepAliveTimeHow long non-core threads survive idleLets burst threads die after the spike passes
workQueueBuffer for tasks when all core threads are busyBounded = backpressure; Unbounded = OOM risk
threadFactoryCustom thread naming and daemon settingsNamed threads = readable thread dumps
rejectionHandlerWhat happens when pool AND queue are fullCallerRunsPolicy = natural backpressure

Task Submission Flow (Critical for Interviews)​

Task SubmittedIs corePool full?Active < coreSizeIs queue full?BlockingQueue.offer()Is maxPool full?Active < maxSize⚑ Start Core ThreadπŸ“₯ Queue Runnable⚑ Start Non-CoreREJECTYESNOYESNOYESNO

2. Core Pool Boundary Check

Condition Check / Action: Compare active threads with corePoolSize

Overview: If the number of running worker threads is less than corePoolSize, the executor always spawns a new core worker thread to run this task.

  • Task Dispatch Decisions:
    • Spawns a new core thread even if other core threads are currently sitting idle.
    • If active threads >= corePoolSize, bypasses thread creation and tries to queue the task.

πŸ’‘ Click on decision blocks (Core check, Queue check, Max check, Rejection) in the pipeline to analyze execution allocations.

⏳Interactive ThreadPool Timeline (T1–T8)

1 Tasks Sent
Pool Settings: corePoolSize=2, maxPoolSize=4, queueCapacity=3Core Slots (2):T1EmptyQueue Slots (3):Q1Q2Q3Non-Core (2):EmptyEmptyRejection:None

πŸ‘‰ Task T1 arrives: Sinks directly into core thread slot 1 (spawns thread).

πŸ’‘ Step through the timeline using the controls above to see the precise scheduling sequence of Java\'s ThreadPoolExecutor.

Why Executors Factory Methods Are Dangerous
Factory MethodHidden Danger
Executors.newFixedThreadPool(n)Uses unbounded LinkedBlockingQueue β†’ tasks pile up β†’ OOM
Executors.newCachedThreadPool()maximumPoolSize = Integer.MAX_VALUE β†’ creates unlimited threads β†’ OOM
Executors.newSingleThreadExecutor()Unbounded queue β†’ same OOM risk as fixed pool

Always use ThreadPoolExecutor directly with bounded queues in production.

Sizing Thread Pools

CPU-Bound Tasks​

Tasks that compute without waiting (sorting, encryption, JSON parsing):

Optimal threads = CPU_cores + 1

Why +1? Insurance against page faults. If one thread stalls on
a memory page fault, the extra thread keeps the CPU busy.

Example: 8-core server β†’ 9 threads for CPU-bound work

I/O-Bound Tasks​

Tasks that spend most of their time waiting (DB queries, HTTP calls, file reads):

Optimal threads = CPU_cores Γ— target_utilization Γ— (1 + wait_time / compute_time)

Example:
8 cores, 70% CPU target
Average HTTP call: 200ms wait, 2ms compute
= 8 Γ— 0.7 Γ— (1 + 200/2) = 8 Γ— 0.7 Γ— 101 = 565 threads

Rule of thumb: if tasks are 99% waiting, you need ~100Γ— more
threads than cores to keep the CPU busy.

Spring Boot @Async Thread Pool

@Configuration
@EnableAsync
public class AsyncConfig implements AsyncConfigurer {

@Override
public Executor getAsyncExecutor() {
ThreadPoolTaskExecutor executor = new ThreadPoolTaskExecutor();
executor.setCorePoolSize(10);
executor.setMaxPoolSize(50);
executor.setQueueCapacity(200);
executor.setThreadNamePrefix("async-worker-");
executor.setRejectedExecutionHandler(new ThreadPoolExecutor.CallerRunsPolicy());
executor.initialize();
return executor;
}
}

See also: Java Concurrency β€” Thread Pools & Executors for ThreadPoolExecutor constructor walkthrough, starvation math, and ScheduledExecutorService.


2. Tomcat β€” Embedded Server Threads

What is Tomcat?

Apache Tomcat is the default embedded servlet container in Spring Boot. It handles HTTP connections and dispatches requests to your controllers. Tomcat uses a thread-per-request model: each incoming HTTP request gets a dedicated thread from a pool.

Tomcat's Internal Architecture

Tomcat Server BoundaryAcceptor ThreadServerSocket.accept()Poller ThreadNIO epoll selectorWorker Thread Poolmax-threads=200DispatcherServlet@RestController MVC

Tomcat Worker Thread Pool

Thread Context Sizing: Thread-Per-Request Pool (ExecutorService)

Tuning Parameter: server.tomcat.threads.max=200 (Default)

Overview: A pool of platform threads that carry out the actual business logic, request parsing, servlet filtering, and controller invocation.

  • Processing Details:
    • Executes the entire lifecycle of a single request blocking on database or downstream APIs.
    • Saturating the max-threads limit (200) stalls all future incoming connections in the OS backlog.

πŸ’‘ Click on Acceptor, Poller, Worker Pool, or DispatcherServlet above to analyze HTTP request boundaries.

How Tomcat Processes a Request​

STEP 01

Acceptor Thread

ServerSocketChannel.accept()

STEP 02

Poller Thread (NIO)

Selector Multiplexing Loop

STEP 03

Worker Thread (Pool)

DispatcherServlet Execution

Tomcat Pipeline: Acceptor Thread

Overview: A simple blocking loop that accepts incoming TCP socket connections from the OS kernel backlog queue. Instantly sets the socket to non-blocking and passes it to the Poller selector.

  • Processing Mechanics:
    • Executes standard TCP handshake completion.
    • Passes socket descriptor (FD) to the next stage immediately.

πŸ’‘ Click on Step 01, 02, or 03 above to trace standard Tomcat request execution stages.

Tomcat Configuration

server:
tomcat:
# === Worker Thread Pool ===
threads:
max: 200 # Maximum worker threads (default: 200)
min-spare: 10 # Minimum idle threads kept warm (default: 10)

# === Connection Limits ===
max-connections: 8192 # Max simultaneous connections the Poller can track (default: 8192)
accept-count: 100 # OS-level TCP backlog queue when max-connections is reached (default: 100)

# === Timeouts ===
connection-timeout: 20000 # ms to wait for first byte after TCP connect (default: 20s)

# === Keep-Alive ===
keep-alive-timeout: 20000 # ms to keep an idle connection open for reuse
max-keep-alive-requests: 100 # requests per keep-alive connection before closing

Understanding the Numbers​

max-connections (8192) β†’ Poller can track this many sockets
↓
max-threads (200) β†’ Only 200 can be actively processed
↓
accept-count (100) β†’ OS queues 100 more when max-connections hit
↓
Beyond that β†’ TCP RST (connection refused)

In steady state with short requests:
200 threads Γ— 10ms avg response = 20,000 requests/sec throughput

With slow requests (2s average):
200 threads Γ— 2000ms = 200 concurrent users max
User #201 waits in the Poller queue

Tomcat and Direct Memory: The Temporary Cache

Like all Java network libraries, Tomcat relies on Direct Memory (off-heap memory) to transmit data between the OS socket buffers and the JVM. Because the Operating System's read() system call requires a static, absolute physical memory address, it cannot write directly to Java heap objects (which are constantly relocated by the Garbage Collector during compaction). A static, off-heap native buffer acts as the intermediate landing pad.

Tomcat's I/O pipeline uses two copy operations:

NIC BufferEthernet RJ-45OS Socket BufferKernel SpaceDirect Buffer (Off-Heap)sun.nio.ch.Util CacheHeap Byte ArrayJVM GC Managed

Temporary Direct Buffer (Off-Heap)

Memory Domain: JVM Native Memory Space

Overview: Tomcat allocates off-heap Direct Buffers to bridge the gap. Data is copied from kernel space directly to this stable native memory block.

  • Direct Memory copy pipeline implications:
    • Eliminates Garbage Collector relocation blockages during socket reads.
    • ⚠️ THE THREAD-LOCAL TRAP: sun.nio.ch.Util caches the largest buffer ever used by a thread. A 50MB file read leaves a persistent 50MB native cache per worker thread, causing container OOM crashes!
    • Remediation: Configure -Djdk.nio.maxCachedBufferSize=262144 (256KB) to enforce clean buffer eviction.

πŸ’‘ Click on NIC, OS Socket, Off-Heap, or Heap boxes above to explore off-heap memory leak gotchas.

The Per-Thread Cache Footprint​

Tomcat manages this direct memory using a temporary thread-local cache inside the JDK (sun.nio.ch.Util).

  • Every worker thread is assigned a single direct buffer.
  • This buffer is initially sized to Tomcat's default buffer size (configured by appReadBufSize, usually 8KB).
  • When a request is read, Tomcat reuses the same 8KB direct buffer repeatedly instead of allocating new blocks or releasing it to the OS.
  • For 200 default worker threads, this consumes only 200 Γ— 8KB = 1.6MB of direct memory, which is negligible and virtually immune to out-of-memory issues.

The Thread-Local Cache Trap (e.g., The Ehcache Trap)​

Although Tomcat's HTTP reading is safe, a major memory leak trap occurs if your business logic uses the same worker thread to perform a large file or socket channel operation:

  • The JDK NIO utility (sun.nio.ch.Util) caches the largest buffer ever requested by that thread.
  • If your business code reads a 50MB file into a heap byte array using a FileChannel on a Tomcat worker thread, the JIT/NIO allocates and caches a 50MB native direct buffer on that thread's local storage.
  • The buffer remains bound to the thread forever and is never resized down. If 100 worker threads run this code path once, your application will silently bleed 5GB of off-heap native memory, leading to container-level OOMKilled crashes while heap usage looks perfectly normal.
  • Remediation: Set the JVM flag -Djdk.nio.maxCachedBufferSize=262144 (e.g., 256KB) to restrict the maximum size of cached thread-local direct buffers, forcing large buffers to be discarded instead of cached.

Tomcat vs Jetty vs Undertow

FeatureTomcatJettyUndertow
Default in Spring Bootβœ… YesNoNo
Threading modelThread-per-request (NIO)Thread-per-request (NIO)XNIO (non-blocking)
WebSocket supportβœ…βœ…βœ…
HTTP/2βœ…βœ…βœ…
Memory footprintMediumLowerLowest
Best forGeneral purposeLightweight appsHigh-performance, reactive

See also: Spring Boot Internals β€” Embedded Server Architecture for how Spring Boot auto-configures the servlet container.


3. Netty β€” Event Loop Architecture

What is Netty?

Netty is an asynchronous, event-driven network application framework for building high-performance protocol servers and clients. Unlike Tomcat's thread-per-request model, Netty uses a small number of event loop threads to handle thousands of connections simultaneously.

Netty is the engine behind: Spring WebFlux, gRPC-Java, Cassandra Driver, Elasticsearch transport, Vert.x, and Kafka clients.

How Netty Works Internally

🧡 Threading Layout: Tomcat (Blocking)

Conn 1 (Active)Conn 2 (Blocked)Conn 3 (Blocked)Thread-1[read] β†’ [process] β†’ [write]Thread-2🚨 BLOCKED on DB I/O waitThread-3🚨 BLOCKED on Socket read wait

Traditional Tomcat Model: 1 Thread = 1 Connection

Concurrency Model: Thread-Per-Request blocking model.

Required OS Threads: High (e.g. 200 Threads)

Connection Scaling: Capped at Thread Pool Size (typically 200 concurrent)

  • Architecture Details:
    • Each incoming HTTP connection is allocated a dedicated platform thread from the executor pool.
    • The thread is physically blocked during socket reads, writes, and database query roundtrips.
    • Under heavy load, threads consume ~1MB of heap stack each, leading to severe RAM exhaustions and OS context-switching overheads.

πŸ’‘ Switch between Tomcat Model and Netty Model tabs to analyze how thread blockages compare.

Netty Architecture (Boss-Worker Model)

Boss EventLoopSelector.select() β†’ Accept TCPWorker EventLoopGroupLoop 1[fd1, fd4]Loop 2[fd2, fd5]Loop 3[fd3, fd6]ChannelPipeline (Handler Chain)Decoderbytes $\rightarrow$ reqEncoderresp $\rightarrow$ bytesIdleStateHeartbeatBusinessLogic

Worker EventLoop Group

Capacity / Sizing: Default = 2 Γ— CPU Core Count

Core Concept: Non-blocking multi-connection multiplexer pool.

  • Architecture Guidelines:
    • Each Worker EventLoop runs a single thread executing a Selector.select() loop in an infinite cycle.
    • A single Worker EventLoop manages many Channels concurrently (Thread affinity keeps L1/L2 caches hot).
    • Fires read and write socket events and triggers the ChannelPipeline.

πŸ’‘ Click on Boss Group, Worker Group, or the ChannelPipeline block to inspect Netty socket routing internals.

Key Netty Concepts

ConceptWhat It IsAnalogy
ChannelAn open connection (socket)A phone line
EventLoopA single thread running an infinite select() loopA switchboard operator
EventLoopGroupA pool of EventLoopsThe operator team
ChannelPipelineChain of handlers processing dataAssembly line
ChannelHandlerA processing step (decode, encode, business logic)A station on the assembly line
ByteBufNetty's buffer (replaces java.nio.ByteBuffer)A smarter byte array with read/write indexes

Netty Configuration

EventLoopGroup bossGroup = new NioEventLoopGroup(1); // 1 thread for accepting
EventLoopGroup workerGroup = new NioEventLoopGroup(); // default: 2 Γ— CPU cores

ServerBootstrap b = new ServerBootstrap();
b.group(bossGroup, workerGroup)
.channel(NioServerSocketChannel.class)

// TCP backlog β€” OS-level queue for pending connections
.option(ChannelOption.SO_BACKLOG, 1024)

// Child channel options (per-connection settings)
.childOption(ChannelOption.TCP_NODELAY, true) // Disable Nagle's algorithm
.childOption(ChannelOption.SO_KEEPALIVE, true) // Enable TCP keep-alive
.childOption(ChannelOption.ALLOCATOR, PooledByteBufAllocator.DEFAULT) // Pooled memory

.childHandler(new ChannelInitializer<SocketChannel>() {
@Override
protected void initChannel(SocketChannel ch) {
ch.pipeline()
.addLast(new HttpServerCodec()) // HTTP encode/decode
.addLast(new HttpObjectAggregator(65536)) // Aggregate HTTP chunks
.addLast(new IdleStateHandler(60, 30, 0)) // Detect idle connections
.addLast(new MyBusinessHandler()); // Your logic
}
});
The Golden Rule of Netty

NEVER block an EventLoop thread. If your handler does blocking I/O (JDBC query, synchronous HTTP call, Thread.sleep()), you block that EventLoop β€” and ALL channels assigned to it are frozen.

// ❌ NEVER DO THIS in a ChannelHandler
@Override
public void channelRead(ChannelHandlerContext ctx, Object msg) {
// This blocks the EventLoop β†’ freezes hundreds of connections
String result = jdbcTemplate.queryForObject("SELECT ...", String.class);
ctx.writeAndFlush(result);
}

// βœ… Offload blocking work to a separate thread pool
private final EventExecutorGroup blockingGroup =
new DefaultEventExecutorGroup(16); // dedicated pool for blocking ops

ch.pipeline().addLast(blockingGroup, new MyBlockingHandler());

Netty and Direct Memory: Pooled Chunks & Reference Counting

Netty optimizes network performance by bypassing the secondary copy operation of Tomcat:

Netty Copy Pipeline (Zero-Copy Heap Path)NIC BufferHardware RingOS Socket BufferKernel TCP bufferPooled ByteBuf (Off-Heap)4MB Chunk AllocatorJava ObjectJVM Heap Space

Pooled Direct Memory (Off-Heap)

Memory Scope: JVM Off-Heap Native Memory

Overview: Pooled off-heap ByteBuf blocks managed by PooledByteBufAllocator.

  • Off-heap Optimization Implications:
    • Eliminates the intermediate copy to heap arrays, saving significant CPU cycles and GC load (Zero-Copy parser path).
    • Leases tiny slices (e.g. 64KB) from massive 4MB native chunks to handle socket payloads.
    • ⚠️ GC BLIND SPOT: Must be released manually using byteBuf.release(). Forgetting to release leaves chunks pinned off-heap forever, leading to container OOM kills.

πŸ’‘ Click on NIC Buffer, OS Socket, Pooled ByteBuf, or Java Object above to analyze Netty off-heap allocations.

By parsing packets directly on the off-heap direct buffer, Netty saves CPU cycles and prevents heap garbage accumulation, resulting in higher throughput and smoother p99 latency. However, because Netty maintains permanent residency on direct memory, it introduces critical off-heap management responsibilities.

Netty's Pooled Allocator (PooledByteBufAllocator)​

Direct memory allocations via the OS are expensive system operations. To avoid doing this on every request, Netty grabs large blocks of native memory called chunks (defaulting to 4MB a chunk since Netty 4.1.75) and manages them using a custom allocator.

  • Each HTTP/gRPC request is leased a tiny slice (e.g., 64KB) of a 4MB chunk.
  • Unlike Tomcat (which scales direct memory with thread count), Netty scales direct memory with the volume of concurrent active data across all connections at a given moment.
  • If client connections slow down or backpressure builds up, data accumulates in these direct buffers, causing the allocator to request more 4MB chunks from the OS.

The GC Blind Spot & Reference Counting Leak​

Because the Garbage Collector cannot see or clean off-heap chunks, Netty implements manual Reference Counting:

  • Each ByteBuf wrapper on the heap holds a reference counter. When you call .release(), the reference count drops. When it hits zero, the off-heap memory slice is immediately returned to the pool.
  • The Wrapper Leak Trap: If your code forgets to call release() (or fails to call it inside a finally block), the heap wrapper object eventually goes out of scope and is garbage collected. The heap remains completely clean, but the off-heap slice is never returned to the Netty pool.
  • The Pinning Effect: Netty can only return a 4MB chunk to the OS when every single slice leased from it has been released. A leak of a single 64KB slice is enough to pin the entire 4MB chunk in physical memory forever.
  • Remediation: Always run with Netty's leak detection enabled in development and staging:
    -Dio.netty.leakDetection.level=ADVANCED
    This tracks resource allocations and prints detailed stack traces when a ByteBuf is garbage-collected without being released.

See also: Socket Programming & I/O Models for epoll, the Reactor pattern, and how Netty uses them under the hood.


4. HikariCP β€” Database Connection Pooling

What is HikariCP?

HikariCP is the fastest JVM connection pool and the default in Spring Boot 2.x+. It manages a cache of pre-opened, pre-authenticated database connections that threads borrow and return β€” eliminating the 10–100ms overhead of establishing a new connection per request.

How HikariCP Works Internally

STEP 01

Thread-Local

⚑ ~250 nanoseconds

STEP 02

Shared List Steal

⚑ ~500 nanoseconds

STEP 03

Handoff Queue

⏳ Waits up to connection-timeout (Default: 30s)

ConcurrentBag: Step 1: Thread-Local Borrow (Fast Path)

Latency Overhead: ⚑ ~250 nanoseconds

Thread Contention: Zero contention (lock-free)

Overview: ConcurrentBag checks the calling thread's local list of connections. If a thread previously borrowed and returned a connection (e.g. C1), it acquires it immediately.

  • ConcurrentBag Mechanics:
    • Bypasses all volatile read/write fences and synchronized locks.
    • Saves high-concurrency CPU cycles by keeping local connection cache structures thread-confined.

πŸ’‘ Click on Step 01, 02, or 03 above to trace HikariCP connection acquisition internals.

HikariCP Configuration & Sizing

For details on how to configure HikariCP for production, fixed-size vs dynamic pools, and the baseline database sizing metrics, see the comprehensive Database Connection Pooling Guide.

For the step-by-step mathematical calculation of Tomcat thread pools vs. HikariCP pool sizes and database cores, see the Production Sizing Guide below.


5. How They All Relate

The Full Request Flow

Understanding how thread pools, Tomcat, Netty, and HikariCP interact in a single HTTP request is the key to diagnosing performance issues.

πŸ”„ Concurrency Model: Spring MVC (Tomcat)

HTTP ClientTomcat Worker Thread🚨 BLOCKED on DatabaseDatabase Server

Spring MVC: Thread-Per-Request (Blocking)

Thread Layout Sizing: 1 OS Thread per HTTP Connection

Memory Footprint: High Memory (~1MB per Thread stack)

Tuning Sizing: Inefficient under high concurrent I/O wait times.

  • Execution flow details:
    • An HTTP request borrows a dedicated worker thread (e.g. http-nio-8080-exec-42).
    • The worker thread is physically BLOCKED during database database SQL queries, remote API calls, or file reads.
    • Saturating the pool prevents the application from accepting future connections, causing high latency or downtime.

πŸ’‘ Toggle between Spring MVC, Spring WebFlux, and Java 21 Virtual Threads tabs to compare thread scheduling block overheads.

The Relationship Diagram

HTTP Web ServersTomcat (BIO/NIO)Netty (EventLoop)Database Connection PoolsHikariCP (JDBC)R2DBC Pool (Reactive)Application Thread Pools (TaskExecutor & Schedulers)@Async Executor@Scheduled PoolForkJoinPool (Streams)Virtual Thread Executor

HTTP Web Server Thread Pools

Overview: The front door of your web application. Responsible for handling incoming network requests (TCP connections), reading data, and dispatching requests into Spring controller mappings.

  • Configured Pools:
    • Tomcat Thread Pool (Default: max=200)
    • Netty EventLoopGroup (Default: 2 Γ— Cores)
  • Threading Impact:
    • Tomcat: Thread-per-request model. Blocks spare threads during I/O blockages.
    • Netty: Non-blocking EventLoop model. Uses few threads to manage thousands of active requests.

πŸ’‘ Click on HTTP Web Servers, Database Connection Pools, or Application Thread Pools above to trace JVM threading boundaries.

Concurrency Model Comparison

AspectTomcat (BIO/NIO)NettyVirtual Threads
ModelThread-per-requestEvent loopVirtual-thread-per-request
Threads needed~200 for 200 concurrent~8 for 10,000+ concurrentMillions possible
Blocking I/OBlocks a worker thread❌ Must never blockβœ… Safe β€” unmounts from carrier
Code styleSimple imperativeCallback/reactiveSimple imperative
Memory per connection~1MB (thread stack)~1KB (channel state)~1KB (heap continuation)
Spring integrationSpring MVCSpring WebFluxSpring MVC (3.2+)
DB accessJDBC + HikariCPR2DBC (reactive)JDBC + HikariCP
Best forTraditional CRUD APIsHigh-connection serversI/O-heavy APIs on Java 21+

6. Production Sizing Guide

Sizing thread pools and connection pools is not about choosing arbitrary numbers. It is a mathematical chain of constraints extending from your user request rates down to your physical database cores.


1. Tomcat Thread Pool Sizing

Tomcat's default pool size of 200 worker threads is often an anti-pattern when running in modern containerized environments (Kubernetes pods or Docker containers) capped at 1 or 2 vCPUs. Too many active threads cause context-switching overhead, CPU throttling, and cache invalidation.

To calculate the optimal worker thread count, use the Brian Goetz formula (from Java Concurrency in Practice):

Optimal ThreadsΒ =Β Available CoresΒ Γ—Β Target CPU UtilizationΒ Γ—Β ( 1 + Wait TimeCompute Time )
  • Available Cores: The CPU core limits allocated to your container (e.g. resources.limits.cpu in Kubernetes).
  • Target CPU Utilization: The desired average CPU load (typically 0.8 or 80% to leave headroom for GC, serialization, and traffic spikes).
  • Wait Time / Compute Time (Blocking Coefficient): The ratio of time a request spends waiting for off-thread I/O (database, cache, HTTP downstream calls) vs. actually processing on the CPU.

Sizing Example​

Suppose a Spring Boot microservice is deployed inside a container with 2 vCPUs. Performance profiling shows that a typical request takes 55ms total: 50ms waiting for the database (Wait Time) and 5ms computing JSON serialization and business logic (Compute Time).

Tomcat ThreadsΒ =Β 2Β Γ—Β 0.8Β Γ—Β ( 1 + 505 )Β =Β 1.6Β Γ—Β 11Β =Β 17.6Β &approx;Β 18 threads

Instead of the default 200 threads, setting server.tomcat.threads.max = 18 is the correct, mathematically derived starting point for performance testing.

Calculating Throughput Capacity (Little's Law)​

Using Little's Law (L = Ξ» Γ— W), we can calculate the theoretical throughput of a single instance:

  • L (Number of concurrent requests in-flight) = 18 (our max threads).
  • W (Latency per request) = 55ms (0.055 seconds).
  • Ξ» (Throughput / Requests Per Second) = L / W.
ThroughputΒ =Β 180.055Β &approx;Β 327 RPS per instance

If your business requirement is to handle 1,600 RPS, you will need to scale out to at least 5 application instances (1600 / 327).

⚠️ Container Core-Counting Warning​

In Java versions prior to 8u191, the JVM was unaware of container limits and read the host's physical cores (e.g. 64 cores). This caused the JVM to spawn too many internal threads (GC, JIT, ForkJoinPool), leading to massive CPU throttling.

  • Fix: Use modern JVMs (Java 11, 17, 21, or 25) which are container-aware by default and properly read cgroups memory and CPU limits.

⚑ Cloud Run Concurrency Aligning​

If you deploy to serverless containers like Google Cloud Run, there is a concurrency setting (max concurrent requests routed per instance).

  • If you set concurrency = 80 (default) but configure Tomcat threads.max = 18, Cloud Run will route up to 80 requests to a single instance. Tomcat will process 18, and the remaining 62 will sit in Tomcat's TaskQueue, causing request latency to explode.
  • Fix: Keep concurrency aligned closely with your Tomcat threads.max (e.g., 20). This forces Cloud Run's native load balancer to scale-out horizontally to a new pod immediately rather than queuing requests internally.

2. Database Connection Pool Sizing (HikariCP)

Once the application threads are sized, the connection pool must be sized to support them. In a thread-per-request architecture, a worker thread only needs a database connection during the database execution phase of the request, not for the entire request lifecycle (e.g. not during CPU-bound serialization or external HTTP calls).

To calculate the connection pool size per instance, use the formula:

Hikari Pool SizeΒ =Β max-threadsΒ Γ—Β Connection Hold TimeTotal Request Time

Using our 2 vCPU example (18 threads, total request 55ms, connection held for 50ms):

Hikari Pool SizeΒ =Β 18Β Γ—Β 5055Β =Β 16.3Β &approx;Β 17 connections

Set maximum-pool-size: 17 and minimum-idle: 17 to keep the pool warm and avoid connection handshake latency during traffic spikes.

πŸ“‰ Why Smaller Pools are Faster​

A common trap is assuming that more connections equal higher throughput. The database engine can only process queries in parallel up to its hardware limits. Excess active connections result in disk spindle thrashing, lock contention, and OS thread context switching, which degrades throughput and causes latency spikes.

The optimal connection limit for a database is defined by the PostgreSQL/HikariCP formula:

Optimal Executing ConnectionsΒ =Β ( Database CPU CoresΒ Γ—Β 2 )Β +Β Spindle
  • Spindle: The number of physical hard disks. On modern SSDs or NVMe drives where the working set fits in cache, this is essentially 0.
  • A 4-Core DB running on SSDs can only run (4 Γ— 2) + 0 = 8 queries in parallel optimally.

Resolving the Mismatch​

If 5 application instances each open 17 connections, the database has 85 total connections open. How does this align with the DB's optimal limit of 8-9 executing queries?

The key is distinguishing between open connections (idle/waiting on network) and executing queries (utilizing database CPU).

  • Out of the 50ms database phase, the database CPU might only spend 5ms executing the query. The other 45ms is network transit, connection checkout, and client-side data buffering.
  • Out of the 85 connections open from the cluster, the concurrent active executing queries are:
Active executing queriesΒ =Β 85Β Γ—Β 5ms50msΒ =Β 8.5 queries

This matches the 4-core database capacity perfectly!

Complete Sizing Chain Example​

To support 1,600 RPS under the 55ms total latency profile:

  1. Instances: 1600 / 327 β‰ˆ 5 instances.
  2. Hikari Pool size per instance: 18 Γ— (50 / 55) β‰ˆ 17. (Total cluster connections = 17 Γ— 5 = 85).
  3. Database execution load: 85 Γ— (5ms / 50ms) = 8.5 active executing connections.
  4. Database sizing: Cores = Executing Connections / 2 = 8.5 / 2 β‰ˆ 4.25 β‰ˆ 4 to 8 cores.

Recommendation: Set the database max_connections limit significantly higher (e.g., 150 to 200) to provide headroom for administrator logins, indexing jobs, and monitoring metrics, even though the cluster pool only checks out 85.

⏳ Connection Hold Time Leaks​

If connection usage spikes in production while the database CPU is idle, check for hold time leaks:

  • The Transaction Trap: Placing @Transactional annotations on outer service methods that call slow external REST APIs keeps the database connection checked out doing absolutely nothing while waiting for the network call.
  • Fix: Keep transactions short. Only hold connections during database operations. Use Hikari's leak-detection-threshold to log warnings for connections held longer than a specific limit (e.g. 5 seconds).

3. Virtual Threads (Project Loom) Sizing Impact

When virtual threads are enabled (spring.threads.virtual.enabled=true), the Tomcat thread pool bottleneck disappears because virtual threads do not require a 1MB native OS stack.

  • The Trap: If you have 5,000 concurrent requests, Spring will spawn 5,000 virtual threads. However, your database connection pool does not scale.
  • The Result: All 5,000 virtual threads will block at the gates of the HikariCP pool waiting for a connection, leading to connection timeouts. Virtual threads shift the application concurrency bottleneck entirely down to the database connection layer.
  • Fix: Use semaphores, rate limiters, or Spring's @ConcurrencyLimit annotations to cap downstream resource access, preventing database pool exhaustion under Loom.

The Mismatch Deadlock

A critical failure mode when thread pool and connection pool sizes don't match:

🚫Deadlock Simulator: Step 1: 200 Requests Arrive

Step 1 of 8
Tomcat Worker ThreadsThread 1Threads 2-10Threads 11-200 (Blocked)HikariCP Connection Pool10 Checked-Out ConnsRemaining: 0 Available

Step 1: 200 Requests Arrive

Tomcat accepts all requests into its network buffers. The first 10 worker threads instantly start executing.

πŸ’‘ Step through the simulator using the controls above to understand how sub-requests cause starvation deadlocks.


7. Troubleshooting & Common Failures

Symptoms β†’ Diagnosis β†’ Fix

SymptomLikely CauseDiagnosisFix
Response times spike under loadPool starvation (thread or connection)Check hikaricp.connections.pending > 0 or high thread countReduce connection-timeout, fix slow queries
SQLTransientConnectionExceptionAll connections borrowed, timeout expiredhikaricp.connections.timeout counter increasingIncrease pool size OR fix connection hold time
RejectedExecutionExceptionThread pool + queue both fullThread dump shows all threads busyIncrease queue or threads; fix slow handlers
CPU at 100% with no useful workToo many threads β†’ context switchingvmstat shows high cs (context switch) rateReduce thread count
Memory growing (OOM)Unbounded queue or thread countHeap dump shows many task objects or threadsUse bounded queues; explicit ThreadPoolExecutor
Tomcat stops accepting requestsAll 200 worker threads blockedThread dump shows all threads in WAITING on HikariCPFix connection leak; reduce connection-timeout
Netty EventLoop blockedBlocking call in a ChannelHandlerSlow channel handlers, increasing event loop latencyOffload blocking work to separate executor

Timeout Exceptions Deep Dive

When connections start failing under load, clients will log network exceptions. Rather than treating them as unrelated glitches, recognize that Connection refused, Connect timed out, Read timed out, and Connection reset are different phases of the same congestion problem.

CLOCK 01

Connect Timeout

SYN Handshake

CLOCK 02

Read Timeout

Active Wait

CLOCK 03

Connection Reset

Server Idle Timeout

Exception Phase: Connect timed out / Connection refused

Clock Pipeline: Connect Timeout Clock (SYN Handshake phase)

Root Cause: Server TCP accept queue (backlog) or max-connections fully saturated.

Overview: Occurs during the initial TCP 3-way handshake. The client sends a SYN packet, but cannot get a response back.

  • Deep-Dive Details:
    • Connection refused: The server actively sends a RST (Reset) packet. Indicates no server process is listening on the port.
    • Connect timed out: Sockets beyond server.tomcat.max-connections (8192) queue up in the OS backlog (accept-count, default 100). When the queue overflows, the kernel silently drops SYN packets, forcing client clocks to timeout.

πŸ’‘ Toggle between Connect Timeout, Read Timeout, and Connection Reset above to trace network failure phases.

Understanding these clocks and server settings allows you to pinpoint precisely where a request fails:

1. Connection refused vs. Connect timed out (TCP Handshake)​

  • What it means: The client attempts to initiate the 3-way TCP handshake (sends a SYN packet) but cannot complete the connection.
  • The Root Cause:
    • Connection refused (ConnectException): The server kernel actively rejects the connection by replying with a RST (Reset) packet. This happens if the target port has no process listening on it, or if the server process has completely shut down.
    • Connect timed out (SocketTimeoutException: connect timed out): The target port is open, but Tomcat's connection capacity is exceeded.
  • The Tomcat Mechanism:
    • Tomcat accepts up to server.tomcat.max-connections (default 8192) active sockets.
    • Sockets beyond this are queued in the OS Kernel TCP Accept Queue, sized by server.tomcat.accept-count (default 100).
    • Request 8293+: When both the 8192 active slots and the 100 queue spots are full, the Linux kernel silent drops incoming SYN packets. The client receives no response, retries the handshake, and eventually gives up when its client-side connectTimeout clock expires.
  • Troubleshooting:
    • If logging Connect timed out, check if your cluster is undersized (RPS is exceeding cluster capacity).
    • Trap: Increasing accept-count to a massive number (e.g., 5000) just creates a longer queue, which eventually converts into Read timed out exceptions as clients wait too long for their turn in the queue.

2. Read timed out (Socket Wait State)​

  • What it means: The TCP handshake completed successfully, the connection was checked out, and the client sent the HTTP request payload. However, the client-side readTimeout expired before the server sent back a response.
  • The Root Cause: The Tomcat thread pool (threads.max, default 200) is fully saturated. Threads are blocked waiting on slow downstream microservices, unindexed database queries, or connection pools.
  • The Mechanism:
    • Tomcat accepted the connection into its max-connections buffer, but no worker thread is free to parse or process the HTTP headers. The request sits idle.
    • The client waits, its readTimeout expires, and the client throws SocketTimeoutException: Read timed out and terminates the socket.
    • Nghα»‹ch lΓ½: The client reports timeouts, but Tomcat logs remain blank and CPU utilization is low. The client aborted the request, but the blocked Tomcat thread is still running the query in the background, unaware the client has departed.
  • Troubleshooting:
    • Do not blindly increase the client readTimeout. This simply holds resources (sockets and client threads) blocked for longer.
    • Locate the thread bottleneck: take a thread dump (jstack) and inspect why Tomcat threads are in WAITING or BLOCKED states.

3. Connection reset / Broken pipe (Server-Side Eviction)​

  • What it means: The TCP socket was open, but the server unilaterally closed the connection, causing the client's next write operation to fail.
  • The Root Cause: The client opened a socket but did not send any bytes within Tomcat's configured server.tomcat.connection-timeout limit (default 20000ms / 20 seconds).
  • The Mechanism:
    • Tomcat keeps open idle TCP connections to support Keep-Alive. However, if a client holds a socket open but doesn't send HTTP headers (e.g., slow clients, network hiccups, port scanning scripts), Tomcat closes the socket to free up resources.
    • If the client tries to send data on this closed socket, the OS returns a RST packet, throwing IOException: Connection reset by peer or Broken pipe on the client.
  • Troubleshooting:
    • Verify if clients are experiencing high network latency or sending headers slowly. If keep-alive connections are being recycled too quickly, tune server.tomcat.connection-timeout carefully.

Essential Metrics to Monitor

# Spring Boot Actuator + Micrometer

# Tomcat thread pool
tomcat.threads.current # Current thread count
tomcat.threads.busy # Threads actively processing requests
tomcat.threads.config.max # Maximum configured threads

# HikariCP connection pool
hikaricp.connections.active # Connections currently in use
hikaricp.connections.idle # Connections sitting idle
hikaricp.connections.pending # Threads waiting for a connection ← ALERT if > 0
hikaricp.connections.timeout # Connection borrow timeouts (cumulative)
hikaricp.connections.usage # Connection hold time histogram

# JVM threads
jvm.threads.live # Total live threads
jvm.threads.peak # Peak thread count since startup
jvm.threads.daemon # Daemon threads

Thread Dump Analysis

# Get a thread dump of your Java process
jstack <pid> > threaddump.txt

# What to look for:
# 1. Many threads in WAITING state on HikariCP
"http-nio-8080-exec-42" WAITING
at com.zaxxer.hikari.pool.HikariPool.getConnection(HikariPool.java:162)
β†’ Pool starvation β€” all connections are borrowed

# 2. Many threads BLOCKED on synchronized
"http-nio-8080-exec-15" BLOCKED
at com.example.LegacyService.criticalSection(LegacyService.java:42)
β†’ Lock contention β€” single synchronized method is a bottleneck

# 3. Deadlock detected
"Found one Java-level deadlock:"
β†’ Thread A holds Lock 1, waits for Lock 2
β†’ Thread B holds Lock 2, waits for Lock 1

Senior Deep Dive: The RUNNABLE Database Call Illusion

When database queries slow down, you might take a thread dump to diagnose the issue, only to find a paradox: dozens of threads are blocked waiting for database results, yet their JVM state is reported as RUNNABLE rather than WAITING or BLOCKED.

1. Why JVM States Mismatch Reality​

To understand why a waiting thread reports as runnable, we must look at where thread states are managed:

  • JVM-Managed States (BLOCKED, WAITING, TIMED_WAITING): These states represent synchronization queues controlled entirely inside the JVM's memory boundary.
    • BLOCKED means a thread is waiting to acquire a Java monitor lock (to enter a synchronized block).
    • WAITING / TIMED_WAITING means the thread is parked inside the JVM via Object.wait(), Thread.join(), or LockSupport.park(), waiting for another Java thread to wake it up.
  • OS-Level Blocking (The JVM Blind Spot): When a Java thread issues a blocking database query via JDBC, the JVM execution engine drops down into native code to perform an OS kernel system call (syscall) to read from a TCP socket (seen in thread dumps as socketRead0).
    • The thread is blocked at the Operating System kernel level, waiting for the network card to receive TCP database packets.
    • Because the wait occurs outside the JVM's synchronization structures, the JVM cannot track it portably across different operating systems.
    • Therefore, the JVM maintains the thread state as RUNNABLE (defined by the Java spec as "executing in the JVM but may be waiting for other resources from the operating system").

The Timesheet Analogy: Imagine an office check-in sheet. When an employee quets their card, they are marked "In Office / Working" on the timesheet. If they sit at their desk staring at a loading screen waiting for an external vendor to email them files, the HR department (JVM) still marks them as "Working" because they haven't checked out. They are only marked "Away" when in an official internal state (e.g. locked out of a meeting room β€” BLOCKED, or waiting for a colleague β€” WAITING).

2. The Traditional Platform Thread Bottleneck​

In classic Java, each platform thread maps 1:1 to a physical OS thread (each consuming a fixed 1MB native stack).

  • When a database query blocks, the underlying OS thread is pinned in the kernel. It cannot do any other work.
  • An I/O-bound microservice spending 90% of its time waiting on network round-trips will quickly exhaust its Tomcat thread pool (threads.max = 200).
  • The Symptom: You observe 200 threads in RUNNABLE (actually blocked in socketRead0 syscalls), CPU utilization is idle at 10-20%, but the application is starved, throwing connection timeouts.

3. How Java 21+ Virtual Threads Resolve the Illusion​

Virtual Threads (Project Loom) decouple logical threads from OS threads, running thousands of virtual threads on a small pool of platform carrier threads:

  • Unmounting on I/O: When a virtual thread executes a blocking socket read (like database queries), the JDK intercepts the call. Instead of blocking the carrier OS thread, the JDK unmounts the virtual thread, serializes its stack frame onto the Java Heap as a continuation, and frees the carrier thread to run other virtual threads.
  • Accurate States: Because the JVM scheduler now manages the block, the virtual thread's state changes to WAITING (and Thread.getState() correctly returns WAITING).
  • Pinning Limit in Java 21/23: If a virtual thread blocks inside a synchronized block or native method, it gets pinned to the carrier thread, reverting to the old 1:1 behavior.
  • Java 24+ Fix (JEP 491): Java 24 completely resolves pinning inside synchronized blocks, allowing virtual threads to unmount freely.

4. Diagnostic Caveat: The jstack Blind Spot​

Traditional profiling tools like jstack only display platform/carrier threads. If your service uses virtual threads and hangs, a standard jstack output will show a completely idle, clean JVM.

  • To dump virtual threads: Run the following jcmd command to output all virtual thread stacks:
    jcmd <PID> Thread.dump_to_file -format=json threads.json
    # Or plain text format:
    jcmd <PID> Thread.dump_to_file threads.txt

Interview Questions

Q: Explain the relationship between Tomcat's thread pool and HikariCP's connection pool.

A: Tomcat's thread pool handles HTTP requests β€” each request gets a worker thread. When that request needs the database, the worker thread borrows a connection from HikariCP. The Tomcat thread is blocked until the DB query completes and the connection is returned. If HikariCP has fewer connections than Tomcat has threads (common: 200 threads vs 10–20 connections), excess threads queue. This is fine for fast queries (under 5ms) but dangerous for slow queries β€” threads pile up waiting, leading to pool starvation and cascading timeouts.

Q: Why does Netty need far fewer threads than Tomcat?

A: Tomcat uses thread-per-request: each thread blocks during I/O. With 200 threads, you handle 200 concurrent requests max. Netty uses the Reactor pattern: a few EventLoop threads multiplex thousands of connections via epoll/kqueue. When data isn't ready on a socket, the EventLoop serves another channel instead of blocking. This means 8 threads can handle 10,000+ concurrent connections β€” but you must never block an EventLoop thread, or all its channels freeze.

Q: What happens when you enable virtual threads in Spring Boot 3.2+?

A: Setting spring.threads.virtual.enabled=true makes Tomcat use virtual threads instead of platform threads for request handling. Each request gets its own virtual thread (not from a fixed pool). When the virtual thread blocks on I/O (JDBC query, HTTP call), it unmounts from the carrier thread β€” the carrier is free to run other virtual threads. This gives Tomcat-like simplicity (blocking code) with Netty-like efficiency (threads aren't wasted during I/O). The new bottleneck shifts from threads to connection pools β€” you must size HikariCP and use Semaphores to prevent 100K virtual threads from overwhelming the database.

Q: How would you size a HikariCP pool for a 4-pod deployment against an 8-core RDS instance?

A: Use the formula: connections = (CPU_cores Γ— 2) + 1 = 17. Round to 20 total connections. With 4 pods: 20 / 4 = 5 connections per pod. Set maximum-pool-size: 5 and minimum-idle: 5 (fixed pool). If you scale to 8 pods without adjusting, you'd get 40 total connections β€” overloading the DB. Either reduce per-pod pool size or add PgBouncer/RDS Proxy as a connection multiplexer.

Q: Why is Executors.newFixedThreadPool() considered dangerous?

A: It uses an unbounded LinkedBlockingQueue. If tasks arrive faster than threads can process them, the queue grows without limit β€” consuming heap memory until OutOfMemoryError. In production, always use ThreadPoolExecutor directly with a bounded ArrayBlockingQueue and a rejection policy like CallerRunsPolicy for backpressure.

Q: How do you diagnose pool starvation?

A: Monitor hikaricp.connections.pending (threads waiting for connections) and hikaricp.connections.timeout (failed borrows). Take a thread dump β€” if many threads are in WAITING state at HikariPool.getConnection(), the pool is starved. Root causes: slow queries (N+1, missing indexes), connections held during non-DB work (@Transactional wrapping HTTP calls), or pool too small for the workload. Fix the query first; increase pool size only as a last resort.


Cross-References

TopicLink
Concurrency vs. Parallelism beginner guideConcurrency vs. Parallelism
ThreadPoolExecutor & Fork/Join detailsJava Concurrency
Virtual Threads deep diveVirtual Threads (Project Loom)
HikariCP anti-patterns & PgBouncerDatabase Connection Pooling
Netty, epoll, and the Reactor patternSocket Programming & I/O Models
Tomcat embedded server internalsSpring Boot Internals
JVM memory & thread stacksJVM Memory Architecture
Spring Boot server tuningSpring Boot Advanced
πŸ“–
Track Page Progress0 / 635 Read
Knowledge Base Completion0%