backenddrills

Database connections

Diagnosing HikariCP connection pool exhaustion

By the BackendDrills editorial team · Published and checked October 6, 2026 · 6-minute read

Before you start: Basic HTTP requests and database connections. Find this reading in a study path →

Pool exhaustion means callers cannot obtain a connection within the configured wait budget. It does not, by itself, prove that the database needs more CPU or that the pool should be enlarged. Compare acquisition wait, connection hold time and database work before changing capacity.

Build a timeline for one slow request

Consider an illustrative reservation service with a 16-connection pool. Sixteen requests hold connections while waiting on a supplier API; another 45 requests queue for a connection. SQL takes 18 milliseconds, but the supplier request takes 4 seconds. A quiet database can coexist with a saturated application pool.

Illustrative measurements, not an actual incident:
pool active=16 idle=0 pending=45
connection acquisition p95=1.8 s
connection hold p95=4.2 s
database statement p95=18 ms
supplier HTTP p95=4.0 s

Instrument the time before obtaining a connection, after acquisition, and at release. Add trace spans for database calls and remote dependencies. Do not infer hold time from statement duration: the application may keep a connection between statements or across unrelated work.

Separate the competing explanations

EvidenceHypothesis to test
Long hold time; short SQLRemote I/O or application work inside the resource boundary
Long SQL and lock waitsA database query, lock or transaction problem
Connections never return after requests finishA resource cleanup failure
Saturation tracks a traffic burstAdmission and concurrency exceed sustainable throughput

HikariCP's connectionTimeout bounds the wait to acquire a connection; it is not a query timeout. maximumPoolSize bounds pooled connections. A leak-detection warning identifies a connection held beyond its threshold, so investigate its stack and duration before declaring a permanent leak.

Choose a bounded containment action

For the supplier example, cap concurrent reservation work and give remote calls a finite timeout. Check that failed or cancelled requests release connections. Moving supplier work outside the transaction may help, but first decide how to handle supplier success followed by a database failure. Resource efficiency is useful only if the reservation invariant survives.

If measurements justify a larger pool, account for every replica and other database clients. A pool of 16 on ten replicas can request 160 database connections. More connections can move the queue into the database rather than reduce end-to-end latency.

Verify recovery under the same pressure

  1. Reproduce the slow dependency with a controllable stub.
  2. Measure successful business operations, acquisition wait and hold duration at the same arrival rate.
  3. Cancel some requests and check that active connections return to baseline.
  4. Remove the delay and confirm the backlog drains within a stated time.

Include both throughput and correctness. A lower timeout count is insufficient if the service has silently dropped reservations. Record which measurement would disprove your leading diagnosis before increasing any limit.

Check the underlying behavior

Original illustrative examples, prepared with AI assistance and checked against the linked primary documentation. No customer incident or vendor endorsement is claimed. Our editorial approach.

Practice the next decision

Try a complete free backend drill: inspect evidence, make three decisions and review the reasoning. No account or card required.

Try a free incident drill →

Explore 1008 scenarios · $28 one-time

Selected practice from the study paths

Free readings need no account. Full edition drills require verified access; opening a paid link does not expose its answers.

Recommended next readings