Database connections
Diagnosing HikariCP connection pool exhaustion
By the BackendDrills editorial team · Published and checked October 6, 2026 · 6-minute read
Before you start: Basic HTTP requests and database connections. Find this reading in a study path →
Pool exhaustion means callers cannot obtain a connection within the configured wait budget. It does not, by itself, prove that the database needs more CPU or that the pool should be enlarged. Compare acquisition wait, connection hold time and database work before changing capacity.
Build a timeline for one slow request
Consider an illustrative reservation service with a 16-connection pool. Sixteen requests hold connections while waiting on a supplier API; another 45 requests queue for a connection. SQL takes 18 milliseconds, but the supplier request takes 4 seconds. A quiet database can coexist with a saturated application pool.
Illustrative measurements, not an actual incident:
pool active=16 idle=0 pending=45
connection acquisition p95=1.8 s
connection hold p95=4.2 s
database statement p95=18 ms
supplier HTTP p95=4.0 s
Instrument the time before obtaining a connection, after acquisition, and at release. Add trace spans for database calls and remote dependencies. Do not infer hold time from statement duration: the application may keep a connection between statements or across unrelated work.
Separate the competing explanations
| Evidence | Hypothesis to test |
|---|---|
| Long hold time; short SQL | Remote I/O or application work inside the resource boundary |
| Long SQL and lock waits | A database query, lock or transaction problem |
| Connections never return after requests finish | A resource cleanup failure |
| Saturation tracks a traffic burst | Admission and concurrency exceed sustainable throughput |
HikariCP's connectionTimeout bounds the wait to acquire a connection; it is not a query timeout. maximumPoolSize bounds pooled connections. A leak-detection warning identifies a connection held beyond its threshold, so investigate its stack and duration before declaring a permanent leak.
Choose a bounded containment action
For the supplier example, cap concurrent reservation work and give remote calls a finite timeout. Check that failed or cancelled requests release connections. Moving supplier work outside the transaction may help, but first decide how to handle supplier success followed by a database failure. Resource efficiency is useful only if the reservation invariant survives.
If measurements justify a larger pool, account for every replica and other database clients. A pool of 16 on ten replicas can request 160 database connections. More connections can move the queue into the database rather than reduce end-to-end latency.
Verify recovery under the same pressure
- Reproduce the slow dependency with a controllable stub.
- Measure successful business operations, acquisition wait and hold duration at the same arrival rate.
- Cancel some requests and check that active connections return to baseline.
- Remove the delay and confirm the backlog drains within a stated time.
Include both throughput and correctness. A lower timeout count is insufficient if the service has silently dropped reservations. Record which measurement would disprove your leading diagnosis before increasing any limit.
Check the underlying behavior
Original illustrative examples, prepared with AI assistance and checked against the linked primary documentation. No customer incident or vendor endorsement is claimed. Our editorial approach.
Practice the next decision
Try a complete free backend drill: inspect evidence, make three decisions and review the reasoning. No account or card required.
Try a free incident drill →Selected practice from the study paths
- The database is quiet. The pool is full. · Complete free drill
- The cache expires. The database melts. · Full edition drill
- The average hides the slow customers · Full edition drill
Free readings need no account. Full edition drills require verified access; opening a paid link does not expose its answers.