KOEN
All notes
SERVER & OPERATIONS

The CPU is idle. The request is still waiting.

An imagined order request shows why waiting for a database connection and holding one need separate clocks.

Paper request tickets queue at a narrow brass gate beside an idle workshop machine

Quiet inside

The CPU graph looks quiet. Someone presses the order button and never gets to the next screen. In this hypothetical situation, ‘Why is it slow when the server is not busy?’ already needs a better definition of busy. Waiting takes time without doing much computation.

The server monitoring guide asks why database connections are being returned slowly before increasing the pool size. Adding concurrent work to an already slow database can make matters worse.

That detail makes me want to follow one imaginary order API. It borrows a database connection, saves an order, notifies a delivery service and sends a response.

The customer experiences one wait. Inside the application, we can give it several timestamps: request arrival, connection requested, connection acquired, database work finished, connection returned.

Which gap is growing?

What happens while the key is borrowed?

Assume connection acquisition is taking longer in this example. That alone does not prove the query is slow. First look at what requests holding connections are doing.

Suppose the code saves the order, then waits for the delivery service before returning the connection. The database has finished its work, but the next request cannot borrow a connection. An idle CPU would be quite compatible with that scene.

This is an invented example, not an incident report. It explains why acquisition time and holding time deserve separate clocks. A slow query, a lock and an external call made while holding a connection require different investigations.

A dashboard showing only ‘DB time’ leaves a question unanswered. Does that mean query execution, or does it include getting a connection? Two measurements can share a unit while measuring entirely different waits.

Adding connections might briefly make this hypothetical order endpoint faster without repairing the code. When the external service slows again, more requests could stand around holding connections. That is a reason to inspect the return path before choosing a pool size.

Nor would I blindly shorten the transaction. The order still needs a defined set of changes that commit together, and a separate plan for retrying a failed external notification.

A failure that finishes quickly

Change the scene once more. Waiting requests start failing early because they cannot obtain a connection. Average response time might look better. The customer still cannot place an order.

Google’s SRE book explains why successful and failed request latency should be separated: fast errors can distort the combined number. Before appreciating a lower value, ask what response produced it.

For the imagined order service, I would start with successful order latency and the error rate. Then put arrival volume, acquisition waits and connection holding time beside them for the same period. If the charts cover different requests or intervals, show that difference.

Some part of the wait may remain invisible. Coloring an uninstrumented segment as healthy makes the picture tidier without explaining it. Marking the gap as unknown gives the next investigation somewhere to begin.

This is a thought experiment about dividing a request’s elapsed time, not a reproduction of an outage or a prescription for the right pool size.

There is a queue outside and an empty chair inside. Before adding another counter, I would like to know who still has the key.

Sources