Anatomy of a Web Request Primer
curl https://api.example.com/user/42 takes 137 milliseconds. Figures you can drive prove three things about them: the twelve stages are strictly serial, so the total is a sum; the server owns 1.03 ms of the 137; and 101 ms is setup a reused connection deletes outright. Twenty primers are linked at the stage each one owns.
One command, twelve stages, 137 milliseconds
A JSON object lands on stdout in about a seventh of a second. Almost none of that time belongs to the server.
The command is curl https://api.example.com/user/42, over a30 ms path, with a cold DNS cache and no connection open. At the far end, nginx fronts a process that does one primary-key lookup in PostgreSQL and returns 2 KiB of JSON. Nothing about it is exotic, which is why it is worth measuring.
The twelve stages run strictly one after another, because each consumes what the one before produced: you cannot open a socket to an address you have not resolved. The stage under the clock is amber and everything behind it is teal. Press play, or drag the clock across the picture:
Notice that the stages already paid for pile up in one row. 5 of the 12 bars sit in the wire lane and carry 131 ms of the 137 ms, while the 5 server-side stages together are a sliver too thin to have a width. The total is a sum rather than a maximum precisely because the stage under the clock is always the only one running.
That is worth reading lane by lane rather than taking on trust. Walk the slider through the five lanes and the one you select reports how many stages it owns and what share of the request it costs:
The server application — parse, query, serialise — comes to 960 µs. Add both kernels and the whole machine-side cost is 1.03 ms: well under one percent of what the caller waited for. Whatever you do inside that handler, you are tuning 0.7% of the number your user experiences.
The reason is arithmetic, not engineering. One 30 ms round trip is so large a unit that everything inside a machine vanishes against it — drag the marker down to each smaller latency and read how many of them fit inside one:
375,000 main-memory accesses fit inside a single round trip, and 30,000,000 L1 hits do. The cache miss a CPU engineer calls catastrophic is under three millionths of the cheapest thing a network can do. That ratio is why the rest of this page is mostly a story about round trips.
It holds at every distance, which is the part that surprises people. Drag the round trip from a same-rack 1 ms out to a cross-continent 61 ms and watch the wire and the server move independently:
The server's bar never moves. It is 1.03 ms at every setting: nothing the server does depends on distance, and the wire's bar is the only thing that grows. So the page collapses into one cost model: total ≈ 5 ms of process start + DNS + 3 × RTT + 2.4 ms. At a 1 ms round trip that is 50.4 ms; at 61 ms it is 230 ms. The rest of this page argues about which of those four terms you can delete.
Before there is a packet
45 ms of the 137 ms are gone before the client even has an address to connect to, and neither culprit is the network.
Pressing Enter makes the shell fork and then execve a binary that is not in memory yet. The kernel builds an address space, maps the ELF and its six shared objects, and hands control to the dynamic linker — see Process & Thread and Virtual Memory.
Almost none of that 5 ms is the fork. Walk the slider through the five phases and watch the clock accumulate on the two nobody expects, with everything already done behind it:
Notice where the mass sits. Relocation — the linker patching every symbol reference in six libraries — is 2.1 ms, and demand-paging the text is another 1.6 ms, at about 1,340 minor faults. Both are the price of starting, not of working, which is the whole argument for a long-lived client over a per-request one. The allocator behind it is in Memory Allocation.
Then getaddrinfo, which costs 20 µs or 40 ms — a swing of 2,000× decided entirely by who already knows the answer. Drag the slider up the delegation and watch the hop that answers move away from you, with the hops in front of it all having to be walked:
The chain empties from the far end, and the order matters: an authoritative answer expires long before the .com delegation, whose NS records carry a two-day TTL. So the common cold lookup is not the 40 ms walk from the root — it is the 25.1 ms one that starts at .com, with the rest of the delegation still cached. DNS & HTTP has the record types.
The failure mode is worse than the cost, because DNS rides UDP and nothing retransmits it but the caller. Drag the resolver's timeout and the number of nameservers in resolv.conf, and read how long the calling thread is blocked by attempts nobody answers:
At the glibc defaults — timeout:5, attempts:2, two nameservers — a black-holed query blocks the caller for 20 seconds and returns EAI_AGAIN. It fails loudly, which is the good case, but it fails after every timeout above it has already fired. Synchronous DNS inside a request handler is a bug, not a performance issue.
Two handshakes before one byte of HTTP
The name is resolved and nothing useful has been said yet. Two negotiations stand between us and the request.
socket() costs a kernel allocation and connect() costs one TX descriptor — 60 µs for the pair. Everything after that is waiting for a peer to answer, and the peer is 30 ms away. Both handshakes are covered in depth in TCP Deep Dive and TLS & Security.
Six messages cross before GET /user/42 can go out. Step through them and watch the message in flight take half a round trip each time, while the ones already delivered stay put:
Watch which half of each round trip is idle. The client sends the ACK and the ClientHello back to back, so TLS 1.3 costs one round trip and not two — the key share rides with the first message rather than waiting for a negotiated cipher. The clock reaches 61.2 ms: 2 round trips plus 1.2 ms of certificate verification.
Against a 137 ms request that is a specific, drivable fraction. Drag the round trip and read what the handshakes take of the whole:
At 30 ms the handshakes are 45% of the request. At a trans-Pacific 150 ms they are 61%, because two of the three round trips are pure setup and only one carries data. This is the term worth deleting, and it is deletable: nothing in it depends on the request itself.
Four ways to start a connection, then, drawn in one stage so they can be compared rather than remembered. Switch between them and read what each one has already paid for:
Read the round-trip column, not the intuition. Session resumption replays a pre-shared key from the last connection and still costs 2 round trips, the same as the full handshake: TLS 1.3 is already one round trip, so there is not a second one to remove. What resumption removes is the certificate — 61.2 ms becomes 60 ms, and the 1.2 ms of chain parsing and two signature checks is the whole saving. The 2-to-1 folklore is TLS 1.2's arithmetic. 0-RTT is the one that removes a round trip (30 ms), by sending application data in the first flight — which is why it is unsafe for anything non-idempotent: an attacker who captured that flight can replay it, and the server has no handshake state yet with which to notice. It is still 1 round trip and not zero, because the TCP handshake in front of it has not gone anywhere; only QUIC folds that in. A reused connection pays nothing at all.
One more thing has to hold before any of it counts: the certificate must chain to a root the client already trusts. Break one link at a time and read what the verifier actually reports:
The trip-up is the third case. unable to get local issuer certificate reads like an untrusted root and is almost never one — it means the server did not send its intermediate, so the client has a valid leaf, a valid root, and nothing joining them. It also fails intermittently across clients, because a browser that has cached that intermediate from another site will succeed where curl fails.
Down the stack and back up
The request is 312 bytes of text. Getting it onto a wire and back off one is where the layers earn their keep — and where they drop things.
Every layer wraps what the layer above handed it and adds a header of its own: a TLS record, a TCP segment, an IP packet, an Ethernet frame, and the preamble the PHY clocks out before any of it. That is the subject of Network Stack, and the byte layouts are in Binary & Number Systems.
The bar below is always one frame, drawn to the same width whatever is in it. Shrink the payload and watch the 112 bytes of framing take over the picture:
At a full 1,426-byte payload the framing is 7.28% and nobody thinks about it. At one byte — an ACK, a keepalive, a chatty RPC that sends a field at a time — goodput is 0.9% and you pay 113 bytes to move one. This is the whole reason Nagle's algorithm exists, and the whole reason turning it off is a decision rather than an optimisation.
Above one segment's worth, TCP splits. The maximum segment size is the 1,500-byte MTU minus 40 bytes of IP and TCP header minus 12 bytes of timestamp option — drag the response body and watch the count step:
Notice what the last segment costs. Our 2 KiB of JSON is two: a full one and a 622-byte remainder in a frame sized for 1,426. Both still cross in half a round trip, because they go out back to back and the window is far wider than two — segment count costs bandwidth, not latency, until a loss makes the second wait for the first.
On the server the NIC writes each frame straight into a ring buffer by DMA and raises one interrupt; a softirq then drains the ring in batches of 64. Raise the arrival rate past what one core can retire:
Past the drain rate the ring fills and the NIC has nowhere to put the next frame, so it increments rx_missed_errors and throws it away. There is no backpressure to apply — the sender is not listening. The drop is silent at every level a normal service can see: TCP retransmits and the request eventually succeeds, slowly. See Syscalls & Interrupts for the softirq budget that sets that rate.
Once the bytes are in the socket buffer, a worker parked in epoll_wait has to become the running thread. Add threads already runnable on that core and watch which of the six steps stretches:
Because the first four steps are just list manipulation, the wake is 22 µs on an idle core, and the run queue is the only step load touches at all. 6 runnable threads ahead of you turn that 22 µs into 9 ms — a p99 no profiler pointed at your code will ever explain. See I/O Models, CPU Scheduling and Concurrency Primitives.
What the server actually does
One millisecond and three hundredths, for the whole handler. It is worth knowing where even that goes, because two of its terms have no ceiling.
Parsing the request line and eight headers is60 µs of pointer arithmetic over a buffer that never leaves L1 — Memory Hierarchy for why that matters, and Cache Coherence for what happens once the connection table is written from several cores. Then SELECT * FROM users WHERE id = 42.
The index is a B+tree, and its height is what decides how many pages the lookup touches. Drag the row count across seven orders of magnitude and watch the tree gain a level only when the fanout runs out:
Notice how rarely it grows. An 8 KiB page holds about 367 bigint index entries, so a hundred million rows need four levels: 367³ is only 49 million, and 367⁴ is 18 billion — one more than a ten-million-row table needs. Adding a zero to the table usually adds nothing to the lookup — that flatness is the whole reason a B+tree is shaped this way rather than as a binary tree. Database Storage has the node layout.
Four index pages and one heap page, then. Whether that costs 550 µs or 5.55 ms depends only on how many of the 5 the buffer pool already holds. Take pages out of the pool, and change the device:
A completely cold descent on NVMe costs 1 ms — 5 4 KiB reads at about 90 µs each, which is why the page cache under it barely matters (File System) and why Disk Storage spends its time on queue depth. On cloud block storage the same descent is 5.55 ms, an order of magnitude, from the same query plan.
None of that is the term with no ceiling. A pooled connection is a finite resource, and the moment demand exceeds it the wait is not service time at all. Raise the requests in flight past the pool size, then lengthen how long each one holds a slot:
As soon as demand passes the pool size, the queue is the entire latency — below it the wait is exactly zero and the picture is boring. 200 in flight against a pool of 20 held for 5 ms each is a 45 ms wait added to a 550 µs query. The ceiling is Little's law and nothing else — 20 slots ÷ 5 ms is 4,000 requests per second, and offering more does not get you more, it gets you a queue.
For a write the last term is durability. COMMIT does not return until the write-ahead log record is on media that survives power loss, and batching is the only lever. Raise how many transactions share one flush:
At one commit per flush the device is the throughput: 1,538 per second on NVMe at 650 µs each, 124 on a cloud volume. Group commit amortises one flush over sixteen transactions and buys an order of magnitude back for a few milliseconds of added latency each. Turning fsync off buys two more orders and silently gives up the D in ACID — see Transactions for what COMMIT is really promising.
Where the 137 ms goes
Every stage has now been measured. Sorted by cost, the same twelve numbers say something the trace does not.
The waterfall in §01 is drawn in the order things happen, which is right for understanding the mechanism and wrong for deciding what to fix. Sorting the same twelve stages by cost puts the argument in a different shape.
Walk the slider down the stages, in cost order this time, and read what each one is as a share of the whole, with the top three held apart from the rest:
Those three — DNS, TLS, TCP — are 101 ms, or 74% of the request, and every one is setup: not one depends on which user we asked for. The bottom 6 together are 1.09 ms. This is a profile on which a flame graph shows nothing at all, because none of the time is on a CPU.
Because the stages are serial, deleting one has to remove exactly its own length and nothing else — which is a claim you can test rather than accept. Pick a stage to delete and watch everything after it slide left into the gap:
Every choice moves the end by exactly the length of the bar that vanished. That is the invariant restated as arithmetic: no stage overlaps another, so the total is the sum, and any millisecond you can stop spending is a millisecond off the end. It also says which deletions are worth chasing — and the three biggest are all setup.
Setup is the part you actually get to delete. Switch between four connection strategies and watch the bars that disappear:
Watch which bars survive. Caching the DNS answer alone takes 137 ms to 97.5 ms. Keeping the connection open as well takes it to 36.2 ms. A long-lived client that also skips the process start — which is what any server-to-server call is — lands at 31.2 ms, a 4.4× cut, and 30 ms of that is the one round trip carrying the request and the reply. There is no fifth strategy: that is the floor.
The saving is per request, so it compounds. Drag how many requests the run makes, and compare one pooled connection against a fresh one every time:
50 sequential calls cost 1.67 s pooled and 6.87 s unpooled — the same work, 4.1× the wall clock, and the difference is entirely handshakes that carried no data. This is also where CPU Architecture and Assembly & ISA stop being the lever: the constant factor you can win in a parse loop is microseconds against a term measured in round trips.
What breaks, and where
Symptoms travel up the layers. The cause lives in exactly one, and the trace is the map from the first to the second.
A 500 is almost never the whole story, and neither is "the API is slow". Each of the ten failures below leaves a signature the layer above cannot leave, which is what makes the map worth memorising.
Walk the slider through them and read where each failure actually lives, and what the caller sees when it fires:
Two of the ten sit on the same stage, and they are the pair people confuse: a slow query and an exhausted pool both show up as "the database is slow". Shape tells them apart — a slow query lengthens some requests, a starved pool lengthens every request equally, including the ones that never touch the database.
The second failure in that list deserves its own picture, because its symptom is a number people quote without knowing its source. Step through the retransmissions of a SYN that nothing ever answers:
Because the retransmission timer doubles, the elapsed time goes 1, 3, 7, 15 seconds — so the "three-second hang" in your logs is one dropped SYN and nothing else. Linux gives up after tcp_syn_retries doublings, 127 seconds by default: long enough that every timeout above it fires first. TCP Deep Dive has the state machine.
Which raises the question of what those timeouts should be. Drag the client's budget down past what a cold request costs and watch what the retry policy then does to the offered load:
Below 137 ms every first call times out and is retried, so the service receives three times the load and completes none of it — the classic retry storm, caused by a timeout set from the warm number instead of the cold one. A budget has to cover the slowest path a caller can legitimately take, or it converts a slow dependency into a dead one.
The last failure mode is not a failure of anything. It is arithmetic that appears only at scale. Raise how many backend calls one user request fans out to:
Notice how fast it compounds. With ten calls, one user request in ten already contains a p99 call (9.6%); with a hundred, 63% do. A service whose own p99 is a healthy 20 ms composes into a user-facing p50 that is not, and no service is at fault: Dean and Barroso's tail at scale (CACM 56:2, 2013).
So: the stages are serial, so the total is a sum; failures are silent at the layer you watch and loud two down; and the trip-up is optimising the 0.7% a profiler can see. Three lines hold the cost:
ms spawn 5.00 DNS 0.02-40.00
TCP 1xRTT TLS 1xRTT data 1xRTT
server 1.03 = 5 + DNS + 3xRTT + 2.4