DNS & HTTP Primer

Run curl https://api.example.com/user/42. Three claims, each provable on a figure you can drag: a cold DNS lookup is four round trips, and the world keeps the old answer for a full TTL after you change it; HTTP/1.1 carries one request per connection, which is where the six connections and the repeated cookie come from; and HTTP/2 fixed the framing but stayed TCP's prisoner, which is why QUIC exists.

01

The name has to become an address

Nothing in curl can open a socket to a name. Before the first byte, api.example.com has to become 203.0.113.42, and that walk is where most “intermittent” outages actually live.

getaddrinfo reads /etc/nsswitch.conf, tries /etc/hosts, then sends one UDP datagram to the address in /etc/resolv.conf. That forwarder is a recursive resolver, and on a cold cache it walks for us: the query in flight goes out four times, and each hop that answered stays on screen. Press play, or scrub one hop at a time:

hop 0 of 4 — 0 ms elapsed

Notice which hop costs the most. The root and .com servers are anycast onto hundreds of sites, so they answer in 12 ms and 18 ms; the authoritative server is wherever the domain's owner put it, and it takes 34 ms. Four hops, 72 ms — and every one is a referral, not an answer. The root never knows where api.example.com lives, only who runs .com.

Almost no lookup pays that. Every layer between the program and the authoritative server keeps a copy, so the same walk gets cut short at whichever one still holds it. Drag the slider down the cache stack and watch the walk that never happens shrink:

browser cache — 0.05 ms

The spread is three orders of magnitude: 0.05 ms out of the browser's own map, 8.0 ms from a resolver that already has it, when nobody does. This is why a first request to a new host feels slow and the second feels instant — and why a latency graph of your service has a fat tail that has nothing to do with your service.

An address is only one of the things a name can carry. The record type in the question decides what comes back in the answer section. Drag the slider through the eight worth knowing:

A — 203.0.113.42

Stop on SOA. Its last field is the one the next two figures are about: the negative TTL, which is how long the world may remember that a name does not exist. Note also that CAA is checked by certificate authorities before they issue, so an out-of-date one breaks renewals rather than traffic — the failure arrives sixty days later.

The price of that caching is that you do not control when the world stops believing the old answer. Every copy carries a TTL, and every resolver started its own clock whenever it happened to ask. Drag the slider through the seconds after a change and count the resolvers still serving the old address:

16 / 16 — 0 s after the change

Watch where the last stale copy flips. Convergence takes a full TTL after the change, not after the last refill you watched, because the resolver that asked one second before you edited the zone holds its copy for the whole window. That is the invariant: every resolver serves the old record for up to one TTL past the moment you changed it, and there is no way to call it back. So lower the TTL before the change, not with it.

Failure runs on its own clock, and it is usually the longer one. Fix a typo and drag the seconds afterwards to count the resolvers still answering NXDOMAIN against the SOA minimum:

16 / 16 — 0 s after the fix

Because the SOA minimum is a property of the zone rather than of the record, a name that never existed is cached with a number nobody set for it — often 1 h. Ten seconds to fix the typo, an hour before every resolver forgets it.

Aliases add their own round trips. A CNAME says “ask again under this other name”, so each one is another authoritative lookup before an address comes back. Drag the slider to lengthen the chain and read the cold cost off the bottom row:

0 aliases — 42 ms cold

Three aliases turn a 42 ms lookup into , and each hop is a fresh authoritative query with its own TTL. The dashed row is the part that catches teams out: the zone apex already carries SOA and NS records, and RFC 1034 forbids a CNAME from coexisting with anything at the same name. So example.com itself cannot be a CNAME; you need your DNS host's ALIAS, ANAME, or flattening, which resolves the chain server-side and hands out an A record.

Size is the other constraint the protocol never let go of. A classic DNS response is one UDP datagram, and the original limit is 512 B. Add A records with the slider and push the answer past the budget:

60 B — one datagram

Past 512 B — — the server sets the TC bit and the client retries the whole query over TCP, for another round trip. EDNS(0) lets the client advertise a bigger buffer instead, which is why almost nothing falls back today. But bigger is not unlimited: past 1.2 KB, which is the IPv6 minimum MTU less its headers, the datagram fragments, and enough middleboxes drop IP fragments that DNS Flag Day 2020 asked the whole industry to settle on exactly that number. is the real ceiling.

One more thing hides behind “the resolver is fast”. 1.1.1.1 is not a server; it is an address announced by BGP from hundreds of sites at once, and your packets stop at whichever is nearest. Drag the client itself along the route and compare the nearest site with a single server:

12 ms to the nearest site. drag the client along the route; the arrow keys move it 400 km at a time, Home returns it.
Tokyo — 12 ms

Because light in fibre moves at about two-thirds of c, a round trip costs roughly one millisecond per 100 km, and no protocol beats that. Anycast does not make packets faster; it moves the destination. The same trick is why the root zone answers in milliseconds from anywhere despite having thirteen named servers.

Which brings us to the failure that actually pages people. A record moves, everything that respects the TTL follows it, and one process that cached the address forever keeps talking to a machine nobody serves from. Drag the slider through the ten minutes after the move:

0 s — 0 failed requests

The client that re-resolves recovers after one TTL. The pinned one is still failing at ten minutes, 24,000 requests deep, and it fails silently: no DNS error to log, only connections to a machine that moved. A JVM is the classic offender — networkaddress.cache.ttl defaults to −1, cache forever, under a security manager. Set it to the record's TTL.

02

One request at a time, in text

HTTP/1.1 is a line-oriented text protocol over TCP with exactly one outstanding request per connection. Every workaround the web grew for twenty years follows from that one sentence.

With an address in hand, curl opens a TCP connection, completes the TLS handshake, and starts talking. What it says is ASCII ending in a blank line — no length prefix, no framing header.

Press play to send the request one line at a time, and watch the bytes already on the wire accumulate in the bar underneath:

0 of 7 lines — 0 B sent

Notice the last line. The blank \r\n is the only thing that tells the server the headers are over — which is why parsing HTTP/1.1 means scanning for a delimiter rather than reading a length. A body with a known size gets Content-Length: N; a body whose size nobody knows yet gets framed a chunk at a time. Step through one:

0 of 3 chunks

Each chunk carries its own hex length in front of it, and a zero-length chunk is the end of the body. That is how a server-sent-events stream, a long poll or a query result set gets out of the door before its size is known — and it is why a proxy that buffers whole responses turns a streaming endpoint into a batch one without anybody changing a line of code.

That 149 B is cheap. What is not cheap is getting to the point where it can be sent. Drag the round-trip slider and switch the handshake to see when the first response byte lands:

first byte 150 ms

At 50 ms the first byte arrives at 150 ms on a fresh TLS 1.3 connection, 200 ms on TLS 1.2, and 50 ms on one already open. That is the whole argument for keep-alive: two of the three round trips are setup, and setup is reusable.

HTTP/1.0 closed the connection after every response. HTTP/1.1 made keeping it open the default. Drag the request-count slider and compare a fresh connection each time with one connection held open:

1 request — 170 ms kept open, 170 ms fresh

Six requests cost 1.02 s the old way and 520 ms the new one. Servers cap the reuse: an idle timeout, usually 60–75 seconds, and a maximum number of requests per connection. This is also why curl https://a/x https://a/y beats running curl twice — one process, one connection.

Keep-alive amortises the handshake but not the round trip. Pipelining was meant to fix that: send request 2 before response 1 comes back. The catch is that responses must come back in order. Drag the right edge of the first response and watch the two behind it:

60 ms on the first response. drag the first response's right edge; the arrow keys move it 10 ms at a time, Home restores it.
60 ms on the first response

Watch response 3. It is ready almost immediately and cannot leave until the slow one ahead of it does. That is head-of-line blocking, and it is the phrase the rest of this page is about. Pipelining also met a second wall: transparent proxies that buffered, reordered or dropped pipelined responses. Browsers enabled it in the mid-2000s and every one of them rolled it back.

So the only concurrency left is more connections. Drag the number of parallel connections and find the number every browser stops at:

1 connection — 2.20 s

Because the curve is a hyperbola in the number of connections, the first six buy and the next six buy . Six is the knee, and it is a compromise: enough parallelism to matter, few enough that a page with a hundred assets does not look like an attack. The old trick to beat the cap was domain sharding — serve assets from img1, img2, … so the browser opens six per name. Under HTTP/2 that same trick is actively harmful, and the advice outlived the protocol it was written for.

The last cost is the one nobody sees. Every one of those 30 requests carries the same headers, and the cookie is usually the biggest of them. Drag the slider up to the 4 KB an over-eager session library will hand you:

3.6 KB per page load

of request headers, uploaded, for one page — on the narrow half of a residential link, and before a single byte of content. Nothing in HTTP/1.1 can help: it has no way to say “the same headers as last time”. Fixing that is the first thing HTTP/2 does.

03

One connection, many streams

HTTP/2 keeps every HTTP semantic and replaces the wire format: binary frames tagged with a stream id, a shared header table, and one connection instead of six. It is still TCP's prisoner.

Text framing is what forced one request per connection: with no length in front of a message there is no way to interleave two of them. So HTTP/2 puts a length in front of everything. Every message becomes a sequence of frames, and every frame starts with the same nine bytes.

Step through the header fields and read the width of each off the row above it:

length — 24 bits

72 bits, and the last 31 are the whole idea: a stream id on every frame. Length says where the frame ends, type says whether it is HEADERS, DATA, SETTINGS, WINDOW_UPDATE, RST_STREAM, PING or GOAWAY, and the id says which conversation it belongs to. Client-initiated streams are odd, server-initiated ones even, and id 0 is the connection itself.

With that tag, the sender can put frames from different requests on the wire in any order it likes. Play the connection frame by frame and watch the frame going out now land in whichever stream owns it:

frame 0 of 9

Notice that no stream ever waits for another to finish. That is what retires the 6-connection workaround: one connection now carries everything, so one congestion window is being probed instead of six competing ones, and domain sharding — a virtue under HTTP/1.1 — becomes a way to throw the benefit away.

The framing layer also made a feature possible that turned out to be a mistake. A server that knows the page needs app.css can push it on its own stream before the browser asks. Drag how much of it the browser already had:

0% cached — 0 B wasted

The server cannot see the browser's cache, so on a repeat visit almost all of 120 KB is sent for nothing — and it is sent ahead of the HTML the browser is actually waiting for. Chrome removed server push in 2022. Link: rel=preload in the response headers does the useful half: it tells the browser what to fetch and lets the browser decide whether it needs it.

Framing also makes the header problem solvable. HPACK gives both ends a synchronised table: send a header once, and afterwards send its index instead. Drag the request-count slider and compare what HTTP/1.1 would have sent with what goes on the wire now:

1 request — 121 B vs 149 B, 19% saved

The first request is barely cheaper — it still carries every header, as a Huffman-coded literal, and installs it in the table. By the second it is 55% smaller, and by the eighth, against 1.2 KB: 82% saved. A static table covers the common ones so even a first request finds indices. The sharp edge: compression makes ciphertext length depend on plaintext, so an attacker who injects a header beside a secret one can read the secret off the size — the CRIME attack. The answer is the never-indexed literal, a flag on anything sensitive that forbids any intermediary from tabling it.

Multiplexing needs its own back-pressure, because one greedy stream could starve five others sharing the socket. Every stream gets a receive window. Drag the window along the byte axis and watch the sender stall between round trips:

64 KB window. drag the window's right edge; the arrow keys widen it a step at a time, Home restores the default.
64 KB — 1.25 MB/s

The default is 64 KB — that is RFC 9113 §6.9.2 — and one window can be in flight per round trip, so at 50 ms a single stream is capped at 1.25 MB/s no matter how fat the pipe is. Raise it to and the same stream reaches 320 MB/s. This is the number behind “HTTP/2 is slow for large downloads”: it is almost always an unraised SETTINGS_INITIAL_WINDOW_SIZE.

Concurrency has a ceiling too, and it is not the one people assume. Drag the slider past what the server advertised:

1 issued — 0 queued

Past 128 — nginx's default http2_max_concurrent_streams — the extra requests are not rejected, they are queued in the client until a stream id frees up. Fire and 172 of them are waiting before they ever reach the network. “HTTP/2 makes concurrency free” is false; it makes it cheap up to a number the server chooses.

And now the flaw that made HTTP/3 necessary. TCP delivers one ordered byte stream, so a lost segment holds back everything behind it — including bytes for streams that already arrived. Raise the loss slider and watch the gap appear in every lane at once:

0.0% loss — 170 ms

Because the streams share one ordering, each waits for the same retransmission: a clean link finishes in 170 ms, 1% loss takes . HTTP/2 fixed application-level head-of-line blocking, inherited the transport-level kind, and by collapsing six connections into one made a single loss more expensive than it had been.

04

Moving the streams below HTTP

Every remaining problem in HTTP/2 is a TCP problem, and TCP lives in the kernel. QUIC rebuilds the transport in user space over UDP, with streams, TLS 1.3 and connection identity in one design.

A new transport protocol number would be dropped by half the middleboxes on the internet, so QUIC rides UDP — which every NAT and firewall already forwards — and rebuilds the transport above it. Switch the segmented control and watch the five jobs move out of the kernel and into the process:

TCP + TLS — the process

All five in user space is the trade. It buys deployability — a QUIC fix ships with the application instead of with a kernel upgrade — and it costs CPU, because every packet is encrypted and acknowledged by code that cannot use the kernel's segmentation offloads. It also means the connection is the library's, not the socket's, which is what makes the next three figures possible.

The first thing that buys is the handshake. QUIC carries TLS 1.3's cryptographic exchange and its own transport parameters in the same flight. Drag the round-trip slider and switch the transport to see when the first byte arrives:

first byte 100 ms

At 50 ms, TCP with TLS 1.3 needs 150 ms and QUIC needs 100 ms — one round trip saved, because there is no separate TCP handshake to finish first. To a server we have talked to before, 0-RTT puts the request in the very first packet: 50 ms, the floor set by the speed of light.

Before any of that, the server has to know the client's address is real, or QUIC would be a reflection amplifier pointed at whoever the source address named. So a server may return at most three times what it received until the address is validated. Drag the client's first datagram down until what the server may send no longer reaches the dashed line:

1,200 B in — 3,600 B out

Notice where the budget actually runs out: at 1,000 bytes in, not at 1,200. The 1,200-byte minimum datagram is a separate rule — RFC 9000 sets it so a QUIC path is never narrower than that — and the 3× ceiling turns it into a 3,600-byte budget, 600 bytes clear of the 3,000-byte certificate flight. Drag below 1,000 and the lower bar falls short of the dashed line: the server sends what it may, stops mid-flight, and waits for another datagram before it can finish — an extra round trip. The padding is the headroom, not the threshold.

The 0-RTT floor has its own hole. That flight is encrypted under a key derived from the previous session and carries no proof of freshness — so anyone who captures it can send it again. Drag the replays and read the ledger:

0 replays — charged 1×

Watch the charge execute once per copy. This is why 0-RTT is only safe for idempotent requests: a GET replayed is a wasted response, a POST /charge replayed is a second charge. Servers must either restrict early data to safe methods or carry their own replay defence — a single-use ticket, or an idempotency key the application checks.

The second thing QUIC buys is the one HTTP/2 could not fix. Streams have their own sequence numbers, so a lost packet holds back only its own stream. Raise the loss slider and compare where this page finishes with where HTTP/2 finished:

0.0% loss — 120 ms

At 1% loss the page lands at 124 ms against HTTP/2's 290 ms; at 3%, against 530 ms. The mechanism is the whole difference: five of the six lanes never notice the retransmission, because nothing about their bytes depends on the missing packet arriving first. On a clean wired link the gain rounds to nothing — the win is loss, not bandwidth.

Independent streams break HPACK, though. HPACK assumes both ends process table updates in the order they were sent, and QUIC deliberately does not guarantee that. QPACK's answer is a budget. Drag how many streams may block on a table entry that has not landed yet:

0 of 6 — 726 B of headers

At zero — the default value of SETTINGS_QPACK_BLOCKED_STREAMS — the encoder may never reference an entry it cannot prove has arrived, so it falls back to literals: 726 B for six requests instead of 84 B. Raising the budget buys compression and pays in head-of-line blocking, one stream at a time. Almost nothing tunes it, which is why real QPACK savings sit below HPACK's.

The last thing QUIC changes is what a connection is. TCP identifies one by the four-tuple, so changing network kills it. QUIC puts a connection ID in every packet. Drag the client from Wi-Fi to LTE and watch the four-tuple stop matching while the connection ID still does:

on Wi-Fi. drag the client from Wi-Fi to LTE; the arrow keys move it, Home puts it back on Wi-Fi.
on Wi-Fi

Because the server keys the session on the ID and not the address, the video keeps playing across the handover instead of stalling while a new connection and a new TLS session are built. The IDs are rotated and unlinkable by design, so migration does not hand a passive observer a way to follow a device between networks.

Which leaves the question of how a client that speaks three protocols ever ends up on this one. TLS's ALPN extension settles HTTP/1.1 against HTTP/2 inside the handshake; HTTP/3 cannot be chosen that way, because choosing it means not making a TCP connection at all. Step through the visits:

visit 1 — HTTP/2 over TCP

The first visit is ordinary HTTP/2 over TCP, and its response carries Alt-Svc: h3=":443"; ma=86400 — “there is a QUIC endpoint here, remember it for a day”. The next visit races QUIC against TCP and keeps whichever finishes first. About 30% of the HTTPS requests Cloudflare served in 2024 came over HTTP/3, essentially all at the CDN edge. Inside a datacentre, HTTP/2 is still the sensible default.

05

What the whole request costs

Four questions worth answering cold, priced with the same models the figures above are drawn from, and five things to catch in a review.

The first question is the one an interviewer actually asks: where does the time go between typing a URL and the first byte of JSON? Almost all of it is round trips, and which ones you pay depends entirely on the transport. Switch it, and drag the round-trip slider:

total 158 ms

At 50 ms with a warm resolver, TCP costs 158 ms, HTTP/3 costs 108 ms, and 0-RTT costs 58 ms; a cold cache adds 72 ms to all three. None of it is bandwidth. It is latency, and the only levers are fewer round trips and shorter ones.

The second is head-of-line blocking, which means a different thing per layer: application-level in HTTP/1.1 pipelining, transport-level in HTTP/2 because TCP is one ordered byte stream, and none at the transport in HTTP/3 — where QPACK can put it back at the header layer if you raise the blocked-streams budget. A browser opened 6 connections for the first and opens one for the second, which carries 128 streams by default.

The third is when HTTP/3 is worth deploying. The honest answer is a crossover, not a verdict — drag the loss slider and watch HTTP/2 lose first to HTTP/3 and then to six HTTP/1.1 connections:

0.0% loss — HTTP/1.1 450 ms, HTTP/2 170 ms, HTTP/3 120 ms

The model is one sentence: every lost packet costs one round trip of stall to whatever shares its transport ordering. At 1% that is 124 ms for HTTP/3, 290 ms for HTTP/2, 470 ms for six HTTP/1.1 connections; past the single connection is worse than the six it replaced. Lossy mobile links and handovers are where HTTP/3 pays.

  • networkaddress.cache.ttl=-1 — pinned to one address for the life of the JVM.
  • Domain sharding on an HTTP/2 origin — six congestion windows where one would do.
  • Depending on HTTP/2 server push — removed from Chrome in 2022; use Link: rel=preload.
  • A non-idempotent handler over 0-RTT — early data is replayable by design.