TCP Deep Dive Primer
A connection to api.example.com:443 is two machines agreeing on numbers. Three claims, each one drivable: the handshake needs three packets, not two; the sender obeys two windows at once, one advertised and one guessed; and every expensive TCP failure in production is silent.
Three packets, and why not two
curl has typed Enter. Before one byte of the request leaves, the kernel has to agree two numbers with a machine it has never met.
Every connection opens with SYN, SYN-ACK, ACK. Three rather than two is not ceremony: each side picks a random initial sequence number, and a number the other side has never acknowledged is one neither of them can trust.
Both sides start knowing nothing about the other's numbering, and each packet moves exactly one fact across. Press play, or take the scrubber, and read what each side knows:
Notice that after two packets the picture is already lopsided. The client knows both numbers; the server has sent y and has no idea whether anyone received it. The third packet is not an acknowledgement of the connection — it acknowledges y, and it is the only thing that makes the server's own numbering trustworthy.
Insisting on it costs a full round trip in which nothing useful crosses. Cut the handshake back to two packets and step an old duplicate SYN through it:
Watch the server's state on the right. On two packets it reaches ESTABLISHED on a connection nobody opened, and it will hold a socket and a buffer there until something times out. On three, the client answers RST: a client that never sent x cannot acknowledge y. The third packet is what lets a stale copy be rejected instead of honoured.
One round trip is the price, and the price is whatever the link charges. Drag the round-trip time, then switch what the crypto layer puts on top of the handshake:
The crypto layer needs round trips of its own, so the cost multiplies rather than adds. At plain TCP is 140 ms, TLS 1.3 is 280 ms and TLS 1.2 is 420 ms — all before the request is written. That is the whole argument for connection pooling, and no server tuning touches it.
Once the third packet lands the kernel has to put the connection somewhere, and there are two queues rather than one. Raise the arrival rate past what the application accepts:
The two queues fail differently, and only one of them fails loudly. A full SYN queue has a defence — tcp_syncookies stops allocating state and encodes the server's ISN as a hash of the 4-tuple, so a SYN flood costs no memory at all. A full accept queue has none: the kernel drops the third ACK, logs nothing, and the client goes on believing it is connected.
So the failure reaches the client as a wait rather than an error, and the wait has a shape you only need to see once. Step the retransmissions:
Watch the tick spacing double. tcp_syn_retries = 6 puts them at , 3 s, 7 s, 15 s, 31 s and 63 s, and connect() gives up at 127 s. A connect-latency histogram spiking at one and three seconds is not a slow server; it is a dropped SYN or a full accept queue.
Numbering every byte
IP loses packets, duplicates them, reorders them, and says nothing. Everything TCP adds is built from one 32-bit counter.
Every byte has a sequence number — the field counts bytes, not packets. The receiver acknowledges a position, and the number it returns is the next byte it expects, not the last one it got.
That single choice decides everything downstream. Drag the acknowledgement point along the tape and watch what is confirmed divide from what is still outstanding:
Notice that the ACK number can only ever say one thing: a prefix arrived. Nothing in the base header can say “I have 1461–2920 and 5841–7300 but not the middle”, so the cumulative ACK stalls at the first hole and stays there — which means one lost segment makes every subsequent ACK identical.
Identical ACKs are information, and TCP spends it. Step the loss and its recovery, and watch the duplicates accumulate on the sender's side:
The third duplicate is the trigger, so the sender never waits for a clock — it retransmits on the receiver's own evidence. Three is not arbitrary: one or two duplicates are exactly what ordinary reordering produces, so a lower threshold would resend healthy data every time two packets arrived swapped.
What the sender may resend, though, depends on what the receiver is allowed to say. Drag the hole along the tape, then turn the selective-acknowledgement option off:
Without SACK the sender knows where the hole is and nothing else, so it resends the hole and everything behind it: with the loss at segment 6 that is 10 KB back on the wire, of which 8.6 KB had already arrived, against 1.4 KB — the hole itself and nothing else. SACK (RFC 2018) is negotiated once, in the SYN, and turns the repair into exactly one segment.
A hole the receiver never reports is a different problem, and that is what the timer is for. Set the smoothed round trip and its variance:
Watch the floor. RFC 6298 computes RTO = SRTT + max(G, 4·RTTVAR), but Linux clamps it at 200 ms, so on a same-rack path with a 0.2 ms round trip the timeout is a thousand times the RTT. That is deliberate: an early retransmission adds load to a network already in trouble.
When the timeouts do start firing they do not fire at a constant rate. Walk the count up and watch the elapsed span:
Notice how little of the total the first few timeouts are. The fifteenth retransmission — the default tcp_retries2 — goes out at 805 s, and the kernel kills the connection one TCP_RTO_MAX later: tcp_model_timeout() in net/ipv4/tcp_timer.c evaluates ((2<<9)−1)·200 ms + 6·120 s = , so a connection to a vanished peer reports an error after 15 minutes, not 13. ip-sysctl gives the range as 13 to 30 minutes, because the real RTO is rarely the floor. A service with no deadline of its own holds the request, and the thread, for all of it.
One loss stays invisible to all of this: the last one. Arm the tail-loss probe and step the same exchange again:
A lost tail has no later segment to complain about, so no duplicate ACK is ever generated and only the timer is left — 200 ms on a 20 ms path. The probe fires at 2·SRTT, turns the tail into an ordinary duplicate-ACK case, and saves 160 ms. That is why RACK-TLP (RFC 8985) is now Linux's default.
The window the receiver advertises
The sender is bounded twice over. The first bound is the only one the receiver has any say in, and it rides in every ACK.
Every ACK carries a receive window: the free space in the receiver's socket buffer at that instant. The sender never puts more unacknowledged bytes on the wire than that, which makes flow control lossless — the receiver is never sent more than it can hold.
The window is not a constant anybody chose; it is arithmetic on the buffer, and the application's read pointer is the other end of it. Drag the pointer behind and watch the advertised window close:
Notice that none of this is about the network. The receive window shrinks because the application is slow and reopens the instant it reads, so a receiver-side stall reaches the sender looking exactly like a bandwidth limit. No packet capture on the wire can tell the two apart.
Push the read pointer all the way behind and the window reaches zero, which is a state the protocol has to be able to leave. Step the persist timer:
A window update is itself an ACK, and ACKs are never retransmitted, so a lost update would deadlock the connection forever. The persist timer is the escape hatch: the sender keeps poking, doubling the interval up to 120 s. No error, no timeout, no log line — a stuck consumer stalls a whole pipeline in silence.
The window has one more problem, and this one is arithmetic rather than behaviour. Raise the scale shift negotiated in the SYN and watch how far the raw field falls short:
Sixteen bits cap the advertised window at 64 KB, which was generous in 1981 and is now three orders of magnitude short. RFC 7323 multiplies it by up to 214, reaching 1 GB. The option appears only in the SYN, so a middlebox that strips it caps the connection at 64 KB for life, silently.
Whether 64 KB is a cap or a comfort depends entirely on the link. Set the bandwidth and the round trip, and read the volume of the pipe against what the window fills:
Because throughput is window ÷ RTT, an unscaled 64 KB window reaches 7.5 Mbps at 70 ms whatever the link was sold as: a path at that latency needs 8.3 MB in flight and delivers under 1% of its rated speed. Check this before anything gets blamed on congestion.
The window the sender guesses
The second bound is one nobody advertises. The sender keeps it alone, from evidence, about a network it cannot see.
The congestion window is the sender's private estimate of how much the path will absorb. It is never in a header and never negotiated, and bytes in flight are capped by the smaller of the two windows, not their sum.
Which one binds is the first thing to establish when a connection under-delivers, because it decides where you go looking. Move rwnd and cwnd against each other:
Notice how different the two diagnoses are. If rwnd binds, the receiver's buffer or its application is the problem; if cwnd binds, the network is, and no buffer tuning helps. ss -ti prints both, which makes this a ten-second question rather than a guess.
A new connection has no evidence at all, so it starts small and doubles. Step the round trips and watch cwnd reach for the pipe:
cwnd doubles once per round trip, so “slow” start is the fastest phase there is: 10 segments (RFC 6928, 14 KB) become 1,280 in seven rounds. The dashed rule is where a 100 Mbps path at 70 ms fills, at 854 KB — six rounds, 420 ms of ramp before the link is used at all.
Doubling forever would be catastrophic, so it stops the first time anything is lost. Walk the rounds through two losses:
Watch the shape after the first one. The congestion window halves and then climbs one segment per round trip — additive increase, multiplicative decrease — which is what makes many senders sharing a bottleneck converge on a fair split instead of oscillating. It also makes recovery time proportional to the window.
That proportionality is the whole reason Reno was replaced. Drag the cursor along the recovery of a 2,000-segment window on a 100 ms path and compare Reno with CUBIC:
Reno adds one segment per round trip, so it needs a thousand of them — here — to recover a window it lost to one packet. CUBIC grows as a cubic in wall-clock time rather than in round trips, so it is back in 11 s whatever the RTT. Past the old peak that cubic keeps accelerating, and what stops it is not the algorithm: it runs into the receive window at 30 s and flattens there, because a sender obeys min(rwnd, cwnd) however fast its own window grows. Linux has defaulted to it since 2006.
Both share one assumption that a modern path often breaks: that a lost packet means congestion. Slide the link's random loss rate up:
The fall is steep, and the two loss-based curves fall differently. Reno's ceiling is about MSS ÷ (RTT·√p) — Mathis et al., 1997 — so at on a 140 ms path it tops out at 3.2 Mbps of a gigabit link. CUBIC has a response function of its own (RFC 8312 §5.1) that falls off as p−3/4 rather than as p−1/2, which is worth 113 Mbps against Reno's 32 at — and worth nothing above 0.15%, where RFC 8312's TCP-friendly region hands CUBIC back to the Reno estimate and the two become one line. BBR estimates the bottleneck rate and the minimum RTT and paces to their product, so random loss costs it only the bytes actually lost.
The reason a loss-based sender has to see loss at all is that it is filling something first. Drag the bottleneck buffer deeper and watch the throughput not move:
A loss-based sender fills whatever buffer sits in front of it before it backs off, so on a 100 Mbps link pushes the round trip from 12 ms to 396 ms and delivers not one extra bit. The fixes are active queue management at the bottleneck (fq_codel is the Linux default) or a sender that watches delay.
Closing, and the parts that bite
Opening a connection is symmetric and quick. Closing it is neither, and most production TCP problems live here.
A connection is two independent byte streams, so closing it is two independent events. Each side sends a FIN when it has finished writing and acknowledges the other's — four packets, and the halves can be minutes apart.
The interesting part is not the packets but the state they leave behind. Step the close and watch the two state columns diverge:
Notice which side is left holding something. Whoever calls close() first ends in TIME_WAIT and keeps the 4-tuple for 2·MSL. On Linux that is TCP_TIMEWAIT_LEN in include/net/tcp.h — a compile-time 60 s, not tcp_fin_timeout, which governs FIN_WAIT_2. Lowering that sysctl to shorten TIME_WAIT is popular and does nothing.
Sixty seconds is free until one client opens connections faster than the port range recycles them. Raise the connect rate to a single peer:
A 4-tuple to a fixed destination varies only in the local port, so ip_local_port_range gives 28,232 of them and the exhaustion rate is 471 a second. At connect() starts returning EADDRNOTAVAIL. The cure is connection reuse; tcp_tw_recycle was removed in Linux 4.12 because it broke every client behind a NAT.
The other classic is not about closing at all: two optimisations that are each correct alone. Step a request written as two small write() calls:
Watch the gap in the first mode. Nagle holds the second small write until the first is acknowledged; delayed ACK holds that acknowledgement for up to 40 ms hoping to piggyback it on a reply the server cannot produce, because it has not seen the whole request. TCP_NODELAY removes the stall and puts two 40-byte headers on the wire; one writev() removes it and puts one segment on the wire.
Nothing so far notices a peer that has simply vanished — no FIN, no RST, just silence. Shorten the idle timer and read the detection time:
tcp_keepalive_time ships at 7,200 s, so with nine probes 75 s apart a dead peer is noticed after 2.2 hours. Dropping the idle timer alone to still leaves 12 minutes, because the probes dominate: 60 s idle, 10 s interval and 3 probes is 90 s, and all three have to move. gRPC ships its own keepalive rather than trust anyone to have set them.
Quick reference
Four questions worth answering cold, three figures that hold the answers, and five red flags.
What makes TCP reliable?
One invariant, and it is drawable. Below the ACK number, every byte is delivered in order exactly once; at or above it, the sender still holds a copy. Drag the edge and watch the two rows trade:
The sender may forget a byte only once the ACK number has passed it, so its buffer is the price of reliability and the receive window is the receiver's half of the same bargain. Retransmission, SACK and every timer are machinery for advancing that one number without breaking the rule.
What does the waiting actually cost?
Nine orders of magnitude separate the cheapest wait on this page from the most expensive one. Walk the rung being priced down the scale:
The step worth memorising is the sixth to the eighth. Everything up to is a property of the path; TIME_WAIT at 60 s and keepalive at 2.2 h are properties of a configuration file, and they are the two you can actually change.
rwnd or cwnd — which one is holding me back?
Read both from ss -ti. A small rwnd means the receiving application, or tcp_rmem, is slow; a small cwnd means loss, slow start, or a long path. Compare either against BDP = bandwidth × RTT to know whether the window is the limit at all — and if the connection is fresh, the answer is often neither:
On a same-region link a fresh connection spends 24 ms of its 84 ms before the request is even sent. the same 200 KB response takes 980 ms, of which 280 ms is setup and 560 ms is slow start ramping the response.
Which failures are silent?
Three, and none writes a log line. Cause each one on a 5 MB download and watch what it adds to the same budget:
Every one of them is a wait, not an error. A full accept queue costs the first SYN retry, 1.0 s; a zero window costs a persist interval, 1.6 s; and on a trans-Pacific path costs 9.9 s, because the transfer stops being congestion-limited and starts being window-limited. Read them with ss -ti and nstat -az TcpExtListenOverflows.
The last of the five flags below is the one you can feel from here. Open the pipeline — put more than one request on the wire before its response is back — and watch the rate move while the link does not change at all:
- Reaching for
TCP_NODELAYfirst. It removes the 40 ms stall and leaves ten 40-byte headers per request. Coalescing in user space removes both. tcp_tw_recycle = 1in a sysctl file. Removed in Linux 4.12; it silently dropped SYNs from anything behind a NAT. The knob you meant istcp_tw_reuse.
The next flag loses data rather than time. Load the send buffer, then close the socket both ways and watch what the peer never reads:
- Lowering
tcp_fin_timeoutto shorten TIME_WAIT. TIME_WAIT is a compile-time 60 s;tcp_fin_timeoutgoverns FIN_WAIT_2, a different state. SO_LINGERwith a zero timeout on live traffic. It forces a RST instead of a FIN, so in-flight data is discarded and the peer sees a connection reset.- A protocol that ping-pongs one message at a time. Throughput is bounded by payload ÷ RTT whatever the bandwidth: at that is 14 requests a second per connection, and 57 with four in flight.