Network Stack Primer
Our client runs curl https://api.example.com/user/42. Five layers stand between that process and the one on the far end. Three claims about them, each proved by the figure beside it: every layer costs bytes you can count, every ceiling on the path is a number, and the failures are silent by design.
What every layer adds
Each of the five prepends a header going down and strips it coming up. The bytes that reach the wire are never the bytes the application wrote.
Universities teach OSI's seven. Production reasons about five, because nothing that shipped implemented the session and presentation layers: HTTP at L7, TCP at L4, IPv4 at L3, Ethernet at L2, line coding at L1.
Our request is 500 bytes of HTTP. Drag the slider one layer at a time and watch the payload stay exactly the size it started while each header is prepended in front of it:
Notice that nothing rewrites the payload. TCP prepends 20 bytes, IPv4 another 20, Ethernet 14 in front and a 4-byte frame check sequence behind — 58 bytes of framing around 500 bytes of request, 558 on the wire. That is the invariant the whole stack rests on: every layer owns its own header and treats everything above it as opaque bytes.
Fifty-eight bytes is cheap on a 500-byte request and ruinous on a small one. Drag the marker along the curve — the whole drawing is a drag surface — to read the framing share against the payload it wraps:
At 500 bytes the framing is 10.4% of the wire. At it is 98.3% — 59 bytes moved to deliver one. This is why chatty protocols lose to batched ones long before the link is busy, and it is the entire argument for Nagle's algorithm, which holds a small write back until the outstanding ACK arrives so the next one can ride with it.
Coming up, the same frame is peeled, and every header's real job is to name the layer above it. Press play and watch each demultiplexing key hand what is left to exactly one upstairs neighbour:
Watch the keys rather than the sizes. EtherType 0x0800 says IPv4; IP's protocol byte 6 says TCP; destination port 443 says which struct sock. A layer never searches — it reads one field and dispatches. Receiving a packet is three table lookups, not three scans, and that is why the receive path is O(1) in the number of open connections.
The wire will not carry an arbitrary frame. Ethernet's payload ceiling is 1500 bytes, so with 20 of IP and 20 of TCP one segment carries at most 1460. Drag the write past that and watch TCP split it:
As soon as the write crosses , the second segment carries a single byte and costs 58 more bytes of framing to deliver it. Because TCP is a byte stream the application never sees the split, so a protocol whose record is one byte over the MSS quietly doubles its packet count and doubles its exposure to loss.
The ceiling moves when the addresses do. An IPv6 header is 40 bytes, not 20, so the same 1500-byte frame carries 1440 bytes of payload instead of 1460. Switch versions with the same write in place:
A 1,460-byte write is one segment over IPv4 and two over IPv6 — the most common way a service gets slower after a dual-stack rollout that changed nothing else. Tunnels do it too, and they stack: 4 bytes for a VLAN tag, 24 for GRE, up to 50-odd for IPsec, every one of them subtracted from the same 1500.
The medium then adds a tax no header accounts for: 7 bytes of preamble, 1 of start-of-frame delimiter, and a 12-byte inter-frame gap between every pair of frames. Drag across the frame size and read what the link actually delivers:
At the minimum frame a 10 Gbps link runs at 14.88 Mpps for 714 Mbps of goodput — 93% of the link spent on framing and gap; at the 1,518-byte maximum, 813 Kpps and 9.49 Gbps. Both come back in section 04.
The socket a flow hangs off
A file descriptor points at a kernel object owning two buffers, the TCP state machine, and the four numbers that identify the flow.
socket() allocates a struct socket wrapping a struct sock — for TCP a struct tcp_sock, which embeds it as its first member. The descriptor is an ordinary VFS handle, which is why read works on it at all.
The four numbers are the source address and port and the destination address and port. Step through the arriving segments and watch the tuple select exactly one socket:
Notice the invariant the whole demultiplexer depends on: any two established connections differ in at least one of the four fields. Two of the segments above share a source IP and two share a source port, and neither pair collides. A tuple that no established socket claims falls through to the listener; a tuple no socket claims at all earns an RST.
That rule is generous to a server and cruel to a client. A server on (*, 443) holds millions of connections because the client side varies; a client to one destination varies only its own port. Raise the connection count and watch the ephemeral range fill:
Linux ships net.ipv4.ip_local_port_range at 32768 60999 — , and therefore 28,232 simultaneous connections from one source IP to one (dst IP, dst port). Past that, connect() returns EADDRNOTAVAIL. That one fails loudly, which is the good case; the silent version is the same range consumed by sockets sitting in TIME_WAIT for 60 seconds after they closed.
Ports are one ceiling; the window is the other. A flow can only have as many bytes unacknowledged as the receiver has advertised, so throughput is window ÷ round trip, never more. Raise the round trip and see where a stock window lands:
At 100 ms a 1 Gbps link needs 12.5 MB in flight to stay full, and a 200 KB window delivers 16.4 Mbps — 1.6% of the link, on a link that is not congested and shows no loss. This is why Linux auto-tunes tcp_rmem, and why setting SO_RCVBUF by hand is almost always a mistake: it pins the window and switches the auto-tuner off.
On the accept path there are two more queues with two more ceilings. listen(fd, backlog) sizes the second; the first is tcp_max_syn_backlog. Raise the arrival count with nobody calling accept():
Watch what happens after the twenty-fourth arrival. The kernel does not refuse the connection and the application logs nothing — the SYN-ACK is simply retransmitted and eventually dropped, so the client sees a timeout and the server sees a healthy process. The counter that tells you is nstat -az TcpExtListenOverflows, and the depth is Recv-Q in ss -lnt.
Once a connection is established the same pressure appears inside it. The receive buffer holds what has arrived and not been read, and the window the kernel advertises is whatever is left. Slow the reader down:
Below roughly 40% of the arrival rate the buffer fills and the advertised window reaches zero — the sender stops, and the only thing that restarts it is the reader draining bytes. That is end-to-end backpressure, and it is the reason a slow consumer shows up as latency at the producer rather than as memory growth in the kernel.
A server holding tens of thousands of these does not poll each one. epoll gives back only the descriptors that changed. Raise the descriptor count and switch between the two interfaces:
select() copies and scans the whole set every call, so its cost is the number of descriptors watched; epoll_wait() returns the ready list, so its cost is the number of descriptors ready. And select() stops entirely at — a compile-time constant, not a tunable.
UDP has the same struct with the state machine removed: one sendto is one packet, one recvfrom is one packet, and there is no window to close. Its receive buffer still fills, and it charges for something other than what you sent:
The buffer accounts skb->truesize — the whole allocation, about 768 bytes for a small datagram — not the 100 on the wire. So a 212,992-byte rmem_default holds 277 datagrams, not 2,130, and the rest are dropped with no signal: no RST, no ICMP, no error at the sender.
How the packet finds its way
Leaving the socket, a packet needs three things: an interface, a next-hop address, and a destination MAC. Routing answers the first two, ARP the third, and the path MTU bounds the result.
The routing table is not a list walked in order. It is a longest-prefix match: every route whose prefix covers the destination is a candidate, and the one fixing the most bits wins.
Walk through eight destinations and watch which prefixes cover each one. The address is under your hand; the winning route is the longest bar that still covers it:
Notice that the default route 0.0.0.0/0 covers every one of the eight — it fixes zero bits, so it always matches and always loses to anything else. A metric only breaks ties between prefixes of equal length; a /24 with a terrible metric still beats a /0 with a perfect one, which is why adding a route "with a high metric so it won't be used" usually uses it.
The route gives a next hop, and the next hop is not the destination. Step through the four hops and compare the two addresses in the headers — the IP one and the MAC one:
This is the layering made concrete: L3 is end to end and L2 is hop by hop. The destination IP is written once by the sender and never changed; the destination MAC is rewritten by every router, because it only ever names the next device on this segment. The TTL falling by one per hop is what makes traceroute possible at all.
Filling in that MAC needs a lookup, and the answer is cached in the neighbour table with a life of its own. Drag time forward on one entry that keeps carrying traffic:
At 30 seconds — base_reachable_time_ms, which the kernel randomises over half to one and a half times its value — the entry goes STALE, and because traffic is still flowing it enters DELAY immediately. The thing worth knowing is that packets keep flowing on the stale address the whole time. Revalidation happens beside the traffic, not in front of it: three unicast probes a second apart, and only then FAILED.
Nothing in that exchange is authenticated. Any host on the segment may answer for any address, and Linux will believe an unsolicited reply that updates an entry it already has. Step through what that buys an attacker on the same L2:
The victim's routing table is untouched, its ARP cache holds a syntactically perfect entry, and every packet to the internet now transits a third party that forwards them on. There is no L2 fix; the answers are dynamic ARP inspection on the switch, 802.1X, or simply not putting untrusted hosts on the same segment — which is what a cloud VPC does for you by giving every tenant its own virtual L2.
With an interface, a next hop and a MAC, the last question is how big the frame may be. The answer is the smallest MTU on the path, and the path includes links nobody told you about. Drag the tunnel in the middle down:
A takes 24 bytes, PPPoE 8, IPsec 50-odd — and the sender, which has never heard of them, keeps emitting 1500-byte packets with DF set. The router that cannot forward one drops it and sends back ICMP type 3 code 4 carrying the MTU it can do, and the sender clamps its MSS. That is Path MTU Discovery, and it is one ICMP message away from working.
So block ICMP "for security" and watch what the same response does. Small responses are unaffected; drag the size up with the path's ICMP allowed, then blocked:
Here is the shape every engineer eventually debugs: work perfectly and hang. The sender is not told anything, so it retransmits the same too-large segment until the connection dies — the handshake succeeded, the headers arrived, and only the bulk is a black hole. Allow ICMP type 3, or turn on net.ipv4.tcp_mtu_probing, which infers the ceiling from loss instead (PLPMTUD, RFC 4821).
The older answer was to fragment instead of clamping, and it has an arithmetic problem. A datagram survives only if every fragment does, so its survival is (1 − p) to the power of the fragment count. Raise both:
At 1% link loss a 6-fragment datagram is lost 5.9% of the time — 5.9× the loss of the link it crossed — and IP has no partial retransmit, so one lost fragment costs the whole datagram. This is why TCP sets DF and clamps instead.
The bottom of the stack
Below IP is the part that decides whether a box tops out at a hundred thousand packets a second or ten million.
The NIC and the kernel meet at a pair of ring buffers in main memory. The kernel posts descriptors pointing at empty 2 KB buffers; the card DMAs an arriving frame into the next one and advances its tail. Nobody copies, and nobody waits.
That works exactly as long as the kernel takes descriptors back faster than the card fills them. Set the arrival rate above the drain rate and watch the ring:
Notice what the ring is actually for. A 512-descriptor ring absorbs a 300 Kpps overload for 1.71 ms and then drops everything above the drain rate — ethtool -G eth0 rx 4096 buys 13.7 ms instead of 1.71 ms. Ring size buys burst tolerance, never throughput, and rising rx_missed_errors in ethtool -S means the CPU, not the card, is the problem.
Draining is where the CPU goes. Until 2003 every frame raised a hardware interrupt; section 01 already told us a 10 Gbps link runs at with full frames. Switch between the two notification schemes at that rate:
An interrupt costs on the order of 1.5 µs once entry, exit and the cache it leaves cold are paid for, so 813 Kpps is 1.22 cores of pure interrupt — the machine is saturated before a single byte reaches a socket. NAPI's answer is to disable the queue's interrupt on the first packet and poll up to a budget of 64 frames per call, so under load the interrupt cost falls by that factor and, while packets keep arriving, the loop never re-enables it at all.
One core polling is still one core. A modern NIC has several receive queues, hashes each packet's 4-tuple, and interrupts the CPU that owns the chosen queue. Raise the flow count, then switch the traffic shape:
Because the hash is over the 4-tuple, every packet of one flow lands on one queue, one core, and one warm cache — no reordering and no cross-core locking. It also means one elephant flow cannot be split: a single 8 Gbps transfer pins one core at 100% %si while the other seven idle, and /proc/interrupts shows it as one column that will not flatten.
Every stage between the wire and the application has its own queue, its own limit, and its own counter — plus one more on the way back out, and none of them is netstat -s alone. Step down the chain:
Learn this ladder rather than the individual commands. A drop at the ring means the CPU is behind; at the per-CPU backlog it means one core is behind (net.core.netdev_max_backlog); at the socket buffer it means the application is behind; at the accept queue it means the application is not calling accept. The fix is different at every rung, and a packet dropped at the wrong rung is invisible to the metric you were watching.
For traffic you intend to discard, the cheapest place to discard it is before any of that exists. XDP runs an eBPF program inside the driver on the raw DMA buffer, before an sk_buff is allocated. Compare the two drop points under :
Dropping in iptables costs a few microseconds per packet because the packet has already been given a socket buffer and walked a chain; XDP_DROP costs around 50 ns because none of that has happened yet. At 10 Mpps that is 25 cores against 0.5, which is the entire reason Cloudflare's L3 scrubber and Meta's Katran load balancer live in XDP rather than in netfilter.
Past that the exits leave the kernel altogether. Set a target rate and read the cores each data path needs to sustain it — the kernel stack against the bypass paths:
At 1 Mpps the kernel stack needs 1.0 cores and DPDK 0.07; at 14.88 Mpps — 64-byte line rate from section 01 — DPDK needs 1.0 pinned cores and the kernel 14.9. The bill is that you now own TCP, retransmission, congestion control and ARP.
Quick reference
Two questions worth answering cold, and five red flags worth catching in a review.
Where does the time in one HTTPS request actually go?
Almost none of it is transmission. A cold request pays one round trip to resolve, one to shake hands, one for TLS 1.3, and one more to ask and be answered. Drag the round trip and read the server's own share of the result:
At the answer arrives 225 ms after the request, and only 25 ms of that was the server. Reusing the connection removes three of the four round trips and lands at 75 ms — which is why keep-alive, TLS session resumption and HTTP/2 multiplexing all pay off far more than any amount of server-side optimisation at that RTT.
Why does a busy client run out of ports before it runs out of anything else?
Because a closed connection holds its port for the whole TIME_WAIT — 60 seconds on Linux, hard-coded — so the ports in use are the request rate times that, not the concurrency. Raise the rate against the 28,232 the range holds:
Without reuse the range runs dry at to one backend. With keep-alive the same rate needs twelve open connections: a connection is held for the 25 ms it is in use, not for 60 seconds after.
SO_RCVBUFset by hand. Pins the window, kills auto-tuning.- Tens of thousands of connections to one backend. 28,232 ports per source IP. Pool, or fan out.
- All ICMP blocked at the edge. PMTUD dies silently; large responses hang forever.
- UDP read "until I have N bytes". Each
recvfromis exactly one datagram. - Unfiltered
tcpdumpon a hot path. Pass a BPF expression so the kernel filters first.