I/O Models Primer

Three claims, each with a number and a figure you can drive. A blocked thread costs 25.5 KiB of kernel memory and two 1.3 µs context switches. Readiness turns O(n) into O(ready) — and stops helping the moment everything is ready. Completion deletes the per-operation syscall: worth nothing at a batch of one, 2.6× at thirty-two.

01

The wait, and who pays for it

Every I/O call either waits for its answer or refuses to. Every ring, every readiness API and every copy on this page follows from that fork.

A socket read has two jobs: wait until bytes are in the receive queue, then copy them out. The copy is a few microseconds of honest work. The wait is unbounded, and it is the entire design problem.

So watch the wait itself. Drag the slider to move the moment the first byte lands, and watch the thread that is not running stretch to meet it:

read() 406 µs · off CPU 397 µs · park + wake 2 × 1.3 µs

Notice that the parked stretch is not the kernel doing anything. The 406 µs the call takes at the shipped setting is 6 µs of copying and 400 µs of nothing, and the thread paid two context switches — 2.6 µs — for the privilege of sleeping through it.

Non-blocking is the same call with the wait deleted. Flip the mode and the syscall returns EAGAIN immediately, and the time the thread keeps is exactly what the parked stretch was:

blocking · read() 406 µs · off CPU 397 µs · park + wake 2 × 1.3 µs

Watch what the mode switch actually changes: nothing about the data, nothing about the socket, only who owns the gap. Blocking hands it to the kernel and gets simplicity back. Non-blocking keeps it, and now owes an answer to the question the kernel used to answer — when is this descriptor worth asking again?

Blocking is not free just because it is easy. Each waiting connection needs its own thread, and each thread is 16 KiB of kernel stack plus a task_struct. Raise the connections and watch the fixed bill:

1,000 threads · kernel 25 MiB · VA 7.8 GiB · 34 KiB each

At that is 249 MiB of kernel memory before a single byte moves, and 78 GiB of address space reserved by glibc's 8 MiB default stacks — only the pages a handler touches are real. Push it to and the threads want 2.4 GiB that never swaps, 3.2 GiB resident all told — a fifth of a 16 GiB machine. Both halves of that 25.5 KiB are countable rather than estimated: 16 KiB is THREAD_SIZE on x86-64, four pages of kernel stack, and 9.5 KiB is what the task_struct slab reports in /proc/slabinfo on a 6.x kernel.

Memory is the bill you can see. The one that hides is the switch: 1.3 µs of scheduler and cold cache, twice per event. Raise the event rate and watch the curve cross the line that is one whole core:

21K/s events
42K/s switches · 5.4% of a core

Because two switches are 2.6 µs, the crossing lands at — a rate one connection can reach in a benchmark. Past it, a core is spent changing whose registers are loaded, and no application tuning touches it.

Non-blocking on its own does not fix that, because a loop with nothing to wait on has to re-ask. Tighten the interval between asks and watch the two bars trade against each other:

954K/s EAGAIN calls · 24% of a core · 52 µs mean delay

Neither end of that slider is a server. At the loop burns 125% of a core answering EAGAIN; at it answers with 2.5 ms of mean delay. What is missing is a way to be told, and the next section is twenty-five years of that one idea.

02

One thread, many sockets

Readiness is one thread asking the kernel a single question about thousands of descriptors at once. Three generations answered it.

select answered it in 1983 with three bitmaps. You set a bit per descriptor you care about, the kernel hands them back with the ready bits set, and you scan. The bitmap is a fixed-size struct, which is where the trouble starts:

descriptor 900
descriptor 900 · 113 B × 3 sets × 2 ways · 678 B a call

Drag the marker and nothing complains. FD_SET is a macro over an array with no bounds check, so a descriptor past the end writes into whatever the caller put next on the stack. No error, no crash at the call — just a corrupted local, later, somewhere else.

poll fixed the ceiling with a caller-sized array, and left the real cost alone. Set the watched set against how many of them are ready, and read the two walks the kernel could have done:

select 1,000 walked · 10 ready · 990 wasted · epoll 10 walked

At a thousand watched and ten ready, 990 of the kernel's 1,000 checks were spent on descriptors with nothing to say — and then userspace walked the same array again to find which ten. Both walks are O(n) in the size of the set, not in the size of the answer.

epoll, in Linux 2.5.44, changed one thing: the set lives in the kernel between calls. Step the transport through the three syscalls and watch what comes back in each:

step 1 of 4 — epoll_create

Because the interest list stays registered, the wakeup does not have to describe the set — it only reports the descriptors that fired. The invariant this buys is worth stating exactly: at the top of every loop iteration each descriptor is either in the kernel's interest list or in the ready set the loop is currently draining. One in neither is invisible.

Keeping the set in the kernel is also where epoll stops behaving the way the descriptor number suggests. The interest list keys on the open file description, not on the integer. Step through a dup and a close of the number you registered:

1 entry · keyed on the description

Notice that the entry outlives the descriptor. Events keep arriving for a socket you believe you closed, and EPOLL_CTL_DEL on the closed number answers EBADF — you can no longer name the thing you registered. fork is the same trap from the other side: the child inherits the epoll fd, and both processes are now woken for one socket. Delete before you close, every time.

Set to answer size instead of set size is a different curve, not a better constant. Grow the watched set and watch the two lines separate:

1,000 watched
1,000 watched · 10 ready · 25 µs vs 500 ns · 51×

At with one percent ready, one wakeup is 250 µs against 2.8 µs — ninety-one times over. Turn the second slider up, though, and the gap closes: at fifty percent ready epoll is doing almost the same walk, because O(ready) and O(n) are the same when everything is ready.

One shared descriptor breaks the same way. Register a listening socket in n worker processes, then raise the worker count and turn EPOLLEXCLUSIVE on and off under one arriving connection:

8 workers · 8 woken · 7 wasted · 18 µs burnt

Without the flag every waiter is woken; one wins the accept and the rest get EAGAIN and park again, each having paid two context switches for nothing. That is the accept thundering herd, and EPOLLEXCLUSIVE (Linux 4.5) wakes one waiter instead of all of them. It is level-triggered only, and it is not the general fix: SO_REUSEPORT, which gives each worker its own accept queue, is the one that scales.

epoll offers the answer in two shapes, and one of them is a trap. Level-triggered reports anything still readable; edge-triggered reports the arrival, once. Shorten the read and switch modes:

level-triggered · 1 wakeup · 0 B stranded

Watch the edge-triggered case at a 16 KiB read: 48 KiB stay in the socket, the loop parks, and no further wakeup is coming for bytes that already arrived. The connection hangs and the logs say nothing. That is the failure mode of this whole section — the invariant broken by a descriptor in neither list, and it fails silently.

03

Stop asking, start handing over

epoll tells you a descriptor is ready and then you do the work yourself, one syscall at a time. io_uring deletes the second half.

Count the boundary crossings a batch of ready operations costs. Two hundred and fifty nanoseconds each with KPTI on, and the readiness loop pays one for the wakeup plus one per operation, where the ring pays one for the batch:

32 ops · 32/33/1 crossings · 2.9 µs/339 ns/129 ns an op

Notice that at all three models cost about the same — 2.9 µs, 3.1 µs and 2.9 µs an operation — because the two context switches around the wakeup dominate everything else. The win is not the mechanism, it is the batch: at thirty-two, io_uring pays one crossing where epoll pays 33. The top row is worth reading twice, because it is the assumption the other two are measured against: a thread per connection has no batch at all. Thirty-two ready operations are thirty-two threads, each woken and each parked again, so the two switches are charged per operation and not per batch — which is why it costs 2.9 µs an operation while epoll, paying one crossing more, costs 339 ns.

It buys that with shared memory. The submission ring is mapped into both address spaces at once, so writing a request is a store, not a call. Drag the tail along the ring:

0 entries pushed. drag the tail along the ring; the arrow keys move it one slot at a time and Home empties it
tail 0 · head 0 · 0 pending · 0 syscalls

Notice the readout: entries are queued and no syscall has happened. Only two indices matter — the tail you advance and the head the kernel advances — and their difference is the occupancy. Push to and the ring is full; submission now blocks on the kernel, which is the worst single operation this model has.

The full cycle is four beats, and only one of them crosses. Step the transport through it:

step 1 of 4 — write the entries

Because the completions come back into mapped memory too, reaping is another load — the one crossing in the middle is the whole bill. This is the shift from readiness to completion: you no longer ask what is ready, you describe work and are told when it is done, which is also why io_uring covers regular files where epoll never could.

There is a mode that removes even that one call. With IORING_SETUP_SQPOLL a kernel thread spins on the ring, so the crossing never happens. Switch the poller on and read what it costs:

io_uring_enter 7.0K/s · 0.2% of a core

Watch both bars at once: the application's share of the boundary goes to zero, and a whole core is held whether the ring is busy or empty. Jens Axboe's 2019 paper measured 1.7 M IOPS this way against 608 K for libaio on the same device — a 2.8× that is real and that you rent a core to get.

04

Moving the bytes

Knowing which descriptor is ready says nothing about what it costs to move the payload. That bill is paid in memory bandwidth.

Serving a file from the page cache to a socket is a read into a buffer and a write out of it. Step the transport through one chunk and watch where the payload actually goes:

step 1 of 5 — read()

Notice that the bytes are already in the kernel at step one and end up back in the kernel at step four. The user buffer is a round trip that the data did not need, and each leg of it is the CPU moving every word by hand at memory bandwidth.

Each kernel interface removes exactly one of those legs. Walk the slider from the loop to MSG_ZEROCOPY and watch a hop the CPU never makes take its place:

read + write · 2 CPU copies · 2 syscalls · 128 KiB by CPU

sendfile deletes the user buffer, so the page cache feeds the socket directly: one copy instead of two, one syscall instead of two. With scatter-gather DMA the card reads the page cache itself and the CPU never touches the payload at all. This is what nginx serves static files with.

Put a size on it. At 10 GB/s per core, drag the file up and read the three bills:

100 MiB · CPU 22 ms / 11 ms / 401 µs · 3,206 syscalls

A 100 MiB file costs 22 ms of CPU under the loop, 11 ms under sendfile and 401 µs with scatter-gather — 54 times, and all of it is copying, because the 3,206 syscalls are only 0.8 ms of the total. At this size the boundary is noise and the bandwidth is everything — and the 54 does not move as you drag, because every term in it is per-byte. Only the bills change.

Which is exactly why it inverts for small messages. Shrink the message until the two lines cross:

64 KiB
64 KiB · write 6.8 µs · MSG_ZEROCOPY 1.6 µs

Because MSG_ZEROCOPY has to pin the pages and reap a completion off the socket's error queue, it carries a flat 1,650 ns that a small message cannot amortise. The crossover is 13.7 KiB, close to the ~10 KiB the kernel documentation gives — under it, zero-copy is slower than the copy it removed, and at it is 1.3 µs slower.

05

Where each one actually wins

Three models, three axes. None is fastest everywhere, and the two axes that decide most arguments have nothing to do with speed.

The first axis is memory. An epoll registration is about 200 bytes; a thread is 25.5 KiB of kernel memory that never swaps and never shrinks.

Grow the connection count and watch one descriptor each stay flat under a line that is climbing away from it:

1,000 connections
1,000 connections · 25 MiB vs 200 KiB vs 125 MiB · 128×

The ratio is flat at 128×; the absolute number is what bites. At threads want 249 MiB against 2.0 MiB. C10K was a named problem in 1999 and is not one now, and the fix was never a faster thread — it was not having one.

Neither coloured line is what caps the box, which is the follow-up worth having ready. The grey one above both is socket buffers: Linux's default tcp_rmem is 4 KiB / 128 KiB / 6 MiB, and both models pay it per connection. Ten thousand active connections is 1.2 GiB of receive buffer — five times the thread cost. Dropping threads buys you the axis below it, not that one.

The second axis is the one people skip. An event loop is a single thread, so anything that blocks inside it blocks everyone. Drag the blocking call out from zero:

199 behind · 0 ns stalled · p99 +0 ns

Because the loop cannot preempt itself, all 199 other ready connections wait out the whole call — a getaddrinfo is 10 ms on every one of their p99s. A goroutine is the same shape: the runtime parks you on epoll for the I/O it knows about, and pins a kernel thread for everything it does not.

The third axis is whether you may run it at all. io_uring is the fastest thing here and the least available: Docker's default seccomp profile blocks io_uring_setup, Google turned it off across ChromeOS and Android in 2023, and kernels from 6.6 ship kernel.io_uring_disabled, which some distributions set to 2. Write the epoll path first.

06

Quick reference

Three questions worth answering cold, the bug that actually ships, and five things to stop in review.

Is io_uring faster?

Only if you batch. Set the operations submitted together and read the ceiling one core could hold on boundary work alone:

batch 32 · 351K/s / 2.9M/s / 7.7M/s · 22×

At the three bars are the same length — between 320 K and 355 K a second — because two context switches dominate. At thirty-two, completion reaches 7.7 M against 2.9 M, and a thread each has not moved, because nothing about it batches.

What is the bug this section is really about?

Edge-triggered epoll reports arrival once. A handler that does one read and returns leaves whatever did not fit in the socket, and no further wakeup is coming for it: on a 64 KiB burst into a 16 KiB buffer the upper line below strands 48 KiB forever, while the lower one drains to EAGAIN:

/* EPOLLET, one read: 48 KiB stranded */
n = read(fd, buf, sizeof buf);

/* correct: drain until EAGAIN */
while ((n = read(fd, buf, sizeof buf)) > 0)
  handle(buf, n);

That upper line is the bug that ships: legal C, no warning, correct in every test where the peer sends less than one buffer. Overhead is only worth arguing about next to the work beside it, so set what one request does:

20 µs of work · overhead 12% / 1.7% / 0.6% · 19×

At 20 µs of work — a cache lookup, a small parse — a thread each spends 12% of the request on overhead against 0.64%. Drag it out to and every model is under 0.3%: pick the one your team can debug at 3 a.m.

Why is epoll faster than poll, and when is it not?

poll is O(n) twice: the kernel walks the whole array and so does your loop. epoll keeps the set registered and returns only what fired. When most of the set is ready every wakeup that is the same walk — read the two costs off one axis, with the top of the axis under your second hand:

syscall, KPTI on · 250 ns

The step to memorise is a syscall at 250 ns against : 5.2×, which is why every model here is really a strategy for not switching. Copying 64 KiB, at 6.55 µs, is bigger than all of them.

Where do the numbers come from?

250 ns and 60 ns are Gregg's 2018 KPTI measurements, mitigated and mitigations=off. 1.3 µs is an lmbench-style two-thread ping-pong on x86-64, the middle of a 1.2–1.5 µs range. 10 GB/s is one DDR4-3200 core copying past L3. 1.7 M against 608 K IOPS is Axboe's io_uring paper. Two are derived, not measured, and the figures that use them say so on the axis: 25 ns per descriptor the kernel polls, and 40 ns per ring entry.

  • select in new code. FD_SETSIZE is 1024 and the macro does not check it.
  • Edge-triggered without a drain loop. Silent data loss; use level-triggered unless you have measured a reason.
  • A blocking call inside an event loop. DNS, a file open, a lock — every ready connection pays for it.
  • read plus write to serve a static file. Two copies and twice the syscalls; sendfile is one line.
  • MSG_ZEROCOPY on small messages. Under about 10 KiB the pinning costs more than the copy.