Syscalls & Interrupts Primer

A core running your code cannot touch a disk, a socket, or another process's memory. Everything it does about the outside world happens because control crossed into the kernel and came back. Three claims, with numbers you can drag: the doorway is chosen by whoever caused it; the price spans four orders of magnitude, from a 2 ns function call to a 26 µs cold-cache tail; and the expensive part is never the instruction — it is the flushed TLB, the cold cache, and the woken thread.

01

Three doorways across the line

Control leaves user mode three ways. What separates them is not what the kernel does next, but who chose the moment.

Your code runs in ring 3, where a privileged instruction faults and the kernel's pages are not mapped. Reaching ring 0 is not a jump to an address: the hardware offers three doorways and no fourth.

The syscall instruction is one of yours. A page fault is nailed to the load that raised it. An interrupt lands wherever the device happened to catch the pipeline — pick a doorway, then drag the violet marker along the stream and watch which of the three follows your hand:

between instruction 8 and 9. drag left and right; the arrow keys move it one step at a time and Home restores the opening pose
syscall — at the syscall instruction — index 7

Notice that only the interrupt moves. Drag with the syscall selected and the dashed violet marker travels the whole stream while the crossing stays on instruction 7; select the exception and it sits on instruction 4 just as stubbornly. That is the whole of synchronous versus asynchronous: a syscall and a fault are consequences of the instruction stream, an interrupt is only a coincidence with it.

Where they differ more is on the way back. Press play and watch the marker leave the strip, cross the kernel, and drop back down — then switch doorways and watch the landing cell move:

syscall — running · ring 3

The syscall lands on instruction 8, the one after it. The interrupt lands on 9, exactly where it left. The exception lands on 4 — the same load, run again, because the handler's whole job was to make it succeed this time.

All three pay the same hardware bill on the way in, and since 2018 that bill has surcharges. Add them one at a time and watch a null getpid() climb:

0 on — 50 ns per null syscall

Notice that every rung starts where the one above it ended, so the number at a bar's end is the running total and the step between two bars is what that mitigation cost: , on hardware that has PCID, . The middle rung is the one to stare at: the same two CR3 writes cost 36 ns when the hardware can tag them and 77 ns more when it cannot, because an untagged write throws away every translation the thread had.

Nanoseconds only get interesting multiplied. Set a call rate on the first slider and a mitigation level on the second, and read off the share of one core that goes on nothing but entering and leaving:

100 k/s — 0.86% of one core

Watch the bar fill as the rate climbs. At — an ordinary web server — it is 0.86%, and nobody profiles it. At it is 8.6%. At it is 97%: a core fully occupied crossing a line, executing none of your instructions. That is why io_uring and epoll exist, and why the interesting optimisation is never a faster syscall.

02

What the hardware does on the way in

Four things happen before the first kernel instruction runs, and every one of them is a place a cost or a bug can hide.

Raising the ring is worth nothing without an address to raise it to. For interrupts and exceptions that address comes out of a table the kernel filled in at boot and the CPU indexes itself, in hardware, with no software in the path.

Every one of the 256 vectors holds a 16-byte gate — handler address, target ring, stack hint. Walk the vector number and watch what the kernel put there:

vector 14 — CPU exception

The table is 4,096 bytes: 256 gates of 16, which is exactly one page, and that is not a coincidence. The first 32 entries are fixed by the architecture — vector 14 is the page fault whatever the operating system thinks — and everything above 32 is the kernel's to assign as devices are probed.

The gate also says which stack to switch to, because the user stack cannot be trusted. Push the trap frame down onto the kernel's own, one slot at a time:

0 slots pushed — 0 bytes

Watch the first five go down before the others: those are pushed by the CPU itself, in microcode, before a single kernel instruction executes. The rest are the entry stub's work. Twenty-one slots, 168 bytes, on a 16 KiB stack — 1% of it, which is why kernel stack overflows come from deep call chains and not from the frame.

A syscall does not use that table at all. SYSCALL jumps straight to the address in the LSTAR MSR — and pays for the shortcut in a register:

C call — arg 4 travels in rcx

The C convention passes the fourth argument in rcx. The kernel takes it in r10, because SYSCALL writes the return address into rcx and RFLAGS into r11 before your code has stopped running. Exactly one cell differs, and it is the one every hand-written syscall stub gets wrong once.

The last piece is KPTI, which gives the kernel its own page tables — so two CR3 writes per crossing, and whether they hurt depends on one bit. Set a working set and switch the tag off:

nothing flushed

Because a tagged write leaves the entries in place, all 64 survive and there is nothing to walk., the same write flushes them: 64 pages costs 533 ns of walking after the return, and pays 12.8 µs — spread over the next few thousand instructions, invisible to any timer around the syscall itself.

03

Why the top half has to stay tiny

A hardware handler runs with its own line masked. That single fact shapes every device driver Linux has.

The masking is not politeness — it is what stops a handler being re-entered by the interrupt it is already handling. But it also means the device cannot say anything else until that handler returns.

So the length of the handler is a ceiling on the whole line. Drag its right edge and watch how many arrivals fall inside the deaf window:

1.5 µs in the handler. drag left and right; the arrow keys move it one step at a time and Home restores the opening pose
1.5 µs masked — ceiling 667 k/s

Notice how little it takes. A 1.5 µs handler caps the line at 667 k/s, already under 10 GbE's 813 k packets a second at full-size frames. Ten microseconds — a perfectly reasonable amount of code — caps it at 100 k/s, eight times short of the wire.

Linux's answer is to cut the work in two. Move it out of the masked window and into a softirq that runs with interrupts back on:

0.00% deferred — line ceiling 100 k/s

Watch what does and does not move. The line ceiling goes from 100 k/s to 2.50 M/s, because the deaf window shrank from 10 µs to 0.4. The core ceiling stays at 100 k/s the whole way, because the deferred work still costs 10 µs of somebody's time. The split buys availability, not throughput — and confusing the two is how a driver gets "fixed" and the box still drops packets.

A softirq is not allowed to run forever either. It drains what it can and hands the rest to ksoftirqd, an ordinary scheduled thread that competes with yours. Fill the ring and find the wall:

150 waiting — 0 to ksoftirqd

The wall is at 200 packets, not at the 300 of netdev_budget that everybody tunes. net_rx_action stops at 300 packets or netdev_budget_usecs, and 2,000 µs of wall clock runs out after 200 packets of 10 µs work. Raising the packet budget on that box changes nothing at all.

Which leaves the interrupt itself, still one per packet. NAPI kills that too: take the first one, then poll. Grow the burst one poll drains and watch the per-packet cost collapse:

1 packet waiting. drag left and right; the arrow keys move it one step at a time and Home restores the opening pose
1 packet — 2.0 µs each

One packet a poll and the interrupt costs 2.0 µs of the 10 the packet is worth. At the 64 of NAPI_POLL_WEIGHT it is 31 ns, and at 300 it is 6.7 ns — under a percent. The tail of that curve is why an idle link has low latency and a busy one has high throughput on the same driver, with no knob turned in between.

04

What a context switch really costs

The number everybody quotes is the one a microbenchmark can see, and it is the smaller half of the bill.

A switch is not a doorway — it happens inside the kernel, after one of the three. But it is what a blocking syscall actually buys, so its price belongs on the same page.

Start with the part that is easy to measure. Choose a pair of tasks and watch which lines of the direct cost get charged:

two threads — 700 ns

Notice that the rows this pair does not pay are drawn but not filled. Two threads of one process share an address space, so there is no CR3 write: 700 ns. Two processes need one: 1.2 µs. Under KPTI the entry and exit either side each swap tables as well: 1.9 µs. Everything on this ladder is a register move or a table pointer, and all of it is what lat_ctx reports.

And then the new thread starts running, into caches that are full of someone else's data. Drag the working set it has to pull back and compare the tail with the switch:

32 KiB of working set. drag left and right; the arrow keys move it one step at a time and Home restores the opening pose
3.3 µs of tail — 2.7× the switch

At — one L1 — the tail is 3.3 µs, already 2.7× the switch. At it is 26 µs and 21.8×. This is the cost Li, Ding and Shen measured in 2007, and the reason a benchmark that switches between two do-nothing threads tells you almost nothing about your server.

Which is what turns a cheap syscall expensive. A read the page cache answers never leaves the CPU; one that blocks pays two switches and a wakeup. Give the device some latency and compare:

8.6 µs for the blocking read

1.2 µs against 8.6 µs — seven times, with the device answering instantly. The kernel does not spend that time idle: somebody else runs in the gap, which is the whole reason blocking exists. It is your latency that pays, not the machine's throughput.

The scheduler is choosing how often to pay this, and it cannot win both ways. Move the quantum and watch the two curves cross each other's territory:

0.04% overhead — 21 ms worst wait

Watch the two costs trade places. At CFS's own , switching costs 0.04% and the eighth runnable task waits 21 ms. Drag down to and the wait falls to 350 µs while switching climbs to 2.3%; at it is 11% and the machine is mostly changing its mind. Neither curve crosses the other — they live in different units — but there is no setting that is good at both, which is why the answer to a latency problem is usually fewer runnable threads rather than a smaller slice.

05

Signals — the way back out

The kernel has no doorway of its own into your process. It has to wait for one of yours, and everything strange about signals follows from that.

kill(pid, SIGTERM) does not call anything. It sets a bit in the target's pending mask and returns. The bit is read on the target's next return from kernel to user mode — and only there.

So the delay is not the kernel's scheduling, it is how long the target stays inside. Set a sleep state and stretch the time it spends there:

running — delivered at once

A running target and an interruptible sleep both take it at once — the sleep is aborted for exactly this purpose. An uninterruptible sleep, the D state, is not on a path that checks: the signal waits the full 6 ms, and SIGKILL waits with it. That is why a process stuck on a dead NFS mount cannot be killed, and why kill -9 is not the trump card it is believed to be.

When the bit is finally read, the kernel does not call your handler directly. It builds a frame on your stack and returns into it. Pick a register set and shrink the stack you gave it:

952 bytes — it fits

Watch the frame outgrow the space reserved for it. On SSE it is 952 bytes; with AVX-512's xsave area it is 3,128, over the old 2,048-byte MINSIGSTKSZ by 1,080. That constant was a compile-time number for thirty years and had to become a runtime one — sysconf(_SC_MINSIGSTKSZ), glibc 2.34 — because the ISA outgrew it.

And your handler runs wherever the signal caught you. Step through a delivery that lands inside malloc, then change what the handler does:

the handler calls malloc — in malloc · lock free

Because the interrupted code holds the arena lock, a handler that calls malloc blocks on a lock its own thread already owns, and nothing will ever release it. The safe handler writes one sig_atomic_t and returns. The full list of what is legal inside a handler is in signal-safety(7), and printf is not on it.

There is one more consequence, and it is the one that reaches ordinary code. Delivering a signal to a thread sitting in a slow syscall has to end that syscall somehow. Move where the signal lands in the transfer:

before any byte arrives. drag left and right; the arrow keys move it one step at a time and Home restores the opening pose
read returns −1 with EINTR

Before the first byte, read returns −1 with EINTR — or, with SA_RESTART, the kernel rewinds rip and nobody notices. After the first byte, both flags return a short count and no error at all, because there is a result to report and it cannot be un-reported. A loop that treats a short read as the end of the stream is correct until the day a signal arrives, and then it is quietly, unreproducibly wrong.

06

Quick reference

Three questions worth answering cold, the handler shape that is always right, and five red flags.

How expensive is a crossing, really?

It depends by four orders of magnitude on which crossing, so the only useful answer is the whole ladder at once. Everything this page has measured, against a plain function call:

function call — 2.0 ns

Notice how much of the axis the top four rungs share. is 43 function calls — annoying, not fatal. is 4,300, and behind a switch is 13,107. The rungs worth engineering around are all at the bottom of the axis, and none is the syscall instruction.

Where does a request's time go?

Count the crossings rather than timing them. Set how many a request makes and read the split against 50 µs of your own work:

59 µs — 16% crossing

Watch the middle segment grow as you add crossings. is 59 µs a request, 16% of it crossing. is 88 µs and 43%. The fix is readv or io_uring, never a faster syscall.

What is the safe shape of a signal handler?

One store, and a loop that expects to be interrupted:

static volatile sig_atomic_t stop;
void on_term(int s) { stop = 1; }  /* all */

while (!stop) {
  n = read(fd, buf, len);
  if (n < 0 && errno == EINTR) continue;
  if (n <= 0) break;   /* real error, or EOF */
  consume(buf, n);     /* n may be short */
}

Which leaves one question: safe by what test? Walk the eight calls either side of the line:

write — safe

The safe side is all bare syscalls. The unsafe side reaches for a user-space lock first — which is why exit is unsafe and _exit is not, and why the rule generalises to any library you did not write.

Five red flags in review

  • A handler that does more than one store. §05's deadlock, waiting for the wrong millisecond.
  • A read loop with no EINTR branch. Wrong from the first handler installed, and only intermittently.
  • One syscall per field or row. 86 ns each; batch with readv or io_uring.
  • A latency fix that lowers the quantum. Halving the slice halves the wait but doubles the switch: 1.2% of a core at 100 µs, 2.3% at 50, 11% at 10. Cut runnable threads instead.
  • ksoftirqd on your worker's CPU. Move the IRQ with /proc/irq/N/smp_affinity.