Assembly & ISA Primer

When curl puts bytes on the network it calls send(), and that call compiles to a handful of instructions ending in syscall — the one opcode that hands the machine to the kernel. This primer follows the path down: what a compiler emits and why it fights to keep values in registers; the two decisions that make x86 and ARM different targets; the calling convention and the three ways it fails silently; and the crossing itself, at roughly 250 ns.

01

What the compiler emits

Source is for humans; the CPU decodes bytes. Between them sits one text form that maps one-to-one onto those bytes, and a register file small enough to run out of.

A compiler turns source into assembly — mnemonics like mov and call — and an assembler turns that into the bytes the CPU decodes. Take int add(int a, int b) at -O2. Nothing has been emitted yet, so the figure holds only the C statement and two empty slots; press play, or walk the scrubber right, and each step emits one assembly instruction and fills in the machine-code bytes underneath it:

C source — 0 bytes

Notice there is no memory in it. The arguments arrived in registers and the answer leaves in one, so the whole function is four bytes: lea is three and ret is one. That lea is doing the addition — it computes an address and stores it without touching the flags. The figure below opens on the x86-64 encoding; tap the architecture under it for ARM64 — the same source, the same shape, different bytes — and walk the instruction slider to light one instruction and its encoding at a time:

x86-64 · lea eax, [rdi + rsi] · 3 bytes

Watch the byte count as you switch. x86-64 spends four bytes here and ARM64 spends eight — two fixed four-byte instructions for the same two operations.

Registers are what makes both this small: reading one costs nothing. There are sixteen on x86-64, and a loop with more live values than that has to put the rest somewhere. Drag the live-value slider — a value in a register is a filled cell in the top bank, and past sixteen the row underneath starts taking the overflow:

2 live values · 0 spilled

Past sixteen the compiler spills. Each spilled value gets a stack slot, and reading it becomes an L1 load — five cycles of load-to-use latency instead of the zero a register costs. At you are watching register allocation fail, which is what an optimising compiler spends most of its time avoiding. The bill is per read, every iteration — drag the spilled-read slider from zero and watch the row of free operands turn into five-cycle loads:

0 spilled reads · 0 cycles

is forty cycles — about 13 ns at 3 GHz — on a body whose register operands cost nothing at all. That is why register pressure is a real thing and not an academic one. Registers are also not all an instruction can name: one x86 address is a base plus a scaled index plus a displacement, in a single instruction. Drive the three sliders — base, index and scale — and read the address off the tape:

0x1028 · ARM64 2

Set the displacement slider to zero and ARM64 does it in one instruction too — ldr x0, [x1, x2, lsl #3]. Put the displacement back and ARM needs an add first: A64 has no form that adds an immediate on top of a scaled index. The address carries a second cost neither ISA controls, because memory arrives in 64-byte lines. Drag the eight-byte load along the two lines below and take it across 0x40:

0x18 → one line. drag the load along the tape; the arrow keys move it one byte and Home puts it back
0x18 → one line

The line, not the byte, is the unit of transfer, so dragging the load across 0x40 turns one line into two: two tag lookups and two chances to miss. Nothing in the source says which you get. Keeping values in registers, addresses in one instruction and loads inside a line is the compiler's whole job here — and the assembly is where you check.

02

Where x86 and ARM actually differ

Two decisions taken in the 1980s — how long an instruction is, and how many registers there are — still decide what your laptop and your cloud bill look like.

An x86-64 instruction is 1 to 15 bytes long; every A64 instruction is exactly 4. Almost everything else follows from that one line of each manual. Below is a real seven-instruction x86 block, eighteen bytes end to end: each instruction ends where the next begins, and the byte the decoder starts on is under your hand. Drag the marker off a boundary:

0 · boundary. drag the start marker along the byte tape; the arrow keys move it one byte and Home returns it to zero
0 · boundary

Watch what the decoder produces at byte 2. It does not fault: almost every byte is a legal opcode, so it decodes an instruction nobody wrote and re-syncs with the real stream a few bytes later. That is how a mis-set jump target becomes working-but-wrong code, and how ROP gadgets are found inside instructions the compiler emitted.

A tape with one instruction length has no such position — drag the same marker along the A64 tape, which holds the same seven instructions in twenty-eight bytes:

0 · boundary

There is only one length, so an offset that is not a multiple of four is not a wrong instruction on ARM64 — it is a PC alignment fault, and the core refuses. That guarantee is paid for in the front end.

Below, both tapes hold the same seven instructions on one byte scale, and only the A64 side knows where its instructions start: walk the cycle slider and watch the x86 boundaries resolve one per cycle, left to right, out of a run of bytes the decoder cannot yet cut up:

0 / 7 boundary · x86-64 9 · ARM64 1 cycle

The length scan is the part with no ARM equivalent, and it is a first-pass cost, not a steady-state one: switch the figure to warm and all seven boundaries are there at cycle zero, because a predecode stage has written length marks into the instruction cache and the uop cache skips the length problem entirely. That is why a hot loop on x86 does not run at a seventh of ARM's speed — and it is also why a cold branch target does.

The silicon that makes the warm case free is what ARM never has to build. What x86 buys with it is density — drag the instruction count and read the two bars against each other:

7 instructions · x86-64 18 · ARM64 28 bytes

Notice the ratio: eighteen bytes against twenty-eight for the same seven instructions — the two tapes you just drove — so ARM64 is 56% larger on this block. It is not a constant, though. Push the count to and the gap moves, because the answer is an average over whichever mix of one-byte rets and five-byte calls the block happens to hold: a fixed four bytes beats x86's call as often as it loses to its ret.

The register file is where the trade runs the other way — drag the live-value count and watch which side starts spilling first, with ARM64's thirty-one against x86-64's sixteen:

12 live values · x86-64 0 · ARM64 0 spilled

ARM64 has 31 general-purpose registers against x86-64's 16, so the live-value count that starts spilling roughly doubles. At twenty, x86 is paying for four stack slots while ARM has paid nothing. Register-hungry code — JIT-compiled JavaScript, interpreter loops — feels this directly.

The difference that breaks programs is neither of these. Below, the flag store and the data store sit on the order they become visible in; leave the model on TSO and drag the skew slider, which tries to let the flag run ahead:

x86-64 · TSO · data = 42 · holds

Notice that TSO refuses: the ghost shows where the store wanted to go and the arrow puts it back, so the reader always sees 42. Switch to ARM's weak model and pull — now the flag becomes visible first and the reader sees zero, with no change to the code.

This is the most common cross-architecture porting bug and it fails silently: the test passes on your laptop and the Graviton node returns garbage under load. The fix is a release store and an acquire load — on both architectures.

03

The calling convention

One agreement about registers and stack layout lets code from any compiler, in any language, call code from any other. Every stack trace and every FFI binding depends on it holding.

On Linux and macOS x86-64 that agreement is the System V AMD64 ABI; Windows has its own and Apple's ARM64 is a variant of AAPCS64. They differ in details and agree on the shape: SysV gives the first six integer arguments a register each, in a fixed order, and puts everything after that on the stack. Drag the argument count past six:

3 arguments · 0 stack

Notice what changes at seven. Up to six, the call touches no memory: every argument is already where the callee looks. costs a store before the call and a load after it, and each one after that eight more bytes of stack — which is why a hot function with nine parameters loses to the same function taking a struct pointer. Memory arrives anyway the moment the callee needs locals.

At rest, below, nothing has been pushed and the callee owns none of the stack; press play, or step the scrubber, and watch the frame grow downward, each slot as tall as the bytes it takes:

before the call · rsp = 0

call pushes the return address; the prologue pushes rbp and subtracts a frame; the epilogue undoes both and ret pops and jumps. Every frame on your stack looks like this, which is exactly why a debugger can walk it. The subtraction is not free-form — the ABI requires rsp to be 16-byte aligned at every call site. Drag the locals slider off a multiple of sixteen and watch the padding appear on the end of the frame:

24 locals · sub rsp, 32

Notice that the padding is never more than fifteen bytes and never less than zero: at it vanishes, because ninety-six is already a multiple of sixteen. Those bytes are what keeps movaps from faulting on a stack-allocated vector three calls deeper: hand-written assembly that forgets the alignment crashes in a function it never called. The other half of the contract is which registers a call may destroy — walk the slider along the register file and read each one's promise: restored by the callee, or the caller's problem:

rax · caller-saved

Seven registers survive a call because the callee promises to restore them; the other nine may be destroyed, so a caller that wants one across a call saves it itself. The split exists because each side knows something the other does not — the caller knows which values it still wants, the callee knows which registers it will touch.

Leaf functions get a further shortcut, 128 bytes of it — drag the scratch size and take it past the edge of the red zone:

64 bytes · fits below rsp. drag to grow the scratch area; the arrow keys move it eight bytes and Home restores it
64 bytes · fits below rsp

Under 128 bytes, the red zone lets a leaf use scratch without moving rsp at all — two instructions saved per call. Drag past it and the prologue reappears. Signal handlers must not touch that region, which is one of the few places signal-safe code differs from ordinary code.

All of this assumes both sides agree. Below, the callee is reading the low bits it was promised; drop the declared width and put something in the high half:

64 bits · 0x2a · holds

The ABI puts a narrow argument in the low bits and leaves the high bits undefined, so declaring int where the C function takes long makes the callee read whatever the previous instruction left above bit 31 — zero on your machine today, 0xdead after an unrelated change tomorrow. It does not crash; it returns a wrong number.

A broken convention has one other hiding place, equally silent — drag the stack depth and watch how far the walker gets:

1 / 1 frame · holds

By the walk has hit a frame built with -fomit-frame-pointer — the default at -O2 — so the chain ends and the rest of the stack is gone. DWARF data describes every frame either way, at a table lookup each: that is why Fedora and Ubuntu turned frame pointers back on in 2023–24, since a truncated stack looks like a shallow call rather than a broken unwinder.

04

The one instruction that changes privilege

Out of hundreds of opcodes, exactly one hands the machine to the kernel. Everything a program does to the outside world goes through it, and it is not cheap.

On x86-64 that instruction is syscall; on ARM64 it is svc #0. Every other instruction runs in user mode, where the hardware allows only what the kernel has mapped for this process. The figure below opens with every register still holding its user value; press play, or step the scrubber once, and watch the top five flip to their kernel values together — that is the whole instruction, and the atomicity is the point, because there is no moment when user code holds kernel privilege:

user mode, ring 3 — nothing has moved

Notice what is not in that first step. syscallwrites rcx, r11, rflags, the segment selectors and rip, and nothing else: when it retires the CPU is in ring 0 and still running on the user stack, under the user page tables. The three steps after it are ordinary instructions in entry_SYSCALL_64, and the middle one is SWITCH_TO_KERNEL_CR3 — before Meltdown that macro expanded to nothing, because the kernel was mapped into every process and the tables never changed on entry. KPTI is why it is not nothing any more, and most of why the crossing got dearer in 2018.

Once inside, the kernel still has to learn which call this was — drag the slider along the table and watch the number in rax pick an entry:

rax = 0 → sys_read

The number in rax indexes sys_call_table, an array of function pointers; the arguments arrive in rdi, rsi, rdx, r10, r8, r9 — r10 where the ordinary convention uses rcx, because syscall itself puts the return address in rcx and the flags in r11 — the two writes you watched it make. Those numbers are ABI forever, which is why read is still 0.

How often you pay for the crossing is a design decision — drag the chunk size down and watch the transition take the bar away from the copy:

64 KB · 16 calls · 109 µs

At a megabyte becomes 1,048,576 crossings and about 262 ms of pure transition against roughly 105 µs of actual copying — the overhead is 99.96% of the bill. At 64 KB it is sixteen calls and 4 µs, and the copy dominates instead. That is the entire reason buffered I/O exists.

The per-crossing price also depends on how the kernel booted: below, one core's second is a hundred cells, and at a million calls a second the crossings have taken a quarter of them. Drive the rate, then flip the isolation off:

1,000,000 / s · 250 ms · 25.0%

Isolation costs roughly 250 ns a crossing against 70 ns without it, so a service making 100,000 syscalls a second spends about 2.5% of a core on the boundary alone, and one making a million spends the twenty-five cells you are looking at. Push it to and the core does nothing else at all.

Turning isolation off is not the fix — crossing less often is. Below, the submission ring is full and one io_uring_enter carries all sixty-four; drag the batch size down and watch the row of crossings underneath grow back:

batch 64 · 16 calls

Batching sixty-four operations per io_uring_enter turns 1,024 crossings into 16; at the row fills and you are back to a syscall per element. Switch to SQPOLL and the crossings go to zero outright: a kernel thread polls the ring, so userspace only writes to shared memory. That costs a core spinning, a trade worth making at high queue depth and not at low.

There is one more way to pay nothing, and it is a matter of which side of a line the data sits on — drag the call rate and watch the path that stays in user mode against the one that crosses:

100,000 / s · 25 ms → 2.5 ms

Because the vDSO page is mapped into every process, the call goes sideways instead of down: it never crosses the ring boundary, so it costs about 25 ns instead of 250. Anything that timestamps in a loop — a tracer, a metrics library, a database taking a commit timestamp per row — is living on this, usually without knowing it, and the day it stops (a clocksource the kernel will not export, a seccomp filter, a container without the page) throughput drops by a factor of ten with no code change.

05

Quick reference

The whole path in one figure, every cost on one axis, and five red flags worth catching in review.

What actually happens between send() and the wire?

Six steps, and only one of them changes privilege: libc loads the arguments, puts 44 in eax, and executes one instruction, after which everything runs in the kernel. Step it down:

send(fd, buf, len, 0)

Notice how little of it is user code: six instructions, none worth measuring. The crossing costs about 250 ns before the kernel copies a byte, and the error check afterwards is the two lines every libc wrapper ends with:

mov  eax, 44        ; sendto
syscall             ; ring 3 -> 0
cmp  rax, -4095     ; error?
jae  __syscall_err
ret                 ; rax = bytes sent

Which registers carry arguments, and how do syscalls differ?

Ordinary calls take rdi, rsi, rdx, rcx, r8, r9, floats in xmm0–7, and return in rax; callee-saved are rbx, rbp, rsp, r12–r15 — seven, and rsp is one of them, because a function that returns with the stack pointer somewhere else has destroyed its caller rather than merely clobbered a register. Syscalls use the same six argument registers with r10 for rcx, because syscall clobbers it. Drag the slider along every cost this page has named, on one axis:

0.3 ns · 1×

Watch the gap between the vDSO and — a factor of ten for the same answer, decided entirely by whether the boundary is crossed. Below that gap the compiler arranges things; above it you do.

When does reading assembly pay for itself?

Three. perf report attributes time to instructions, not lines; whether a loop vectorised is answered only by the emitted code; and memory-ordering bugs are invisible above the assembly.

  • Inline asm without clobbers. Silent, and only when the allocator picks that register.
  • Lock-free code with no explicit ordering. Passes on TSO, breaks on ARM. Name the memory_order.
  • FFI declarations with the wrong width. Generate bindings; do not hand-write them.
  • A syscall per element. 262 ms of transition against 105 µs of copying. Buffer, or readv.
  • Single-architecture images. docker buildx --platform linux/amd64,linux/arm64.