Virtual Memory Primer
Every pointer your curl process touches is a virtual address, translated to a physical one on every access, billions of times a second. Five sections build the picture: why it exists, page tables, the TLB, page faults, and a quick reference.
Why virtual memory exists
Every process sees a private 64-bit address space with 128 TB usable. There isn't 128 TB of RAM in the machine. That illusion buys three separate things.
Without it, each process would address physical RAM directly, and the first thing to break would be isolation. Below, process A's eight virtual pages are scattered across the machine's frames, process B's live in frames A has no name for, and A's last two pages are mapped nowhere. Drag the slider through A's pages, then past the end of its map:
Notice what A cannot express. No virtual address it can form names one of B's frames, because no entry of A's table points there. Isolation is not a check the kernel runs — it is the absence of a mapping, which the hardware enforces for free.
The unmapped read fails loudly: no region covers the address, so the process takes SIGSEGV at the instruction that did it. Loud is the good case. The quiet one lands inside a mapping the process does own — the hardware has nothing to object to, and the corruption surfaces elsewhere.
The second thing is address-space planning. Sharing one physical space, every program would need an address no other program wanted, and two that both wanted 0x400000 could not run together. Here both are linked there anyway — program A and program B — and the slider moves where the loader drops B:
Because the address B is compiled for is virtual, the loader never rewrites the binary: it fills in a page table and jumps. That is what makes ASLR nearly free, and what lets two copies of one binary share a single set of read-only frames while both believe they own 0x400000.
The third is overcommit. Programs reserve far more than they touch: a 200 GB JVM heap on a 64 GB machine is ordinary. Only the touched part of the reservation costs a physical frame. Drag the slider up and watch the reservation stay exactly as wide as it started:
Watch what happens past 72 GB — RAM plus swap. Nothing failed at malloc time, because malloc only promised address space; the failure arrives later, at a fault the kernel cannot satisfy, and lands on whichever process the OOM killer picks rather than the one that over-reserved. That is the price of vm.overcommit_memory=0.
None of it is free. The access itself is about 1 ns out of L1; the translation overhead is whatever fraction of accesses must walk the page table, at roughly 100 cycles each. Raise how many of every thousand accesses miss:
Notice how fast the overhead overtakes the work. At 32 misses per thousand — a 3.2% miss rate, unremarkable for a large working set — translation costs as much as the access it is translating. That curve is the whole reason the TLB exists. The page tables themselves add about 0.2% of the mapped size in kernel memory, which is the small half of the bill.
Page tables — the translation hardware reads
The kernel keeps, per process, a tree of tables mapping virtual pages to physical frames. The CPU walks it on every access whose translation isn't already cached.
Memory is managed in pages — fixed-size aligned chunks, 4 KB on every common CPU. Translation happens at page granularity, so the bottom twelve bits pass through untouched while the four nine-bit index fields above them choose the page. Drag one byte at a time and watch what actually moves:
Notice that the frame does not change while the offset does. Four thousand and ninety-six consecutive addresses share one translation, which is why translation is affordable at all: the hardware does this once per page, not once per byte. Cross the 4 KB boundary and the PT index steps, and the frame with it.
So why a tree? A flat table with an entry per virtual page needs 236 entries at 8 bytes each. Widen the flat table against the four-level tree mapping the same 4 MB:
At 48 bits the flat table is 512 GiB per process — and it is 512 GiB whether the process maps four gigabytes or four kilobytes, because a flat array reserves a slot for every address it could be asked about. The tree allocates a node only where something is mapped: five 4 KB pages here, one PML4, one PDPT, one PD and two PTs.
x86-64 spends the 36 index bits as four nine-bit fields, and the shape is not arbitrary: 29 entries × 8 bytes is exactly 4096, so every table is itself one page. Step the track and spend one level per notch:
Watch the load counter. Four dependent loads — each address comes out of the entry the level before it returned, so no memory-level parallelism helps. Linux gained a fifth level on Ice Lake for 57-bit addresses, used only above 128 TB, at the cost of a fifth dependent load.
Each entry is eight bytes: forty of frame number, the rest permission and state. The bits that permit and the bit that denies are checked in parallel with the walk. Switch the access and drive this entry into a fault:
Notice that the fault carries an error code for what was tried, not what was wrong: bit 0 that the page was present, bit 1 that it was a write, bit 2 that it came from user mode — 0x7. The kernel decides from that whether this is a bug or a copy-on-write to service silently. Accessed and Dirty are set by the CPU, and page reclaim reads them when it picks a victim.
A huge page is an entry with the PS bit set, telling the CPU it is already looking at the final page. Set that bit one level higher each time and watch the level the walk stops at climb:
Two things improve together: one fewer dependent load, and one TLB entry covering 2 MB instead of 4 KB. The bill is in the last readout. A 2 MB page backing a 100 KB allocation wastes 1.9 MB, and cannot be copied-on-write in pieces, so one byte written after fork() copies the whole 2 MB — which is why Linux ships Transparent Huge Pages in madvise mode.
One register decides which tree the CPU walks: CR3 on x86, TTBR0_EL1 on ARM64. It holds the physical address of the top-level table, so changing it changes every translation at once. Move which process holds the CPU and watch the same virtual address land elsewhere:
Swapping the tree under CR3 is what a context switch does to memory, and what happens on every kernel entry under post-Meltdown KPTI, which gives each process two trees — one with the kernel mapped, one without — and switches on every syscall. That doubling is where much of the 2018 syscall regression went.
The TLB — why translation is fast
Without it, every memory access would be five: one for the data, four for the walk. The TLB is the small cache that makes the scheme affordable.
The Translation Lookaside Buffer caches recent page-to-frame mappings and is consulted in parallel with the L1 lookup, so a hit costs essentially nothing while a miss pays for the walk and installs what it found. Step the track through sixteen accesses against eight entries:
Notice which misses are unavoidable. A first touch must miss — nothing could have installed it — but the late miss on page 17 is different: it was resident until page 70 pushed it out, because eight slots cannot hold the nine pages this loop touches. Compulsory misses shrink as a program runs; capacity misses do not, and they decide throughput.
Real hardware answers in tiers: a small L1 dTLB on every load and store, a larger L2 behind it, and only then the hardware page-walker, which itself reads through the data cache. Move the slider down the tiers:
Those are Intel Skylake-through-Golden-Cove numbers for each tier: ~64 L1 dTLB entries for 4 KB pages, ~1500 in L2, ~100 cycles for a walk whose tables are cached and several hundred when they are not. Apple M-series ships ~3000 L2 entries, part of why those cores tolerate large working sets.
Entries times page size is coverage — how much memory the TLB can describe at once. Push the working set past the coverage line and the miss rate does not degrade gently; it goes over a cliff:
Once the working set exceeds coverage, a random access has only coverage/working-set chance of hitting — the curve is 1 − cov/ws, already at 50% at twice the coverage. At 1500 × 4 KB that is just under 6 MB, smaller than the L2 data cache on the same chip: a workload can fit in cache and still spend a third of its time walking page tables.
Huge pages attack the coverage term directly: the entry count never changes, only what each entry describes. Switch the page size under the same cliff:
The same 1500 entries put the coverage line at about 6 MB, 3 GB and 1.5 TB. Redis, PostgreSQL's shared buffers, JVM heaps and HPC arrays are not asking for faster memory; they are asking for the cliff to move past their working set. Check with perf stat -e dTLB-load-misses,dTLB-loads: above ~1% on a workload that claims to fit in cache, the cliff is what you are on.
The TLB caches translations from one tree, so anything that invalidates a translation must invalidate the copy — INVLPG for one, a CR3 reload for all. Before ASID tagging, every context switch threw it all away. Drag the switch count and compare the untagged TLB with the tagged one:
With address-space identifiers — Intel's PCIDs, shipped in Westmere in 2010 — each surviving entry carries the tag of the tree it came from, so a switch stops using the other process's entries rather than deleting them. What is left is capacity pressure. This is the biggest single reason context switches got cheaper through the 2010s, and why KPTI hurt so much on CPUs without PCID.
Page faults — and the uses Linux makes of them
A page fault is what happens when the CPU asks for a translation the table does not have. Linux treats it as a hook, not an error path.
On a missing present bit, or a permission violation, the CPU traps. The kernel reads the faulting address out of CR2, looks it up in the process's list of virtual memory areas, and either repairs it and re-runs the instruction or gives up. Switch the kind of fault and watch what it costs:
The three outcomes differ by a factor of fifty thousand. A minor fault never touches storage — find or zero a frame and return, a few thousand cycles. A major fault has to read, so its cost is the device's. An invalid one is not repaired: no area covers the address, so the process ends in SIGSEGV.
Demand paging is built on this. malloc of a gigabyte marks the range allowed and returns; nothing is backed until touched. Walk through the allocation and watch the resident set grow one fault at a time behind the page being touched:
Notice the fault count rather than the megabytes: 256 minor faults per megabyte, one per 4 KB page. A service that allocates and immediately fills 1 GB pays about 262,000 kernel entries to do it — which is why allocators reuse arenas, why MAP_POPULATE exists, and why RSS rather than VSZ is the number worth alerting on.
mmap is the same mechanism pointed at a file. The mapping starts empty, so the first read of each page faults out to storage while every read after it is ordinary memory. Step through twelve reads of an eight-page file, the last four revisiting pages already read:
Because the first touch does the I/O, the cost is not where the code says: there is no read() in the source, only a dereference that occasionally takes 78 µs. That makes mmap excellent for repeated random access to a hot file and poor for a single streaming pass, where read() with readahead wins.
fork() copies the page table, not the pages: every frame goes read-only in both processes and its reference count rises. Drag the write count and watch the shared frames split, one at a time, into a private copy:
Only the written pages cost anything, and everything still shared costs nothing — which is why forking a multi-gigabyte process is instant. The mirror image is the failure: a parent writing across its whole heap during the child's lifetime duplicates the whole heap. And because the copy happens on a write fault, the memory is committed then — a fork can OOM long after it returned.
When frames run out the kernel reclaims: pick the coldest pages, write them to swap, mark the entries not-present, free the frames. Drag the working set past the size of RAM and watch the average access cost move:
Watch the multiplier, not the picture. One access in three landing in swap does not make the program a third slower — it makes it four hundred times slower, because the two costs are 80 ns and 100 µs. That is why swap is a safety valve, and why the symptom is seconds of unresponsiveness rather than a gentle slowdown.
Past the end of swap the kernel stops choosing pages and starts choosing processes. The OOM killer scores by footprint and oom_score_adj, and it will pick the database over the log shipper unless you have said otherwise.
Quick reference
Three questions worth answering cold, and five red flags.
What does one memory access actually cost?
Entirely on where the answer is found, and the spread is six orders of magnitude. Every rung below is one memory access, and the slider walks the one being paid for:
The step worth memorising is the fourth rung to the fifth: a TLB miss is still nanoseconds, a fault is microseconds. Above the gap is a hardware problem you tune with page sizes; below it is an I/O problem you tune with memory budgets.
What does mmap actually do?
It creates a virtual memory area backed by a file or by anonymous zeroed memory, and marks the range not-present. The first access to each page faults, reads the data into a frame and wires up the translation. Used for large files, shared libraries, shared memory, and large malloc allocations.
When do huge pages help, and when do they hurt?
They help when TLB coverage, not bandwidth, is the bottleneck. They hurt on small allocations, and the slider shows why — every request under 2 MB pays for a whole page, and the remainder is waste:
That waste is not the only cost: they also hurt fork-heavy processes, where a one-byte write copies 2 MB, and hosts too fragmented to find a contiguous 2 MB block.
- A huge reservation never touched. No RAM, but page-table memory and TLB pressure across the range.
- Repeated mmap/munmap of small regions. Each is a syscall, a page-table edit, and a TLB shootdown per core.
mlockon a large range. Locks pages out of reclaim. For secrets, not speed.fork()after a large write burst. Every recently written page is a copy waiting on the next write.- Ignoring
vm.overcommit_memory. The default turns an allocation failure into a later OOM kill of some other process.