Process & Thread Primer
Our curl command runs as a process: a page table, a descriptor table and a few kilobytes of kernel bookkeeping, with one or many threads inside it. Five sections prove that fork copies the index and not the data, that a thread is the same syscall with more flags set, and that a switch costs a hundred times what the kernel's counter reports.
What a process is, and what fork actually copies
A running program is not the file on disk. It is a page table, a descriptor table, and a few kilobytes of kernel bookkeeping the scheduler can name.
When the shell runs curl it calls fork() then execve(), and the kernel allocates a task_struct, an mm_struct for the address space, a descriptor table and a PID. The largest of those, and by far the emptiest, opens on the stack — drag it along the 128 TB user half and pull the zoom back until the mappings vanish into it:
Notice how little there is to find. Nine mappings — text, rodata, data, the heap, three shared libraries, the stack and the vDSO — total 5.5 MB, four millionths of one percent of the space they sit in. An address space is not memory. It is a 47-bit namespace with a few needles in it, and the page table is the index that says where they are.
That index is exactly what fork() copies. Not the data — parent and child go on pointing at the same physical frames. Drag the parent's resident size and watch the page table grow while the data stays put:
At a 1 GB parent the kernel walks 262,144 leaf entries into 2 MB of fresh page tables and copies zero bytes of data. That is where the 1.1 ms comes from, and why fork is linear in RSS rather than free: at the same walk is 128 MB of tables and 67 ms — the pause a Redis background save has to hide.
Both halves of that bargain are conditional. Every shared frame is mapped read-only, so the child's first write to one traps into the kernel. Drag the writes across the parent's gigabyte and watch the copies pile up:
Watch the bill rather than the row. Each fault costs 1.2 µs — two crossings, a frame allocation and a 4 KB copy — and there are 262,144 pages, so a child that writes pays 325 ms in faults on top of the 1.1 ms fork. Copy-on-write is not cheap. It is deferred, and a child that touches everything gets the whole invoice.
The descriptor table is copied too — but only the table. The open file description underneath it, which is where the offset lives, is shared. Alternate writes between parent and child, then switch to two separate open() calls:
Because both descriptors index the same open file description, f_pos advances once per write and the twelve 4-byte records land end to end. Give each process its own open() and both start at offset 0: eleven of the twelve writes are overwritten and the file is 4 bytes long.
execve() then throws the address space away and keeps the task. Step through it, and turn O_CLOEXEC on to decide whether the descriptor goes with it:
As soon as the new image is mapped, all four mappings are gone and all four task fields are still there — same PID, same cwd, same credentials, same fd 3. That last one is a leak: without O_CLOEXEC every descriptor you happen to be holding is handed to the program you exec, which is how a server's listening socket ends up inside a child it spawned — one flag apart, and the compiler cannot see the difference:
int fd = open(p, O_RDWR); /* leaks */
int fd = open(p, O_RDWR | O_CLOEXEC); /* does not */Both lines compile and both pass tests. The same shape sits one level up as well, in the C library's own buffer: printf writes into process memory, so fork duplicates it. Print a few lines, then fork:
Both processes flush the same buffered bytes at exit, so four printfs before a fork put eight lines on the pipe. fflush(stdout) first leaves an empty buffer and nothing to duplicate. Nothing in the fork man page says so and no compiler warns you, which is why this one outlives every code review.
Threads are the same syscall with more flags
A thread is what the scheduler hands a CPU to. On Linux it is created by the same call as a process, with more of the sharing bits set.
pthread_create does not call a thread syscall. It calls clone() — the one fork() calls — with a different flag word. Everything people mean by "process" and "thread" is one dial between two extremes, and the kernel has no opinion about where you stop turning it.
That dial has five detents, and the familiar names only appear at the ends. Add the sharing flags one at a time and watch the child stop keeping its own copy of each resource:
With no flags set you have fork(): five private copies. With you have a thread — one address space, one working directory, one descriptor table, one signal disposition, one thread group. Everything between is legal, and some of it is useful: CLONE_FILES without CLONE_VM shares descriptors with a child that has its own memory.
What that thread costs is two different numbers and only one of them is real. Slide the thread count and watch the reserved address space and the resident memory pull apart:
Notice that at 1,024 threads the process has reserved 8 GB and is using 8 MB of it. The 8 MB stack is ulimit -s, a reservation; a fresh thread touches two pages of it. The cost that is really there is 25 KB of kernel memory per thread — a 16 KB kernel stack plus a 9 KB task_struct — and that one is not in your RSS at all.
Which is why the ceiling is set by the kernel's memory rather than yours, and why a second, unrelated ceiling usually gets there first. Drag the machine's RAM and watch the two limits change places:
threads-max is total RAM over 32 pages, so a 16 GB machine allows 131,072 threads. But pid_max defaults to 32,768 and every thread takes one, so 32,768 is the real answer — the same on a 16 GB box and a 1 TB one until somebody raises the sysctl. At the memory ceiling binds instead, at 16,384.
Sharing the address space has a consequence at the language level: a global is one cell no matter how many threads read it. Pick a thread, then switch the declaration to thread_local:
With static int n all four threads increment the same word, so it reads 4 and every increment is a race. Declaring it thread_local makes the compiler emit %fs:-0x8 instead of an absolute address, and each thread reads its own 1. Every per-thread allocator cache and per-CPU counter in the wild is this trick.
One kernel thread per concurrent task stops working long before the memory does. Drag the task count and compare what a runtime's own tasks cost against a thread each:
At goroutines at 4 KB apiece cost 4 GB and threads at 33 KB apiece cost 33 GB — but the thread bar turned rose at 32,768, because pid_max stopped you long before the memory did. That is the argument for M:N: not that threads are slow, but that there is a hard count of them.
The bill for that arrives at the first blocking syscall, because a task inside the kernel is holding the OS thread it was running on. Raise the share of blocked tasks, then turn the hand-off on:
Drive it to 100% with no hand-off and the runtime has zero usable cores while eight OS threads sit in read(). Go's sysmon fixes it by retaking the P after 20 µs and starting a replacement thread; Java's virtual threads unmount at the blocking point instead. Both fail the same way on a block the runtime cannot see — a page fault on a mapped file, a CGO call, dlopen.
A task's stack is the other thing that is cheap only until it is not. Go's starts at 2 KB and doubles when a frame will not fit, copying every live frame into the new one. Take the call depth up:
At 1,024 frames deep the stack has doubled five times to 64 KB and copied 62 KB getting there — a cost a pthread never pays, because it reserved 8 MB up front. So "goroutines are cheap" is a claim about memory and scheduling, not about syscalls and not about deep recursion: a runtime can only multiplex what it can observe.
The privilege boundary, and what crossing it costs
Two privilege levels, enforced by the CPU. Which one you are in decides what an address means and what an instruction is allowed to do.
x86-64 has four rings and uses two: ring 3 for applications, ring 0 for the kernel. ARM64 calls them EL0 and EL1. The kernel's half of the address space is described by the same page table as yours, and the hardware — not the kernel — is what stops ring 3 reading it.
So the boundary is a property of the address, not of a separate memory. Drag the address across all 64 bits and watch the two verdicts disagree:
Notice that the middle of the space is not memory at all. Only the bottom 128 TB and the top 128 TB are canonical — bits 48 to 63 must copy bit 47 — and an address in between faults on its own shape before any page table is consulted. The kernel half is mapped and ring 3 still cannot read it: that is the U/S bit in the page table entry, checked by the MMU on every access.
Getting into ring 0 legitimately takes eight steps, and the CPU does the first of them atomically so nothing can be observed half-done. Step through one syscall:
Watch where the page table changes. Step 5 is KPTI: since Meltdown the user page table does not contain the kernel, so entering means writing CR3 and leaving means writing it back. Two of the eight steps exist only because of a 2018 hardware bug, and they are the two that cost the most.
How much they cost depends on what your kernel has turned on. Switch between the mitigation configurations and raise the syscall rate until the core bar moves:
A bare crossing is 55 ns of instructions. KPTI with PCID adds 50 ns for the two CR3 writes; KPTI without PCID adds 190, because each write flushes the whole TLB. At 100,000 syscalls a second — an ordinary busy server — that is 1.1% of a core with PCID and 2.5% without; at it is 11% of a core with PCID and 25% without.
The cheapest crossing is the one that does not happen. clock_gettime reads a page the kernel maps into every process, so it never leaves ring 3. Raise the call rate:
At a million calls a second — a timestamp per request in a hot loop is exactly that — the vDSO costs 2.5% of a core and the syscall path costs 34%, fourteen times more. There is no trick in it: the kernel writes the clock into a shared page and the vDSO does the arithmetic in user space. io_uring is the same idea applied to I/O.
Faults cross the boundary too, and they do not all cost the same. Drag the pointer across sixteen pages and read what each access actually costs:
The spread is five orders of magnitude. A TLB hit is a nanosecond, a minor fault is 900 ns, a major fault is 100 µs because the page has to come off the device, and a SIGSEGV is 2.2 µs and then the process is gone. The cheapest thing on that list is the only one that ends the program — which is the good case, because the expensive ones fail silently and keep going.
Which is why a server that wants throughput does not make syscalls faster — it makes fewer of them. Hold the work at a million operations a second and submit them a few at a time:
Because the crossing is per submission and not per operation, an io_uring queue 128 deep turns a million crossings a second into 7.8K — 0.105 of a core down to 0.0008. The work did not get cheaper and the ring change is still 55 ns; there are simply 128 times fewer of them.
The context switch, and the bill it leaves behind
Two threads cannot run on one core, so the kernel swaps them. The direct cost is under a microsecond. What it leaves behind is a hundred times that.
A switch happens when the timer expires the running thread's slice, when it blocks on I/O, or when something more urgent wakes up. The kernel saves its registers, picks the next runnable thread, and loads that one's. What varies is whether the incoming thread belongs to the same process.
That one question decides whether one of the six steps happens at all. Step through the switch, then change who the incoming thread is:
Step 4 is switch_mm, and a same-process switch skips it: the incoming thread already has the right page table, so CR3 is never written. That is the entire direct difference — 730 ns for a thread, 930 ns for a process. If it were the whole story, process-per-request would be 27% worse and nobody would have written a runtime about it.
It is not the whole story, because CR3 is what the TLB is keyed on. Grow the incoming thread's working set and watch what it pays to re-walk translations it used to have:
Notice where the curve stops climbing. The L2 STLB holds 1,536 entries, so a working set past 6 MB was never fully cached, and from upward the flush cannot cost more than 38 µs of page walks. That ceiling is the good news: 38 µs is forty-one times the direct cost of the switch that caused it.
Which is why the TLB stopped being flushed. Every entry carries a process id, and Linux keeps six of them per CPU. Add processes to one core until it runs out of tags:
TLB_NR_DYN_ASIDS is 6. Up to six processes round-robining on a CPU keep their translations across every switch and pay 930 ns. The seventh evicts somebody, and from then on each switch costs 930 ns plus the re-walk — 13 µs for a 2 MB working set. PCID is not "the flush is gone"; it is "the flush is gone if your runqueue is short".
The caches have no such tagging at all. Grow the working set again and watch the direct cost disappear into the two things the switch dragged behind it:
At the direct 930 ns is a hairline at the left of the bar. 6.4 µs of page walks and 87 µs of refilling L2 at 12 GB/s put the real cost at 95 µs — a hundred and two times the number the kernel's own counter reports, and it tops out at 127 µs. Li, Ding and Shen measured the same shape in 2007: indirect costs from a few microseconds to over a millisecond.
So the scheduler's job is to switch as rarely as it can get away with. Shrink the time slice and watch the share of the core that switching takes:
At the shipped 0.75 ms default, switching costs 0.097% of a core — invisible. Pull the slice down to and it is 12.7%, with every cache effect above now happening a hundred and fifty times more often. Preemption granularity is not a fairness knob you turn freely. It is a tax rate.
Which is what a thread pool is really choosing. Grow the pool for a request that spends 200 µs on CPU and 2 ms waiting, and watch throughput and latency come apart:
Throughput saturates at 88 threads — eight cores times the elevenfold ratio of wall time to CPU time — and then stays flat all the way to 4,096. Latency does not: 2.2 ms at the knee, 13 ms at 512 threads, . Every thread past the knee buys nothing and adds its own queueing delay to every request in flight.
The invariant worth carrying out of this section is that a runnable thread which cannot get a core is not waiting — it is making everyone else wait. Pool size is not a capacity dial. It is the length of a queue you have decided to keep inside your own process instead of in front of it, where you could have shed it.
Quick reference
Three questions worth answering cold, two of them with a slider on them, and five red flags.
How much is context switching costing me right now?
Read the cs column out of vmstat 1 and multiply by 730 ns. The eight cells are eight cores, and the slider walks what the rate takes off them:
At 100,000 switches a second — an ordinary busy server — 0.07 of a core is gone before your code runs, which is nothing. At 4.1 cores of eight are gone and the other 3.9 are running with cold caches. The number to alarm on is not the rate but the rate per unit of work: two switches per request is healthy, twenty is a design problem.
Process, thread or task — how do I actually choose?
By counting the concurrency you need against the memory and the PIDs you have. Drag the concurrency and watch which model runs out first:
Notice which bar turns rose first. At process-per-request has eaten 24 GB. Thread-per-request survives to 32,768 and is then stopped by pid_max, not by memory. Tasks reach 4.2 million on the same box. Process-per-request is not wrong — nginx and Postgres both use it — it is wrong per request.
And how expensive is it to create one?
Enough that the creation rate is its own budget line. Raise it and watch the cores go:
fork plus execve is 355 µs — 55 µs of page table for a 10 MB parent, then 300 µs to tear the address space down and map a new image. pthread_create is 25 µs and go func() is 300 ns. At a thousand a second that is 0.36 of a core, 0.03, and 0.0003.
- A process per request. 3 MB of private RSS each, and a fork whose cost is linear in the parent's page table.
- An unbounded thread pool. The ceiling is
pid_maxat 32,768, and the knee was at 88. - A blocking call inside an async task. It pins the OS thread underneath and starves everything queued on it.
- Shared mutable state with no synchronisation. Threads share the address space; nothing in the language reminds you.
- Tuning the scheduler's granularity for latency. Below about 50 µs of slice you are paying more than 1% of every core for it.