File System Primer
A server appends one line to access.log. The write() returns in under a microsecond, and the line is nowhere a power cut cannot reach it. Twenty-five figures you can drive, and one argument: write() promises RAM and nothing else.
A file is a name, an inode, and some blocks
A file is not a primitive. It is a name in a directory, an inode the name points at, and a set of blocks the inode points at. Every operation rearranges those three.
Bottom up: blocks are fixed-size chunks of the device, 4 KB on ext4 to match the page size; an inode is one file's metadata plus the map from file offset to block; and a directory is an inode whose data is a table of names. Start with the path as the kernel first sees it — nothing but a chain of names, none of them read yet, and the slider walking one link at a time:
Nothing there costs anything, and that is the point of the pose: a path is a sequence of lookups, not one. Each lookup has to read the directory's own data block to find the name in it, so the same walk with that one box added is:
Five names, five directory blocks. But a name in a directory is only a number — it says which inode, not what is in it — so every one of those lookups needs a second read of the inode it names. Add that box, walk the components again, then switch the cache from cold to warm:
Notice where the cost is: not in the file. Cold, five components are ten block reads before a single byte of the log is touched. Warm, the count is zero — the dentry cache answers every lookup from RAM. That is why open() on a deep path is expensive once and free afterwards.
Inside the inode the map is fifteen slots: twelve direct, then single, double and triple indirect. Where a byte sits decides what reading it costs. Drag the offset, or pull the handle along the axis:
Watch the chain grow. With 4 KB blocks and 4-byte block numbers one indirect block holds 1,024 pointers, so the tiers reach 48 KB, 4 MB, 4 GB and 4 TB. A byte at costs three block reads, and one at four, three of them pure map. The deeper the file, the more expensive its tail.
Extents replace the list with a run. One ext4_extent is 12 bytes and describes up to 128 MB of contiguous blocks, and four of them fit inside the inode itself. Drag the fragmentation up and watch the extent map climb toward the pointer list:
A gigabyte is 262,144 blocks. As pointers that is 1 MB of map, plus the indirect blocks to hold it; as eight extents it is 96 bytes — and yet even that does not fit. ee_len has 15 usable bits, so one extent covers at most 32,768 blocks, and a perfectly contiguous gigabyte still needs eight of them; i_block is 60 bytes, a 12-byte header and 12-byte records, so the fifth spills the map into an extent-tree leaf block — one extra read, not thirty-two.
The name and the file are separate objects, and the inode counts its names. Drag the count to zero, then switch to a symlink and do it again — the names, the inode, what is left:
Because the inode counts references, rm is unlink: it removes a name and decrements. The blocks are freed at zero — and not even then if a process still holds the file open, which is how a deleted log keeps filling a disk that rm cannot free. A symlink counts nothing: delete its target and the link survives, pointing at ENOENT.
Directories are the other place cost hides. ext4 indexes them with an htree, a hashed B-tree one or two levels deep; without it a lookup is a linear scan of every block. Drag the entry count up the log axis and watch the two separate:
They separate immediately. At the htree is three block reads where the linear scan averages 2,942. The index is why nobody notices — until something calls readdir and then stat on every entry, which no index helps with.
All of this is laid down at mkfs time and sized then. Drag the inode density, then drag the size of the files you plan to store below it:
The default of one inode per 16 KB spends 1.6% of the device — 64 GB of a 4 TB volume, empty, before you write anything. (KB here is 1,024 bytes throughout, the way the tools report it: 4 TB / 16 KB = 268,435,456 inodes × 256 B = 64 GB, exactly.) The count is fixed at format: with on a filesystem the inodes run out at 6.3% of capacity, and you get ENOSPC while df still shows the disk nearly empty. Only df -i says why.
The page cache, and what write() actually promises
Between your code and the device sits a cache of file pages in RAM. Almost every read hits it, every write lands in it first, and that second fact is where lost data comes from.
The page cache is the kernel's copy of recently touched file data, indexed by (inode, offset) — buff/cache in free -h, routinely half of RAM on a server, handed back the moment anything else wants it. On a read the kernel looks there first: a resident page is a memcpy, a missing one is a trip to the device. Walk the slider across the pages, then change the device underneath:
Notice the gap. A hit is 0.7 µs — a syscall, a lookup, and 4 KB of memcpy. The same read from NVMe is 80 µs, 114× more; from a 7,200 rpm disk it is 10.2 ms — 6 ms of mean seek plus 4.17 ms of mean rotation at 7,200 rpm — . Nothing in your code changed. The only thing that changed is whether the page happened to be there.
A write is different: the kernel copies into the cache, marks the page dirty, and returns. Getting it onto the media is somebody else's job, later. Drag the power cut along the timeline and find the moment the page stops being yours to lose:
Because dirty_expire_centisecs is 3000 and the flusher wakes every 5 s, a page written now is only eligible for writeback after 30 s and can sit in RAM for 35. A successful write() means the bytes are in kernel memory. It does not mean they are anywhere a power cut cannot reach.
Dirty pages are capped, and the cap is enforced on whoever is writing. dirty_background_ratio is 10% of RAM and dirty_ratio is 20%. Below both, a write is a memcpy. Push the dirty share past them and watch what write() starts costing:
Watch the step. Under 10% a write is a memcpy; past 20% the caller is made to do the writeback itself and the cost becomes the device's — 321 µs instead of 0.7 µs at . That figure is a model, not a measurement: above the threshold balance_dirty_pages runs a feedback loop that sleeps the writer in proportion to the excess, and the drawing prices one sleep as a couple of device writes. The height of the curve is kernel-version specific; the step at 20% is not. This is the mechanism behind “the process froze for two seconds” during a large copy. Nothing froze: one write() paid everyone's backlog.
Reads get help that writes do not. When accesses look sequential the kernel opens a readahead window and doubles it up to the 128 KB in read_ahead_kb. The strip below is exactly one such window — 32 blocks of a file — and the stride picks which of them you ask for:
At the window is open: one request covers all 32 blocks, and the walk costs 1.1 µs for every kilobyte it uses. At it shuts. Every block you wanted is now its own request — 16 of them for the 64 KB you actually read — and a useful kilobyte costs 20 µs, eighteen times worse, and stays there at . You are not paying for bytes; you are paying for requests, and each one still drags in a whole 4 KB block of which you use a field.
O_DIRECT removes the cache entirely: DMA from your buffer to the device, alignment constraints and all. It is often reached for as an optimisation. Drag the re-read count and watch the gap the cache is opening on it:
It survives exactly one read. From , buffered costs 0.7 µs where O_DIRECT goes back to the device — 81 µs against 159 µs for two reads of one block. O_DIRECT wins only when you have your own cache and better information than the kernel, which is why databases use it and application code usually should not.
Durability — the journal, fsync, and the promise nobody makes
A power cut is the test. After the reboot the filesystem has to still make sense, and your file has to be either the old one or the new one. Those are two different guarantees, and only one of them is free.
A metadata change touches several blocks at once — a bitmap, an inode, a directory block — and interrupted halfway it leaves a filesystem that contradicts itself. Before any of that, though, the bytes have to get somewhere. Four places, and only the last one survives: your buffer, the page cache, the device's own write cache, then the media. Press play, or scrub the sequence by hand:
Notice that fsync is two steps, not one. Flushing the dirty pages to the block layer is not enough: drives acknowledge writes into volatile DRAM, so fsync also issues a cache-flush command and waits for the device to admit it is done. A write that stops at step five is durable exactly as long as the power holds.
Journalling makes the metadata change atomic. The kernel writes the new blocks to a log, then a commit block, and only then touches the filesystem. Move the power cut through the transaction and read the verdict at every point:
The invariant is one sentence: at recovery a transaction with a commit block is and one without is , so the filesystem is always exactly the old one or exactly the new one. And recovery is bounded by the log, not the disk — mke2fs caps the journal at 128 MB, which is 38 ms of sequential NVMe read against 20 s merely to stream a 4 TB volume's inode table. That 20 s is a floor, not an estimate: a real e2fsck also walks the bitmaps and the directory tree, in seeks rather than one sequential pass. The ratio survives it — a bounded log and an unbounded disk.
ext4 offers three modes and the default is a compromise. data=writeback journals metadata and lets the data land whenever; data=ordered writes the same bytes but forces the data down first; data=journal puts the data through the log too. Take the segmented control left to right and grow the transaction with the slider:
Watch the amplification. At data=journal writes 2.01× — every byte twice — for the strongest crash semantics on offer. And watch what data=writeback does not buy: its bar is the same length as data=ordered's, because it writes exactly the same bytes. What it drops is the order, and that is free until it isn't — the metadata may commit before the data lands, so a crash can leave a newly-extended file showing whatever those blocks held before, which may be someone else's deleted data. A mode that costs nothing and can hand you another tenant's bytes is the one to avoid.
fsync is the only POSIX call that promises anything, and it is priced per commit rather than per byte. Pick a device, then put more writers behind one call:
Because the cost is per flush, batching is nearly free throughput. One writer on a 7,200 rpm disk gets 100 durable commits a second — the platter has to come round, and 4.17 ms of that is mean rotational latency at 7,200 rpm. sharing one fsync get 1,581 — sixteen commits per rotation, less the queueing this figure prices at 8 µs per extra writer. Only the rotation is measured; the queueing term is the drawing's model, and it is the only reason the number is not a flat 1,600. That is exactly what a database's group commit is.
Which raises the obvious next question: fsync is 700 µs, so what does fdatasync skip, and why not always use it? Drag the call, then change the condition it runs under:
Notice the second column. forces the data and skips the inode, so it saves the metadata write and its journal commit — 540 µs against 700. But switch to and the discount vanishes: the new size is needed to read the data back, so POSIX makes it part of the data, and an appending writer — a log, a WAL, a segment file — gets nothing from fdatasync at all. Preallocate with fallocate and the discount comes back.
The other three rows are the same breakdown with pieces missing, and the pieces have very different sizes. Take the flush out of the three device writes and look at what is left:
sync_file_range is the row people misread. and forces nothing else — no metadata, and no cache flush — so it is a way to start writeback early, never a way to be durable. Its own manual page says so. And is that same deletion applied to fsync: drop the flush and a commit is 241 µs instead of 700. It is only safe on a device with power-loss protection, which is what the PLP rung on the ladder is; on anything else it is a 2.9× speed-up that loses the last second of writes on a power cut, silently.
And fsync can fail. When the device reports an error the kernel holds a dirty page it cannot write and no way to keep it forever, so it drops it. Step through what the next call reports:
This is fsyncgate, found in PostgreSQL in 2018. The second fsync returns 0 because the error was consumed and the page is already gone — success reported for a write that never happened. Before Linux 4.13 only one file descriptor saw the error at all. The answer was uniform across the industry: treat an fsync failure as unrecoverable, panic, and rebuild from the log.
Which is why replacing a file is a four-step recipe and not a write. Pick a recipe, then move the cut through it and read what survives the reboot:
Only the full recipe is safe at every cut. Writing over the original in place leaves it half old and half new with no copy of either. Renaming without fsync on the temporary file flips the name onto blocks that were never allocated — a zero-length config file, which is what bit thousands of ext4 users in 2009. The wrong version differs from the right one by two lines, and it passes every test you will write for it:
write(f, buf, n) fsync(f) # the wrong version skips this close(f) rename(tmp, dst) fsync(dirfd) # and this
Underneath all four recipes is one hard floor: the device's 4 KB sector is the largest thing that is atomic, and nothing above it is. Move the cut through a page write and widen the page:
An 8 KB page spans two sectors, so one crash point in three leaves it torn — half the new page, half the old, checksum failing. POSIX never promised otherwise. Postgres pays for the gap with full-page writes into the WAL; ZFS and btrfs pay for it by never overwriting a live block at all, which is where §04 starts.
Where ext4, XFS, ZFS and btrfs actually differ
Everything so far is common to all four. What separates them is three decisions: overwrite or copy, checksum the data or trust the device, and how much the allocator knows before it commits.
ext4 and XFS overwrite in place and journal the metadata. ZFS and btrfs never overwrite a live block: a write goes to fresh space and the change is published by rewriting the path up to the root, so the commit is one pointer. Walk a 4 KB overwrite from the leaf to the superblock and watch the new blocks appear beside the old ones rather than on top of them:
Notice that the old tree is complete and mountable at every step until the last. That is the property a journal buys with a log; COW gets it from the shape of the write. The bill is amplification — a 4 KB overwrite touched six blocks — and fragmentation, which is why COW filesystems slow down under random writes and want periodic balance.
The second difference is whether anyone checks. A consumer disk is specified at fewer than one unreadable sector per 1014 bits read. Drag the amount read up to the size of an array rebuild:
At the chance of hitting one is 62%. ext4 and XFS checksum their own metadata and not your data, so a flipped bit in a file is handed back to you as data; ZFS and btrfs checksum every block and can tell you the file is wrong. Detection, not correction — correction needs a second copy, which is what a mirror is for.
The third is what the allocator knows when it runs. ext4 delays allocation until writeback, so it sees the whole run at once instead of one write() at a time. Change how much is buffered before the allocator gets to choose:
At 4 KB granularity a gigabyte becomes 262,144 extents and 3 MB of map. Buffered in 128 MB chunks it is 8 extents and 96 bytes — still two records past the four i_block holds, but one leaf block instead of a three-level tree. Delayed allocation is also why a crash straight after write() can find nothing allocated at all: the trick that makes files contiguous is the same trick that makes them empty.
One more, because most of us now meet a filesystem through a container. overlayfs stacks a read-only lower layer under a writable upper one, and the first write to any lower file copies the whole file up. Drag the file's size:
A one-byte change to a blocks for about 175 ms — the copy runs at ~1.2 GB/s, reading and writing the same device — and then costs 200 MB of the container's writable layer. It is why a database inside an image wants a volume, and why chmod -R on a layer is not the cheap operation it looks like.
Quick reference
Four questions worth answering cold, and five red flags.
What does one file operation cost?
Entirely on where the answer is found, and the spread is four orders of magnitude. Each rung below is one operation, and the slider walks the one being paid for:
The step worth memorising is the third rung to the fourth: is 80 µs, and on the same consumer drive is 700 µs — click either and the ladder brackets it against the page-cache hit it is priced from. Above the gap you buy throughput with the cache; below it you buy it with group commit, and there is nothing else to buy it with.
What does write() promise, and what does fsync() add?
write() promises the bytes are in the page cache and that a later read() — from any process — will see them. It promises nothing about the media: the page can sit dirty for 35 s. fsync() adds the flush to the block layer and the device cache-flush, and returns only when the device acknowledges. If it fails, treat the data as gone.
Why is a million files in one directory a bad idea?
The lookup itself is fine — the htree makes it three block reads. It is everything that walks the directory that hurts. Drive the entry count up the log axis:
Notice the second band. ls -l is one readdir plus one stat per entry, and at that is 68,383 block reads — 5.5 s on NVMe, 697.5 s on a spinning disk. Hash into subdirectories the way git does.
When would you pick XFS over ext4, or ZFS over both?
XFS for very large files and many concurrent writers: allocation groups let cores allocate without contending on one bitmap. Past that it is not a matter of taste — two countable mechanisms decide it. Grow the dataset, then switch the question from what a snapshot costs to what a flipped bit does:
Notice which row stays teal in both modes. ext4 has no reflink, so a snapshot is cp; XFS clones the extent map instead — one 16-byte xfs_bmbt_rec per run, and since its br_blockcount field is 21 bits one record reaches just under 8 GiB, so a contiguous gigabyte clones as one record in a 512-byte inode. ZFS pays one branch — six blocks, 24 KB, and the bar does not move at a terabyte, because the branch it copies is one it already wrote. But only ZFS checksums the data — the other two hand you the flipped byte as a value, and no fsck finds it. ext4 unless one of those rows is your problem.
write()in a loop, then claiming durability. One power cut from never having existed.- Retrying a failed fsync. The page is gone; the retry returns 0.
O_TRUNCover a config file. Temp file,fsync, rename,fsyncthe directory.- A directory as a key-value store. The lookup stays cheap;
readdirplusstatdoes not. O_DIRECTas an optimisation. From the second read, the cache you removed was free.