Disk Storage Primer

Seven sections, three claims. A platter spends 99.8% of a 4 KB read waiting. A flash drive answers one write with six or seven. A million-IOPS drive gives you 28,571 if you ask for one thing at a time. Every number below is read off the figure beside it.

01

The head, and what it is waiting for

A 4 KB random read is the canonical storage benchmark. On a platter it costs 8.34 ms, and almost none of that is reading.

The numbers below are the Seagate Exos X18 product manual's: 7,200 RPM, 4.16 ms average seek, 258 MB/s sustained. Two of the three are the drive doing nothing useful.

The first is geometry. Once the arm is on the right track, the head still has to wait for the sector to arrive, and the arc left to travel is the wait. Drag the platter round:

180° still to turn — 4.17 ms of waiting. drag the platter to turn it; the arrow keys turn it six degrees at a time and Home puts the sector back half a turn away
180° still to turn — 4.17 ms of waiting

Notice there is no way to make it shorter. One revolution at 7,200 RPM takes 8.33 ms, so the wait is uniform between zero and 8.33 and its average is half a turn. Nothing above the head can influence it: a faster CPU, a bigger cache, a better plan — the platter turns at the same speed.

Getting to the track costs too. Send the head across the platter and read the seek off the curve below it:

track 0 — 0.00 ms

Watch where the average lands: a third of the way across. Two random tracks are a third of the stroke apart on average, so 0.5 ms of settle plus a third of 11 ms is 4.17 ms — the 4.16 ms on the datasheet. Each of the twelve rings stands for about forty thousand real ones.

Add the two waits to the transfer and you have the whole cost of one read. The transfer is the only part that moves data; everything else is the head not reading. Grow the read from 4 KB and watch the proportions invert:

4 KB of read

At , 99.8% of the time is waiting: 120 IOPS, 491 KB/s. At the waiting is down to 20% and the transfer finally delivers 205 MB/s. A platter's 258 MB/s rating is only reachable at megabyte-scale I/O; its 4 KB number is a hundredth of it.

Which is why layout beats bandwidth on this device. The same 64 blocks read as scattered 4 KB requests pay 64 waits; read as one run they pay one. Switch the layout under the same block count:

scattered — 491 KB/s

Because the head never moves for the run, 64 blocks cost 9.34 ms together instead of 534 ms apart — 57 times faster for the same bytes. Every write-ahead log and every LSM compaction exists to turn the first pattern into the second.

None of these numbers mean anything until you put them beside what the CPU could have been doing. Walk the slider down from an L1 hit to the platter, and read the multiplier at each rung:

L1 — 1 ns

The step worth memorising is DRAM to NVMe: 80 ns to 35 µs, a factor of 440. The platter is 8.3 million times an L1 hit — the gap every cache above storage exists to hide.

02

Inside the flash

An SSD is not a hard drive with the moving parts removed. It is a very different device wearing a hard drive's interface, and the disguise costs something.

A flash cell is a transistor with an insulated pocket you push electrons into; the charge in the pocket is the stored value. The trick that made flash cheap was splitting that range into more levels.

There is one fixed window of charge to divide, so an extra bit does not get an extra window — it gets half of the margin between neighbours. Step through the four technologies and watch that margin collapse:

SLC — margin 1/1 of the window

Notice the arithmetic. Three bits need eight levels and seven gaps, so a TLC cell has a seventh of an SLC cell's margin and QLC a fifteenth. That is the whole reason QLC is rated for 1,000 program/erase cycles where SLC gets 100,000: nothing about the oxide changed, only how much room the reader has to be wrong in.

And the reader does get less room over time. Every erase pushes electrons through the insulator and damages it, so each level's charge distribution spreads. Wear the cell past its rating and watch the neighbours meet:

0% of the rated endurance used — 0 cycles

The rating is not a cliff. It is the point where the spread has just filled the margin. the levels overlap and reads return the wrong bits — hidden behind ECC until it can't be, and then the block is retired. It is also why a worn drive gets slower: read-retry at shifted reference voltages costs microseconds an attempt.

The same leakage runs when the drive is switched off, and then nothing is refreshing anything. JEDEC's JESD218 sets the floor: a worn client SSD must hold its data a year at 30 °C. Wear the cells and warm the shelf:

60 months unpowered

Because charge loss is thermally activated, retention roughly halves every 9 °C: five years for a fresh drive at 30 °C, and under two months for a worn one in a 55 °C rack. An archive on unpowered SSDs is not an archive.

Above the cell sits the asymmetry that shapes everything else. You can program one page — 4 to 16 KB — but you can only erase a whole block of them, and within a block the program pointer only moves forward. Fill one:

0 of 16 pages programmed

Sixteen pages here; a real TLC block holds 768 to 2,304 of 16 KB each, which is 12 to 36 MB. There is no way back up the row: to put new data at a page you have already written, the whole block has to be erased first.

So what happens when the host rewrites a block it has already written? The device cannot honour that literally. It writes the new copy to a free page somewhere else and marks the old one dead. Try it both ways:

refused — the page is already programmed

Because the physical location moved and the logical address did not, something has to remember where the data went. That table is the flash translation layer. Walk the logical addresses and watch it resolve:

logical block 0 → physical 21

Notice how much this costs: a page-level map needs about 1 GB of controller DRAM per 1 TB of NAND, which is why cheap drives use a coarser map or borrow host memory over HMB — and why enterprise drives carry capacitors. That map is the drive.

03

The writes you did not ask for

Rewriting in place is impossible, so overwrites strand dead pages inside live blocks. Reclaiming them is where an SSD spends its life and its endurance.

Every overwrite leaves a stale page behind. Free space runs out not because the drive is full but because it is littered, and the only broom is the erase.

So the collector has to move the pages that are still live out of a block before it can erase it, and those moves are writes nobody asked for. Step through one cycle:

the block is full: eight pages live, eight already replaced elsewhere

Notice what step four costs. Sixteen free pages came out of that erase, but eight of them were paid for with eight relocations — writes the host never issued and will never see. That ratio is write amplification, and it decides how long the drive lives.

The arithmetic is one line. If the collected block is a fraction u live, erasing it yields only 1 − u free pages, so one host page costs 1 ÷ (1 − u) physical writes. Slide the live count and read the cost off the bar:

1 of 16 pages still live — 1.07×

At a half-live block the cost is 2×; at fifteen of sixteen it is 16×. The collector's whole job is therefore to pick the emptiest block it can find — and its whole problem is that a drive with no spare room has none to pick.

Which is what over-provisioning buys: flash the host is never told about, so there is always somewhere to put a relocation and always a nearly-dead block to collect. Drag the reserve up and watch the amplification fall:

7% held back from the host. drag the marker along the curve; the arrow keys move it one step and Home puts it back
7% held back from the host — 6.57×

Collecting greedily beats collecting blindly, so the solid curve sits well under the dashed one — picking the emptiest block instead of any block takes over-provisioning from 15.3× down to 6.6×, and takes it to 1.6×. The curve is simulated, not fitted — uniform random 4 KB overwrites, greedy victim choice, run to steady state.

The drive can only be greedy about pages it knows are dead, and a filesystem deleting a file changes nothing on the device unless it says so. Turn the discard command off and watch the collector keep copying data nobody wants:

16.00×

Without TRIM every deleted page still looks live, so the collector relocates it forever. That is why fstrim runs weekly, why lsblk --discard is worth checking on a new host, and why a thin volume that swallows discards is a performance bug you will not find in your own code.

The collector has a twin on the write path. A TLC array can be programmed one bit per cell — fast, and needing far less margin — so a drive absorbs a burst into a region in that mode and folds it down later. Write past the end of it:

20 GB written in one burst — 6.9 GB/s

Watch the cliff at 114 GB, the Samsung 990 PRO's dynamic cache: 6.9 GB/s while the burst fits, 2.9 GB/s averaged over and 1.9 over . Every review that copies a 30 GB file measures the left of that curve; every backup job measures the right.

All of it lands on the endurance budget. The cells can absorb about 3,000 TB of NAND writes on a 1 TB TLC drive; what the host gets is that divided by the amplification. Raise it and watch the host's share shrink:

write amplification 1× — 16.4 years

Now look at where the dashed mark sits. Samsung warrants that part for 600 TBW, and 3,000 ÷ 600 is 5 — the rating already assumes an amplification of about five. Run with 500 GB of writes a day and the cells are done in 2.5 years, not the the budget suggests.

The second bill arrives while the drive is still healthy. Relocations are real device writes, so the collector competes with the application for the same channels. Push the write rate up:

6 MB/s of host writes — 36 µs

Watch the latency curve, not the bar. Reads that cost 35 µs on an idle drive cost milliseconds once the collector saturates the channels — exactly when the application is busiest. A p99 that is fine in staging and terrible in production is usually this.

04

The wire, and how many things you ask for at once

A million-IOPS flash array behind a SATA cable is a hundred-thousand-IOPS drive. And a million-IOPS drive asked for one thing at a time is a twenty-eight-thousand-IOPS drive.

Flash is parallel: several dies per channel, each able to work on a different request. Whether you see that parallelism depends on what the wire can carry and how many requests the host has outstanding.

The wire first, because it is the simpler ceiling. What the interface can carry is fixed by the link, and everything above it is flash the host cannot reach. Switch the interface:

600 MB/s

Notice that SATA is not slightly slow, it is a different era: 600 MB/s against PCIe 4.0 x4's 7.9 GB/s, thirteen times narrower, with an array behind it that could saturate the wider one. Everything above the ceiling is flash the host cannot reach.

The wire is not the interesting limit though. AHCI, the controller SATA uses, was specified for a device that could physically do one thing at a time: one command queue, thirty-two entries deep. Offer more commands than that:

1 commands offered

There is nowhere to put command thirty-three, so it waits in the host. NVMe's answer is not a bigger queue but many of them — up to 65,535 submission queues of 65,535 entries, in practice one pair per CPU core, so two cores submitting I/O never touch the same cache line.

A queue only helps if there is something to keep busy, and there is: the rating divided by the service time says at least thirty-five operations must be in progress at once. Raise the depth and count the dies working:

1 / 40 working

Because a die is busy for the whole of its 35 µs, one outstanding request leaves thirty-nine of forty idle. That is the entire mechanism behind the next figure — the ceiling is not the wire and not the controller, it is how many pieces of silicon you have managed to wake up.

Now the part that surprises people. A drive rated for a million IOPS delivers that only if a million requests a second are actually in flight. Raise the number kept outstanding and watch what arrives and what it costs:

queue depth 1. drag the marker along the curve; the arrow keys move it one step and Home puts it back
queue depth 1 — 28,571 IOPS, 35 µs

Watch the knee, and where it is. One 4 KB read takes about 35 µs, so yields 28,571 IOPS — 2.9% of the rating. The curve saturates at exactly depth 35, and past it the extra queue shows up as latency and nothing else: at the drive still does a million IOPS, at 128 µs each.

That number is not folklore, it is Little's Law. The requests in flight are the throughput multiplied by the time each one takes, so the depth you need is an area with IOPS and latency for sides. Move either:

required depth 35

The dashed rectangle is the datasheet — a million IOPS at 35 µs — and its area is 35. That is where "you need QD 32" comes from, and it tells you the shape of the fix: a benchmark reporting IOPS without a latency beside it has reported one side of a rectangle.

Keeping thirty-five requests in flight is the host's job, and a blocking pread cannot do it: it submits one, waits, and crosses the kernel boundary twice per 4 KB. Step through the path, then switch to the shared ring:

the application asks for 4 KB

A syscall round trip costs about 1.5 µs on post-Meltdown x86-64 with page table isolation on, against 0.1 µs before it. At a million IOPS that is 1.5 CPU cores burned on mode switches before any read has touched a byte. io_uring shares its rings with the kernel: one syscall a batch, or none.

05

Surviving the drive that dies

Every redundancy scheme answers one question — survive N device failures without losing bytes — and charges for it in capacity, in write cost, or in the length of the window after a failure.

Start with the layout: a stripe is one row of blocks across a group of drives, and what varies is how many hold data and how many hold something recomputable.

Read the usable share off the bar and the failures survived off the right. Switch levels, and change the group width under each:

usable 6/6 — survives 0 drive losses

Notice that the tolerance shown is the guaranteed one, not the lucky one. Three mirrored pairs survive one loss always and a second only when it misses the first one's partner — a runbook that plans for the lucky number has not planned. RAID 0 is the honest extreme: every drive of capacity, and the group dies with the first one.

The write cost is less visible and bites more often. Parity is a function of the whole stripe, so changing one block means reading the old block and the old parity before either can be replaced. Step through one small write:

the host writes 4 KB

Four device I/Os for one host write on RAID 5, six on RAID 6, two on RAID 10 — which is why OLTP databases sit on mirrors and archives sit on parity. Switch to RAID 10 and watch the two reads disappear: with a mirror there is nothing to recompute, so nothing to read first.

Then there is the window. When a drive dies, the array reconstructs it by reading every surviving drive end to end, throttled so the array still serves traffic. Drag the drive capacity up and read the rebuild time:

4 TB per drive. drag the marker along the curve; the arrow keys move it one step and Home puts it back
4 TB per drive — 7.4 hours

Rebuild time is capacity divided by a throughput that has not grown with capacity, so an drive at 150 MB/s takes 33 hours — and for all 33 the survivors run at 100%, exactly when a marginal drive is most likely to join the first one. On RAID 5 a second failure in that window is total loss.

And a second failure is not the only way to lose. Drives are specified to return one unrecoverable read per 1014 bits, which is one per 12.5 TB — and a rebuild reads far more than that. Set the group and read the odds:

72% · 12%

Even four 4 TB drives are a 72% chance at the consumer rate. — an ordinary RAID 5 shelf — is 99.996%: the rebuild is expected to fail. Enterprise drives are specified ten times better and still lose the coin flip at 63.5%.

A worse version of the same failure returns no error at all. If a drive hands back wrong bytes and reports success, parity does not help — the array has no idea which copy is wrong. Scrub, and switch what the filesystem checks:

served to the application

Watch what changes and what does not: the corruptions happen either way. ZFS and btrfs store a block's checksum in its parent, not beside it, so a scrub detects the mismatch and rebuilds from redundancy. ext4 over MD RAID has nothing to compare against and serves the bad block, because as far as the drive was concerned the read succeeded.

At rack scale the same arithmetic gets a cheaper answer: splitting an object into k data shards and m parity shards survives any m losses for (k+m)/k of the storage. Move either count:

1.18× — survives 3 lost shards

Backblaze's published Vault layout is 17 + 3: three simultaneous losses survived for 1.18× the storage, against 3× for triple replication and only two. The price is CPU on rebuild and a read that touches k machines.

06

What "written" means

A successful write() promises that the kernel has your bytes. It promises nothing whatsoever about the media, and the gap between the two is where committed transactions go missing.

A byte on its way to the media passes through three places that lose their contents when the power goes — the application's buffer, the page cache, and the drive's own DRAM — each a different amount of data at risk.

Walk the write along its path — drag it, or step it — and read what a power cut at that point costs against what survives:

app buffer. drag along the path to move the power cut; the arrow keys move it one stage and Home puts it back
app buffer

Notice that only the last stop is durable. In the page cache Linux writes back on its own schedule — vm.dirty_expire_centisecs is 3000, so a dirty page may sit 30 seconds — and in the drive's cache the data is one capacitor away from gone. fsync walks the byte to the end of the path and does not return until it is there.

Ordering matters as much as timing, and it breaks first where two writes depend on each other. A parity array updates a block and the parity covering it separately. Cut the power between them, then switch to copy-on-write:

the stripe is consistent: parity covers the data

This is the write hole, and notice that nothing detects it: the stripe is self-consistent nonsense, so a later rebuild reconstructs bytes that were never written. ZFS never updates a stripe in place — it writes a whole new one and switches the pointer — so there is no window to be interrupted in.

The same hazard lives one layer up. A journal, a WAL and an LSM manifest all work the same way: write the data, then write a record that points at it. Let the device reorder, and cut the power:

recovery finds both — the transaction replays

As soon as the commit record reaches the media before the data it points at, recovery replays a transaction whose payload was never written — silent corruption in a system that keeps a journal to prevent it. Ordering has to be asked for: a cache flush, or a FUA write, between the two.

And the trap: the correct version and the fast-and-wrong one look almost identical, and the fast one passes every test that keeps the power on.

write(data_fd, payload, n);
write(log_fd, commit_rec, m);    # either order may land

write(data_fd, payload, n);
fdatasync(data_fd);              # payload is on the media
write(log_fd, commit_rec, m);    # only now may the record exist
fdatasync(log_fd);

Two flushes, then — and each one costs whatever the hardware charges. Switch the drive's cache and batch transactions into one flush:

909 / s

On a consumer drive a flush pushes the cache to NAND: 1.1 ms, so 909 durable commits a second. An enterprise drive empties its own cache after the power fails, so it acknowledges in 60 µs — eighteen times faster, and 32 transactions a flush is 29,091.

Three questions are worth asking of any storage stack: how long is the window before this is durable, who may reorder it, and does the hardware acknowledge from volatile memory.

07

Quick reference

Three questions worth answering cold, and five red flags.

Why is dd if=/dev/zero the wrong write benchmark?

Because many controllers compress or deduplicate before writing, so a buffer of zeros never reaches the flash. Raise the share of the buffer that is zeros and watch the printed number leave the NAND behind:

0% of the buffer is zeros

Past 77% zeros only the PCIe link is still limiting the printed number: 14 GB/s reported against 3.2 that reached the flash. Use fio with --refill_buffers, or seed from /dev/urandom.

When is raw NVMe with O_DIRECT worth it over a filesystem?

When the working set has outgrown the page cache, so the cache is copying bytes for hit rates it will never achieve. Slide the working set past RAM and compare the two paths:

100 ns

Below 1× RAM the cache is unbeatable — a hit is 100 ns against 35 µs. Past it the curves converge and the double-caching is pure cost. That is why Postgres, ScyllaDB and Oracle ASM own their buffer pools, and why for everything else XFS on NVMe beats what you re-engineer in a quarter.

Why does a new drive benchmark faster than the one you deploy?

Because every block is erased, so the collector never runs and write amplification is exactly 1. Write the capacity through once and watch the number you would have reported collapse:

400,000 IOPS

, — the same drive, six and a half times apart, and only the collector in between.

  1. Planning capacity from --iodepth=1. That measures syscall overhead, not the drive.
  2. RAID 5 above about 4 TB. 33 hours of rebuild at 18 TB, and 99.996% odds of an unrecoverable read inside it.
  3. Skipping fsync because "it's an SSD." Media never determined durability — 30 seconds are at risk on any device.
  4. Assuming the discard reached the drive. A thin volume can swallow TRIM; the symptom is a p99 that grows over months.