Hardware & Tensors Primer
A tensor is a pointer, a shape and a stride, laid over memory that has exactly one dimension. Everything a GPU is fussy about — contiguity, coalescing, tiling, the shapes it wants your dimensions to be — follows from that one fact. Every number on this page is computed by the figure beside it, so you can drive it anywhere and it stays true.
A tensor is a pointer, a shape and a stride
Memory has one dimension. Every rectangle in deep learning is arithmetic laid over a line.
Ask for A[2, 1] and the machine does not go looking for row 2. There is no row 2. There is a pointer at the start of one buffer, a shape saying how you have agreed to read it, and a stride per axis saying how far to step. Drag the slider along the buffer and watch which cell of the rectangle it turns out to be:
Notice that the walk is a straight line. The buffer stays in slot order the whole way; it is the rectangle that folds, four slots to a row. Row 1 starts at for exactly one reason: there are four columns in front of it.
That “four in front of it” is the whole index formula. The row stride is 4 because a row is four elements long; the column stride is 1 because columns are neighbours; and the offset is what you get when you multiply and add. Set i and j and read the sum off the line:
Watch the teal bracket measure i × 4 and the amber one add j. That is all of offset = i·s₀ + j·s₁, and it generalises: a rank-k tensor has k strides and the offset is the dot product of the index with them. PyTorch stores literally this — storage(), stride(), storage_offset() — and every operation below edits one of the three.
Nothing forces the strides to be (4, 1). NumPy and PyTorch default to that; Fortran, MATLAB and every BLAS descendant default to (1, 3). Switch between them and watch which four slots the view's first row lands on:
Because row-major puts columns beside each other, a row is one contiguous run of the buffer and a column is a strided one; column-major swaps that. It is not a preference. It is why torch.matmul on a transposed operand sometimes copies first, and why cuBLAS, which is column-major, is usually handed B transposed on purpose.
One thing is still missing from the address: how wide a slot is. Strides count elements, the hardware counts bytes, and the dtype is the exchange rate. Step through the three widths a 2026 model actually ships in and watch the same twelve elements shrink:
Notice the buffer shortens while the shape does not move: it is (3, 4) on all three rows and only element_size() changed. Multiply by a real parameter count and it stops being a detail — 70 billion weights are 260.8 GiB in fp32, 130.4 in bf16 and , against the 80 GB an H100 carries.
A view is stride arithmetic
Transpose, slice, broadcast. None of them touch a byte — until one of them has to, and the copy arrives with no warning.
Once the address is a formula, most of what you write edits the formula rather than the data. A transpose moves nothing at all: it swaps two entries in the shape and two in the stride tuple. Flip between A and Aᵀ and watch the storage hold perfectly still:
Notice the wires. Under A the view runs straight down in reading order; under Aᵀ they braid, because reading it in its order jumps four slots at a time. The bytes are identical and the walk is not, and that difference is the whole performance story of the rest of this page. A.t() is O(1) whatever the size of A.
A slice is the same trick with two more knobs. Keeping every step-th column from a start column adds start × s₁ to the offset and multiplies s₁ by step. Move the two sliders and watch the kept columns land on the line:
Because both edits are arithmetic, A[:, 1::2] costs nothing to produce — and it shares the storage with A, so writing into the slice changes A. That aliasing is the price of the free view, and a whole class of bugs lives there: a slice you thought was a copy, and a loop quietly writing through it.
The third edit is the strange one. A stride of zero makes one stored row answer for as many rows as you like, because every index maps to the same address. Raise the row count and watch the storage refuse to grow:
Watch the wires converge. Broadcasting a bias across 4,096 rows costs four slots, not 16,384 — which is why x + b on a (4096, 768) activation allocates nothing for b. The stride is 0, the memory is one row, and the cache makes the 4,096 rereads nearly free.
So when does a view stop being free? Exactly when the walk it forces on the buffer is no longer +1 at every step. That is the definition of contiguous, and it is the condition .view() checks before agreeing to reinterpret a shape. Widen the step until the walk breaks:
At step = 1 the walk is +1 twelve times and .view(-1) is free. At it has gaps, .view() raises, and .reshape() silently allocates and copies instead. That silence is the trap: reshape is the forgiving one, so it is the one people reach for, and it turns an O(1) call into an O(n) allocation inside a loop that runs a million times. Write x.contiguous().view(-1) and the copy is where you put it.
What a walk actually costs
The memory system does not sell you the bytes you asked for. It sells you 32-byte sectors, and it has no smaller size.
A strided walk has cost nothing but a multiply so far. The hardware disagrees, because no memory system on earth delivers four bytes. An NVIDIA GPU services a global-memory request in 32-byte sectors. Walk the slider along the buffer, ask for one fp32, and look at what actually arrives:
Notice the seven elements that came with it. They crossed the memory bus, landed in L2 and were thrown away. On its own that is a curiosity — if the next thing you want is element 1, the sector is already here. The cost only shows up when nothing you ask for is next to anything else.
Which brings in the other half of the machine. A GPU does not issue one load; it issues one per warp — 32 lanes, one instruction, 32 addresses. Consecutive addresses cost the warp four sectors. Strided ones can cost it 32. Open the stride and watch the sectors multiply:
Watch the lanes spread and the sectors fill in behind them. At stride 1 the warp wants 128 bytes and the hardware moves 128. At it still wants 128 and the hardware moves 1,024 — every lane in a sector of its own, seven eighths of the traffic discarded. This is coalescing, and it is the largest single lever on any memory-bound kernel.
The relationship is exact and worth seeing as a curve rather than as two poses. Efficiency is 1/stride until every lane owns a sector, after which there is nothing left to lose — drag the same stride slider and read it off the curve:
Because the fall in efficiency is hyperbolic, the first doubling is the expensive one: stride 1 to 2 costs half the bandwidth, while 4 to 8 costs another eighth. On an H100 that is 3.35 TB/s of delivered bandwidth against of useful bandwidth, on identical arithmetic. Nothing about the kernel changed — only the order in which it asked.
So the same sum, written two ways, is two different programs. Summing a row-major matrix along its rows reads consecutive addresses; summing it down its columns reads one element per sector. Switch the direction and watch the sectors light up:
Because a column of this matrix is eight elements 32 bytes apart, reading it touches all eight sectors to use 32 of the 256 bytes they carry. The row version touches one sector per eight elements. Same loop, same FLOPs, eight times the traffic — and at 4096 × 4096 the column version misses in L2 as well.
Matmul is a memory schedule
The arithmetic is nearly free. Getting A's rows and B's columns to it is the entire engineering problem.
A matrix multiply is the most compute-dense thing a Transformer does, and written naively it is still memory-bound. Every output element is a dot product of one row of A and one column of B. Drag over the output and watch what one element has to read:
Notice that the next output along needs the same row of A again. Written with no reuse, an n × n matmul does 2n³ FLOPs and reads 2n³ operands — at 2 bytes each, that is 0.5 FLOP per byte. The H100's arithmetic units need 295 to stay busy.
That ratio is arithmetic intensity: FLOPs divided by bytes moved from DRAM. Reuse is what raises it. If a thread block holds a T × T block of the output and streams the matching stripes of A and B through it, every operand is read n/T times instead of n. Raise the tile edge and read the intensity off the curve:
Both axes are logarithmic and say so, so a straight line here is a power law rather than a constant slope. Intensity is T/2 at bf16 — each doubling of the edge halves the traffic. Reaching the ridge, where arithmetic finally becomes the limit, would take T = 591, far more than one SM can hold. We come back to why 128 gets away with it.
The ridge is a property of the machine, not of the kernel: divide peak arithmetic by peak bandwidth and you get the intensity below which memory is the ceiling and above which the tensor cores are. Slide the intensity and watch which ceiling you are under:
Below 295 FLOP per byte you are on the sloped part of the roof, and throughput is bandwidth × intensity however many tensor cores the die has. A naive matmul at tops out at 1.7 TFLOP/s on hardware rated for 989 (H100 SXM5, NVIDIA datasheet, 2023). Softmax, LayerNorm and every element-wise op live down there permanently.
Draw the output as blocks and the reuse becomes obvious. Each block reads one horizontal stripe of A and one vertical stripe of B, so the whole matmul reads each operand n/T times. Change the edge and watch the DRAM traffic:
Watch the traffic halve with every doubling: 2.0 GiB at T = 128, 1.0 GiB at 256, 512 MiB at 512 — against the 96 MiB the three matrices actually contain. The gap is what L2 absorbs. An H100 has 50 MB of it, which is why a real 128-tile kernel measures far closer to the ideal than this model predicts.
The tile cannot simply keep growing, and the wall is shared memory. Two T × T operand tiles at 2 bytes each, double-buffered so the next load overlaps the current math, is 8T² bytes against the 228 KiB an SM can allocate. Push the edge past the wall:
At T = 128 a tile pair is 128 KiB and fits with room for the accumulators. At it is 288 KiB and the kernel does not launch at all — an error at launch time, not a slow kernel. The largest square tile that fits is 168, and cuBLAS picks 128 because a power of two divides the tensor-core MMA shape and leaves registers for the accumulator.
Why the GPU wants the shapes it wants
Any dimension that is not a multiple of something is arithmetic you pay for and throw away.
The tile is 128 wide and the hardware cannot compute part of one. An output 512 columns wide is covered exactly; 520 columns means the launch computes 640 and discards 120. Widen the output and watch the padding appear:
Notice that the waste is not proportional to how far past the boundary you went: it is 94% of the last tile at and 6% at N = 632. This is tile quantisation, and it is why production Transformer dimensions are multiples of 64 or 128 — GPT-2's 50,257-token vocabulary padded to 50,304, head dimensions of 64, FFN widths four times the model width.
There is a second quantisation above that one, and it is coarser. Tiles are handed to SMs, and an H100 has 132 of them. 132 tiles is one full wave; 133 is two waves, the second occupying one SM while 131 sit idle. Walk the tile count across a wave boundary:
Watch the cliff. Occupancy falls from 100% to for one extra tile of work, and does not recover until 264. This is wave quantisation, and it is why a shape you benchmark matters as much as the kernel you benchmark. It shrinks with the wave count — at eight waves the same extra tile costs 12% rather than half.
The last shape problem comes from the data rather than the kernel. Sequences have different lengths and a tensor is a rectangle, so a batch is padded to its longest member and a mask throws the extra away afterwards. Cut the batch into length buckets and watch the wasted arithmetic fall:
With one bucket, 55% of this batch's FLOPs compute tokens nobody keeps. Sorting by length and brings that to 12% without touching a kernel. That is length-grouped batching, and it is also why variable-length attention drops the rectangle entirely for a cu_seqlens offset array — a ragged tensor, and one more thing that is really just a stride.
What to hold on to
One invariant, and three constants.
The element at (i, j) is the byte at base + (offset + i·s₀ + j·s₁) × itemsize, and every operation that does not copy edits those four numbers. The constants: the 32-byte sector, the 128-wide tile, the 132 SMs. Coalesce the walk, fill the tile, fill the wave.