Data for Training

Three claims, proved with twenty-five figures you can drive. That a dataset's errors are four different things with four different budgets, and only one is fixed by collecting more. That every way of leaking the test set into training fails silently and upward, which is why it survives review. And that weighting, resampling and thresholding are one lever with three names.

01

The pile you actually have

A model knows what its rows knew. The six sections after this one are repair work on what this one got wrong.

A dataset is a sample: somewhere there is a population you care about, and what you have is whatever a script reached on a Tuesday afternoon. Four separate things go wrong with that, and only the first is fixed by collecting more — drag the slider to grow the rows you kept and watch their mean walk toward the population's:

6 rows collected, sample mean 45.8 minutes

Notice how fast it settles. Six rows put the estimate at 45.8 minutes against a true 39.2; sixty put it at 38.1, and the last hundred and eighty barely move it. Sampling noise shrinks like one over the root of n — it is the error you can buy your way out of.

Selection bias does not shrink, because reachability is rarely independent of what you are measuring. Drag the wall to shrink the frame and watch the mean it reports detach from the real one:

the frame reaches 240 of 240 users. drag to move the edge of the frame; the arrow keys move it one column at a time and Home restores it
the frame reaches 240 users, and its mean is off by 0.0 minutes

Because heavy users are what the frame drops first, halving it to 120 users moves the estimate from 39.2 minutes to 28.2. Ten times as many rows inside that frame buys a more precise number that is just as wrong. This is the one error more data makes worse, by making it look trustworthy.

Coverage is the third. Production traffic has a long tail, and a sample holds only the classes it hit — drag from eight rows up to and watch the eighth class sit at zero:

8 rows drawn, the rarest class appears 0 times

Eight rows in, five of the eight classes have no example at all. The rarest is 0.72% of the traffic, so 128 rows miss it 40% of the time, and 1,024 rows give it ten examples against the first class's 709. Ten examples is a class your model gets wrong in production and is never marked down for.

The fourth is not about which rows you have. Two annotators shown the same row disagree — drag the rate to and watch the labels that are simply wrong cap every score anyone will ever report:

0 labels wrong, so no model measures above 100.0%

Flip 26 of the 240 and a predictor right about every true label still scores 89.2% against the recorded ones. Northcutt and colleagues found errors in at least 3.3% of the test labels across ten of the most-cited benchmarks in the field. Four errors, four budgets: rows, a better frame, targeted collection, better annotation. Only the first is a volume problem, and it is the only one anybody funds.

02

Cleaning is a set of decisions

Rows arrive broken in four ways, and every repair changes a number you are going to quote.

None of this is tidying: each of the four is a choice about what the data means, made once by whoever wrote the loader and inherited by every experiment after. Start with the crawler that does not know it has seen a page before — drag the slider and watch the copies outnumber the documents that are new:

24 documents scraped, 17 of them unique

By only 98 are new: 142 are copies, and the worst offender appears 23 times. Each copy multiplies that document's pull on the gradient — and, as §04 shows, a copy that lands on both sides of the split is the cheapest way there is to fake a good score.

Next, the holes. Twenty rows, four columns, and the highest earners are the ones who decline to state an income — switch the control to see what each repair does to the column with holes in it:

the repaired column reports a mean of 34.9

Notice that the mean never moves: the true mean is 42.0 and all three repairs report 34.9, because the missing rows are the high ones. What moves is the spread. Filling with the mean keeps all twenty rows and reports a standard deviation of 13.3 where the truth is 20.7 — six rows of agreement nobody observed.

Outliers are the opposite problem: too influential rather than absent. A least-squares fit weights a row by how far out it sits — drag the ringed row anywhere in the plot and watch the line follow it:

drag the ringed row anywhere in the plot; the arrow keys move it and Home puts it back on the line
the fitted slope is 0.66

At rest that row is on the line and the fit hardly notices: slope 0.66 with it, 0.67 without. Drag it to the bottom-right corner and the slope falls to 0.42 — one row in twenty-three moving the whole model, because its leverage out there is 0.19 against an average row's 0.087.

The fourth survives every other check. Four scraped strings, meant to be one word — step through trim, casefold and Unicode NFC and watch four distinct byte sequences collapse into one:

after normalisation 4 distinct strings remain

As scraped they are 6, 5, 6 and 5 bytes, and the third spells é as an e plus a combining acute — three bytes where the composed form is two. Every exact-match de-duplicator you ran before this step counted four documents where there is one. Order matters as much as the steps: normalise before you de-duplicate, de-duplicate before you split, and decide what a missing value means before you fill it.

03

Cutting it in three

Three slices, three jobs — and a reported number only as sharp as the smallest of them.

Training rows are what the optimiser sees; validation rows are what you choose between models with; test rows are opened once, at the end. Three and not two, because the moment a validation result changes a hyperparameter that slice has begun leaking. Drag the share held out and watch what the slice that measures you costs the slice that trains you:

1,000 test rows, which measures the score to plus or minus 1.86 points

Watch the right-hand number as you drag. A 10% test slice — 1,000 rows — measures a 90% model to ±1.86 points at 95% confidence. and the interval widens to ±2.63. A test set is not spare capacity; it is the resolution of your own scoreboard.

That interval is not a formality. Here is one model, unchanged, scored on twelve independently drawn test sets — drag the size and watch the twelve measurements collapse toward the accuracy it actually has:

the twelve runs land between 83.1% and 96.7%

As soon as the test set is small the measurement stops meaning anything: at fifty rows the same model reads anywhere from 83.1% to 96.7%, so two teams reporting 84% and 96% off that test set have shipped identical models. Push it to and the spread is 88.9% to 91.1% — still wide enough to swallow most of the improvements in most papers.

Size is not all a slice can get wrong. Twenty positives in a thousand rows, a hundred of them drawn for test — drag through forty draws and watch how many positives the random slice caught wander:

this random slice holds 3 positives where the stratified one holds 2

Across the forty it holds anywhere from 0 to 6 of the twenty, and none at all 12% of the time — a test set that cannot score the only class anyone cares about. The stratified slice holds two every time, because it is built class by class rather than row by row.

And rows are not always independent. Forty patients with six scans each: a model can recognise this patient rather than the disease — switch from splitting by row to splitting by patient and watch the groups straddling the wall disappear:

31 patients have rows on both sides of the wall

Since six rows belong to each patient, splitting by row leaves 31 of the 40 with scans on both sides, so all 60 test rows have a sibling the model already fitted; splitting by patient takes both counts to zero and the reported score with them. That drop is the finding. When 10% of a small dataset is too few rows to measure anything, k-fold cross-validation buys resolution with compute: fit k times, hold out a different k-th, report the mean and the spread.

04

Four ways to show the model its own exam

Leakage never throws. It makes your metrics better, which is exactly why it survives review.

Each failure here is one line of code in the wrong place, and each is invisible to the tests you already run. Start with the one every notebook has — a scaler fitted, then the data split. Drag the test share and compare the mean fitted on every row with the mean fitted on training rows only:

fitted on everything the mean is 51.5, fitted on training rows only it is 48.8

Notice that the two rules do not sit together: 51.5 fitted on everything against 48.8 fitted on training rows. So the test rows stand 1.14 standard deviations above the first and 1.64 above the second, and the scaler that saw them has flattened the drift the model is about to walk into.

The second shape is a column that is the answer wearing a hat: refund_issued in a churn model, discharge_ward in a mortality model — drag how much of the label the column carries and watch the two classes come apart:

AUC 0.756 in validation against 0.756 in production

At validation AUC is 0.967 while the honest features hold 0.756 — and 0.756 is what production gets, because at prediction time that column has not been written. Ask when each column is populated relative to the moment you must predict, and delete anything that answers “after”.

The third is §02's duplicates arriving late: de-duplicate after splitting and copies land on both sides — drag the test rows the model has already read and watch the score you would report lift off the score you would get:

the run reports 81.0% where the honest score is 81.0%

Because a model gets a memorised row right every time, eighteen contaminated rows in sixty turn an honest 81.0% into a reported 86.7%, and half the test set turns it into 90.5%. Nothing in the run looks wrong. This is exactly benchmark contamination in language models: the training corpus contains the benchmark, and the benchmark stops measuring anything.

The fourth is time. If rows are dated, a random split trains on the future and tests on the past — drag the wall along the timeline and compare a split by row with a split by date:

the wall sits at day 180 of 240. drag the wall along the timeline; the arrow keys move it one day at a time and Home restores it
0 training rows are dated after the first test row

Split at random and 176 of the 180 training rows are dated after the first test row: the model is handed next week to predict last week with. Splitting by date takes that to zero. All four share one diagnostic — a validation score much better than you expected, from a change you did not think was that good. Treat it as a bug report, not a result.

05

When the thing you want is rare

Fraud, disease, outages, churn. The interesting class is always the small one, and every default is tuned for the other.

Imbalance is a property of the world, not of the model, and it breaks the metric, the loss and the threshold in three different ways. Two hundred and forty transactions came in overnight — drag the ones that are actually fraud down toward one in a batch and watch what a model that answers “legitimate” to everything scores:

12 cases in 240, and a model that always answers no scores 95.0%

Watch the do-nothing model. Twelve in 240 is 5%, and answering “legitimate” to all of them scores 95.0% while finding nothing at all. Take it to and it scores 99.6%. Accuracy on an imbalanced problem is the majority class's share wearing the costume of a measurement.

So the model gives a score and someone draws a line. Here are the scores fraud gets and the scores everything else gets — drag the line and watch the fraud you catch trade against the alerts that are nothing:

recall 50% · precision 54%. drag the line across the scores; the arrow keys move it and Home restores it
recall 50% at precision 54%

At rest the line catches half the fraud and 54% of what it flags is real. Drag left until recall reaches 84% and precision falls to 22% — one real alert in every 4.6, at 5% prevalence. Which end you want is a costing question: a missed fraud against an hour of an analyst's time.

The two curves people report disagree about how much that matters. One classifier, the same two distributions, drawn as ROC — recall against false-positive rate — and as precision against recall: drag the prevalence and watch only one of them move:

AUROC stays 0.921 while average precision is 0.921

From balanced down to , the ROC curve does not move: AUROC stays 0.921, because it is a property of the two distributions and nothing else. Average precision drops from 0.921 to 0.181. A strong AUROC on a rare-event problem is a number nobody has priced.

The standard fix is to weight the rare class in the loss. Drag the weight and watch where the decision boundary actually goes:

a weight of 1 is the same decision as a threshold of 0.500

For a calibrated score, weighting positives by w gives exactly the decision the unweighted model makes at a threshold of 1/(1+w): is threshold 0.077, recall 78% at 27% precision. Good operating point, no new information — and the probabilities are decalibrated, so everything downstream of them is wrong. Resampling to a ratio r is the same move plus duplicated rows. Weighting, resampling and thresholding are one lever with three names, and none of them adds positives.

06

Rows you did not collect

Augmentation is the only way to add rows without going outside. It rests on a claim about invariance, and it fails where that claim does.

The claim: there is a set of transforms under which the label does not change. If a six stays a six when you shift it, every shifted six is a legal training row and you can manufacture them. Press play, or drag the track, to walk the copies a pipeline makes from the example as collected — each one a new training row:

as collected: overlap 1.00 with the original

Notice how uneven the cost is. Shifting one cell in each direction changes 30 of the 121 cells and drops the overlap with the original to 0.52, while rotating twelve degrees changes 12 and holds 0.80. All five reach the optimiser labelled six — the invariance claim, stated as data.

Every such claim has an edge, and this one is easy to walk over. Drag the rotation past a quarter turn and watch how much of a 6 the copy still is cross how much of a 9 it has become:

rotated 0 degrees; the label still holds

The two numbers converge and swap at 77°, and by the copy is exactly a nine, still carrying the label six. In between is a band where it is neither digit. Past that crossing every copy is a training row with a wrong answer attached — §01's label noise, self-inflicted and generated at loader speed.

Which leaves the question nobody asks aloud: what is a copy worth? Error falls as a power of the row count, so a copy is worth however far it moves your effective one — set the number of copies, then set what one is worth against a real row:

the copies are worth 1.00 times the data, taking error to 3.9%

The horizontal axis is logarithmic, so a fixed multiplier is a fixed distance right at any row count — the reason this trade works, and the reason it runs out. at a quarter of a real row each is 6.75× the data, taking error from 3.9% to 2.2%; set the worth to 0.05 and the same copies are 2.15× and 3.1%. Nobody can hand you that second number: it is how far the transform moves an example against how far the task's own variation does. ImageNet-1k has 1,281,167 real training images, and crop-and-flip is worth a large multiple of that and still not a second ImageNet.

07

Quick reference

The whole page is one ordering question, and here is the order.

Eight steps run between a raw file and a trained model, and exactly one of them is a wall: above it everything sees every row, below it everything sees training rows only — drag the split up and down the ladder and watch the steps that go wrong change:

1 step below the split belongs above it

has nothing red in it. Above the wall belong only the two steps that must see everything to be correct — loading, and dropping duplicates, because a duplicate removed after the split already straddles it.Everything below is fitted.

In code, the commonest leak in the field is one statement in the wrong order, caught by no test and no type:

# wrong - the scaler saw every row
X = scaler.fit_transform(X)
Xtr, Xte = split(X)

# right - the split comes first
Xtr, Xte = split(X)
Xtr = scaler.fit_transform(Xtr)
Xte = scaler.transform(Xte)

One last budget, because §03's interval implies a question nobody asks — drag the gain you want to see and read off the test rows it takes against the thousand you have:

seeing a gain of 1.0 points needs 13,493 test rows per side

A one-point gain at 80% power needs 13,493 test rows a side — thirteen times the slice §03 cut. Every failure here breaks one invariant, quietly and upward: at scoring time, nothing that touched the test set has moved a parameter or a preprocessing constant.

  • The score jumped and you cannot say why. Work §04's four shapes first.
  • One feature nearly as good as the model. Ask when that column is written.
  • Test accuracy above the label-noise ceiling. Nobody beats their annotators.
  • Accuracy quoted on an imbalanced problem. Ask for the base rate.
  • A test score quoted to one decimal. At 1,000 rows the interval is ±1.9 points.