📐 ML Fundamentals
Build a trustworthy train/validation workflow: splits, bias and variance, regularization, calibration and production diagnosis.
On this page
Your training score looks great. New examples still go wrong. Before changing the architecture, find out whether the model learned the wrong pattern, saw information it should not have, or met a different population after deployment.
Before you start: probability and statistics covers uncertainty and base rates. Linear algebra explains vectors and norms; calculus explains penalty gradients. Here we build the complete workflow: define the prediction, split the data, fit a baseline, diagnose errors, then select and evaluate a model.
First pass: follow sections 1–5 and their checks. The later sections develop regularization and probability-based decisions. The practice questions retain a few Bayes exercises as a bridge from the previous foundation lesson.
1. Define the prediction before fitting
A supervised dataset contains inputs x and targets y. A model f_θ(x), with learned parameters θ, predicts y from x. Training minimizes a loss over observed examples. Generalization concerns performance on new examples from the population you care about.
Suppose you predict cancellation within 30 days. Define a prediction time t, features available by t, and a target for (t,t+30 days]. A cancellation-confirmation email is an excellent predictor and an invalid feature: it arrives after the decision. A historical table corrected later can leak information even if its event timestamp looks old.
Start with a baseline: the majority class, a seasonal forecast, or logistic regression. Compare with the actual alternative decision process. The fitting loss and product metric need not be identical.
2. Split first, then learn transformations
Training data fit parameters. Validation data select features, settings and thresholds. A test set assesses the selected procedure. Repeatedly choosing changes from the test score turns it into more validation data.
For a churn example, train on January–June prediction dates, validate on August and test on October. The gaps allow 30-day labels to mature. Fit scalers, imputers, vocabularies, PCA and feature selection on the relevant training partition only, then transform held-out data with those fitted objects.
Match the split to deployment. For future behavior of existing users, preserve time and feature availability. For unseen users, hold out users as groups. If both matter, enforce both. A timestamp column alone does not make every random split invalid; ask which deployment question the split answers.
Check: can you standardize everything before splitting because no labels are involved?
No. Held-out distribution statistics still influence the transformation. Fit on training data and repeat that fit inside each cross-validation fold. The same applies to unsupervised PCA and imputation.
3. Understand what bias and variance measure
For regression with squared-error loss, expected error at a fixed input is squared bias of the average fitted prediction, prediction variance across training sets, and conditional target noise.
Suppose the true conditional mean is 5. Models fitted to two equally likely training sets predict 4 and 8, averaging 6. Squared bias is (6−5)²=1. Prediction variance is [(4−6)²+(8−6)²]/2=4. With independent target noise of variance 2, expected squared error is 1+4+2=7.
This averages over possible training sets and a fresh target. It does not apply unchanged to accuracy or cross-entropy. Prediction variance is not simply the train–validation gap.
Check: what if every fitted model predicted 6?
Squared bias stays 1, prediction variance becomes 0, and noise stays 2, giving expected error 3. In real tasks we usually do not know the true conditional mean and noise variance exactly.
Ensembling and the Tradeoff
Bagging can reduce variance by averaging learners whose errors are imperfectly correlated; random forests combine bootstrap samples and feature subsampling to encourage diversity. Boosting sequentially adds learners to improve an objective and can affect both bias and variance. For losses beyond squared error, it fits pseudo-residuals rather than ordinary target residuals. Stacking learns from out-of-fold base predictions. None guarantees a reduction in exactly one error term.
The Modern Regime: Double Descent
The familiar U-shaped test-error curve is a useful pattern, not a consequence guaranteed by the bias–variance decomposition. In some settings, error rises near the interpolation threshold, where the model can fit the training data, then falls again as capacity increases. This is double descent. It can occur in linear models as well as neural networks; its shape depends on the data, model, optimization and regularization. A large model can generalize well, but adding parameters does not guarantee improvement.
4. Diagnose before prescribing a fix
The decomposition is a thought experiment over training sets. In practice, comparable train/validation metrics and learning curves provide evidence for a diagnosis, not direct measurements of every term.
Diagnosing via the Train-Validation Gap
The single most useful diagnostic is the gap between training and validation error. Underfitting shows up as both errors high with a small gap; overfitting as low training error but a large gap to validation. Two cases people miss:
- Both errors low: A promising fit, or an easy/leaky evaluation. Check independent data, relevant slices and decision costs before release.
- Validation error below training error: First compare like with like. Training may use dropout, harder augmented inputs or a regularization penalty that validation omits. Recompute the same data-loss metric in evaluation mode before suspecting leakage, an easier validation distribution or incorrect batch-normalization statistics.
Learning Curves
Plot training and validation error as a function of training-set size. The shapes tell you which problem you have:
- Possible high bias: Both curves plateau at a high error, close together. More of the same data may have diminishing returns; also check optimization, missing features and label noise before increasing capacity.
- Possible high variance: Training error is low and validation error is higher. More representative data often helps, but the curve is evidence to test that hypothesis, not a guarantee that the gap will close.
- Healthy fit: Validation error converges toward training error as data grows, both at a low value.
Learning curves help narrow the diagnosis. Use comparable losses and repeated splits where practical; one noisy curve cannot uniquely identify the cause.
Concrete Fix Patterns
For poor training fit, first audit labels/features and fit a tiny clean subset to check optimization. Then test reduced excessive regularization, better features, suitable extra capacity or another model family. Train longer only when the loss trajectory supports it.
For low training error but worse validation, audit the split and population before tuning regularization. Compare more representative data, simpler models, early stopping and task-valid augmentation. Extra duplicates do not create new information, and a stronger penalty can underfit.
Beyond Overfitting: Distribution Shift
A production drop can reflect distribution shift, but also leakage, missing features or a difference between training and serving code. Under covariate shift, P(X) changes while P(Y|X) stays the same. Under label shift, P(Y) changes while P(X|Y) stays the same. Concept drift changes P(Y|X). Real changes can mix these patterns. Compare feature pipelines, monitor input and prediction distributions, and measure performance on labelled production slices. Retraining or reweighting may help when the new data and assumptions support it; neither monitoring nor scheduled retraining is a cure on its own.
5. Evaluate the selection procedure
A held-out validation set gives you one estimate of generalization error, which can be high-variance for small datasets and tempting to overfit through repeated hyperparameter tuning. Cross-validation compares the fitting procedure across held-out partitions. Its usefulness still depends on representative splits and disciplined model selection.
k-Fold
Split the data into k roughly equal partitions. For each fold, train on the other k−1 folds and evaluate on the held-out fold, then average the scores. For example, five fits on 1,000 examples each use 800 training examples. Scores 0.74, 0.78, 0.76, 0.79 and 0.73 average 0.76. Common choices are k = 5 or k = 10. Larger k trains on more data per fit but costs more fits; it does not necessarily reduce the variance of the final estimate because the fitted models and errors are correlated. Each example appears in exactly one validation fold. Use a splitting scheme that matches the task's groups and time structure.
Stratified k-Fold
Standard k-fold splits randomly, which can produce class-imbalanced folds when the dataset itself is imbalanced, fatal for small minority classes (a fold might contain zero positive fraud examples). Stratified k-fold preserves the class proportions in every fold by sampling within each class. Use it when the task permits exchangeable examples; group and temporal constraints take priority. Check that minority counts support the chosen number of folds.
Time-Series CV (Walk-Forward / Rolling Window)
When deployment predicts future outcomes, random splitting can let later information influence earlier evaluation. Match the actual forecast origin and horizon. Use forward-chaining splits: train on [1..t], validate on [t+1..t+h], roll forward, repeat. Never include a validation index earlier than a training index. Two common variants: expanding window (training set grows with each split) and sliding window (fixed-size training set rolls forward). A sliding window excludes older regimes but loses data; an expanding window uses more history. Validate which tradeoff fits the task, including seasonal coverage and label delays.
Leave-One-Out (LOOCV)
The extreme case where k = n: train on n−1 examples, evaluate on the held-out one, repeat n times. Each fit uses almost the full dataset, so training-size bias is often small. The estimate can still be variable, especially for unstable learners; overlapping training sets do not give independent error estimates. It costs n fits unless the model has an efficient shortcut, such as linear regression's PRESS statistic. Use 5- or 10-fold as a practical starting point and choose based on the data, learner and compute budget.
Nested CV
Selecting hyperparameters and reporting the same CV score creates selection bias: the winning setting partly fits that evaluation's noise. Nested CV separates the tasks. For each outer training fold, run inner CV to choose settings and fit preprocessing; refit the selected model on that outer training fold and score it on the untouched outer validation fold. The outer scores estimate the whole selection procedure and reduce tuning optimism. They still have sampling uncertainty and a training-size difference from a final full-data fit. Cost grows with both fold counts and the number of candidate settings.
Group k-Fold
When examples are not independent (multiple records per patient, multiple frames per video, multiple search queries per user), random splitting puts correlated examples in both train and validation, leaking information and inflating CV scores. Group k-fold keeps each group entirely on one side of a fold. Choose the group that matches deployment: new patients, new videos or new users. Combine grouping with time constraints when both matter; optional stratification must respect the groups.
Common Pitfalls
- Tuning hyperparameters on the test set. The test set must be used once, at the end. If you tune on it, you have a second validation set, not a test set.
- Leakage via preprocessing. Standardize/normalize inside each fold using only the fold's training data, not across the whole dataset before splitting. Same for imputation, target encoding, and feature selection.
- Forgetting groups. If your data has natural groups (users, patients, time), random k-fold gives overly optimistic scores.
- Random splits on time series. Mentioned above; the most common production failure mode.
- Reading CV scores as a guarantee. They estimate generalization to the same distribution as training. Out-of-distribution behaviour requires separate monitoring and offline shift tests.
6. Regularization changes the preferred solution
Regularization constrains or biases learning to favor some solutions over others. It can improve generalization, especially with limited or noisy data, but too much can underfit. An expressive model may fit noise without it; that does not mean every unregularized model will fail. Weight penalties, early stopping and augmentation impose different preferences, so choose and validate them for the task.
For w=[3,−4], L1 norm is 7 and squared L2 norm is 25. At λ=0.1 their penalty contributions are 0.7 and 2.5. These are not directly comparable strengths: the gradients and units differ. Scaling an input from meters to millimeters changes its coefficient, so fit feature scaling on training data and check the penalty convention.
L1 vs L2 Penalties
The two canonical penalties add a term to the loss:
L_total = L_data + λ · R(w)
L2 (Ridge) uses
R(w) = ||w||_2² = Σ w_i². Its penalty gradient is2λw, producing smooth shrinkage. It does not usually create sparse solutions, though a coefficient can be exactly zero. The constrained region is a circle or sphere. With a negative log-likelihood data loss and compatible scaling, the penalty corresponds to MAP estimation under a zero-mean Gaussian prior.L1 (Lasso) uses
R(w) = ||w||_1 = Σ |w_i|. Away from zero, the penalty's derivative isλ · sign(w_i); at zero its subgradient is the interval [−λ, λ]. This kink allows optima with exactly zero coefficients. Solvers such as coordinate descent or proximal gradient exploit it; ordinary gradient steps merely crossing zero do not explain sparsity. The constrained region is a diamond with corners on the axes. With a negative log-likelihood loss, L1 corresponds to a Laplace prior.ElasticNet mixes both:
λ_1 ||w||_1 + λ_2 ||w||_2². Useful when you want sparsity (L1) but L1 alone is unstable under correlated features (it picks one and drops the others arbitrarily); L2 stabilizes the selection.
Check: a proximal L1 step has u=[0.03,−0.20] and threshold ηλ=0.05. What remains?
Soft-thresholding gives [0,−0.15]. A small coefficient becomes exactly zero. L2 instead shrinks proportionally. This explains sparsity through the optimization rule rather than a gradient merely crossing an axis.
With a negative log-likelihood data objective, a Gaussian prior with variance τ² adds Σw_i²/(2τ²) to the negative log posterior. This gives a MAP estimate, not a full posterior. Averaging instead of summing the data loss changes the corresponding penalty coefficient by sample size.
Weight Decay vs L2 Regularization
These are often conflated, but they differ subtly and importantly. Weight decay is the update rule w ← (1 - η·λ) · w - η · ∇L_data. L2 regularization adds λ ||w||_2² to the loss, giving gradient update w ← w - η · (∇L_data + 2λw). For vanilla SGD, the two are mathematically equivalent (with a factor-of-2 reparameterization of λ).
For adaptive optimizers like Adam, they differ. A coupled L2 term enters both moment estimates, making its effect depend on gradient history. In a simplified view with the second moment v held fixed, scaling by 1/(√v + ε) makes the same penalty contribution larger when v is smaller, not smaller. The full update also includes momentum and the penalty's effect on v. AdamW (Loshchilov & Hutter, 2019) applies a separate proportional shrinkage to the current weights and computes Adam's adaptive step from the data gradient. It decouples decay from that adaptive scaling; it does not make L2-regularized Adam mathematically invalid.
Dropout
Dropout (Srivastava et al., 2014) randomly zeros activations during training with probability p (typically 0.1-0.5). It can be viewed as training an exponential ensemble of subnetworks that share weights, and as injecting multiplicative Bernoulli noise that prevents co-adaptation of features.
Inverted dropout is the standard implementation: scale activations by 1/(1-p) at training time so that no scaling is needed at inference. This keeps the inference path identical to a model without dropout, simplifying deployment.
For a=4 and p=0.25, inverted dropout outputs 0 with probability 0.25 and 5.333… with probability 0.75, so its expectation is 4. Ordinary inference returns 4. Preserving this layer expectation does not make the nonlinear network exactly equal to the mean over every subnetwork. Validate placement/rates and interactions with normalization for the architecture and data.
Early Stopping
Monitor validation performance and stop when it stops improving, with suitable patience. Early stopping limits how far optimization proceeds and can reduce fitting to noise. In linear models it can have a spectral-shrinkage effect related to ridge, but it is not generally identical to using one L2 coefficient. Save the checkpoint selected by your validation criterion and evaluate final performance on separate data.
Label Smoothing
Replace one-hot targets [0, 0, 1, 0] with smoothed targets [ε/K, ε/K, 1−ε+ε/K, ε/K], where K is the number of classes and ε controls the smoothing. This discourages driving the predicted probability of the correct class all the way to one. It can improve generalization or calibration, but check both on held-out data. In distillation, smoothing can remove useful information from the teacher's relative probabilities, so the interaction needs validation.
Data Augmentation as Regularization
Augmentation encodes task assumptions. A horizontal flip may preserve an animal label but reverse a left-turn instruction. Common methods include random crops, flips, color jitter, mixup (convex combinations of inputs and labels), CutOut (zeroing random patches), and learned policies like RandAugment / AutoAugment. Transformations and synthetic targets such as mixup must make sense for the task. Compare their held-out effect rather than assuming stronger augmentation always helps.
Batch / Layer Normalization
Normalization layers regularize as a side effect. BatchNorm computes mean/variance from the mini-batch, so each training example sees slightly different statistics, introducing batch-dependent variation in activations. LayerNorm/RMSNorm don't have this stochasticity (they normalize per-example) but stabilize training, which lets you use higher learning rates and stronger augmentation.
Modern LLM Regularization
Large-model training recipes often combine data diversity, deduplication, learning-rate schedules and AdamW weight decay. Some omit dropout; there is no single recipe for all models and datasets. During fine-tuning on a small dataset, compare limited-rank updates, dropout, early stopping and suitable learning rates using task and capability-retention evaluations. A KL penalty against a reference policy can discourage policy drift during RLHF, but a finite penalty does not prevent all drift.
7. Calibrate probabilities for decisions
A binary classifier is calibrated when examples assigned probability p have positive frequency p in the population. Predicting 0.1 for everybody on a 10%-positive population can be calibrated but offer no useful ranking.
Use reliability plots, log-loss and Brier score (mean squared probability error), with relevant slices and uncertainty. ECE averages binned confidence–frequency gaps; its result depends on bins and sample size. Fit temperature scaling, Platt scaling or isotonic regression on held-out calibration data and assess separately. Positive temperature scaling preserves top class but can change multiclass confidence ranking across examples.
For a simplified binary decision with zero cost for correct outcomes, false-positive cost C_FP and false-negative cost C_FN, predict positive when calibrated p exceeds C_FP/(C_FP+C_FN). Costs 1 and 9 give threshold 0.1. Capacity limits or downstream effects need a richer decision rule.
Check: at 0.1% fraud prevalence, what does always predicting ordinary achieve?
99.9% accuracy and zero fraud recall. Inspect the confusion matrix, false-alarm cost and review capacity. Resampling or weighted training changes the fitting distribution/objective, so calibrate and evaluate on representative data.
8. Put the workflow together
Define the population and prediction time. Split for that deployment question. Fit transformations and a baseline inside the training partition. Diagnose comparable train/validation metrics before changing capacity. Select settings/checkpoints on validation and assess the full procedure independently. After deployment, reproduce actual failures, compare labeled slices as outcomes arrive, and test remedies on representative data.
Continue to classical ML to compare model families, then deep learning. Evaluation develops the comparison process. Return to probability for interval and Bayes calculations and information theory for log-loss.
Further Reading
scikit-learn: common pitfalls — leakage and fitted preprocessing pipelines.
An Introduction to Statistical Learning — resampling and model-assessment exercises.
Mathematics for Machine Learning — probability, distributions and the mathematical foundations used here.
The Elements of Statistical Learning — model assessment, bias–variance and regularization with their assumptions.
Decoupled Weight Decay Regularization — the AdamW paper and its distinction between weight decay and an L2 penalty.
On Calibration of Modern Neural Networks — calibration experiments and temperature scaling; results depend on the model and dataset studied.