🧠 Deep Learning Basics
Trace neural-network updates with worked derivatives and tensor shapes, then reason about activations, normalization, optimizers, losses, regularization, and numerical failures.
On this page
Training a neural network means changing many numbers so that its predictions become more useful on new examples. The central challenge is to make those changes measurable: which operation produced an error, how did that error reach a parameter, and did the update help outside the training batch?
Before you start
Linear Algebra supplies vectors, matrix products, and shapes. Calculus & Optimization introduces derivatives and the chain rule; ML Fundamentals explains why lower training loss is not sufficient evidence of improvement. Follow the scalar example first, then the matrix version. Numerical Computing and Information Theory support the precision and loss sections.
You will learn to trace a forward and backward pass, choose normalization axes, explain an optimizer update, and debug training with evidence. Architecture-specific mechanisms come next in Pre-Transformer Architectures and Transformers.
1. One weight, one prediction, one update
Let the model be ŷ = wx, with input x = 2, target y = 5, and initial weight w = 1. Use half squared error, L = (ŷ−y)²/2.
- Forward:
ŷ = 2;L = 4.5. - Backward: the derivative of loss with respect to prediction is
ŷ−y = −3. Prediction changes by x for a unit change in w. The chain rule givesdL/dw = (ŷ−y)x = −6. - Update: with learning rate
η = 0.1,w ← w−η(dL/dw) = 1.6. The new prediction is 3.2 and loss is 1.62.
The negative gradient says that increasing w locally reduces this loss. It does not say every step size works: using η=1 instead gives w=7, prediction 14, and loss 40.5. Optimization follows local information; the step size matters.
Check yourself: estimate the derivative at w=1 using [L(1+h)−L(1−h)]/(2h) with h=0.001. Solution: the two losses are 4.494002 and 4.506002, giving −6. This is a useful gradient check for smooth points in a small implementation, not an efficient way to train millions of weights.
2. Backpropagation is the same chain rule with shapes
A network composes functions. Its forward pass records intermediate values; reverse-mode automatic differentiation propagates the scalar loss gradient backward through them. A layer from n inputs to m outputs has an m×n Jacobian J. If the output gradient is an m-entry column vector v, the input gradient is Jᵀv. Frameworks compute these products without materializing every full Jacobian. For a scalar loss, the backward pass generally costs a modest multiple of the forward pass, depending on operations and recomputation.
A two-layer classifier
Let B be batch size, d the input width, h the hidden width, and K the number of classes. Store examples as rows:
X: B×d W1: d×h b1: h
Z1 = XW1 + b1 # B×h; broadcast b1 over rows
A1 = ReLU(Z1) # B×h
W2: h×K b2: K
Z2 = A1W2 + b2 # B×K, called logits
P = softmax(Z2, classes) # B×K; each row sums to 1
Y: B×K # one-hot or distribution-valued targets
L = −sum(Y * log(P)) / B
For normalized target rows, softmax plus cross-entropy gives:
dZ2 = (P−Y)/B # B×K
dW2 = A1ᵀ dZ2 # h×K
db2 = sum_rows(dZ2) # K
dA1 = dZ2 W2ᵀ # B×h
dZ1 = dA1 * (Z1 > 0) # B×h, elementwise mask
dW1 = Xᵀ dZ1 # d×h
db1 = sum_rows(dZ1) # h
The choice of mean versus sum loss changes gradient scale; masked token losses should specify whether they average over examples or valid tokens. Compute all gradients before changing weights used by earlier backward steps. At ReLU's nondifferentiable zero, frameworks choose a convention, commonly zero. Away from zero, masking with A1 > 0 is equivalent here to Z1 > 0.
Keep the calculation in a safe numerical range
Softmax is unchanged if the same scalar is subtracted from every logit. For logits [1000,1001], exponentiate [-1,0], obtaining probabilities about [0.269,0.731]. Stable log-sum-exp is m + log(sum(exp(z−m))), where m is the largest logit. For target class 2, loss is about 0.3133 and the logit gradient is [0.269,−0.269].
Use a loss API that takes logits, or implement log-softmax carefully. Taking the log of an already rounded-to-zero probability can produce infinity even after the softmax exponentials themselves were stabilized. Separate operations can be implemented stably; the important distinction is the numerical calculation and API contract. A bounded logit gradient does not bound gradients of every earlier weight.
3. Nonlinearities, initialization, and gradient paths
Without nonlinearities, a stack of affine layers is another affine map. Nonlinearities let intermediate features change how the input space is divided.
| Activation | Mechanism | Boundary |
|---|---|---|
Sigmoid σ(z)=1/(1+e^(−z)) |
Maps to 0–1; derivative σ(z)(1−σ(z)) |
Saturates at large magnitude; useful for probabilities and gates |
| Tanh | Maps to −1–1; derivative 1−tanh²(z) |
Also saturates; useful in recurrent states |
ReLU max(0,z) |
Derivative 1 on the positive branch | Zero derivative on the negative branch; an inactive unit may stay inactive |
| Leaky ReLU / PReLU | Fixed / learned negative slope | Reduces zero-gradient regions; does not guarantee stable training |
GELU zΦ(z) |
Smoothly weights z by a Gaussian CDF | Different cost and derivative from ReLU; quality depends on architecture |
Swish zσ(z) |
Smooth self-gating | Used inside gated feed-forward layers |
ELU has an exponential negative tail. SELU's self-normalizing argument requires specific initialization, architecture, and input assumptions. Mish, z tanh(softplus(z)), is another smooth option; no activation wins on every task.
A SwiGLU hidden layer can be written H = swish(XWg) ⊙ (XWu), followed by HWd. With input/output width d and hidden width h, it has approximately 3dh weights, versus 2dh for an ordinary two-projection feed-forward layer. Matching a conventional h=4d layer gives gated h≈8d/3, ignoring biases and implementation rounding. That ratio is parameter accounting, not a universal quality guarantee.
Why gradients shrink or grow
Backpropagation multiplies Jacobians along paths. If all relevant operator norms are bounded below one, long products can shrink exponentially. Norms above one allow amplification but do not guarantee it: directions, activation masks, and cancellation matter. ReLU removes sigmoid's positive-branch saturation but does not remove weight matrices from the product. Units inactive for all current examples receive no gradient through that branch; upstream changes or optimizer state can sometimes revive them.
A residual block y = x + F(x) contributes Jacobian I + J_F. The identity path can help, but cancellation and unstable branches remain possible. An LSTM similarly offers a direct cell-state path. These are mechanisms for better gradient flow, not proofs that gradients cannot vanish or explode.
Initialize to a sensible scale
For approximately independent zero-mean weights and inputs, preactivation variance scales roughly as fan_in × Var(weight) × E[input²]. Glorot initialization uses weight variance 2/(fan_in+fan_out) to balance forward and backward scales under its assumptions. He initialization uses 2/fan_in for ReLU: a symmetric preactivation distribution loses half its second moment after rectification. This is not exactly a halving of centered variance, because ReLU changes the mean.
Leaky ReLU uses a slope-dependent gain. GELU and multiplicative gates need architecture-appropriate initialization; the ReLU derivation does not automatically transfer unchanged. Orthogonal matrices preserve vector norms in the linear map, but nonlinearities and weight updates change that property. Some deep residual designs scale branch outputs at initialization, for example by 1/√(2N) for N layers under a particular convention. Inspect activation and gradient distributions to test the intended effect.
4. Normalization: name the axes before the layer
Normalization divides by a scale computed from a specified set of values, with ε to avoid division by zero. Learned scale and sometimes shift parameters then restore flexibility.
For a token tensor X: B×T×D, LayerNorm over D computes a mean and variance separately for each (batch, token) pair. It does not average over future tokens. RMSNorm over D divides by sqrt(mean(x²)+ε) without subtracting the mean. For one vector [1,3], ignoring ε and affine parameters, LayerNorm gives [-1,1], while RMSNorm gives [1/√5,3/√5] ≈ [0.447,1.342]. Both rescale; only the former centers.
For an image tensor B×C×H×W, convolutional BatchNorm normally computes statistics over B,H,W for each channel C. It uses batch statistics during training and running estimates at evaluation by default. Small or unrepresentative batches can make those estimates poor. GroupNorm instead computes statistics over channels within each group and spatial positions for each example; InstanceNorm commonly normalizes spatial positions per example and channel.
BatchNorm is not inherently impossible in sequence models. The axes and availability of statistics matter: including future positions in training statistics can leak information into a causal prediction, and different inference statistics can cause mismatch. LayerNorm/RMSNorm over token features avoid this particular dependence on batch composition and future positions. Their own normalization rules are the same in training and evaluation, though dropout and other model components can still differ. RMSNorm has less arithmetic than centered LayerNorm; realized speed, affine parameter count, and quality depend on implementation and design.
5. Optimizers change how the gradient becomes an update
SGD uses θ ← θ−ηg. Momentum maintains v ← βv+g and steps along v; some libraries use an equivalent normalized convention with a correspondingly different learning rate. Nesterov momentum evaluates or approximates a gradient at a look-ahead point. Neither SGD nor momentum is inherently too slow for a particular architecture.
Adagrad accumulates squared gradients, reducing step scales for repeatedly active coordinates; this can help sparse features but make later progress small. RMSprop uses an exponential moving average instead. Adam maintains two averages:
m_t = β1 m_(t−1) + (1−β1) g_t
v_t = β2 v_(t−1) + (1−β2) g_t²
m_hat = m_t / (1−β1^t)
v_hat = v_t / (1−β2^t)
θ ← θ − η m_hat / (sqrt(v_hat)+ε)
The square and division are elementwise. Zero initialization biases the moving averages downward under stationary moments, motivating the corrections. But their ratio need not be small: for a first nonzero gradient, β1=0.9, β2=0.999 and negligible ε, the uncorrected update magnitude is η×0.1/√0.001 ≈ 3.162η; correction gives approximately η.
Adam with an L2 loss penalty puts the penalty gradient into both moment estimates. AdamW computes moments from the data gradient and applies proportional decay separately, for example θ ← θ−η adaptive_step−ηλθ. They are different algorithms; neither choice implies a guaranteed training failure. Learning rate, decay, and which parameters receive decay all need validation.
Adam/AdamW usually store two optimizer-state tensors per parameter; momentum SGD usually stores one. Total memory also includes weights, gradients, precision copies, activations, and temporary buffers. Lion uses a sign-based update and one momentum state. Sophia uses curvature estimates as part of an adaptive update. Such alternatives have workload-dependent quality, implementation cost, and memory benefits; published speedups are not universal constants.
Schedule, clip, and preserve numerical range
Warmup limits early parameter changes; it serves a different purpose from Adam bias correction. After warmup, a cosine schedule can use η(t)=η_min + (η_max−η_min)(1+cos(πt/T))/2, where t counts decay steps and T is their total. At t=0 it is η_max, halfway it is their midpoint, and at t=T it is η_min. Step decay, inverse-square-root, one-cycle, and warmup–stable–decay are alternatives. Choose using the training budget and measured learning behavior.
Global norm clipping uses g ← g min(1,τ/‖g‖₂). For g=[3,4] and τ=2, the result is [1.2,1.6], preserving direction. Coordinate clipping at 2 gives [2,2], a different direction. There is no universal threshold: loss reduction, batch scaling, model, and optimizer matter. Gradient clipping bounds this gradient, not necessarily the entire Adam update or parameter movement.
Mixed precision uses low-precision operations where appropriate and higher precision for sensitive reductions or state. FP16 has 5 exponent and 10 fraction bits; BF16 has 8 exponent and 7 fraction bits. BF16 has a much wider range but coarser spacing. FP16 commonly needs loss scaling to preserve small gradients: scale loss, backpropagate, unscale, check finiteness, clip, then update. BF16 usually needs no loss scaling, yet can still overflow, underflow, or round away small changes. FP8 requires additional format/scaling choices and hardware support. See Numerical Computing.
6. Choose a loss and regularizer for the actual target
Cross-entropy is negative log-likelihood for categorical targets; BCE is its binary counterpart. Squared error estimates a conditional mean and corresponds to a Gaussian observation model with fixed variance when used as a likelihood. MAE targets a conditional median and penalizes large residuals linearly. Huber is quadratic near zero and linear beyond a chosen threshold. Hinge loss max(0,1−yf(x)) rewards a margin for y∈{−1,+1}.
Focal loss uses the true-class probability p_t: −α_t(1−p_t)^γ log(p_t). For γ=2 and no class weight, p_t=0.9 gets a multiplier of 0.01; p_t=0.2 gets 0.64. It emphasizes hard examples, which can include mislabeled examples. Compare it with weighting, sampling, calibration, and operating-threshold changes; no imbalance ratio mandates it.
Contrastive losses learn relationships. InfoNCE treats a positive as the desired choice among competing candidates, with temperature controlling logit scale. Negatives need not be every other batch item; false negatives and sampling matter. Triplet loss max(0,d(a,p)−d(a,n)+margin) enforces relative distances and depends strongly on which triples are selected.
KL divergence is D_KL(P‖Q)=Σ P log(P/Q). It is nonnegative but asymmetric, not a distance metric. H(P,Q)=H(P)+D_KL(P‖Q) explains why minimizing cross-entropy fits Q to fixed P. Missing probability where P is positive has infinite forward-KL cost. Reverse KL has the opposite support requirement. When the approximating family is restricted, these directions often favor covering versus selecting modes, but neither slogan is a universal theorem about fitted models. Jensen–Shannon divergence averages KLs to a mixture; symmetry does not guarantee absence of mode collapse.
Inverted dropout keeps an activation with probability 1−p and scales it by 1/(1−p) during training. For an activation 4 and p=0.25, its expected output is 0.75×(4/0.75)=4. Individual predictions remain noisy; evaluation normally disables it. Deliberate Monte Carlo dropout is a different inference procedure. Stochastic depth drops residual branches. Weight decay penalizes parameter growth through the update. Label smoothing replaces one-hot y by (1−ε)y+ε/K; for three classes and ε=0.1, the target becomes [0.9333,0.0333,0.0333]. It can reduce confidence but does not guarantee improved calibration or accuracy. Early stopping and label-preserving augmentation also require validation.
7. Debug before changing everything
Save enough information to replay a failing step: checkpoint, optimizer/scaler state, batch identifiers, random state, and distributed configuration. Check inputs, labels, masks, finite activations, loss denominators, gradient norms, and updates. A division by zero or invalid label can fail abruptly without a gradual warning. Replay a small case in FP32 or with selected kernels disabled to localize a numerical fault. A loss spike alone does not prove a hardware failure or optimizer defect.
Fix the diagnosed cause and resume from known-good state. Lowering a learning rate, changing clipping, or skipping a batch may be useful experiments, but unexplained data deletion can hide a systematic problem. Validate recovery over enough steps and across data slices. Track validation behavior as well as optimization health.
Where the architectures go next
n-grams use counts and smoothing over bounded context; embeddings and recurrence share parameters across contexts. RNNs carry fixed-width state, introducing sequential dependence and long gradient paths. LSTM/GRU gates help control memory; attention lets a query read multiple stored states. Transformer training parallelizes known token positions with the appropriate mask, while ordinary autoregressive generation remains sequential across new tokens. Convolutional, recurrent, state-space, and attention-based models remain useful under different memory and latency constraints.
Final check: a tiny model fits in a 10 MB keyboard budget. Must producing three candidate next tokens require three forward passes? Solution: no. One vocabulary distribution provides its top three entries. Generating three-token continuations requires additional steps. Budget weights, vocabulary, runtime, and recurrent state or KV cache before comparing measured device latency; no architecture is excluded by its name alone.
Sources and next lessons
- Deep Learning, Chapter 6: Deep Feedforward Networks — computational graphs, backpropagation, and output losses.
- Deep Learning, Chapter 8: Optimization for Training Deep Models — initialization, gradients, and optimizers.
- Decoupled Weight Decay Regularization — the distinction between adaptive L2 regularization and weight decay.
- PyTorch automatic mixed precision examples — loss scaling, unscaling, and clipping order.
Continue with Pre-Transformer Architectures to follow states through time and space, then Transformers to follow queries, keys, values, and masks.