🏛️ Pre-Transformer Architectures
Work through convolution geometry, recurrent states, gated memory, residual paths, and selective state-space models to understand architecture and deployment tradeoffs.
On this page
An image classifier needs to recognize the same pattern in different places. A streaming sequence model needs to update its prediction when one more observation arrives. Convolution and recurrence solve these different problems by reusing parameters: across space in a CNN, and across time in an RNN. Their constraints explain both their strengths and the motivation for attention.
Before you start
Read Linear Algebra for shapes and Deep Learning Basics for the chain rule, nonlinearities, initialization, and normalization. Calculus & Optimization supports the gradient sections. The main route is convolution → recurrence → gated memory → attention; the architecture families and state-space section provide deeper connections.
By the end, you should be able to compute a convolution's output shape and receptive field, trace a recurrent state, explain an LSTM's direct gradient path, and compare architectures using information access and deployment costs.
1. Convolution: reuse one detector across locations
Take a one-dimensional input [1,2,4,7] and a kernel [-1,1], with stride 1 and no padding. Sliding the kernel gives [1,2,3]: each output is the increase between neighboring inputs. The same two parameters detect that relationship at every position.
Deep-learning libraries usually call this operation convolution while implementing cross-correlation, without flipping the kernel. Learned filters can represent either convention, but the distinction matters when checking a hand calculation.
For images stored as X: B×C_in×H×W, a standard kernel has shape C_out×C_in×k_h×k_w. Each output channel sums spatial and input-channel products, then optionally adds a bias. The output has shape B×C_out×H_out×W_out. Ignoring bias, the parameter count is k_h k_w C_in C_out; it does not grow with input image area.
Three useful structural assumptions are:
- Locality: one layer reads a local neighborhood. Stacking layers expands access to distant regions.
- Parameter sharing: the same filter is applied at many positions, reducing the number of independent parameters.
- Translation equivariance: under compatible boundary and sampling conditions, shifting the input shifts the feature map. Stride and finite-image padding limit the exact property.
Equivariance means a corresponding output shift, not an unchanged output. Pooling and global aggregation can reduce sensitivity to position, but ordinary local max pooling does not guarantee invariance to arbitrary translations. These assumptions can help when the task has local repeated structure; they do not guarantee a particular accuracy advantage or data-count threshold.
Output size, stride, padding, and dilation
For one dimension with input length n, symmetric padding p, kernel width k, dilation d, and stride s:
output = floor((n + 2p − d(k−1) − 1)/s) + 1.
Dilation spaces the kernel taps; its effective span is d(k−1)+1. With n=7, k=3, p=1, d=1, s=2, output length is 4. Thus stride 2 does not always yield exactly half the input length. A framework's SAME convention commonly gives ceil(n/s) and may use asymmetric padding; VALID means no padding, with output size still depending on stride and dilation.
Check yourself: use n=7, k=3, d=2, p=0, s=1. Solution: the kernel spans 5 positions and produces 3 outputs. The number of learned taps is still 3.
Receptive field: what can influence one output?
Track the receptive field R and the spacing j between neighboring outputs measured in original-input coordinates. Start at R=1,j=1. Each layer updates:
R_new = R + (k−1)d j
j_new = j s
Five 3×3 stride-1 layers have side-length receptive field 11. With stride 2 in layers 2 and 4, the successive R values are [3,5,9,13,21], and j ends at 4. With stride 1 and dilations [1,2,4,8,16], R reaches 63. Ten ordinary 3×3 layers reach 21. These are theoretical spans; learned influence within the span can be uneven, and dilation patterns can introduce gaps.
2. Build useful CNN blocks without confusing operations
A 1×1 convolution mixes input channels at each position without expanding the spatial receptive field at stride 1. It can reduce or expand width around an expensive spatial layer. A nonlinearity is a separate operation; a bare 1×1 convolution is linear apart from bias.
A depthwise convolution applies a spatial filter independently to each input channel, often with one filter per channel. A following pointwise 1×1 convolution mixes channels. For a 3×3 layer with C_in=32 and C_out=64:
| Design | Weights, ignoring bias |
|---|---|
| Standard | 9×32×64 = 18,432 |
| Depthwise then pointwise | 9×32 + 32×64 = 2,336 |
The ratio is 1/C_out + 1/9 ≈ 0.127, about 7.89 times fewer weights and multiply-accumulates under matching spatial assumptions. One multiply-accumulate is often counted as two FLOPs; state the convention. This factorization constrains the kernels and does not have identical expressive power to an arbitrary standard convolution of the same shape. Actual latency also depends on arithmetic intensity, data movement, layout, fusion, and supported kernels.
Two 3×3 stride-1 layers cover a 5×5 region and introduce an extra activation if one is inserted. They use fewer weights than one 5×5 layer when channel widths make the comparison favorable—for constant width C, 18C² < 25C². With arbitrary intermediate widths, the conclusion can change.
Max pooling selects local maxima; average pooling averages. Global average pooling maps a feature map to one value per channel and reduces head parameters, but loses explicit spatial layout. A learned strided convolution is another downsampling option. For dense predictions such as segmentation, preserve enough spatial information for the output task.
Architecture ideas worth recognizing
| Family | Reusable idea |
|---|---|
| LeNet | Convolution, subsampling, and classification stages |
| AlexNet | Large supervised CNN training with ReLU, augmentation, dropout, and GPU computation |
| VGG | Repeated small spatial kernels and deeper feature hierarchies |
| Inception | Parallel branches at different scales, with 1×1 reductions |
| ResNet | Learn a residual added to a skip path |
| DenseNet | Concatenate features from earlier layers to reuse them |
| MobileNet / Xception | Factor spatial filtering and channel mixing |
| EfficientNet | Coordinate depth, width, and resolution scaling |
| ConvNeXt | Revisit CNN block design and training recipes alongside transformer-era methods |
These are mechanisms, not a ranking of current defaults. Detection heads, classification heads, and segmentation decoders can reuse related backbones while needing different resolution and latency tradeoffs.
Residual learning and U-Net skips serve different roles
A residual block y=x+F(x) can represent an identity mapping by making F zero. This helped address the degradation problem: some deeper plain networks had worse training error, showing an optimization problem rather than simply overfitting. With an identity skip and no activation after addition, the Jacobian is I+J_F; the gradient is (I+J_F)ᵀg. If F(x)=−x the terms cancel, so the identity route is not a stability proof. Projection shortcuts, activation placement, normalization, and initialization change the details.
U-Net uses an encoder to build coarser semantic features and a decoder to restore spatial resolution. At matching resolutions, it commonly concatenates encoder features into the decoder, preserving access to local detail. Upsampling can use transposed convolution or interpolation followed by convolution. Alignment, cropping/padding, and output losses matter. U-Net variants are useful in segmentation and some diffusion models; neither all segmentation nor all diffusion systems require this backbone.
3. Recurrence: update a fixed-width memory over time
For a sequence X: B×T×D, a simple RNN maintains h_t: B×H. Using row vectors:
W_x: D×H W_h: H×H b: H
h_t = tanh(x_t W_x + h_(t−1) W_h + b)
logits_t = h_t W_y + b_y
The same weights apply at every time step. As a toy linear recurrence, use h_t=0.5h_(t−1)+x_t, h_0=0, and inputs [2,0,0]. The states are [2,1,0.5]. The contribution of that first input decays geometrically; after two transitions its derivative is 0.25. A coefficient larger than one would allow growth. Nonlinear vector RNNs are more complicated because their Jacobians vary with input and state.
Backpropagation through time (BPTT) applies ordinary backpropagation to the unrolled graph and sums gradient contributions to shared parameters. Long paths multiply activation and recurrent Jacobians. Individual spectral radii alone do not determine a product of varying, possibly non-normal matrices. Operator norms give useful bounds, while directions and saturation matter in practice.
Truncated BPTT carries a numerical hidden state across chunks but detaches the graph at boundaries. It saves activation memory while cutting gradient credit assignment across those boundaries. Fixed inference state does not imply constant training memory for an unrolled graph. Gradient clipping limits large finite gradients; it does not recover a vanished signal or repair non-finite values. Choose its threshold from observed scales, as discussed in Deep Learning Basics.
4. Gated memory: trace the LSTM and GRU
An LSTM keeps cell state c and exposed hidden state h, each width H. Concatenating the current D-entry input and previous H-entry hidden state gives D+H inputs to four learned affine maps. Three maps produce sigmoid gates; one produces a tanh candidate:
f_t = sigmoid([x_t,h_(t−1)] W_f + b_f) # forget
i_t = sigmoid([x_t,h_(t−1)] W_i + b_i) # input
g_t = tanh([x_t,h_(t−1)] W_g + b_g) # candidate
o_t = sigmoid([x_t,h_(t−1)] W_o + b_o) # output
c_t = f_t ⊙ c_(t−1) + i_t ⊙ g_t
h_t = o_t ⊙ tanh(c_t)
Here each W has shape (D+H)×H; ⊙ means elementwise multiplication. If one cell has previous value 2, forget gate .9, input gate .2, candidate .5, and output gate .8, then its new cell is 1.9 and exposed hidden value is .8 tanh(1.9) ≈ .765.
Holding gates and other paths fixed, the direct derivative from c_(t−1) to c_t is diag(f_t). Over many steps that path multiplies forget gates. A constant .9 retains only .9^100≈0.0000266 after 100 steps, whereas .99^100≈0.366. Values near one help; the path still multiplies values, and the total derivative includes gate dependencies. A positive initial forget bias can encourage retention, but bias 1 gives sigmoid(1)≈.731 before other contributions, not a gate value of one. This mitigates vanishing gradients without eliminating them.
A GRU keeps one state and uses update and reset gates. One convention is:
z_t = sigmoid([x_t,h_(t−1)] W_z + b_z)
r_t = sigmoid([x_t,h_(t−1)] W_r + b_r)
h_candidate = tanh([x_t,r_t ⊙ h_(t−1)] W + b)
h_t = (1−z_t) ⊙ h_(t−1) + z_t ⊙ h_candidate
Under this convention z near zero retains the old state. Some libraries reverse that convention or place the reset multiplication after the recurrent linear map; match the implementation before porting weights. A basic LSTM uses four affine transforms and a GRU three: at equal D,H and matching bias conventions, the GRU has about 25% fewer recurrent-layer parameters, while the LSTM has about 33% more than the GRU. Neither comparison includes embeddings and output heads or guarantees a speed/quality ranking.
5. Attention changes how stored information can be read
An encoder–decoder without attention asks a fixed summary to support every future decoding decision. Attention supplies a separate context for each decoder query. Given encoder states h_i: H and decoder query s_t: S, additive attention can use:
e_(t,i) = vᵀ tanh(W_s s_t + W_h h_i)
α_(t,:) = softmax(e_(t,:)) # across source positions
context_t = Σ_i α_(t,i) h_i # H entries
The query may be a previous decoder state depending on the update convention. Mask padded source positions before softmax. For two encoder values [1,0] and [0,2] with weights [.25,.75], context is [.25,1.5]. This is a weighted read, not a choice of one mandatory alignment. Attention maps may help inspect behavior but do not establish a causal explanation of the output. Dot-product and bilinear scoring are alternatives to the additive network.
Transformer self-attention makes tokens query other allowed tokens directly. With all training tokens known, many positions can be processed in parallel; causal masks prevent reading future targets. Ordinary autoregressive generation still waits for newly sampled tokens. A dense attention layer has token-pair work quadratic in length, while an ordinary RNN has sequential state updates. Shorter attention paths do not make gradient stability automatic. Continue in Transformers.
Vision and state-space connections
A ViT can split a 224×224 RGB image into 16×16 patches: 196 patches, each 768 numbers, projected to model width. Position information and often a class token are added. This patch projection can be implemented as a convolution with kernel and stride 16; ViTs have architectural priors such as patching and positional structure, though different from CNN locality. Training data, augmentation, pretraining, resolution, and compute affect comparisons. No fixed dataset size makes CNNs or ViTs universally better.
A linear state-space model starts from dh/dt=Ah+Bx, y=Ch, then discretizes to h_t=A_bar h_(t−1)+B_bar x_t. Structured models exploit algebra to evaluate sequences efficiently. Mamba introduces input-dependent selection, including B, C, and the discretization step Δ; the underlying learned continuous A is not simply an arbitrary input-generated matrix. The effective discrete transition depends on Δ.
Affine updates h→a_t h+b_t compose associatively when their coefficients are available from inputs. Two updates compose as (a2 a1, a2 b1+b2), enabling a parallel scan. A general nonlinear LSTM does not have this same simple affine closure. Recurrent inference keeps state whose size is fixed with respect to sequence length, though it grows with model dimensions; it compresses the past and need not preserve every retrievable detail. Attention/SSM hybrids inherit memory costs from their attention layers as well as recurrent state. Evaluate recall, quality, kernels, and latency rather than inferring superiority from asymptotic sequence cost.
6. Choose by information access and measured constraints
Bidirectional RNNs combine a forward and backward state. Exact full-sequence outputs need future context, but bounded lookahead or chunked bidirectionality can provide streaming outputs with delay. If a moderation service sees complete messages, a bidirectional model may already fit; if it must decide token by token, define the allowed delay and whether decisions may be revised.
For a mobile classifier, account for model weights, quantization metadata, runtime binary, activations, and peak scratch memory. Compare suitable CNN and small transformer candidates on representative devices and task data. INT8 often reduces raw FP32 weight storage by approximately four times; latency and accuracy gains depend on operators and hardware. Quantization-aware training, structured pruning, and distillation may help, but none supplies a fixed free accuracy or speed improvement. Measure cold/warm latency, tail latency, thermal conditions, and held-out quality after conversion.
Final check: an architecture has one tenth as many multiply-accumulates. Is it necessarily faster? Solution: no. Poor arithmetic intensity, extra transfers, launch overhead, unsupported operations, or unfavorable shapes can dominate. Keep both the mathematical cost model and the deployment measurement.
Sources and next lessons
- Deep Learning, Chapter 9: Convolutional Networks — convolution, pooling, and structural assumptions.
- Deep Learning, Chapter 10: Sequence Modeling — recurrence, BPTT, gated units, and encoder–decoder models.
- Deep Residual Learning for Image Recognition — degradation and residual learning.
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces — selection, hardware-aware scan, and recurrent state.
Next, work through Transformers. Return to Numerical Computing when debugging precision or kernels, and Model Serving for measured inference constraints.