Week 5 · architectures
Training error on MNIST-1D for plain fully connected networks of depth 2, 8 and 32, counting hidden 64-to-64 layers: width 64, SGD at learning rate 0.005, three seeds each, faint lines for single runs, solid lines their mean.
\[ \hidden_1=\vect{f}_1[\vect{x},\params_1],\qquad \hidden_2=\vect{f}_2[\hidden_1,\params_2],\qquad \hidden_3=\vect{f}_3[\hidden_2,\params_3],\qquad \vect{y}=\vect{f}_4[\hidden_3,\params_4] \]
\[ \frac{\partial y}{\partial h_1} = \frac{\partial h_2}{\partial h_1} \frac{\partial h_3}{\partial h_2} \frac{\partial y}{\partial h_3} \]
\[ \params \leftarrow \params-\alpha\nabla_{\params}\loss \]
Width 200, with weights and biases drawn at variance 1/fan-in. Left and centre: dy/dx across the input for one shallow and one 24-layer plain network, each on its own vertical scale. Right: autocorrelation against input distance, averaged over twenty networks, with a 24-layer residual network whose branches are scaled by 1/sqrt(24); dotted line, one half.
\[ \hidden_{\text{out}} = \vect{f}[\hidden_{\text{in}},\params] \]
\[ \hidden_{\text{out}} = \hidden_{\text{in}} + \vect{f}[\hidden_{\text{in}},\params] \]
\[ \hidden_k=\hidden_{k-1}+\vect{f}_k[\hidden_{k-1},\params_k] \]
A residual block. The input travels unchanged along the identity path while the branch computes a learned correction; the two are added elementwise, so they must have the same shape.
\[ \hidden_1=\vect{x}+\vect{f}_1[\vect{x}], \qquad \vect{y}=\hidden_1+\vect{f}_2[\hidden_1] \]
\[ \vect{y}=\vect{x}+\vect{f}_1[\vect{x}]+\vect{f}_2\!\left[\vect{x}+\vect{f}_1[\vect{x}]\right] \]
\[ \frac{\partial \hidden_k}{\partial \hidden_{k-1}} = \mat{I} + \frac{\partial \vect{f}_k}{\partial \hidden_{k-1}} = \mat{I}+\mat{J}_k \tag{1}\]
\[ (\mat{I}+\mat{J}_2)(\mat{I}+\mat{J}_1) = \mat{I}+\mat{J}_1+\mat{J}_2+\mat{J}_2\mat{J}_1 \]
The same two-block residual network drawn four times, one route through it highlighted in each: one panel per term of the expanded product. A path through a branch picks up that branch’s Jacobian; the later branch’s sits on the left.
Left: the eight paths through three residual blocks, one row each, a filled dot where the path takes the learned branch. Right: for a network of 54 blocks, the share of its paths passing through each number of branches, with the central band holding half of all paths highlighted and, shaded, the lengths Veit et al. found carry most of the gradient.
The size of a gradient with respect to the activations, after passing back through k layers from size 1, over 200 draws. Plain: a product of k random ReLU-layer Jacobians J. Residual: a product of I + J, or of I + J/sqrt(50) for small branches. The line is the median; the band runs from the 10th to the 90th percentile.
\[ \hidden \rightarrow \text{normalisation} \rightarrow \text{ReLU} \rightarrow \text{conv} \rightarrow \text{normalisation} \rightarrow \text{ReLU} \rightarrow \text{conv} \]
\[ \var{A+B} = \var{A}+\var{B}+2\operatorname{Cov}\!\left[A,B\right] \]
\[ \var{A+B} \approx \var{A}+\var{B} \]
\[ \var{h+f[h]}\approx 2v \]
\[ v_0=v,\qquad v_1\approx 2v,\qquad v_2\approx 4v,\qquad\cdots,\qquad v_K\approx 2^K v \]
\[ \var{\frac{h+f[h]}{\sqrt{2}}} \approx \frac{2v}{2}=v \]
\[ m_h=\frac{1}{|\set{B}|}\sum_{i\in\set{B}}h_i \]
\[ s_h^2=\frac{1}{|\set{B}|}\sum_{i\in\set{B}}(h_i-m_h)^2, \qquad s_h=\sqrt{s_h^2} \]
\[ \hat h_i= \frac{h_i-m_h}{\sqrt{s_h^2+\epsilon}} \]
\[ \tilde h_i=\gamma \hat h_i+\delta \]
torch.var defaults to the unbiased form, \(6.67\) here; BatchNorm normalises with the population formThe four activations at each stage of batch normalisation, epsilon taken as zero. Centring moves their mean to zero; standardising scales their spread to one; the learned scale and offset then move them to wherever training prefers.
\[ v_{k+1}\approx v_k+1 \quad\Longrightarrow\quad v_K\approx v_0+K, \qquad\text{against}\qquad v_K\approx 2^K v_0 \]
Activation variance after each of 20 residual blocks at initialisation, one random network of width 256 per curve, measured on a batch of 512 random inputs, on a log scale. Dashed: 2^k. Dotted: 1 + k, which the normalised curve lies on. The normalised branch batch-normalises its input, the block the depth-32 runs train.
The MNIST-1D runs again: plain depth 8 and 32, and a depth-32 residual network with BatchNorm at the start of each branch. Lines are means over three seeds, at the same learning rate and epochs as before.
One example’s BatchNorm output, epsilon taken as zero, in 4,000 random minibatches of 8 and of 64. The unit’s activations over the data have mean 1 and standard deviation 2, and the example’s is 2.5. The dotted vertical line is the value standardising with the data’s own mean and standard deviation would give.
\[ \frac{ah_i-am_h}{as_h} = \frac{h_i-m_h}{s_h} \]
\[ \vect{z}_{ij}=\layerbias+\layerweights\transpose\hidden_{ij}, \qquad \layerweights\in\reals^{C_i\times C_o} \]
A basic block’s branch, left, and a bottleneck branch, right, each taking 256 channels back to 256. Each layer’s top and bottom edges widen with its input and output channels, and the skip path carries all 256 around the branch to the sum. Weights exclude biases.
\[ \hidden_k=\operatorname{concat}\left(\hidden_{k-1},\vect{f}_k[\hidden_{k-1}]\right) \]
\(\hidden_{k-1}\) already holds the input and every earlier layer’s output, so each layer sees all of them
Addition keeps the width fixed; concatenation grows it
TRY: \(C_0=64\), \(g=32\). How many channels enter layer 12?
A dense block of 12 layers with growth rate 32, against a residual stage of constant width 64. Left: the channels entering each layer, which concatenation grows by 32 every layer. Right: the weights in that layer’s 3 x 3 convolution; each dense layer outputs 32 channels and each residual layer 64, so the dense layers start cheaper, draw level at layer 3 and cost more after.
\[ \hidden_{\text{decoder}}'= \operatorname{concat} \left(\hidden_{\text{decoder}},\hidden_{\text{encoder}}\right) \]
Last week’s encoder-decoder shapes, with U-Net’s skips added; box height is spatial size and box width is channels. Each intermediate encoder stage’s features are carried across and concatenated with the decoder stage of the same spatial size, which widens it, returning detail at the 112 x 112 and 56 x 56 resolutions that the downsampling discarded.
| Method | Statistics gathered across, for activations \((B,C,H,W)\) |
|---|---|
| BatchNorm | batch and spatial positions, separately per channel |
| LayerNorm | channels and spatial positions within each example |
| GroupNorm | groups of channels and spatial positions within each example |
| InstanceNorm | spatial positions within each channel and example |
An activation tensor drawn with its batch, channel and spatial axes, one faint line per example and per channel, in the order of the table. The shaded values are the ones one mean and variance are computed over. BatchNorm: one channel across the batch and every position. LayerNorm: one example across every channel and position. GroupNorm: a group of channels of one example, across positions. InstanceNorm: one channel of one example.
\[ \hidden_k=\hidden_{k-1}+\vect{f}_k[\hidden_{k-1}] \]
| Choice | Update | Earlier representation | Width |
|---|---|---|---|
| Replace | \(\hidden_k=\vect{f}_k[\hidden_{k-1}]\) | rewritten | set by \(\vect{f}_k\) |
| Add | \(\hidden_k=\hidden_{k-1}+\vect{f}_k[\hidden_{k-1}]\) | kept, mixed with the change | fixed |
| Concatenate | \(\hidden_k=\operatorname{concat}(\hidden_{k-1},\vect{f}_k[\hidden_{k-1}])\) | kept, as separate channels | grows |
eval() does not mean “compute a validation score”; it changes how modules behavemomentum; it’s unrelated to optimiser momentumA synthetic unit whose true activation mean, dotted, drifts from 0 to 2 between steps 100 and 500 as training changes the weights. Dots: the batch mean at each step, noise standard deviation 0.3. Lines: running estimates starting at 0, at two values of rho.
eval() uses the stored estimates from training, so those estimates need enough representative batches to become usefulmomentum; it is unrelated to optimiser momentum| ResNet | DenseNet | U-Net | |
|---|---|---|---|
| Merge | add | concatenate | concatenate |
| Earlier features | mixed with the new | kept as separate channels | carried past the coarsest stage |
| Width | fixed | grows unless controlled | grows at each merge |
| Main purpose | short optimisation path | feature reuse | restore spatial detail |