Week 4 · architectures
\[ 224\times224\times3=150{,}528 \]
A dense hidden layer of the same width needs
\[ 150{,}528^2 \approx 2.27\times10^{10} \]
weights
Nothing in the layer’s structure encodes the fact that the input is an image.
The same two digits before and after one fixed permutation of their 784 pixels. A dense layer whose weight columns are permuted the same way computes exactly the same outputs from the scrambled images, so nothing in its connectivity records which pixels were neighbours. MNIST, LeCun and Cortes, CC BY-SA 3.0.
Left: a dense layer from six inputs to six outputs, one free weight per edge. Right: a convolution with kernel size three over the same units. Each output reads only its neighbours, and every edge of one colour carries the same weight. Inputs past the edge count as zero, so the end outputs lose the edge that would reach past the input.
\[ \vect{f}[\vect{t}[\vect{x}]] = \vect{f}[\vect{x}] \]
Definition 1 A model \(\vect{f}\) is translation invariant if translating its input leaves its output unchanged: \(\vect{f}[\vect{t}[\vect{x}]] = \vect{f}[\vect{x}]\) for every input \(\vect{x}\) and every translation \(\vect{t}\).
\[ \vect{f}[\vect{t}[\vect{x}]] = \vect{t}[\vect{f}[\vect{x}]] \]
Definition 2 A model \(\vect{f}\) is translation equivariant if translating its input translates its output in the same way: \(\vect{f}[\vect{t}[\vect{x}]] = \vect{t}[\vect{f}[\vect{x}]]\) for every input \(\vect{x}\) and every translation \(\vect{t}\).
Left: a short pattern, and the same pattern shifted to the right. Centre: a convolution with a kernel matched to the pattern. Its response shifts by exactly the same amount, which is equivariance. Right: the largest response anywhere in the signal is the same for both, which is invariance.
| Task | Input | Desired output |
|---|---|---|
| Image classification | object moves | class stays the same |
| Semantic segmentation | object moves | mask moves with it |
| Keypoint localisation | object moves | keypoints move with it |
Definition 3 A kernel, or filter, is the vector of weights \(\kernel=[\omega_1,\ldots,\omega_K]\transpose\) applied to every local window of the input. Each weight is one tap, and the number of taps \(K\) is the kernel size.
\[ z_i = \omega_1x_{i-1}+\omega_2x_i+\omega_3x_{i+1} \]
\[ z_i=(-1)(2)+(0)(5)+(1)(4)=2 \]
The kernel [-1, 0, 1] at three neighbouring positions of one signal. Each highlighted output has one connection from each input in its window, and the connections carry the kernel’s weights: the same three, in the same colours, in every panel. The output row fills in one value per position. Outputs exist only where the whole window fits, so none sits above the end inputs. The first panel is the worked example from the previous slide.
\[ \begin{aligned} z_i &= \omega_1x_{i-1}+\omega_2x_i+\omega_3x_{i+1}\\ z_{i+1} &= \omega_1x_i+\omega_2x_{i+1}+\omega_3x_{i+2} \end{aligned} \]
\[ z_i = \sum_{j=1}^{3} \omega_j x_{i+j-2} \]
\(j=1,2,3\) reads \(x_{i-1},x_i,x_{i+1}\): the sum from the previous slides
Mathematical convolution flips one operand first; the unflipped operation is cross-correlation
\[ \text{convolution:}\quad z_i = \sum_{j=1}^{3} \omega_j x_{i-j+2} \]
\[ \begin{aligned} z'_i &= \sum_{j=1}^{3}\omega_j x'_{i+j-2}\\ &= \sum_{j=1}^{3}\omega_j x_{(i-s)+j-2}\\ &= z_{i-s} \end{aligned} \]
\[ h_i=\activation{\beta_i+\sum_{j=1}^{D}\omega_{ij}x_j} \]
\[ h_i= \activation{ \beta+ \omega_1x_{i-1}+ \omega_2x_i+ \omega_3x_{i+1} } \]
\[ \layerweights=\begin{bmatrix} \omega_2 & \omega_3 & 0 & \cdots\\ \omega_1 & \omega_2 & \omega_3 & \ddots\\ 0 & \omega_1 & \omega_2 & \ddots\\ \vdots & \ddots & \ddots & \ddots \end{bmatrix} \]
The convolution with kernel size three, zero padding and eight inputs, written as the dense matrix that computes it. Each colour is one kernel weight, repeated down its diagonal; every other entry is a structural zero.
Definition 4 Padding extends the input by \(P\) positions at each end before the kernel is applied. Zero padding fills those positions with zeros.
\[ \dout = \din-K+1 \]
Definition 5 A valid convolution evaluates the kernel only where every tap falls inside the input. It uses no padding.
Definition 6 The stride \(S\) is the distance, in input positions, between the windows of consecutive outputs.
\[ K_{\text{eff}}=1+\delta(K-1) \]
Definition 7 The dilation \(\delta\) is the distance, in input positions, between adjacent taps of the kernel. An ordinary kernel has \(\delta=1\).
\[ \dout= \left\lfloor \frac{\din+2P-K_{\text{eff}}}{S} \right\rfloor+1, \qquad K_{\text{eff}}=1+\delta(K-1) \]
Seven inputs and a kernel of size three under four settings. Every output’s connections are drawn faintly; one output’s three are in colour, each carrying its weight. Padding adds zeros at the edges; stride spaces the outputs; dilation spaces the taps.
Definition 8 A channel is one array of values over the spatial positions. In a hidden convolutional layer, each output channel is often called a feature map because it records one learned feature across positions.
One 1D signal, top, and the three feature maps that three different kernels produce from it, on the same positions. Each row is one channel: the same signal, a different local question. The rising-edge map is positive where the signal climbs, the falling-edge map is its mirror, and the average is a smoothed copy.
\[ h_{ci}=\activation{\beta_c+\sum_{c'=1}^{C_i}\sum_{j=1}^{3}\omega_{c'cj}\,x_{c',i+j-2}} \]
\[ \layerweights\in\reals^{C_i\times C_o\times K}, \qquad \layerbias\in\reals^{C_o}, \qquad \text{parameters}=C_o(C_iK+1) \]
Definition 9 The receptive field of a hidden unit is the region of the original input that can influence it.
Three convolutional layers with kernel size three. The highlighted unit at the top and every unit that can influence it. Left, stride one: the receptive field grows by two per layer. Right, stride two: 3, 7, then 15 positions, because each layer’s taps sit twice as far apart as the last’s.
\[ r_L = 1 + L(K-1) \]
\[ \begin{aligned} r' &= r + (K-1)\delta J\\ J' &= SJ \end{aligned} \]
| Layer | \(r\) | \(J\) |
|---|---|---|
| input | 1 | 1 |
| 1 | 3 | 2 |
| 2 | 7 | 4 |
| 3 | 15 | 8 |
Our re-run of the chapter’s MNIST-1D comparison: SGD at learning rate 0.01, batch size 100, 100,000 steps, three seeds each. Left: both networks reach zero training error. Right: test error, where the convolutional network ends far lower.
\[ \begin{bmatrix} x_{i-1,j-1} & x_{i-1,j} & x_{i-1,j+1}\\ x_{i,j-1} & x_{i,j} & x_{i,j+1}\\ x_{i+1,j-1} & x_{i+1,j} & x_{i+1,j+1} \end{bmatrix} \]
\[ h_{ij}= \activation{ \beta+ \sum_{m=1}^{3}\sum_{n=1}^{3} \omega_{mn}x_{i+m-2,j+n-2} } \]
A real MNIST digit, left, and what three 3 x 3 kernels compute from it, each kernel’s weights drawn under its output. The edge kernels respond with opposite signs on the two sides of a stroke, in the same colours as the negative and positive weights that produce them. The averaging kernel blurs. MNIST, LeCun and Cortes, CC BY-SA 3.0.
\[ \layerweights\in\reals^{C_i\times C_o\times K\times K}, \qquad \layerbias\in\reals^{C_o} \]
\[ C_o(C_iK^2+1) \]
| Method | Operation on each local region |
|---|---|
| Stride-2 convolution | a learned weighted combination, evaluated at every second position |
| Max pooling | keep the largest value |
| Mean pooling | average the values |
Definition 10 Pooling with block size \(k\) replaces each non-overlapping \(k\times k\) block of a channel by one number, its maximum or its mean. Each channel is pooled separately, so its \(H\times W\) array becomes \(H/k\times W/k\), when \(k\) divides both, and the channel count is unchanged.
\[ \max \begin{bmatrix} 0.2 & 1.7\\ 0.9 & 0.4 \end{bmatrix} =1.7 \]
\[ \operatorname{mean} \begin{bmatrix} 0.2 & 1.7\\ 0.9 & 0.4 \end{bmatrix} =0.8 \]
A 4 x 4 feature map, then the same map shifted right by one and by two, each pooled over the 2 x 2 blocks outlined. Each block’s largest value is underlined. In the shifted rows, a faded pooled value is one the original already had: in the same cell after a shift of one, one cell to the right after a shift of two. Bold is new. Shifted by one, every peak stays in its block, so max pooling is unchanged while mean pooling changes. Shifted by two, the pooling step, both pooled maps move one cell right; only the column filled from the edge is new.
\[ \text{high resolution} \rightarrow \text{compressed representation} \rightarrow \text{high resolution} \]
Left: a stride-2 convolution with kernel size three and zero padding, from eight fine positions to four coarse ones. Right: its transpose, from four back to eight. The same eleven connections and the same sharing pattern, with the arrows reversed; the primed weights are the transposed layer’s own. A fine unit now receives two contributions or one, depending on where it sits.
\[ \vect{z}_{ij}=\layerbias+\layerweights\transpose\hidden_{ij}, \qquad \layerweights\in\reals^{C_i\times C_o} \]
| Task | Output | Meaning |
|---|---|---|
| Classification | one score per class | one class distribution for the image |
| Object detection | classes and boxes at many locations | what is where |
| Semantic segmentation | \(H\times W\) scores per class | one class distribution per pixel |
A schematic segmentation network’s tensor shapes, not the chapter’s network. Height is side length; width grows with channel count, floored so the thinnest stay visible. Three stride-2 convolutions halve the image three times while channels grow; three transposed convolutions restore it, ending in one score per class at every pixel.
| Quantity | Changes |
|---|---|
| Kernel size | taps: local extent, and parameters |
| Stride | output sampling density and spatial size |
| Dilation | spatial extent, without adding taps |
| Number of channels | number of learned response types |
\[ \dout=\left\lfloor\frac{\din+2P-K_{\text{eff}}}{S}\right\rfloor+1, \qquad K_{\text{eff}}=1+\delta(K-1) \]