Week 4 · principles
Definition 1 Regularisation is any mechanism that biases learning toward a subset of the fitting solutions, to improve generalisation.
Three functions that pass exactly through the same seven training points, so all three have zero training loss. Blue is the smallest-norm fit; the other two add a bump between two observations, where no training point can see it. Training loss cannot choose between them.
\[ \widehat{\params} = \argmin_{\params} \sum_{i=1}^{I}\exloss_i[\vect{x}_i,y_i]. \]
\[ \widehat{\params} = \argmin_{\params} \left[ \sum_{i=1}^{I}\exloss_i[\vect{x}_i,y_i] + \lambda g[\params] \right]. \]
The penalty is therefore not another measurement from the data
It is a modelling choice about which fitting solutions we would rather select when data fit alone does not decide
Left: a data-fitting loss over two parameters with three basins; the lowest is marked. Centre: an L2 penalty, smallest at the origin. Right: their sum. The penalty raises the far basin more than the near one, so the global minimum of the sum moves to a different basin.
Small parameter magnitude is a proxy, not a theorem about simple functions
Neural networks can be reparameterised, so the same function can sometimes be represented with different weight scales
Remember how Pythagoras combines movement along perpendicular axes?
For a parameter vector
\[ \params = \begin{bmatrix} \phi_1 & \phi_2 & \cdots & \phi_P \end{bmatrix}^{\top}, \]
a norm is a rule that reduces a whole vector to one non-negative measure of its size
The L2 norm is the familiar Euclidean choice: distance from the origin
\[ \norm{\params}_2 = \sqrt{\phi_1^2+\phi_2^2+\cdots+\phi_P^2}. \]
In two parameter dimensions, the L2 norm is ordinary Euclidean distance from the origin. The vector [3,4] has length 5 by Pythagoras. An L2 penalty therefore charges according to how far the whole parameter vector lies from zero, not according to any one coordinate alone.
Intuition: smaller weights can make the network less eager to produce sharp changes from small upstream differences
Regularisation normally uses
\[ \norm{\params}_2^2 = \sum_{j=1}^{P}\phi_j^2 \]
rather than the square root itself
So L2 regularisation says: fit the data, but pay an increasing cost for moving far from the origin in parameter space
Data term alone, minimised at \(\phi=3\):
\[ \loss_{\text{data}}(\phi)=(\phi-3)^2. \]
Add an L2 penalty with \(\lambda=1\):
\[ \loss_{\text{reg}}(\phi)=(\phi-3)^2+\phi^2. \]
The data-fitting loss is minimised at φ = 3, the L2 penalty at φ = 0, and their sum at φ = 1.5: the regularised optimum gives up some data fit for a smaller parameter.
\[ \frac{d\loss_{\text{reg}}}{d\phi} =2(\phi-3)+2\phi. \]
\[ 2(\phi-3)+2\phi=0 \quad\Longrightarrow\quad \widehat{\phi}=1.5. \]
Regularised loss:
\[ \loss_{\text{reg}}(\params) = \loss(\params)+\lambda g[\params] \]
Its gradient, which the optimiser follows:
\[ \nabla_{\params}\loss_{\text{reg}} = \nabla_{\params}\loss + \lambda\nabla_{\params}g[\params]. \]
For the penalty
\[ \lambda\norm{\params}_2^2, \]
the penalty gradient is
\[ 2\lambda\params. \]
Geometric reading: the L2 penalty adds a smooth bowl centred at the origin, and its gradient pulls the parameter vector inward
Definition 2 L2 regularisation adds \(\lambda\) times the sum of squared parameters to the data loss:
\[ \widehat{\params} = \argmin_{\params} \left[ \sum_{i=1}^{I}\exloss_i[\vect{x}_i,y_i] + \lambda\sum_j\phi_j^2 \right]. \]
\[ \lambda\sum_j\phi_j^2 = \frac{1}{2\sigma_\phi^2}\sum_j\phi_j^2 \quad\Longrightarrow\quad \lambda=\frac{1}{2\sigma_\phi^2} \]
A weight matrix is still just a collection of scalar weights arranged in rows and columns
The resulting matrix norm is the Frobenius norm:
\[ \norm{\layerweights_k}_F = \norm{\operatorname{vec}(\layerweights_k)}_2. \]
For \(\layerweights_k\in\reals^{\dimension_{k+1}\times\dimension_k}\),
\[ \norm{\layerweights_k}_F^2 = \sum_{r=1}^{\dimension_{k+1}} \sum_{c=1}^{\dimension_k} (\Omega_{k,rc})^2. \]
For example,
\[ \layerweights= \begin{bmatrix} 3&4\\ 0&0 \end{bmatrix} \quad\Longrightarrow\quad \norm{\layerweights}_F = \sqrt{3^2+4^2}=5. \]
The same twelve training points and the same fourteen-joint model, fitted with increasing L2 strength λ. With no penalty the fit chases the noise; moderate λ removes the sample-specific bends; large λ removes the underlying shape as well. The thin grey curve is the true mean.
The L2 objective above used a summed data loss:
\[ \loss_{\mathrm{sum}} = \sum_{i=1}^{I}\exloss_i + \lambda_{\mathrm{sum}}\,g[\params]. \]
CrossEntropyLoss takes the mean by default:\[ \loss_{\mathrm{mean}} = \frac{1}{I}\sum_{i=1}^{I}\exloss_i + \lambda_{\mathrm{mean}}\,g[\params]. \]
\[ \lambda_{\mathrm{mean}} = \frac{\lambda_{\mathrm{sum}}}{I}. \]
\[ \loss_{\text{reg}}(\params) = \loss(\params)+\lambda\norm{\params}^2 \]
Definition 3 Weight decay shrinks the parameters by a factor \(1-\lambda'\) in the update itself:
\[ \params \leftarrow (1-\lambda')\params - \alpha\nabla_{\params}\loss. \]
Gradient of the squared penalty:
\[ \nabla_{\params}\left(\lambda\norm{\params}^2\right) =2\lambda\params. \]
\[ \begin{aligned} \params_{t+1} &=\params_t- \alpha\left( \nabla_{\params}\loss +2\lambda\params_t \right)\\ &=(1-2\alpha\lambda)\params_t -\alpha\nabla_{\params}\loss. \end{aligned} \]
\[ \lambda'=2\alpha\lambda. \]
With the convention \(\loss_{\mathrm{reg}}=\loss+\lambda\norm{\params}^2\):
\[ \nabla_{\params}\loss_{\mathrm{reg}} = \nabla_{\params}\loss+2\lambda\params. \]
weight_decay=wd adds \(wd\,\params\) to the gradient\[ wd=2\lambda. \]
\[ \lambda'_t=\alpha_t\,wd \]
Two update rules for an adaptive optimiser. Top, coupled L2: the penalty’s gradient 2λφ joins the data gradient g before the adaptive rescaling, so each coordinate’s penalty is rescaled too. Bottom, decoupled weight decay (AdamW): only g is rescaled, and the shrinkage is applied to φ on a separate branch. Under plain SGD there is no rescaling and the two coincide.
The penalty defines what counts as an expensive parameter vector
Choosing a norm means choosing a geometry of preference: different penalties favour different kinds of solutions
Definition 4 Implicit regularisation is a systematic preference among equally good minima that comes from the training algorithm rather than from any term in the loss.
We want to isolate an effect that is easy to miss:
the optimiser can bias the solution even when we never write a regularisation term
Gradient flow is not a better training algorithm here
It is a mathematical baseline for asking what finite step size changes
\[ \frac{d\params}{dt} = -\nabla_{\params}\loss(\params). \]
\[ \params_{t+1} = \params_t - \alpha\nabla_{\params}\loss(\params_t). \]
The same loss surface and the same starting point, shown separately so the trajectories are easy to compare. Left: continuous gradient flow follows a smooth curved path to one point on the valley of global minima. Right: finite-step gradient descent follows straight chords based on the gradient at the start of each step and reaches a different point on the same zero-loss valley, where the valley is wider.
\[ \widetilde{\loss}_{\mathrm{GD}}(\params) = \loss(\params) + \frac{\alpha}{4} \norm{\nabla_{\params}\loss(\params)}^2. \]
\[ \frac{\alpha}{4} \norm{\nabla_{\params}\loss}^2 \]
\[ \norm{\nabla_{\params}\loss}_2 = \sqrt{ \sum_j \left(\frac{\partial\loss}{\partial\phi_j}\right)^2 } \]
Our rerun of the chapter’s experiment: MNIST-1D, a network with two hidden layers of 100 units, SGD with batches of 100. Each run takes 6000/LR steps, so every run has the same opportunity to travel; every run fits the training set exactly. Dots are single runs from three seeds, the line their mean. Larger learning rates reach lower test error.
\[ \loss = \frac{1}{I} \sum_{i=1}^{I}\exloss_i, \qquad \loss_b = \frac{1}{|\mathcal{B}_b|} \sum_{i\in\mathcal{B}_b}\exloss_i. \]
Two pairs of minibatch gradients along one parameter. Top: both batches say 1, so they agree exactly. Bottom: one says 2 and the other 0. Both pairs average to the same full gradient, 1 (white), so the mean alone cannot tell the two cases apart.
The full-data gradient keeps only the average direction
That average can hide whether minibatches:
Variance answers the missing question:
how far do the individual batch gradients spread around their mean?
A small variance means the batches tell a similar local story; a large variance means different subsets of the data are pulling the parameters differently
\[ \text{A: }\ \frac{(1-1)^2+(1-1)^2}{2}=0, \qquad \text{B: }\ \frac{(2-1)^2+(0-1)^2}{2}=1. \]
\[ \nabla\loss = \frac{1}{B}\sum_{b=1}^{B}\nabla\loss_b. \]
\[ \underbrace{\loss}_{\text{fit the data}} + \underbrace{\frac{\alpha}{4}\norm{\nabla\loss}^2}_{\text{avoid steep regions}} + \underbrace{\frac{\alpha}{4B}\sum_b\norm{\nabla\loss_b-\nabla\loss}^2}_{\text{prefer batch agreement}} \]
The pieces of the modified SGD objective, over the same two parameters, for two minibatches whose losses are the same three basins shifted slightly apart. The mean loss; the gradient-norm term (α/4)‖∇L‖²; the disagreement term, largest on the basin floors where the curvature is highest and the two shifted batches pull furthest apart; and their sum. Unlike the gradient-norm term, the disagreement term is not zero at a minimum, so the minimum of the sum moves.
Our rerun of the chapter’s experiment: the same network and data, SGD at learning rate 0.05, with each run stopped at the first epoch where it fits the training set exactly, so every run is compared at the point of memorisation. Dots are single runs from three seeds, the line their mean. The smallest batch reaches the lowest test error; from 25 upwards, the steps between means are smaller than the spread between seeds.
Definition 5 Early stopping ends optimisation before full convergence and keeps the parameters from the checkpoint with the best validation performance.
Full-batch gradient descent on the fixed-joint model from last Friday, forty joints and fourteen training points, started at zero. Training loss keeps falling towards zero; validation loss, on fresh points from the same process, reaches its minimum and then rises. Early stopping keeps the checkpoint at that minimum.
The same run at four moments: at initialisation, early, at the validation minimum, and at the final step. The coarse shape of the true mean (thin grey) arrives first; the late steps bend the fit through individual noisy points while training loss keeps falling.
\[ 4.6, \qquad 5.4, \qquad 5.0 \]
\[ \frac{4.6+5.4+5.0}{3}=5.0. \]
Three predictions of a target of 5. Top: the errors fall on both sides and the mean lands on the target. Bottom: every error has the same sign, and the mean inherits it.
Definition 6 An ensemble is a group of models whose predictions are combined into one prediction.
Definition 7 Bootstrap aggregating, or bagging, trains each model of an ensemble on its own resample of the training data, drawn with replacement, and combines their predictions.
Bagging on a toy regression with one unusual point (circled). The first three panels are bootstrap samples: marker size shows how many times each point was drawn, and hollow markers were not drawn at all. The last panel overlays all forty bootstrap fits, their average, and the single fit to the full data; at the unusual point the average is pulled less far from the true mean.
A network can rely on fragile partnerships between units:
A fixed smaller network would reduce capacity, but it would not create dropout’s changing-subnetwork pressure
A changing random subnetwork repeatedly breaks those partnerships, forcing useful behaviour to survive under many combinations of active units
Definition 8 Dropout clamps a random subset of hidden units to zero at each training iteration, drawing a fresh subset every time.
p means the same\[ \vect{h} = \begin{bmatrix} 1.2\\0.4\\2.0\\0.7 \end{bmatrix}, \qquad \vect{m} = \begin{bmatrix} 1\\0\\1\\0 \end{bmatrix} \]
\[ \widetilde{\vect{h}} = \vect{m}\odot\vect{h} = \begin{bmatrix} 1.2\\0\\2.0\\0 \end{bmatrix}. \]
One network, four training iterations. Each iteration samples a fresh mask; dropped hidden units (hollow) take their incoming and outgoing connections with them. Every subnetwork uses the same underlying weights, so each update trains part of one shared model.
Left: a fit whose three hidden units at the dashed joints build a bump in a gap between the training points; they cancel exactly everywhere else, so the bump costs no training loss. Centre: drop one of the three and the cancellation fails, and the function moves far from the data. Right: after training with dropout, the bump’s weights have shrunk and the gap is spanned smoothly, at the cost of a looser fit to the data everywhere. The same training without dropout leaves the bump in place.
\[ \widetilde{\vect{h}}=\vect{m}\odot\vect{h}, \qquad \expect{\widetilde h_j}=(1-p)h_j \]
Definition 9 Inverted dropout divides the kept activations by \(1-p\) during training, so evaluation needs no rescaling.
\[ \widetilde{\vect{h}} = \frac{\vect{m}\odot\vect{h}}{1-p}, \qquad \expect{\widetilde{\vect{h}}}=\vect{h} \]
nn.Dropout is the identity operationThe same activation h and drop probability p under both conventions. Top, inference scaling: training only masks, so a unit’s expected output is (1 − p)h, and evaluation multiplies by 1 − p to match. Bottom, inverted dropout, as PyTorch implements it: training masks and divides by 1 − p, so the expectation is already h and evaluation is the identity.
\[ \expect{\widetilde h} =0.5(0)+0.5(4)=2=h. \]
\[ \widetilde{\vect{x}} = \vect{x}+\noise. \]
\[ \mat{J}_{\vect{x}} = \frac{\partial \vmodel{\vect{x}}}{\partial \vect{x}} \in\reals^{\dout\times\din}. \]
\[ \vmodel{\vect{x}+\delta\vect{x}} \approx \vmodel{\vect{x}} + \mat{J}_{\vect{x}}\,\delta\vect{x}. \]
The same twelve observations and the same thirty-joint model, fitted to many perturbed copies of each input. The faint dots are some of those copies. With no noise the fit interpolates; as the noise standard deviation grows, the fitted function gets smoother.
\[ \widetilde{\params} = \params+\noise. \]
Two minima with the same loss. The same parameter perturbation ε raises the loss sharply in the narrow basin and only a little in the wide one. The link from wide minima to generalisation is an intuition with supporting evidence, not a theorem.
Definition 10 Label smoothing replaces the one-hot target with a distribution that keeps most of the probability mass on the labelled class and spreads the rest over the other classes.
A one-hot target says
\[ [0,0,1,0,0] \]
and cross-entropy keeps rewarding a model for pushing the labelled class closer and closer to probability \(1\)
The aim is not to make the model uncertain at random
It is to stop the target from demanding ever larger logit separation after the class is already correct
\[ \widetilde y_c=1-\rho, \qquad \widetilde y_k=\frac{\rho}{K-1} \quad(k\neq c). \]
\[ [0,0,1,0,0] \longrightarrow [0.05,0.05,0.80,0.05,0.05]. \]
\[ \exloss_i = -\sum_{k=1}^{K} \widetilde y_k\log \hat{y}_k. \]
CrossEntropyLoss(label_smoothing=eps), with \(\varepsilon\) = eps
\[ [0,0,1,0,0] \longrightarrow [0.04,0.04,0.84,0.04,0.04]. \]
The target for a five-class example whose true class is 3. Left: one-hot. Centre: the chapter’s smoothing with ρ = 0.2, which spreads ρ over the four wrong classes. Right: PyTorch’s label_smoothing = 0.2, which mixes in a uniform distribution over all five, so the true class keeps more mass.
Bayesian inference keeps a distribution over plausible parameter vectors instead of collapsing immediately to one.
From the loss-function session, the likelihood asks
\[ p(\set{D}\mid\params): \]
if these were the parameters, how compatible would the observed data be?
The likelihood answers a forward question: parameters \(\rightarrow\) possible data
For inference, the question we actually want is
\[ p(\params\mid\set{D}): \]
after seeing these data, how plausible is each parameter vector?
Bayes’ rule tells us how to reverse that conditioning while accounting for what was plausible before the data arrived.
Suppose only two candidate models are possible
| candidate | prior | likelihood of observed data | prior \(\times\) likelihood |
|---|---|---|---|
| A | \(0.75\) | \(0.20\) | \(0.15\) |
| B | \(0.25\) | \(0.80\) | \(0.20\) |
\[ \Pr(A\mid\set{D})\approx0.43, \qquad \Pr(B\mid\set{D})\approx0.57. \]
Bayes updates relative plausibility: prior \(\times\) likelihood, then renormalise
Remember conditional probability:
\[ \Pr(A\mid B) = \frac{\Pr(B\mid A)\Pr(A)}{\Pr(B)}. \]
For parameters and data, read the same structure as
\[ \underbrace{p(\params\mid\set{D})}_{\text{posterior}} \propto \underbrace{p(\set{D}\mid\params)}_{\text{likelihood}} \underbrace{p(\params)}_{\text{prior}}. \]
One scalar parameter φ: the mean of four noisy observations (ticks). The prior is a Gaussian centred on 0; the likelihood peaks near the observations and is rescaled here to share the axis, since it is not a density in φ. Their product, normalised, is the posterior: between the two, and narrower than the prior.
For the full parameter vector \(\params\) and training data
\[ \set{D}=\{(\vect{x}_i,y_i)\}_{i=1}^{I}, \]
the posterior is
\[ p(\params\mid\set{D}) = \frac{ \left[\prod_{i=1}^{I}\Pr(y_i\mid\vect{x}_i,\params)\right]p(\params) }{ \int \left[\prod_{i=1}^{I}\Pr(y_i\mid\vect{x}_i,\params')\right]p(\params')\,d\params' }. \]
Sometimes we still want one parameter vector rather than a full distribution
Definition 11 The maximum a posteriori (MAP) estimate is the parameter vector at the mode of the posterior.
\[ \widehat{\params}_{\mathrm{MAP}} = \argmax_{\params} \left[ p(\params) \prod_{i=1}^{I}\Pr(y_i\mid\vect{x}_i,\params) \right]. \]
Earlier, L2 was introduced geometrically:
prefer a shorter weight vector when data fit is comparable
Now ask the Bayesian version of the same modelling choice:
what prior would express a soft, symmetric preference for weights near zero?
An independent zero-centred Gaussian prior does exactly that
\[ \phi_j\sim\normal(0,\sigma_\phi^2). \]
A zero-centred Gaussian prior and the penalty it implies. Left: smaller prior variance concentrates more density near zero. Right: taking the negative log turns each Gaussian into a quadratic bowl; the narrower prior gives the steeper penalty.
For independent Gaussian weight priors,
\[ p(\params) \propto \exp\left[ -\frac{1}{2\sigma_\phi^2} \sum_j\phi_j^2 \right]. \]
Take the negative log of the MAP objective:
\[ -\log p(\set{D}\mid\params) + \frac{1}{2\sigma_\phi^2}\sum_j\phi_j^2 + \text{constant}. \]
Compare with the L2 objective from earlier:
\[ \loss(\params)+\lambda\sum_j\phi_j^2. \]
\[ \boxed{\lambda=\frac{1}{2\sigma_\phi^2}} \]
Geometric reading
Probabilistic reading
A smaller prior variance means a stronger belief in weights near zero, and therefore a larger L2 coefficient \(\lambda\)
This connection reinterprets L2; it is not needed in order to use L2 regularisation
Suppose two parameter settings both remain plausible after training, but they disagree on a new input
The Bayesian rule is: ask every plausible parameter vector for its prediction, then weight that prediction by its posterior plausibility
For a new input \(\vect{x}\),
\[ \Pr(y\mid\vect{x},\set{D}) = \int \Pr(y\mid\vect{x},\params) \,p(\params\mid\set{D}) \,d\params. \]
Definition 12 The posterior predictive distribution averages the prediction from every parameter vector, weighted by its posterior density.
Bayesian inference for the fixed-joint model with a Gaussian prior on its weights, computed in closed form. Thin lines are the functions given by weight vectors drawn from the posterior; the thick line is the predictive mean and the band two predictive standard deviations. Left, a broad prior: the functions follow the data and spread furthest at the right-hand end. Right, a narrow prior: the functions are held to a smoother trend, agree with one another, and miss the peak.
One visual grammar for three ways to borrow training signal: a shared representation (blue) with task-specific heads. Transfer learning trains on a secondary task, then replaces the head. Multi-task learning trains several heads at once. Generative self-supervision masks part of an unlabelled input and trains the network to fill it in.
Condition: useful only when the correct target stays unchanged
Pick a task: name one transformation that would break its label
Failure cases:
Augmentation asserts an invariance; it does not create truth.
One MNIST digit and six transformed copies. Each copy is a new training input with the same label, 7; together they tell the network which changes should not move its prediction. MNIST © Yann LeCun and Corinna Cortes, CC BY-SA 3.0.
Four broad regularisation mechanisms and representative methods. Methods highlighted in cyan act through more than one mechanism: input noise, label smoothing, ensembling and the Bayesian approach through two, dropout through three.
| observed concern | plausible intervention |
|---|---|
| highly sample-sensitive fit | L2, early stopping, bagging |
| brittle hidden-unit dependencies | dropout |
| sensitivity to small input changes | input noise, adversarial training |
| overconfident class probabilities | label smoothing |
| limited labelled target data | transfer, self-supervision, augmentation |
| uncertainty across plausible fitted models | ensemble, Bayesian treatment |
# L2-style penalty under SGD
optimizer = torch.optim.SGD(
weight_params,
lr=lr,
weight_decay=wd,
)
# Decoupled weight decay
optimizer = torch.optim.AdamW(
weight_params,
lr=lr,
weight_decay=wd,
)
# Inverted dropout: identity at evaluation time
regulariser = torch.nn.Dropout(p=0.5)
# Framework-specific smoothing convention
loss_fn = torch.nn.CrossEntropyLoss(label_smoothing=0.1)| method | what changes? |
|---|---|
| L2 | objective / parameter penalty |
| finite-step GD / SGD | optimisation trajectory |
| early stopping | optimisation duration |
| ensemble | prediction rule across fitted models |
| dropout | stochastic hidden activations |
| input / weight noise | training computation |
| label smoothing | target distribution |
| Bayesian inference | treatment of parameter uncertainty |
| transfer / self-supervision | information available to fitting |
| augmentation | training examples and assumed invariances |
| distinction | do not conflate |
|---|---|
| explicit vs implicit | a written penalty in the objective vs solution-selection bias introduced by optimisation |
| L2 vs weight decay | squared-parameter penalty vs direct parameter shrinkage; equivalent only for particular update rules |
| training vs inference dropout scaling | compensate at inference vs inverted dropout during training; use one convention, not both |
| label-smoothing coefficients | the same nominal coefficient can define different target distributions |
| regularisation vs validation | regularisation supplies candidate biases; validation chooses among their strengths and variants |