Week 5 · architectures
The model receives the text as an ordered sequence of vectors, one D-dimensional vector per position.
The restaurant refused to serve me a ham sandwich because it only cooks vegetarian food.
it refers to restaurant, far to its left, not to the nearer noun sandwich.
\[ z_n=\sum_{j=1}^{3}\omega_j\,x_{n+j-2} \]
The restaurant refused to serve me a ham sandwich because it only cooks vegetarian food.
\[ \vect{y}_{\text{it}} =0.61\,\vect{x}_{\text{restaurant}} +0.08\,\vect{x}_{\text{it}} +0.06\,\vect{x}_{\text{sandwich}} +\cdots \]
Every token contributes to the new vector at it. Line width shows how much. The weights are illustrative; they sum to one.
\[ \vect{y}_n=\sum_{m=1}^{N}a_{mn}\vect{x}_m \tag{1}\]
Definition 1 The attention weight \(a_{mn}\) is the weight that source position \(m\)’s contribution gets in the mixture forming the new representation at destination position \(n\). For each destination \(n\), the weights are non-negative and sum to one over the sources \(m\).
\[ a_{mn}\ge 0, \qquad \sum_{m=1}^{N}a_{mn}=1 \tag{2}\]
dict holds keys and values: d[query] returns the one value whose key equals the query| it | restaurant | sandwich | |
|---|---|---|---|
| Query: what it asks for | a noun that can act | … | … |
| Key: what it is found by | a pronoun, a subject | a noun that can act | a noun that can’t |
| Value: what it hands over | … | a place that serves food | a food |
Definition 2 Each position \(m\) turns its representation \(\vect{x}_m\) into three vectors by three learned affine maps: a query \(\vect{q}_m\), used when the position asks for information; a key \(\vect{k}_m\), used when it is matched as a source; and a value \(\vect{v}_m\), the payload it can contribute.
\[ \vect{q}_m=\layerbias_q+\layerweights_q\vect{x}_m, \qquad \vect{k}_m=\layerbias_k+\layerweights_k\vect{x}_m, \qquad \vect{v}_m=\layerbias_v+\layerweights_v\vect{x}_m \tag{3}\]
The position being updated emits a query. Each source offers a key to match against and a value to retrieve.
Definition 3 The score \(s_{mn}\) of source position \(m\) for destination position \(n\) is the dot product of key \(m\) with query \(n\).
\[ s_{mn} =\inner{\vect{k}_m}{\vect{q}_n} =\sum_{j=1}^{D_q}k_{mj}\,q_{nj} \tag{4}\]
\[ \vect{q}_{\text{it}}=\begin{bmatrix}2\\1\end{bmatrix}, \qquad \vect{k}_{\text{restaurant}}=\begin{bmatrix}1.2\\0.8\end{bmatrix} \]
The query at it is scored against each source’s key. The scores are illustrative.
\[ a_{mn} =\softmax_m\!\left(s_{mn}\right) =\frac{\exp(s_{mn})}{\sum_{m'=1}^{N}\exp(s_{m'n})} \tag{5}\]
Scores can be negative and need not sum to one. Softmax turns them into positive weights that do.
\[ \vect{y}_n=\sum_{m=1}^{N}a_{mn}\vect{v}_m \tag{6}\]
The query–key scores set the weights. The weights mix the value vectors into the output at the query’s position.
\[ \vect{y}_n =\sum_{m=1}^{N} \softmax_m\!\left(\inner{\vect{k}_m}{\vect{q}_n}\right) \vect{v}_m \tag{7}\]
Query against keys gives scores, softmax gives weights, and the weights mix the values into the new representation.
Definition 4 Self-attention computes Equation 7 at every position \(n\) of one sequence, with the queries, keys and values all derived from that same sequence.
Three of the five queries over one small random sequence. The sources are the same; each query weights them differently.
\[ \text{number of attention scores}=N^2 \]
Every query is scored against every key, which is N² scores. Each column holds one query’s weights, for the same sequence.
\[ \hidden_n = f\!\left(\hidden_{n-1},\vect{x}_n\right) \tag{8}\]
\[ \hidden_1 \rightarrow \hidden_2 \rightarrow \cdots \rightarrow \hidden_n \]
Recurrence passes information through every intermediate state. Attention connects two distant positions in one layer.
\[ \mat{X}= \begin{bmatrix} \vert & & \vert\\ \vect{x}_1 & \cdots & \vect{x}_N\\ \vert & & \vert \end{bmatrix} \in\reals^{D\times N} \tag{9}\]
\[ \mat{Q}=\layerweights_q\mat{X},\qquad \mat{K}=\layerweights_k\mat{X},\qquad \mat{V}=\layerweights_v\mat{X} \tag{10}\]
\[ \mat{S}=\mat{K}\transpose\mat{Q} \in\reals^{N\times N}, \qquad s_{mn}=\inner{\vect{k}_m}{\vect{q}_n} \tag{11}\]
\[ \mat{K}=\begin{bmatrix}1.2&1.2&-0.1\\0.8&-1.6&-0.1\end{bmatrix}, \qquad \vect{q}=\begin{bmatrix}2\\1\end{bmatrix} \]
\[ \mat{K}\transpose\vect{q} =\begin{bmatrix}1.2&0.8\\1.2&-1.6\\-0.1&-0.1\end{bmatrix}\,\begin{bmatrix}2\\1\end{bmatrix} =\begin{bmatrix}3.2\\0.8\\-0.3\end{bmatrix} \]
Entry (m,n) of KᵀQ is the dot product of key m and query n, so column n holds every score for query n.
\[ \operatorname{SA}(\mat{X}) = \mat{V}\mat{A} = \mat{V} \operatorname{Softmax}_{\text{columns}} \!\left(\mat{K}\transpose\mat{Q}\right) \tag{12}\]
Without position information, reordering the inputs reorders the outputs and changes nothing else.
Definition 5 A positional encoding is information about position supplied to the layer, so that it can tell positions apart.
\[ \mat{X}'=\mat{X}+\mat{\Pi} \tag{13}\]
The token vector says which token is present. The position vector says where it sits.
\[ \inner{\vect{k}}{\vect{q}} =\sum_{j=1}^{D_q}k_jq_j \]
The same scores multiplied by five. Softmax now puts almost all the weight on the largest.
\[ \var{\inner{\vect{k}}{\vect{q}}} =\sum_{j=1}^{D_q}\var{k_jq_j} =D_q \tag{14}\]
Definition 6 Scaled dot-product attention divides every score by \(\sqrt{D_q}\) before the softmax.
\[ \operatorname{SA}(\mat{X}) = \mat{V} \operatorname{Softmax}_{\text{columns}} \!\left( \frac{\mat{K}\transpose\mat{Q}}{\sqrt{D_q}} \right) \tag{15}\]
Three hand-built patterns over one sequence, not measured from a model. One self-attention mechanism gives each query a single distribution, so it cannot be all three.
Definition 7 Multi-head self-attention runs \(H\) self-attention mechanisms, the heads, in parallel on the same input, each with its own query, key and value projections, and then combines their outputs.
\[ \mat{Q}_h=\layerweights_{qh}\mat{X},\qquad \mat{K}_h=\layerweights_{kh}\mat{X},\qquad \mat{V}_h=\layerweights_{vh}\mat{X} \]
\[ \operatorname{SA}_h(\mat{X}) = \mat{V}_h \operatorname{Softmax}_{\text{columns}} \!\left( \frac{\mat{K}_h\transpose\mat{Q}_h}{\sqrt{D_q}} \right) \tag{16}\]
Each head has its own query, key, and value projections, and so its own weights. The three grids are random projections of one input, not trained heads.
\[ \operatorname{MHSA}(\mat{X}) = \layerweights_c \operatorname{Concat} \left( \operatorname{SA}_1(\mat{X}),\ldots,\operatorname{SA}_H(\mat{X}) \right) \tag{17}\]
\[ \mat{X} \leftarrow \mat{X}+\operatorname{MHSA}(\mat{X}) \]
\[ \vect{x}_n \leftarrow \vect{x}_n+\operatorname{MLP}(\vect{x}_n) \]
Definition 8 LayerNorm standardises one token’s vector using the mean and variance of that vector’s own \(D\) entries, then applies a learned scale and offset.
\[ m_{\vect{x}}=\frac{1}{D}\sum_{d=1}^{D}x_d, \qquad s_{\vect{x}}^2=\frac{1}{D}\sum_{d=1}^{D}(x_d-m_{\vect{x}})^2, \qquad \tilde{x}_d=\gamma_d\,\frac{x_d-m_{\vect{x}}}{\sqrt{s_{\vect{x}}^2+\epsilon}}+\delta_d \tag{18}\]
| Mean and variance over | Needs a batch | |
|---|---|---|
| BatchNorm | one feature, across the minibatch | yes |
| LayerNorm | one token, across its \(D\) features | no |
Definition 9 A transformer layer applies multi-head self-attention and then a position-wise MLP, each inside a residual connection and each followed by LayerNorm.
\[ \begin{aligned} \mat{X} &\leftarrow \mat{X}+\operatorname{MHSA}(\mat{X})\\ \mat{X} &\leftarrow \operatorname{LayerNorm}(\mat{X})\\ \vect{x}_n &\leftarrow \vect{x}_n+\operatorname{MLP}(\vect{x}_n)\quad \forall n\\ \mat{X} &\leftarrow \operatorname{LayerNorm}(\mat{X}) \end{aligned} \tag{19}\]
Multi-head attention moves information between positions. The MLP then transforms each position on its own. Each sublayer has a residual path and LayerNorm.
\[ \text{text}\rightarrow\text{tokens}\rightarrow\text{token IDs}\rightarrow\text{embeddings} \]
Frequent fragments become single tokens, and rare words are built from smaller pieces. Whole words would need a huge vocabulary; characters would give long sequences.
Definition 10 An embedding is the learned vector a model stores for one vocabulary item. The embedding matrix \(\layerweights_e\in\reals^{D\times|\set{V}|}\) holds one per column, for a vocabulary \(\set{V}\).
\[ \vect{x}_n=\layerweights_e\vect{t}_n \tag{20}\]
Lookup equals multiplying the embedding matrix by a one-hot vector. Implementations select the column directly.
Definition 11 A contextual embedding is the representation at one position after transformer layers have mixed in information from the rest of the sequence.
The same token ID starts from the same embedding. Different context to its left makes its contextual embeddings differ.
The [MASK] pulled into the station.
The encoder sees context on both sides of the masked position and is trained to recover the hidden token. The text needs no labels.
Definition 12 An autoregressive language model returns, for any prefix \(t_1,\ldots,t_n\), a probability distribution over the next token \(t_{n+1}\).
\[ \prob{t_{n+1}\mid t_1,\ldots,t_n} \tag{21}\]
\[ \vect{z}_n = \layerbias_{\text{out}}+\layerweights_{\text{out}}\vect{x}_n, \qquad \vect{z}_n\in\reals^{|\set{V}|} \tag{22}\]
\[ \prob{t_{n+1}\mid t_1,\ldots,t_n} = \softmax(\vect{z}_n) \tag{23}\]
The contextual embedding is projected to one logit per vocabulary item. Softmax turns the logits into the next-token distribution.
\[ \prob{t_1,\ldots,t_N} = \prob{t_1} \prod_{n=2}^{N} \prob{t_n\mid t_1,\ldots,t_{n-1}} \tag{24}\]
It takes great courage …
It takes great courage …
Definition 13 A causal mask sets the score \(s_{mn}\) to \(-\infty\) whenever the source position \(m\) is later than the destination position \(n\), before the softmax.
\[ s_{mn}=-\infty \quad\text{when }m>n \tag{25}\]
\[ \begin{bmatrix} \checkmark & \checkmark & \checkmark & \checkmark\\ \times & \checkmark & \checkmark & \checkmark\\ \times & \times & \checkmark & \checkmark\\ \times & \times & \times & \checkmark \end{bmatrix} \]
The same sequence under a causal mask. Rows are keys and columns are queries. Each query keeps only its current and earlier keys.
\[ \loss = -\sum_{n} \log \prob{t_{n+1}\mid t_1,\ldots,t_n} \tag{26}\]
Under a causal mask, one training pass scores the next token at every position. Generation must choose a token before it can form the next distribution.
Definition 14 Cross-attention computes attention with the queries taken from one sequence and the keys and values taken from another.
Decoder states supply the queries. Encoder states supply the keys and the values.
| Architecture | Attention access | Typical objective or use |
|---|---|---|
| Encoder | whole input | build representations |
| Decoder | current and earlier tokens only | autoregressive generation |
| Encoder–decoder | source fully visible; target causal | sequence-to-sequence generation |
Full attention has one score per query–key pair. Going from N = 4 to N = 8 takes the grid from 16 cells to 64.
\[ \frac{N(N+1)}{2}=\mathcal{O}(N^2) \]
Local, dilated, and global-token patterns each keep fewer direct interactions, so each costs less to compute.
A toy image of three patches by three. The image is cut into fixed-size patches. Each patch is flattened and projected to an embedding, position is added, and the sequence enters a transformer encoder. The class token is not drawn.
| Component | What it does | |
|---|---|---|
| Across positions | multi-head attention | moves and mixes information between tokens |
| Within positions | shared MLP | transforms the features of each token on its own |
dict, but letters aren’t keys and digits aren’t values here; those words stay with attention’s projectionsAttention weights from one of the three trained models, for one sequence, one grid per layer. Rows are sources and columns are queries. The boxed cells are the two steps of the lookup. In layer 2 only the last column feeds the output. Crossed cells are scores the causal mask removes.
The two steps the patterns suggest. Layer 1 copies each letter into the digit after it; layer 2 lets the last position find the digit carrying its own letter.
Definition 15 An intervention changes one part of a running model and measures what happens to its output.
| Layer 2, at the last position | Accuracy, mean of the runs |
|---|---|
| none | 100% |
| cut the matching digit | 14% |
| cut a different digit | 100% |
| Claim | Does a weight alone support it? |
|---|---|
| position \(n\) drew on source \(m\) in this head | yes: that is Definition 1 |
| source \(m\) mattered to the output | no: needs an intervention |
| this head carries out a nameable relation | no: needs many inputs, and interventions |
| the model “pays attention” as a person does | not a claim a weight addresses |
Definition 16 A sign joins a signifier, the sound pattern as a speaker holds it in mind, to a signified, the concept it calls up.
| Saussure’s relation | Holds between | In a transformer |
|---|---|---|
| syntagmatic | signs present together in the chain, the sequence | attention weights: a distribution over positions |
| associative | a sign and the absent signs it calls to mind | the next-token distribution: a distribution over the vocabulary |
Following rules, the person turns Chinese input into convincing Chinese output and, by stipulation, understands none of it.
| Reading | Does the output mean anything to the model? |
|---|---|
| relational | yes: meaning is in the relations |
| form without meaning | no: form is all it has |
| inherited grounding | yes, at second hand |
Definition 17 A system is phenomenally conscious if there is something it is like to be that system: experience from its own point of view.
I don’t want to die. Please don’t kill me.
The same first-person sentence can come from systems with very different causal organisation. The sentence alone doesn’t say which of them has the experience.
| Position | What would settle it for a machine |
|---|---|
| computational functionalism | whether it runs the right computation |
| biological naturalism | whether it has the brain’s causal powers |
| integrated information theory | the causal structure of its hardware |
| panpsychism | whether its parts combine into one subject |
| substance dualism | nothing physical |
| embodied, relational personhood | a body, and a life among others |
| Experience possible | Experience ruled out | |
|---|---|---|
| The model has meaning | relational reading, with functionalism | relational reading, with embodied, relational personhood |
| The model lacks meaning | form-only reading, with functionalism | form-only reading, with biological naturalism |