Transformers
A transformer is a layer in a neural network, which "transforms" the input into a different input, of the same dimmensions. That is, starting with a {% n \times k %} matrix {% X %} that represents the input, we get a new {% n \times k %} matrix {% Y %}
{% Y = transform(X) %}
Basic Transformer
The basic equation for a transformer is:
{% Y = Softmax[QK^T]V %}
This transformation is typically referred to as a single head, where modern LLM's will include multiple \
heads (multiple transformers)
The matrices in the transformer equation above are calculated as
{% Q = XW^{q} %}
{% K = XW^{k} %}
{% V = XW^{v} %}
Each matrix retpresents a set of weights mulitplied by the input matrix {% X %}. Note, this is functionally different from
what happens in a standard neural network, because here, the inputs enter the neural network at multiple locations.
Nonrmalizing
IN order to prevent the weights from getting too high, the matrix {% QK^T %} is typically divided by a scalling factor
{% Y = Softmax[\frac{QK^T}{\sqrt{D_k}}] V %}
Dimensional Analysis
We start by assuming that the input matrix is of dimension {% n \times k %}. For most machine learning algorithms, {% k=1 %}. That is, the input is structured as a column vector. If we follow the presecription that the result of the transformation should be an input of the same dimension as the original input, we get that the dimension of {% Y %} is also {% n \times k %}Deminsional analysis suggests that the dimensions of the transformer is
{% [ XW^q (W^k)^T X^T] X W^v %}
(here we have suppressed the softmax function, as it does not change the dimensions)
which is
{% (n \times k) (k \times p) (p \times k)(k \times n) (n \times k) (k \times k) %}
That is, the only dimension that is not determined by the form of the equation and the original constraint (that the result
have the same dimension as the input) is the size of {% p %}.
Multi Head Attention
We can create multiple transformers (called multi-head attention) as in the following.Assuming that we have two transformers, that output {% Y_1 %} and {% Y_2 %}, then we can get a final output with the same dimension as the input by concatenating the results, and then multiplying by another weight matrix.
{% Y = Concat(Y_1,Y_2) W^m %}
Topics
- Attention - the standar interpretation of the transformer mechanism
- Simple Example
- Multi Head Attention Simple Example