Attention
The attention mechanism is a type of transformation of the inputs to a neural network (transformation takes the inputs and creates anew set of inputs of the same dimension as the orginal set) and produces a new set.The idea of attention is to transform an input token by combining it with other similar tokens. That is, we compute a set of similarity scores, here labeled {% a_{ij} %} that represent the similarity of {% token_{i} %} to {% token_j %}. Then, the new set of tokens are computed by summing these together.
{% \vec{y}_i = \sum a_{ij}\vec{x}_{ij} %}
Self Attention
IN simple self attention, the similarity scores are computed simply as the dot product of one token with another.
{% a_{ij} = \frac{exp(\vec{x}_i^T \vec{x}_j)}{\sum_p exp(\vec{x}_i^T \vec{x}_p )} %}
which can be summarized as
{% Y = Softmax[XX^T]X %}
Projection
The self attention model can be extended to include an additional set of weights as follows. Define the following three matrices
{% Q = XW^{q} %}
{% K = XW^{k} %}
{% V = XW^{v} %}
and then compute the attention as
{% Y = Softmax[\frac{QK^T}{\sqrt{D_k}}] V %}
That is, the set of inputs is transformed linearly prior to running the self attention. The gemoetric interpretation of
this linear transformation is projecting the input tokens into a (typically low dimensional) space prior to computing the
dot product similarity scores.
Positional Encoding
The algorithm above does not utilize the position of a token when computing the new set of tokens. That is {% \sum a_{ij}x_j %} does not consider position. However, position is often an important consideration in sequence data. As such, it is often useful to encode position into the input tokens.A simple way to do this is append an element onto each vector that containes the token position. While this coud be just an integer index, it is probably better to normalize all the position indices by dividing by the total number of positions. (i.e. converting a position to a float)
For large models, in particular, large language models this often represents an unacceptable increase in the resources needed to process the model. The typical way to encode position is simply to calculate a position encoding and add it back to the original input token.