Перейти к содержимому

s01-p05: Positional Encoding & Relative Position Information | CS4CV

Computer Science for Computer Vision

0:00 / 0:00

s01-p05: Positional Encoding & Relative Position Information | CS4CV

8 просмотров · 8 дней назад
Computer Science for Computer Vision
9 подписчиков
8 просмотров · 8 дней назад
Course: Computer Science for Computer Vision Channel: @CS4CV (   / @cs4cv  ) Playlist: Section 1 – Attention Mechanisms & Transformers (   • CS4CV Section 01: Attention Mechanisms & T...  ) In the fifth video of Section 1 – Attention Mechanisms & Transformers of Computer Science for Computer Vision (CS4CV), we study how positional information is incorporated into self-attention models. We begin by explaining why self-attention alone does not inherently encode the order of tokens. Unlike RNNs, which process inputs sequentially, self-attention performs computations in parallel, so information about token position must be added explicitly. We then introduce positional encoding by adding a positional embedding matrix to the input token embeddings. Different possible strategies are discussed, including normalized position values and simple integer-based position indices, along with their limitations for variable-length sequences and generalization. Next, we define the desirable properties of a positional encoding scheme: it should provide a unique representation for each position, preserve meaningful distances between positions, generalize to longer sequences, remain bounded, and be deterministic. We then introduce the sinusoidal positional encoding used in the Transformer architecture, where different dimensions use sine functions with different frequencies. The intuition is also compared with binary representations, where different bits change at different rates. The video also covers relative positional information, showing how the encoding of a shifted position can be related to the original position through a linear transformation that is independent of the absolute position. Finally, we summarize the key properties of multi-head attention, self-attention, computational complexity, shortest path length, and the role of positional encoding in sequence modeling. Topics covered: Why self-attention needs positional information Positional encoding Absolute position information Limitations of naive positional representations Sinusoidal positional encoding Frequency-based positional representations Relative positional information Self-attention summary #PositionalEncoding #SelfAttention #Transformers #Attention #DeepLearning #ComputerVision #CS4CV