Перейти к содержимому

s01-p07: Transformer Decoder, Inference & Attention Visualization | CS4CV

Computer Science for Computer Vision

0:00 / 0:00

s01-p07: Transformer Decoder, Inference & Attention Visualization | CS4CV

21 просмотр · 9 дней назад
Computer Science for Computer Vision
9 подписчиков
21 просмотр · 9 дней назад
Course: Computer Science for Computer Vision Channel: @CS4CV (   / @cs4cv  ) Playlist: Section 1 – Attention Mechanisms & Transformers (   • CS4CV Section 01: Attention Mechanisms & T...  ) In the seventh video of Section 1 – Attention Mechanisms & Transformers of Computer Science for Computer Vision (CS4CV), we continue our study of the Transformer decoder and examine how attention behaves inside a trained Transformer model. We begin with the structure of the Transformer decoder. Each decoder block contains three main components: Masked Multi-Head Attention Encoder–Decoder Multi-Head Attention Position-wise Feed-Forward Network The masked self-attention layer ensures that, during sequence generation, each token can only attend to previously generated tokens. The second attention layer uses queries from the decoder while receiving keys and values from the encoder. We then look at how a trained Transformer is used for sequence generation and follow the flow of information through the encoder and decoder. The final part of the video focuses on visualizing attention weights. We first inspect attention inside the Transformer encoder, including the effect of masking padded tokens. Next, we visualize the decoder attention weights. Since tokens are generated one at a time, the first decoder attention mechanism only attends to previously generated tokens. We then examine the attention weights of the decoder’s second attention layer, which connects the decoder to information produced by the encoder. Topics covered: Transformer decoder architecture Masked multi-head attention Encoder–decoder attention Autoregressive sequence generation Transformer inference Encoder attention visualization Padding masks Decoder attention visualization #Transformers #TransformerDecoder #AttentionVisualization #MaskedAttention #MultiHeadAttention #DeepLearning #CS4CV