Transformers Explained: Attention Is All You Need, from the 2017 Paper to GPT and BERT
AI Coding Arena
0:00 / 0:00
Transformers Explained: Attention Is All You Need, from the 2017 Paper to GPT and BERT
9 просмотров · 8 дн. назад
AI Coding Arena
1 подписчик
9 просмотров · 8 дн. назад
Nearly every modern AI model runs on the transformer. This explainer covers the June 2017 Google paper "Attention Is All You Need" and its eight authors, why recurrent networks and LSTMs hit a wall, how attention works with queries, keys and values, the 8 parallel attention heads, positional encoding, the 6-layer encoder and decoder with cross-attention, the record 28.4 and 41.8 BLEU translation scores after 3.5 days on 8 GPUs, and how the design split into decoder-only GPT (OpenAI, June 2018) and encoder-only BERT (Google, October 2018).
Chapters:
0:00 Attention Is All You Need
0:37 Before transformers: RNNs and LSTMs
1:05 Throwing away recurrence
1:11 Attention: query, key and value
1:41 Multi-head attention
1:59 Positional encoding
2:05 The encoder and decoder
2:37 The translation results
2:59 Two families: GPT and BERT
4:07 How the transformer took over
Part of AI Model Families Explained - one hand-drawn explainer per AI lab.
#Transformers #AttentionIsAllYouNeed #AIExplained