Перейти к содержимому

Transformer Architecture Explained | Attention Is All You Need — Part 1

AI & ML with Sanjay Chouhan

0:00 / 0:00

Transformer Architecture Explained | Attention Is All You Need — Part 1

182 просмотра · 2 нед. назад
AI & ML with Sanjay Chouhan
117 подписчиков
182 просмотра · 2 нед. назад
In this video, we start our journey to understand the *Transformer architecture* introduced in the landmark paper “Attention Is All You Need” (2017). Before diving into Self-Attention and the mathematical details, we first build an intuitive understanding of the **overall Transformer architecture**. In this video, we cover: • Encoder and Decoder stacks • Why the original Transformer has 6 Encoder blocks and 6 Decoder blocks • How contextual representations flow from the Encoder to the Decoder • The two inputs involved in the Transformer • How the Decoder starts generation with a start token • How a Transformer generates text *one token at a time* • A simple example of autoregressive text generation This is *Part 1* of my Transformer series, where we'll break down the architecture step by step, starting from the basics and eventually understanding the calculations behind Self-Attention, Multi-Head Attention, and the Feed-Forward Network. 📄 Paper: Attention Is All You Need — Vaswani et al., 2017 📚 Transformer Series Part 1 — Transformer Architecture Part 2 — Tokenization Part 3 — Embeddings, Positional Encoding & Softmax ...and more Transformer Playlist:    • Transformer Explained | Attention Is All Y...   Subscribe to follow the complete series and learn **Transformers from scratch**. #Transformer #AttentionIsAllYouNeed #TransformerArchitecture #DeepLearning #NLP #LLM #GenerativeAI #MachineLearning