Transformer Architecture Explained | Attention Is All You Need — Part 1
AI & ML with Sanjay Chouhan
0:00 / 0:00
Transformer Architecture Explained | Attention Is All You Need — Part 1
182 просмотра · 2 нед. назад
AI & ML with Sanjay Chouhan
117 подписчиков
182 просмотра · 2 нед. назад
In this video, we start our journey to understand the *Transformer architecture* introduced in the landmark paper “Attention Is All You Need” (2017).
Before diving into Self-Attention and the mathematical details, we first build an intuitive understanding of the **overall Transformer architecture**.
In this video, we cover:
• Encoder and Decoder stacks
• Why the original Transformer has 6 Encoder blocks and 6 Decoder blocks
• How contextual representations flow from the Encoder to the Decoder
• The two inputs involved in the Transformer
• How the Decoder starts generation with a start token
• How a Transformer generates text *one token at a time*
• A simple example of autoregressive text generation
This is *Part 1* of my Transformer series, where we'll break down the architecture step by step, starting from the basics and eventually understanding the calculations behind Self-Attention, Multi-Head Attention, and the Feed-Forward Network.
📄 Paper: Attention Is All You Need — Vaswani et al., 2017
📚 Transformer Series
Part 1 — Transformer Architecture
Part 2 — Tokenization
Part 3 — Embeddings, Positional Encoding & Softmax
...and more
Transformer Playlist: • Transformer Explained | Attention Is All Y...
Subscribe to follow the complete series and learn **Transformers from scratch**.
#Transformer #AttentionIsAllYouNeed #TransformerArchitecture #DeepLearning #NLP #LLM #GenerativeAI #MachineLearning