Fine-Tuning Explained: SFT, LoRA, QLoRA and RLHF - When to Fine-Tune an AI Model
AI Coding Arena
0:00 / 0:00
Fine-Tuning Explained: SFT, LoRA, QLoRA and RLHF - When to Fine-Tune an AI Model
4 просмотра · 8 дней назад
AI Coding Arena
1 подписчик
4 просмотра · 8 дней назад
Fine-tuning trains a general AI model a little more, on your data, for your specific job. This explainer covers how pretraining differs from fine-tuning, supervised fine-tuning (SFT) with example pairs, what fine-tuning teaches well (style, tone, format) and what it doesn't (new facts), LoRA from Microsoft in 2021 (10,000x fewer trainable parameters), QLoRA from Tim Dettmers' team in 2023 (a 65B model on one 48GB GPU), RLHF and RLAIF, and how to choose between prompting, retrieval and fine-tuning - plus pitfalls like catastrophic forgetting.
Chapters:
0:00 What fine-tuning is
0:19 Pretraining vs fine-tuning
0:42 Supervised fine-tuning (SFT)
1:16 What it teaches, and what it doesn't
1:44 LoRA (2021)
2:30 QLoRA (2023)
2:58 RLHF and RLAIF
3:36 Prompting vs retrieval vs fine-tuning
4:05 Pitfalls and catastrophic forgetting
4:21 Generalist to specialist
Part of AI Model Families Explained - one hand-drawn explainer per AI lab.
#FineTuning #LoRA #AIExplained