Diffusion Models Explained (Hindi) — Read The Paper With Me | DDPM & Score Fields
Vishesh Yadav
0:00 / 0:00
Diffusion Models Explained (Hindi) — Read The Paper With Me | DDPM & Score Fields
31 просмотр · 1 день назад
Vishesh Yadav
176 подписчиков
31 просмотр · 1 день назад
=== DESCRIPTION (Hinglish) ===
Is episode mein hum 2020 ka landmark paper "Denoising Diffusion Probabilistic Models" (DDPM — Ho, Jain & Abbeel) — poori tarah, line by line, ek deep ~50-minute dive mein, saath baith kar padhte hain. Ye wahi paper hai jisne Stable Diffusion, DALL·E 2, Midjourney aur Imagen ki neev rakhi — pure noise se nayi images banane ka jaadu.
Aur is baar visuals ko naye level par le gaye hain: dense particle clouds, flowing score-field streamlines, smooth cinematic motion — sab kuch asli paper par (real cropped PDF), aur har idea 3Blue1Brown-style geometric animation se. On-screen text English mein hai; narration Hindi.
Kya-kya cover hota hai (line by line):
Generative modelling & the data manifold
Forward process: gradually adding Gaussian noise (the beta schedule)
The closed form: xₜ = √ᾱₜ·x₀ + √(1−ᾱₜ)·ε — a smooth interpolation between image and noise
Reverse process: learned Gaussian denoising, step by step
The SCORE FIELD (∇ log p) — arrows flowing toward the data; denoising = following the field
Connection to Langevin dynamics & denoising score matching
Predicting the noise ε (the reparameterization) instead of the mean
The variational bound (ELBO) → collapses to L_simple = ‖ε − ε_θ‖² (just an MSE!)
Algorithm 1 (training) & Algorithm 2 (sampling)
Noise → image emerging, step by step
The U-Net + skip connections (ResNet tie) + time embedding
Results: FID 3.17 on CIFAR-10, beating GANs; progressive generation; latent interpolation
Diffusion vs VAE vs GAN
Legacy: Stable Diffusion, DALL·E 2, Midjourney, Imagen, text-to-image
Paper:
"Denoising Diffusion Probabilistic Models" (Ho, Jain & Abbeel, 2020) — arXiv:2006.11239
Agar aapko ye "Read The Paper With Me" series pasand aa rahi hai — Episode 1 (Attention), Episode 2 (LoRA), Episode 3 (Word2Vec), Episode 4 (Batch Normalization), aur Episode 5 (ResNet) bhi dekhiye. Channel subscribe karna mat bhooliye.
#Diffusion #DDPM #StableDiffusion #GenerativeAI #DeepLearning #MachineLearning #Hindi #3Blue1Brown #AI #ScoreMatching
0:00 Intro
0:05 A face from pure noise
0:44 Abstract
1:19 The generative problem
2:00 The core idea: add noise, learn to reverse
2:39 Why so many tiny steps
3:16 Two processes: forward & reverse
3:53 Forward: destroying structure
4:32 Reverse: creating structure
5:09 The Markov chain (Fig 2)
5:47 T = 1000 steps
6:20 Everything ends as pure noise
6:56 The ink-in-water analogy
7:36 What we must learn
8:12 The forward process q
8:48 Shrink a little, wander a little
9:24 The β noise schedule
9:55 A random walk to noise
10:33 The closed-form shortcut
11:09 Eq 4: the closed form
11:46 Image ↔ noise interpolation
12:30 ᾱₜ — the signal that remains
13:10 Signal vs noise tug-of-war
13:49 x_T is pure Gaussian
14:27 Why the closed form matters
15:05 Forward has no learning
15:45 The reverse process pθ
16:22 Small steps → a simple Gaussian
17:04 The denoising network
17:44 What the network sees
18:24 The score field ∇log p
19:05 Arrows pointing to the data
19:47 Denoising = following the score
20:28 Score at high vs low noise
21:11 Langevin dynamics
21:54 = denoising score matching
22:35 Predicting the mean
23:12 The reparameterization
23:54 Predict the noise ε
24:34 Noise = the score, geometrically
25:14 Why ε-prediction is elegant
25:54 The training goal
26:34 The variational bound (ELBO)
27:17 KL: matching distributions
27:55 Matching two Gaussians
28:37 It collapses to noise prediction
29:13 Dropping the coefficients
29:52 L_simple = ‖ε − ε_θ‖²
30:36 Just an MSE — vs GAN instability
31:17 Why the simple loss works
31:57 Algorithm 1: training
32:37 One training step, walked through
33:19 No per-step targets needed
34:00 Training recap
34:43 Algorithm 2: sampling
35:24 Start from pure noise
36:04 The sampling step
36:43 Why add noise back (Langevin)
37:26 Noise → image emerges
38:07 The opening, explained
38:46 Slow sampling — the honest cost
39:25 The U-Net denoiser
40:03 Skip connections (ResNet tie)
40:45 Time embedding
41:24 The whole architecture
42:07 A noise value per pixel
42:47 Sampling recap
43:28 Results — the samples
44:08 FID 3.17, beating GANs
44:51 Progressive coarse→fine
45:37 Latent interpolation
46:25 Diffusion vs VAE
47:10 Diffusion vs GAN
47:53 VAE vs GAN vs Diffusion
48:36 Legacy begins
49:11 Stable Diffusion, DALL·E, Midjourney
49:53 Text-to-image (Attention tie)
50:35 Recap
51:18 The big lesson