Перейти к содержимому

Diffusion Models Explained (Hindi) — Read The Paper With Me | DDPM & Score Fields

Vishesh Yadav

0:00 / 0:00

Diffusion Models Explained (Hindi) — Read The Paper With Me | DDPM & Score Fields

31 просмотр · 1 день назад
Vishesh Yadav
176 подписчиков
31 просмотр · 1 день назад
=== DESCRIPTION (Hinglish) === Is episode mein hum 2020 ka landmark paper "Denoising Diffusion Probabilistic Models" (DDPM — Ho, Jain & Abbeel) — poori tarah, line by line, ek deep ~50-minute dive mein, saath baith kar padhte hain. Ye wahi paper hai jisne Stable Diffusion, DALL·E 2, Midjourney aur Imagen ki neev rakhi — pure noise se nayi images banane ka jaadu. Aur is baar visuals ko naye level par le gaye hain: dense particle clouds, flowing score-field streamlines, smooth cinematic motion — sab kuch asli paper par (real cropped PDF), aur har idea 3Blue1Brown-style geometric animation se. On-screen text English mein hai; narration Hindi. Kya-kya cover hota hai (line by line): Generative modelling & the data manifold Forward process: gradually adding Gaussian noise (the beta schedule) The closed form: xₜ = √ᾱₜ·x₀ + √(1−ᾱₜ)·ε — a smooth interpolation between image and noise Reverse process: learned Gaussian denoising, step by step The SCORE FIELD (∇ log p) — arrows flowing toward the data; denoising = following the field Connection to Langevin dynamics & denoising score matching Predicting the noise ε (the reparameterization) instead of the mean The variational bound (ELBO) → collapses to L_simple = ‖ε − ε_θ‖² (just an MSE!) Algorithm 1 (training) & Algorithm 2 (sampling) Noise → image emerging, step by step The U-Net + skip connections (ResNet tie) + time embedding Results: FID 3.17 on CIFAR-10, beating GANs; progressive generation; latent interpolation Diffusion vs VAE vs GAN Legacy: Stable Diffusion, DALL·E 2, Midjourney, Imagen, text-to-image Paper: "Denoising Diffusion Probabilistic Models" (Ho, Jain & Abbeel, 2020) — arXiv:2006.11239 Agar aapko ye "Read The Paper With Me" series pasand aa rahi hai — Episode 1 (Attention), Episode 2 (LoRA), Episode 3 (Word2Vec), Episode 4 (Batch Normalization), aur Episode 5 (ResNet) bhi dekhiye. Channel subscribe karna mat bhooliye. #Diffusion #DDPM #StableDiffusion #GenerativeAI #DeepLearning #MachineLearning #Hindi #3Blue1Brown #AI #ScoreMatching 0:00 Intro 0:05 A face from pure noise 0:44 Abstract 1:19 The generative problem 2:00 The core idea: add noise, learn to reverse 2:39 Why so many tiny steps 3:16 Two processes: forward & reverse 3:53 Forward: destroying structure 4:32 Reverse: creating structure 5:09 The Markov chain (Fig 2) 5:47 T = 1000 steps 6:20 Everything ends as pure noise 6:56 The ink-in-water analogy 7:36 What we must learn 8:12 The forward process q 8:48 Shrink a little, wander a little 9:24 The β noise schedule 9:55 A random walk to noise 10:33 The closed-form shortcut 11:09 Eq 4: the closed form 11:46 Image ↔ noise interpolation 12:30 ᾱₜ — the signal that remains 13:10 Signal vs noise tug-of-war 13:49 x_T is pure Gaussian 14:27 Why the closed form matters 15:05 Forward has no learning 15:45 The reverse process pθ 16:22 Small steps → a simple Gaussian 17:04 The denoising network 17:44 What the network sees 18:24 The score field ∇log p 19:05 Arrows pointing to the data 19:47 Denoising = following the score 20:28 Score at high vs low noise 21:11 Langevin dynamics 21:54 = denoising score matching 22:35 Predicting the mean 23:12 The reparameterization 23:54 Predict the noise ε 24:34 Noise = the score, geometrically 25:14 Why ε-prediction is elegant 25:54 The training goal 26:34 The variational bound (ELBO) 27:17 KL: matching distributions 27:55 Matching two Gaussians 28:37 It collapses to noise prediction 29:13 Dropping the coefficients 29:52 L_simple = ‖ε − ε_θ‖² 30:36 Just an MSE — vs GAN instability 31:17 Why the simple loss works 31:57 Algorithm 1: training 32:37 One training step, walked through 33:19 No per-step targets needed 34:00 Training recap 34:43 Algorithm 2: sampling 35:24 Start from pure noise 36:04 The sampling step 36:43 Why add noise back (Langevin) 37:26 Noise → image emerges 38:07 The opening, explained 38:46 Slow sampling — the honest cost 39:25 The U-Net denoiser 40:03 Skip connections (ResNet tie) 40:45 Time embedding 41:24 The whole architecture 42:07 A noise value per pixel 42:47 Sampling recap 43:28 Results — the samples 44:08 FID 3.17, beating GANs 44:51 Progressive coarse→fine 45:37 Latent interpolation 46:25 Diffusion vs VAE 47:10 Diffusion vs GAN 47:53 VAE vs GAN vs Diffusion 48:36 Legacy begins 49:11 Stable Diffusion, DALL·E, Midjourney 49:53 Text-to-image (Attention tie) 50:35 Recap 51:18 The big lesson