Перейти к содержимому

Sparse Autoencoders Unlearn Knowledge in LLMs | A Paper-Based Walkthrough

Papers Are Wonderful

0:00 / 0:00

Sparse Autoencoders Unlearn Knowledge in LLMs | A Paper-Based Walkthrough

4 701 просмотр · 1 г. назад
Papers Are Wonderful
1,5 тыс. подписчиков
4 701 просмотр · 1 г. назад
I made a video about one of my favorite papers! I hope you enjoy :) ===Summary=== "Applying Sparse Autoencoders to Unlearn Knowledge in Language Models" investigates using SAEs—tools that peer into the inside of LLMs—to remove undesirable capabilities from language models. In this video, I walk through the motivation of this work, the methods used, and the interesting results the authors found. I highly recommend you read it for yourself here: https://arxiv.org/pdf/2410.19278#page... ===My other videos on Sparse Autoencoders=== Matroshkya SAEs:    • Matryoshka (Nested) Sparse Autoencoders Ex...   SAEs from the Ground Up:    • A Window  Into LLMs | Sparse Autoencoders ...   ===Video Chapters=== 0:00 Intro 0:14 Context/Motivation 0:46 SAE Negative Clamping 1:01 Feature Identification 1:35 Experimental Setup 1:49 Single-Feature Steering 2:24 Multi-Feature Steering 3:29 Investigating RMU Hypothesis