Перейти к содержимому

EZ撸paper: DeepSeek-R1 论文详解 part 3:GPT发展史 | scaling law | 训练范式 | emergent ability #deepseek

EZ.Encoder Academy

0:00 / 0:00

EZ撸paper: DeepSeek-R1 论文详解 part 3:GPT发展史 | scaling law | 训练范式 | emergent ability #deepseek

12 174 просмотра · 1 год назад
EZ.Encoder Academy
17,6 тыс. подписчиков
12 174 просмотра · 1 год назад
DeepSeek-R1 是DeepSeek 发布的开源推理模型,其性能可以媲美 OpenAI 的 o1,在数学、编程和逻辑推理任务上表现突出。而更令人惊讶的是,它的运行成本仅为 OpenAI 的 2%!最重要的是,它是一个完全开源的模型,任何人都可以自由使用其模型权重,进行训练和开发。 在本视频中,我将GPT发展史和背后的思想. ------------------------------------------------------ 视频中使用的手写板: HUION Inspiroy H1060P Graphics Drawing Tablet https://amzn.to/41DvOiW ------------------------------------------------- DeepSeek R1 paper: https://arxiv.org/pdf/2501.12948 DeepSeekMath paper: https://arxiv.org/pdf/2402.03300 Kimi k1.5 paper: https://arxiv.org/pdf/2501.12599v1 Kimi k1.5 作者思路: 英文版本: https://x.com/Kimi_Moonshot/status/18... 中文版本: https://www.zhihu.com/people/flood-su... ViT paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale https://arxiv.org/pdf/2010.11929 Rich Sutton: The Bitter Lesson http://www.incompleteideas.net/IncIde... Rich Sutton: AI Succession (WAIC keynote) http://www.incompleteideas.net/Talks/... Hyung Won Chung (OpenAI): Shaping the future of AI from the history of Transformer https://docs.google.com/presentation/... LLM 发展树: https://x.com/tariqkrim/status/165185... Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond https://arxiv.org/abs/2304.13712 GPT1: Improving Language Understanding by Generative Pre-Training https://cdn.openai.com/research-cover... GPT2: Language Models are Unsupervised Multitask Learners https://cdn.openai.com/better-languag... GPT3: Language Models are Few-Shot Learners https://arxiv.org/pdf/2005.14165 PaLM: Scaling Language Modeling with Pathways https://arxiv.org/abs/2204.02311 Scaling Laws for Neural Language Models https://arxiv.org/abs/2001.08361 Open Pretrained Transformers - Susan Zhang | Stanford MLSys    • Open Pretrained Transformers - Susan Zhang...   2 OLMo 2 Furious https://arxiv.org/abs/2501.00656 Deep reinforcement learning from human preferences https://arxiv.org/abs/1706.03741 InstructGPT: Training language models to follow instructions with human feedback https://arxiv.org/abs/2203.02155 Learning to summarize from human feedback https://arxiv.org/abs/2009.01325 The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale https://arxiv.org/abs/2406.17557 DPO: Direct Preference Optimization: Your Language Model is Secretly a Reward Model https://arxiv.org/abs/2305.18290 Emergent Abilities of Large Language Models https://arxiv.org/abs/2206.07682 Are Emergent Abilities of Large Language Models a Mirage? https://arxiv.org/abs/2304.15004 ------------------------------------------------- DeepSeek-R1 技术报告详细解读 part1:    • EZ撸paper: DeepSeek-R1 论文详解 part 1:比肩 OpenA...   DeepSeek-R1 技术报告详细解读 part2:    • EZ撸paper: DeepSeek-R1 论文详解 part 2:AGI是什么? ...   DeepSeek-R1 技术报告详细解读 part3:    • EZ撸paper: DeepSeek-R1 论文详解 part 3:GPT发展史 |...   ------------------------------------------------- DeepSeek-V3 技术报告详细解读 part1:    • EZ撸paper: DeepSeek-V3 技术报告详细解读 part1 | 开源最...   DeepSeek-V3 技术报告详细解读 part2:    • EZ撸paper: DeepSeek-V3 技术报告详细解读 part2 | 开源最...   DeepSeek-V3 技术报告详细解读 part3:    • EZ撸paper: DeepSeek-V3 论文中的隐藏细节 (part 3):你不...   DeepSeek-V3 技术报告详细解读 part4:    • EZ撸paper: DeepSeek-V3 论文中的隐藏细节 (part 4):从入...   ------------------------------------------------- 我正在写的书, 我的经历和我的胡思乱想: https://ez-encoder-academy.gitbook.io... 我希望能结交更多朋友, 想和我云聊: https://calendly.com/chvlyl/coffee_ch...