EZ撸paper: DeepSeek-R1 论文详解 part 3:GPT发展史 | scaling law | 训练范式 | emergent ability #deepseek
EZ.Encoder Academy
0:00 / 0:00
EZ撸paper: DeepSeek-R1 论文详解 part 3:GPT发展史 | scaling law | 训练范式 | emergent ability #deepseek
12 174 просмотра · 1 год назад
EZ.Encoder Academy
17,6 тыс. подписчиков
12 174 просмотра · 1 год назад
DeepSeek-R1 是DeepSeek 发布的开源推理模型,其性能可以媲美 OpenAI 的 o1,在数学、编程和逻辑推理任务上表现突出。而更令人惊讶的是,它的运行成本仅为 OpenAI 的 2%!最重要的是,它是一个完全开源的模型,任何人都可以自由使用其模型权重,进行训练和开发。
在本视频中,我将GPT发展史和背后的思想.
------------------------------------------------------
视频中使用的手写板:
HUION Inspiroy H1060P Graphics Drawing Tablet
https://amzn.to/41DvOiW
-------------------------------------------------
DeepSeek R1 paper:
https://arxiv.org/pdf/2501.12948
DeepSeekMath paper:
https://arxiv.org/pdf/2402.03300
Kimi k1.5 paper:
https://arxiv.org/pdf/2501.12599v1
Kimi k1.5 作者思路:
英文版本: https://x.com/Kimi_Moonshot/status/18...
中文版本: https://www.zhihu.com/people/flood-su...
ViT paper:
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
https://arxiv.org/pdf/2010.11929
Rich Sutton: The Bitter Lesson
http://www.incompleteideas.net/IncIde...
Rich Sutton: AI Succession (WAIC keynote)
http://www.incompleteideas.net/Talks/...
Hyung Won Chung (OpenAI): Shaping the future of AI from the history of Transformer
https://docs.google.com/presentation/...
LLM 发展树:
https://x.com/tariqkrim/status/165185...
Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond
https://arxiv.org/abs/2304.13712
GPT1: Improving Language Understanding by Generative Pre-Training
https://cdn.openai.com/research-cover...
GPT2: Language Models are Unsupervised Multitask Learners
https://cdn.openai.com/better-languag...
GPT3: Language Models are Few-Shot Learners
https://arxiv.org/pdf/2005.14165
PaLM: Scaling Language Modeling with Pathways
https://arxiv.org/abs/2204.02311
Scaling Laws for Neural Language Models
https://arxiv.org/abs/2001.08361
Open Pretrained Transformers - Susan Zhang | Stanford MLSys
• Open Pretrained Transformers - Susan Zhang...
2 OLMo 2 Furious
https://arxiv.org/abs/2501.00656
Deep reinforcement learning from human preferences
https://arxiv.org/abs/1706.03741
InstructGPT: Training language models to follow instructions with human feedback
https://arxiv.org/abs/2203.02155
Learning to summarize from human feedback
https://arxiv.org/abs/2009.01325
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
https://arxiv.org/abs/2406.17557
DPO: Direct Preference Optimization: Your Language Model is Secretly a Reward Model
https://arxiv.org/abs/2305.18290
Emergent Abilities of Large Language Models
https://arxiv.org/abs/2206.07682
Are Emergent Abilities of Large Language Models a Mirage?
https://arxiv.org/abs/2304.15004
-------------------------------------------------
DeepSeek-R1 技术报告详细解读 part1:
• EZ撸paper: DeepSeek-R1 论文详解 part 1:比肩 OpenA...
DeepSeek-R1 技术报告详细解读 part2:
• EZ撸paper: DeepSeek-R1 论文详解 part 2:AGI是什么? ...
DeepSeek-R1 技术报告详细解读 part3:
• EZ撸paper: DeepSeek-R1 论文详解 part 3:GPT发展史 |...
-------------------------------------------------
DeepSeek-V3 技术报告详细解读 part1:
• EZ撸paper: DeepSeek-V3 技术报告详细解读 part1 | 开源最...
DeepSeek-V3 技术报告详细解读 part2:
• EZ撸paper: DeepSeek-V3 技术报告详细解读 part2 | 开源最...
DeepSeek-V3 技术报告详细解读 part3:
• EZ撸paper: DeepSeek-V3 论文中的隐藏细节 (part 3):你不...
DeepSeek-V3 技术报告详细解读 part4:
• EZ撸paper: DeepSeek-V3 论文中的隐藏细节 (part 4):从入...
-------------------------------------------------
我正在写的书, 我的经历和我的胡思乱想:
https://ez-encoder-academy.gitbook.io...
我希望能结交更多朋友, 想和我云聊:
https://calendly.com/chvlyl/coffee_ch...