构筑 AI 理论体系:深度学习 100 篇论文解读

第 0 篇:总纲——系列导航与学习地图

💡 核心命题:在原典中理解智能的构造

深度学习并非黑箱魔法,而是人类数十年间基于数学、神经科学信息论一步步构造起来的工程奇迹

本系列旨在打破“调参侠”的局限,带你回到深度学习历史上 100 篇最具开创性、里程碑式的论文现场。我们的目标不是简单复述摘要,而是还原论文的动机、核心数学推导和历史上下文,理解每个算法是如何解决前一代的瓶颈。

看完这 100 篇论文,你将拥有一个坚不可摧的深度学习知识体系,并能:

  • 理解底层逻辑: 明白为什么 Transformer 能够战胜 RNN,为什么 Diffusion Model 正在超越 GAN
  • 掌握设计哲学: 具备快速评估并改进当前最先进模型的能力。
  • 预测未来趋势: 从历史中推导出下一个十年深度学习的方向。

🗺️ 方法论:主题-影响力-时间 混合排序

本系列采用 “主题分块,块内按影响力与时间排序” 的混合策略。我们不会严格按照时间线跳跃,而是系统地完成一个技术模块的学习,确保知识体系的连贯性和逻辑的流畅性。

全系列分为五个核心技术模块

部/模块 主题 核心目标 论文数量
第一部 基石与感知机原点 (Foundations) 奠基: 掌握 BP 算法、优化器、正则化等一切深度学习的计算前提。 15 篇
第二部 卷积网络与视觉复兴 (CNN & Vision) 感知: 理解图像处理的结构(局部连接、权值共享),掌握深度残差设计。 25 篇
第三部 序列模型与 Transformer 革命 (RNN & NLP) 语言: 理解长序列处理的困难,掌握门控机制和自注意力机制的范式转变。 28 篇
第四部 生成模型与表征学习 (Generative & Representation) 创造/表征: 掌握 VAE、GAN、Diffusion 以及 自监督学习,学习数据的分布和高效特征。 22 篇
第五部 决策与前沿 (RL & GNN) 决策: 掌握 AI 的决策能力,包括强化学习理论、连续控制和处理非结构化数据的图网络。 10 篇

📜 100 篇论文全景图(Deep Learning 100)

第一部:基石与感知机原点 (15 篇)
序号 论文名称 核心主题
1 A Logical Calculus… (McCulloch & Pitts, 1943) 神经网络的数学起点:M-P 模型
2 The Perceptron… (Rosenblatt, 1957) 第一个可学习的神经模型:感知机
3 Learning Internal Representations… (Rumelhart et al., 1986) 深度学习的计算基石:反向传播 (BP)
4 Efficient Backprop (LeCun et al., 1998) 早期 BP 算法的工程优化与技巧
5 A Fast Learning Algorithm for Deep Belief Nets (Hinton et al., 2006) 解决 BP 瓶颈 1:DBN 概率预训练
6 Reducing the Dimensionality… (Hinton & Salakhutdinov, 2006) 解决 BP 瓶颈 2:堆叠自编码器重构预训练
7 Rectified Linear Units Better Feature Representation (Nair & Hinton, 2010) 终结预训练: ReLU 激活函数革命
8 Adam: A Method for Stochastic Optimization (Kingma & Ba, 2014) 优化加速: Adam 自适应优化器
9 Dropout: A Simple Way to Prevent Overfitting (Srivastava et al., 2014) 泛化基石: Dropout 正则化
10 Batch Normalization… (Ioffe & Szegedy, 2015) 训练稳定: 批量归一化 (BN)
11 Understanding the difficulty of training deep feedforward neural networks (Xavier/Glorot, 2010) 初始化基石:Xavier 初始化
12 Deep Learning (LeCun, Bengio & Hinton, 2015) 深度学习的综述与未来展望
13 Practical Recommendations for Gradient-Based Training (LeCun, 2012) 梯度训练的工程实践建议
14 Delving Deep into Rectifiers… (He et al., 2015) Kaiming/He 初始化与深入研究 ReLU
15 Weight Normalization: A Simple Reparameterization (Salimans & Kingma, 2016) 另一种归一化方法

第二部:卷积网络与视觉复兴 (25 篇)
  1. Gradient-Based Learning Applied to Document Recognition (LeCun et al., 1998) (LeNet-5/CNN 基础)
  2. ImageNet Classification with Deep Convolutional Neural Networks (Krizhevsky et al., 2012) (AlexNet)
  3. Visualizing and Understanding Convolutional Networks (Zeiler & Fergus, 2014) (ZF Net)
  4. Very Deep Convolutional Networks for Large-Scale Image Recognition (Simonyan & Zisserman, 2014) (VGG)
  5. Going Deeper with Convolutions (Szegedy et al., 2015) (GoogLeNet)
  6. Deep Residual Learning for Image Recognition (He et al., 2016) (ResNet)
  7. Densely Connected Convolutional Networks (Huang et al., 2017) (DenseNet)
  8. Squeeze-and-Excitation Networks (Hu et al., 2018) (SENet)
  9. Aggregated Residual Transformations… (Xie et al., 2017) (ResNeXt)
  10. Rethinking the Inception Architecture… (Szegedy et al., 2016) (Inception v3)
  11. MobileNets: Efficient Convolutional Neural Networks (Howard et al., 2017)
  12. ShuffleNet: An Extremely Efficient CNN (Zhang et al., 2018)
  13. Xception: Deep Learning with Depthwise Separable Convolutions (Chollet, 2017)
  14. Fully Convolutional Networks for Semantic Segmentation (Long et al., 2015) (FCN)
  15. U-Net: Convolutional Networks for Biomedical Image Segmentation (Ronneberger et al., 2015)
  16. Fast R-CNN (Girshick, 2015)
  17. You Only Look Once: Unified, Real-Time Object Detection (Redmon et al., 2016) (YOLO v1)
  18. Feature Pyramid Networks for Object Detection (Lin et al., 2017) (FPN)
  19. Mask R-CNN (He et al., 2017)
  20. Focal Loss for Dense Object Detection (Lin et al., 2017)
  21. DeepLab: Semantic Image Segmentation… (Chen et al., 2017)
  22. Neural Style Transfer (Gatys et al., 2016)
  23. Learning Deep Features for Discriminative Localization (Zhou et al., 2016) (CAM)
  24. High-resolution image synthesis… with conditional GANs (Isola et al., 2017) (Pix2Pix)
  25. Deformable Convolutional Networks (Dai et al., 2017)

第三部:序列模型与 Transformer 革命 (28 篇)
  1. Long Short-Term Memory (Hochreiter & Schmidhuber, 1997) (LSTM)
  2. Sequence to Sequence Learning with Neural Networks (Sutskever et al., 2014) (Seq2Seq)
  3. Learning Phrase Representations using RNN Encoder–Decoder (Cho et al., 2014) (GRU)
  4. Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau et al., 2015) (Attention 基础)
  5. Efficient Estimation of Word Representations in Vector Space (Mikolov et al., 2013) (Word2Vec)
  6. GloVe: Global Vectors for Word Representation (Pennington et al., 2014)
  7. Deep contextualized word representations (Peters et al., 2018) (ELMo)
  8. Attention Is All You Need (Vaswani et al., 2017) (Transformer)
  9. BERT: Pre-training of Deep Bidirectional Transformers (Devlin et al., 2019)
  10. Generative Pre-Training of a Large Language Model (Radford et al., 2018) (GPT-1)
  11. Language Models are Unsupervised Multitask Learners (Radford et al., 2019) (GPT-2)
  12. Language Models are Few-Shot Learners (Brown et al., 2020) (GPT-3)
  13. RoBERTa: A Robustly Optimized BERT Pretraining Approach (Liu et al., 2019)
  14. XLNet: Generalized Autoregressive Pretraining (Yang et al., 2019)
  15. Exploring the Limits of Transfer Learning with T5 (Raffel et al., 2020) (T5)
  16. ELECTRA: Pre-training Text Encoders as Discriminators (Clark et al., 2020)
  17. Vision Transformer (ViT) for Image Recognition (Dosovitskiy et al., 2021)
  18. Swin Transformer: Hierarchical Vision Transformer (Liu et al., 2021)
  19. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context (Dai et al., 2019)
  20. Training Verifiable Neural Networks (Weng et al., 2018)
  21. Exploring the Limits of Language Modeling (Jozefowicz et al., 2016)
  22. Improving Language Understanding by Generative Pre-Training (Radford et al., 2018)
  23. Universal Language Model Fine-tuning… (Howard & Ruder, 2018) (ULMFiT)
  24. PALM: Scaling Language Modeling with Pathways (Chowdhery et al., 2022)
  25. In-Context Learning (ICL) in Large Language Models (Min et al., 2022)
  26. Recurrent Neural Network Regularization (Zaremba et al., 2014)
  27. Character-Aware Neural Language Models (Kim et al., 2016)
  28. Zero-shot Text-to-Image Generation (Ramesh et al., 2021) (DALL-E)

第四部:生成模型与表征学习 (22 篇)
  1. Auto-Encoding Variational Bayes (Kingma & Welling, 2014) (VAE)
  2. Generative Adversarial Nets (Goodfellow et al., 2014) (GAN)
  3. Wasserstein GAN (Arjovsky et al., 2017) (WGAN)
  4. Improved Training of Wasserstein GANs (Gulrajani et al., 2017) (WGAN-GP)
  5. Progressive Growing of GANs (Karras et al., 2018) (ProGAN)
  6. Style-Based Synthesis of High-Resolution Images (Karras et al., 2019) (StyleGAN)
  7. Denoising Diffusion Probabilistic Models (Ho et al., 2020) (DDPM)
  8. Diffusion Models Beat GANs on Image Synthesis (Dhariwal & Nichol, 2021) (ADM)
  9. High-Resolution Image Synthesis with Latent Diffusion Models (Rombach et al., 2022) (LDM/Stable Diffusion)
  10. Adversarial Autoencoders (Makhzani et al., 2016)
  11. Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks (Zhu et al., 2017) (CycleGAN)
  12. Semi-Supervised Learning with GANs (Odena et al., 2016)
  13. InfoGAN: Interpretable Representation Learning (Chen et al., 2016)
  14. A Simple Framework for Contrastive Learning of Visual Representations (SimCLR, Chen et al., 2020) (自监督)
  15. Self-Training with Noisy Student improves ImageNet classification (Xie et al., 2020) (自监督)
  16. Deep Unsupervised Learning using Nonequilibrium Thermodynamics (Sohl-Dickstein et al., 2015)
  17. Classifier-Free Diffusion Guidance (Ho & Salimans, 2022)
  18. Consistency Models (Song et al., 2023)
  19. GLIDE: Towards Photorealistic Generation (Nichol et al., 2022)
  20. Image Generation with Conditional Diffusion Models (Ho et al., 2021)
  21. Learning Transferable Visual Models From Natural Language Supervision (Radford et al., 2021) (CLIP)
  22. Scaling Up Visual and Vision-Language Representation Learning (Jia et al., 2021) (ALIGN)

第五部:决策与前沿 (10 篇)
  1. Playing Atari with Deep Reinforcement Learning (Mnih et al., 2013) (DQN)
  2. Asynchronous Methods for Deep Reinforcement Learning (Mnih et al., 2016) (A3C)
  3. Continuous control with deep reinforcement learning (Lillicrap et al., 2016) (DDPG) (连续控制)
  4. Proximal Policy Optimization Algorithms (Schulman et al., 2017) (PPO)
  5. Mastering the game of Go with deep neural networks… (Silver et al., 2016) (AlphaGo)
  6. Mastering the game of Go without human knowledge (Silver et al., 2017) (AlphaGo Zero)
  7. MuZero: Mastering Games without Rules (Schrittwieser et al., 2020) (MuZero)
  8. Semi-Supervised Classification with Graph Convolutional Networks (Kipf & Welling, 2017) (GCN)
  9. Inductive Representation Learning on Large Graphs (Hamilton et al., 2017) (GraphSAGE)
  10. Graph Attention Networks (Velickovic et al., 2018) (GAT)

🚀 结语

本系列将是您攀登深度学习技术树的最坚实阶梯第 1 篇即将开始对 M-P 模型的深入解析。

期待与您一同,重走智能之路。

更多推荐