# Paper List for AI beginners
**作者**: Xiuyu Li
**日期**: 2026-04-18T04:55:23.000Z
**来源**: [https://x.com/sheriyuo/status/2045365552848482680](https://x.com/sheriyuo/status/2045365552848482680)
---

This list comes from a reading guide by my supervisor Prof. Mingyang Yi, aimed at helping second-year CS / Math undergraduates get started with ML & RL.
Foundations
- Deep Residual Learning for Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Layer Normalization
- Attention Is All You Need
AI Infrastructure
- An Introduction to Variational Autoencoders
- Language Models are Unsupervised Multitask Learners
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Transformers for Image Recognition at Scale
- Learning Transferable Visual Models From Natural Language Supervision
- LoRA: Low-Rank Adaptation of Large Language Models
- Let's Verify Step by Step
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Inference Acceleration
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Linformer: Self-Attention with Linear Complexity
- Fast Inference from Transformers via Speculative Decoding
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reinforcement Learning
- Training language models to follow instructions with human feedback
- Token-level Direct Preference Optimization
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- Trust Region Policy Optimization
- Proximal Policy Optimization Algorithms
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Asynchronous Methods for Deep Reinforcement Learning
- It Takes Two: Your GRPO Is Secretly DPO
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
## 相关链接
- [Xiuyu Li](https://x.com/sheriyuo)
- [@sheriyuo](https://x.com/sheriyuo)
- [4.1K](https://x.com/sheriyuo/status/2045365552848482680/analytics)
- [Mingyang Yi](https://mingyangyi.github.io/)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [12:55 PM · Apr 18, 2026](https://x.com/sheriyuo/status/2045365552848482680)
- [4,151 Views](https://x.com/sheriyuo/status/2045365552848482680/analytics)
- [View quotes](https://x.com/sheriyuo/status/2045365552848482680/quotes)
---
*导出时间: 2026/4/18 23:38:35*
---
## 中文翻译
# AI 初学者论文清单
**作者**: Xiuyu Li
**日期**: 2026-04-18T04:55:23.000Z
**来源**: [https://x.com/sheriyuo/status/2045365552848482680](https://x.com/sheriyuo/status/2045365552848482680)
---

这份清单源自我的导师 Mingyang Yi 教授的一份阅读指南,旨在帮助计算机科学或数学专业的大二学生入门机器学习(ML)与强化学习(RL)。
基础
- Deep Residual Learning for Image Recognition(用于图像识别的深度残差学习)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift(批归一化:通过减少内部协变量偏移加速深度网络训练)
- Layer Normalization(层归一化)
- Attention Is All You Need(注意力机制就是你所需的一切)
AI 基础设施
- An Introduction to Variational Autoencoders(变分自编码器介绍)
- Language Models are Unsupervised Multitask Learners(语言模型是无监督的多任务学习者)
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding(BERT:用于语言理解的深度双向变换器预训练)
- Transformers for Image Recognition at Scale(大规模图像识别中的变换器)
- Learning Transferable Visual Models From Natural Language Supervision(从自然语言监督中学习可迁移的视觉模型)
- LoRA: Low-Rank Adaptation of Large Language Models(LoRA:大型语言模型的低秩适应)
- Let's Verify Step by Step(让我们逐步验证)
- Reflexion: Language Agents with Verbal Reinforcement Learning(Reflexion:具有口头强化学习的语言智能体)
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection(Self-RAG:通过自我反思学习检索、生成和评判)
推理加速
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention(Transformers 就是 RNN:具有线性注意力的快速自回归变换器)
- Linformer: Self-Attention with Linear Complexity(Linformer:具有线性复杂度的自注意力)
- Fast Inference from Transformers via Speculative Decoding(通过推测解码实现 Transformer 的快速推理)
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty(EAGLE:推测采样需要重新思考特征不确定性)
强化学习
- Training language models to follow instructions with human feedback(使用人类反馈训练语言模型以遵循指令)
- Token-level Direct Preference Optimization(Token 级直接偏好优化)
- A General Theoretical Paradigm to Understand Learning from Human Preferences(理解人类反馈学习的通用理论范式)
- Trust Region Policy Optimization(信任区域策略优化)
- Proximal Policy Optimization Algorithms(近端策略优化算法)
- High-Dimensional Continuous Control Using Generalized Advantage Estimation(使用广义优势估计的高维连续控制)
- Asynchronous Methods for Deep Reinforcement Learning(深度强化学习的异步方法)
- It Takes Two: Your GRPO Is Secretly DPO(两个都少不了:你的 GRPO 其实就是 DPO)
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning(Search-R1:训练 LLM 利用强化学习进行推理并借助搜索引擎)
## 相关链接
- [Xiuyu Li](https://x.com/sheriyuo)
- [@sheriyuo](https://x.com/sheriyuo)
- [4.1K](https://x.com/sheriyuo/status/2045365552848482680/analytics)
- [Mingyang Yi](https://mingyangyi.github.io/)
- [升级到 Premium](https://x.com/i/premium_sign_up)
- [2026年4月18日 下午 12:55](https://x.com/sheriyuo/status/2045365552848482680)
- [4,151 次查看](https://x.com/sheriyuo/status/2045365552848482680/analytics)
- [查看引用](https://x.com/sheriyuo/status/2045365552848482680/quotes)
---
*导出时间: 2026/4/18 23:38:35*