# The 1,000-Reading AI Curriculum — Audited Discovery Edition

**Status:** This is a progressive discovery curriculum, not a verified bibliography. A direct link records a supplied or audit-corrected destination; it does not by itself prove peer review, authorship, document type or republication rights. Any entry marked **UNVERIFIED** is explicitly **not verified/citable** until its identity is repaired.

**Audit basis:** curriculum identity and structure audit `paper-curriculum-2026-07-22` · audit date 2026-07-22 · source rows 1000 · link cards 1000.

**Flag key:** **VAGUE AUTHOR** = unnamed attribution; **NON-PAPER** = book, report, documentation, policy, product/topic or multi-work record; **UNVERIFIED** = no unique public work established; **DUPLICATE** = another curriculum row appears to represent the same work; **WRONG SUPPLIED MAPPING** = the original direct link named a different work and is replaced here with the audited correction.

## 1. Mathematical & Statistical Foundations

1. **A Mathematical Theory of Communication** — Shannon · 1948 · **Type:** Research-paper candidate. [Supplied direct source](<https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf>)

2. **Computing Machinery and Intelligence** — Turing · 1950 · **Type:** Research-paper candidate. [Supplied direct source](<https://academic.oup.com/mind/article/LIX/236/433/986238>)

3. **A Logical Calculus of Ideas Immanent in Nervous Activity** — McCulloch & Pitts · 1943 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.cs.cmu.edu/~./epxing/Class/10715/reading/McCulloch.and.Pitts.pdf>)

4. **An Inductive Inference Machine** — Solomonoff · 1957 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=An%20Inductive%20Inference%20Machine&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=An%20Inductive%20Inference%20Machine%20Solomonoff%201957>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=An%20Inductive%20Inference%20Machine&sort=relevance>)

5. **On Computable Numbers** — Turing · 1936 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.cs.virginia.edu/~robins/Turing_Paper_1936.pdf>)

6. **Maximum Likelihood from Incomplete Data via the EM Algorithm** — Dempster, Laird, Rubin · 1977 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.jstor.org/stable/2984875>)

7. **Regression Shrinkage and Selection via the Lasso** — Tibshirani · 1996 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.jstor.org/stable/2346178>)

8. **Bayesian Reasoning and Machine Learning** — Barber · 2012 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=Bayesian%20Reasoning%20and%20Machine%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Bayesian%20Reasoning%20and%20Machine%20Learning%20Barber%202012>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Bayesian%20Reasoning%20and%20Machine%20Learning&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper.

9. **Elements of Information Theory** — Cover & Thomas · 1991 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=Elements%20of%20Information%20Theory&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Elements%20of%20Information%20Theory%20Cover%201991>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Elements%20of%20Information%20Theory&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper.

10. **Monte Carlo Sampling Methods Using Markov Chains** — Hastings · 1970 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Monte%20Carlo%20Sampling%20Methods%20Using%20Markov%20Chains&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Monte%20Carlo%20Sampling%20Methods%20Using%20Markov%20Chains%20Hastings%201970>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Monte%20Carlo%20Sampling%20Methods%20Using%20Markov%20Chains&sort=relevance>)

11. **Stochastic Relaxation, Gibbs Distributions and Bayesian Restoration of Images** — Geman & Geman · 1984 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Stochastic%20Relaxation%2C%20Gibbs%20Distributions%20and%20Bayesian%20Restoration%20of%20Images&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Stochastic%20Relaxation%2C%20Gibbs%20Distributions%20and%20Bayesian%20Restoration%20of%20Images%20Geman%201984>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Stochastic%20Relaxation%2C%20Gibbs%20Distributions%20and%20Bayesian%20Restoration%20of%20Images&sort=relevance>)

12. **An Introduction to the Bootstrap** — Efron & Tibshirani · 1993 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=An%20Introduction%20to%20the%20Bootstrap&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=An%20Introduction%20to%20the%20Bootstrap%20Efron%201993>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=An%20Introduction%20to%20the%20Bootstrap&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper.

13. **The Kernel Trick** — Aizerman, Braverman, Rozonoer · 1964 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Kernel%20Trick&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Kernel%20Trick%20Aizerman%201964>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Kernel%20Trick&sort=relevance>)

14. **Singular Value Decomposition and Least Squares Solutions** — Golub & Reinsch · 1970 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Singular%20Value%20Decomposition%20and%20Least%20Squares%20Solutions&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Singular%20Value%20Decomposition%20and%20Least%20Squares%20Solutions%20Golub%201970>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Singular%20Value%20Decomposition%20and%20Least%20Squares%20Solutions&sort=relevance>)

15. **Random Matrices** — Mehta · 1967 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Random%20Matrices&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Random%20Matrices%20Mehta%201967>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Random%20Matrices&sort=relevance>)

## 2. Classical Machine Learning Foundations

16. **The Perceptron: A Probabilistic Model for Information Storage and Organization** — Rosenblatt · 1958 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Perceptron%3A%20A%20Probabilistic%20Model%20for%20Information%20Storage%20and%20Organization&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Perceptron%3A%20A%20Probabilistic%20Model%20for%20Information%20Storage%20and%20Organization%20Rosenblatt%201958>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Perceptron%3A%20A%20Probabilistic%20Model%20for%20Information%20Storage%20and%20Organization&sort=relevance>)

17. **Perceptrons** — Minsky & Papert · 1969 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=Perceptrons&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Perceptrons%20Minsky%201969>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Perceptrons&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper.

18. **Nearest Neighbor Pattern Classification** — Cover & Hart · 1967 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Nearest%20Neighbor%20Pattern%20Classification&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Nearest%20Neighbor%20Pattern%20Classification%20Cover%201967>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Nearest%20Neighbor%20Pattern%20Classification&sort=relevance>)

19. **A Theory of the Learnable** — Valiant · 1984 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=A%20Theory%20of%20the%20Learnable&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Theory%20of%20the%20Learnable%20Valiant%201984>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Theory%20of%20the%20Learnable&sort=relevance>)

20. **Induction of Decision Trees** — Quinlan · 1986 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Induction%20of%20Decision%20Trees&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Induction%20of%20Decision%20Trees%20Quinlan%201986>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Induction%20of%20Decision%20Trees&sort=relevance>)

21. **A Training Algorithm for Optimal Margin Classifiers** — Boser, Guyon, Vapnik · 1992 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=A%20Training%20Algorithm%20for%20Optimal%20Margin%20Classifiers&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Training%20Algorithm%20for%20Optimal%20Margin%20Classifiers%20Boser%201992>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Training%20Algorithm%20for%20Optimal%20Margin%20Classifiers&sort=relevance>)

22. **Support-Vector Networks** — Cortes & Vapnik · 1995 · **Type:** Research-paper candidate. [Supplied direct source](<https://link.springer.com/article/10.1007/BF00994018>)

23. **Statistical Learning Theory** — Vapnik · 1998 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=Statistical%20Learning%20Theory&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Statistical%20Learning%20Theory%20Vapnik%201998>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Statistical%20Learning%20Theory&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper.

24. **Bagging Predictors** — Breiman · 1996 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Bagging%20Predictors&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Bagging%20Predictors%20Breiman%201996>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Bagging%20Predictors&sort=relevance>)

25. **Random Forests** — Breiman · 2001 · **Type:** Research-paper candidate. [Supplied direct source](<https://link.springer.com/article/10.1023/A:1010933404324>)

26. **A Decision-Theoretic Generalization of On-Line Learning (AdaBoost)** — Freund & Schapire · 1997 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=A%20Decision-Theoretic%20Generalization%20of%20On-Line%20Learning%20%28AdaBoost%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Decision-Theoretic%20Generalization%20of%20On-Line%20Learning%20%28AdaBoost%29%20Freund%201997>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Decision-Theoretic%20Generalization%20of%20On-Line%20Learning%20%28AdaBoost%29&sort=relevance>)

27. **XGBoost: A Scalable Tree Boosting System** — Chen & Guestrin · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1603.02754>) · [Direct PDF](<https://arxiv.org/pdf/1603.02754.pdf>)

28. **Greedy Function Approximation: A Gradient Boosting Machine** — Friedman · 2001 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Greedy%20Function%20Approximation%3A%20A%20Gradient%20Boosting%20Machine&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Greedy%20Function%20Approximation%3A%20A%20Gradient%20Boosting%20Machine%20Friedman%202001>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Greedy%20Function%20Approximation%3A%20A%20Gradient%20Boosting%20Machine&sort=relevance>)

29. **An Introduction to Variable and Feature Selection** — Guyon & Elisseeff · 2003 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=An%20Introduction%20to%20Variable%20and%20Feature%20Selection&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=An%20Introduction%20to%20Variable%20and%20Feature%20Selection%20Guyon%202003>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=An%20Introduction%20to%20Variable%20and%20Feature%20Selection&sort=relevance>)

30. **Latent Dirichlet Allocation** — Blei, Ng, Jordan · 2003 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.jmlr.org/papers/volume3/blei03a/blei03a.pdf>)

31. **A Tutorial on Principal Component Analysis** — Shlens · 2014 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=A%20Tutorial%20on%20Principal%20Component%20Analysis&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Tutorial%20on%20Principal%20Component%20Analysis%20Shlens%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Tutorial%20on%20Principal%20Component%20Analysis&sort=relevance>)

32. **Reducing the Dimensionality of Data with Neural Networks** — Hinton & Salakhutdinov · 2006 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Reducing%20the%20Dimensionality%20of%20Data%20with%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Reducing%20the%20Dimensionality%20of%20Data%20with%20Neural%20Networks%20Hinton%202006>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Reducing%20the%20Dimensionality%20of%20Data%20with%20Neural%20Networks&sort=relevance>)

33. **K-Means++: The Advantages of Careful Seeding** — Arthur & Vassilvitskii · 2007 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=K-Means%2B%2B%3A%20The%20Advantages%20of%20Careful%20Seeding&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=K-Means%2B%2B%3A%20The%20Advantages%20of%20Careful%20Seeding%20Arthur%202007>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=K-Means%2B%2B%3A%20The%20Advantages%20of%20Careful%20Seeding&sort=relevance>)

34. **Gaussian Processes for Machine Learning** — Rasmussen & Williams · 2006 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=Gaussian%20Processes%20for%20Machine%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Gaussian%20Processes%20for%20Machine%20Learning%20Rasmussen%202006>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Gaussian%20Processes%20for%20Machine%20Learning&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper.

35. **No Free Lunch Theorems for Optimization** — Wolpert & Macready · 1997 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=No%20Free%20Lunch%20Theorems%20for%20Optimization&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=No%20Free%20Lunch%20Theorems%20for%20Optimization%20Wolpert%201997>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=No%20Free%20Lunch%20Theorems%20for%20Optimization&sort=relevance>)

## 3. Neural Network Foundations

36. **Learning Representations by Back-Propagating Errors** — Rumelhart, Hinton, Williams · 1986 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/323533a0>)

37. **Multilayer Feedforward Networks are Universal Approximators** — Hornik, Stinchcombe, White · 1989 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Multilayer%20Feedforward%20Networks%20are%20Universal%20Approximators&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Multilayer%20Feedforward%20Networks%20are%20Universal%20Approximators%20Hornik%201989>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Multilayer%20Feedforward%20Networks%20are%20Universal%20Approximators&sort=relevance>)

38. **Learning Internal Representations by Error Propagation** — Rumelhart, Hinton, Williams · 1985 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Learning%20Internal%20Representations%20by%20Error%20Propagation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Learning%20Internal%20Representations%20by%20Error%20Propagation%20Rumelhart%201985>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Learning%20Internal%20Representations%20by%20Error%20Propagation&sort=relevance>)

39. **Connectionist Learning Procedures** — Hinton · 1989 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Connectionist%20Learning%20Procedures&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Connectionist%20Learning%20Procedures%20Hinton%201989>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Connectionist%20Learning%20Procedures&sort=relevance>)

40. **A Fast Learning Algorithm for Deep Belief Nets** — Hinton, Osindero, Teh · 2006 · **Type:** Research-paper candidate. [Corrected direct source](<https://www.cs.toronto.edu/~hinton/absps/fastnc.pdf>)
   - **Audit:** WRONG SUPPLIED MAPPING — The supplied link named a different work and is corrected in this edition. Original target: “A collaborative framework to exchange and share product information within a supply chain context”.

41. **Greedy Layer-Wise Training of Deep Networks** — Bengio et al. · 2007 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Greedy%20Layer-Wise%20Training%20of%20Deep%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Greedy%20Layer-Wise%20Training%20of%20Deep%20Networks%20Bengio%202007>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Greedy%20Layer-Wise%20Training%20of%20Deep%20Networks&sort=relevance>)

42. **Understanding the Difficulty of Training Deep Feedforward Neural Networks** — Glorot & Bengio · 2010 · **Type:** Research-paper candidate. [Supplied direct source](<https://proceedings.mlr.press/v9/glorot10a.html>)

43. **Rectified Linear Units Improve Restricted Boltzmann Machines** — Nair & Hinton · 2010 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Rectified%20Linear%20Units%20Improve%20Restricted%20Boltzmann%20Machines&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Rectified%20Linear%20Units%20Improve%20Restricted%20Boltzmann%20Machines%20Nair%202010>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Rectified%20Linear%20Units%20Improve%20Restricted%20Boltzmann%20Machines&sort=relevance>)

44. **Deep Sparse Rectifier Neural Networks** — Glorot, Bordes, Bengio · 2011 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Deep%20Sparse%20Rectifier%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Deep%20Sparse%20Rectifier%20Neural%20Networks%20Glorot%202011>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Deep%20Sparse%20Rectifier%20Neural%20Networks&sort=relevance>)

45. **Maxout Networks** — Goodfellow et al. · 2013 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Maxout%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Maxout%20Networks%20Goodfellow%202013>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Maxout%20Networks&sort=relevance>)

46. **Improving Neural Networks by Preventing Co-Adaptation of Feature Detectors** — Hinton et al. · 2012 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Improving%20Neural%20Networks%20by%20Preventing%20Co-Adaptation%20of%20Feature%20Detectors&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Improving%20Neural%20Networks%20by%20Preventing%20Co-Adaptation%20of%20Feature%20Detectors%20Hinton%202012>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Improving%20Neural%20Networks%20by%20Preventing%20Co-Adaptation%20of%20Feature%20Detectors&sort=relevance>)

47. **Dropout: A Simple Way to Prevent Neural Networks from Overfitting** — Srivastava et al. · 2014 · **Type:** Research-paper candidate. [Supplied direct source](<https://jmlr.org/papers/v15/srivastava14a.html>)

48. **Batch Normalization: Accelerating Deep Network Training** — Ioffe & Szegedy · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1502.03167>) · [Direct PDF](<https://arxiv.org/pdf/1502.03167.pdf>)

49. **Layer Normalization** — Ba, Kiros, Hinton · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1607.06450>) · [Direct PDF](<https://arxiv.org/pdf/1607.06450.pdf>)

50. **Group Normalization** — Wu & He · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1803.08494>) · [Direct PDF](<https://arxiv.org/pdf/1803.08494.pdf>)

51. **Deep Learning** — LeCun, Bengio, Hinton · 2015 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/nature14539>)

52. **Representation Learning: A Review** — Bengio, Courville, Vincent · 2013 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Representation%20Learning%3A%20A%20Review&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Representation%20Learning%3A%20A%20Review%20Bengio%202013>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Representation%20Learning%3A%20A%20Review&sort=relevance>)

53. **On the Number of Linear Regions of Deep Neural Networks** — Montufar et al. · 2014 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=On%20the%20Number%20of%20Linear%20Regions%20of%20Deep%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=On%20the%20Number%20of%20Linear%20Regions%20of%20Deep%20Neural%20Networks%20Montufar%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=On%20the%20Number%20of%20Linear%20Regions%20of%20Deep%20Neural%20Networks&sort=relevance>)

54. **Residual Learning (Identity Mappings in Deep Residual Networks)** — He et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1603.05027>) · [Direct PDF](<https://arxiv.org/pdf/1603.05027.pdf>)

55. **The Lottery Ticket Hypothesis** — Frankle & Carlin · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1803.03635>) · [Direct PDF](<https://arxiv.org/pdf/1803.03635.pdf>)

## 4. Optimization & Training Techniques

56. **A Stochastic Approximation Method** — Robbins & Monro · 1951 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=A%20Stochastic%20Approximation%20Method&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Stochastic%20Approximation%20Method%20Robbins%201951>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Stochastic%20Approximation%20Method&sort=relevance>)

57. **On the Importance of Initialization and Momentum in Deep Learning** — Sutskever et al. · 2013 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=On%20the%20Importance%20of%20Initialization%20and%20Momentum%20in%20Deep%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=On%20the%20Importance%20of%20Initialization%20and%20Momentum%20in%20Deep%20Learning%20Sutskever%202013>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=On%20the%20Importance%20of%20Initialization%20and%20Momentum%20in%20Deep%20Learning&sort=relevance>)

58. **ADADELTA: An Adaptive Learning Rate Method** — Zeiler · 2012 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ADADELTA%3A%20An%20Adaptive%20Learning%20Rate%20Method&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ADADELTA%3A%20An%20Adaptive%20Learning%20Rate%20Method%20Zeiler%202012>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ADADELTA%3A%20An%20Adaptive%20Learning%20Rate%20Method&sort=relevance>)

59. **Adam: A Method for Stochastic Optimization** — Kingma & Ba · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1412.6980>) · [Direct PDF](<https://arxiv.org/pdf/1412.6980.pdf>)

60. **Decoupled Weight Decay Regularization (AdamW)** — Loshchilov & Hutter · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1711.05101>) · [Direct PDF](<https://arxiv.org/pdf/1711.05101.pdf>)

61. **SGDR: Stochastic Gradient Descent with Warm Restarts** — Loshchilov & Hutter · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1608.03983>) · [Direct PDF](<https://arxiv.org/pdf/1608.03983.pdf>)

62. **Cyclical Learning Rates for Training Neural Networks** — Smith · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Cyclical%20Learning%20Rates%20for%20Training%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Cyclical%20Learning%20Rates%20for%20Training%20Neural%20Networks%20Smith%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Cyclical%20Learning%20Rates%20for%20Training%20Neural%20Networks&sort=relevance>)

63. **Super-Convergence** — Smith & Topin · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Super-Convergence&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Super-Convergence%20Smith%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Super-Convergence&sort=relevance>)

64. **Large Batch Training of Convolutional Networks (LARS)** — You et al. · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Large%20Batch%20Training%20of%20Convolutional%20Networks%20%28LARS%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Large%20Batch%20Training%20of%20Convolutional%20Networks%20%28LARS%29%20You%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Large%20Batch%20Training%20of%20Convolutional%20Networks%20%28LARS%29&sort=relevance>)

65. **Large Batch Optimization (LAMB)** — You et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Large%20Batch%20Optimization%20%28LAMB%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Large%20Batch%20Optimization%20%28LAMB%29%20You%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Large%20Batch%20Optimization%20%28LAMB%29&sort=relevance>)

66. **Sharpness-Aware Minimization (SAM)** — Foret et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2010.01412>) · [Direct PDF](<https://arxiv.org/pdf/2010.01412.pdf>)

67. **An Overview of Gradient Descent Optimization Algorithms** — Ruder · 2016 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=An%20Overview%20of%20Gradient%20Descent%20Optimization%20Algorithms&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=An%20Overview%20of%20Gradient%20Descent%20Optimization%20Algorithms%20Ruder%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=An%20Overview%20of%20Gradient%20Descent%20Optimization%20Algorithms&sort=relevance>)

68. **Fixing Weight Decay Regularization in Adam** — Loshchilov & Hutter · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Fixing%20Weight%20Decay%20Regularization%20in%20Adam&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Fixing%20Weight%20Decay%20Regularization%20in%20Adam%20Loshchilov%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Fixing%20Weight%20Decay%20Regularization%20in%20Adam&sort=relevance>)

69. **Mixed Precision Training** — Micikevicius et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1710.03740>) · [Direct PDF](<https://arxiv.org/pdf/1710.03740.pdf>)

70. **Automatic Mixed Precision** — NVIDIA · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Automatic%20Mixed%20Precision&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Automatic%20Mixed%20Precision%20NVIDIA%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Automatic%20Mixed%20Precision&sort=relevance>)

71. **Gradient Checkpointing** — Chen et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Gradient%20Checkpointing&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Gradient%20Checkpointing%20Chen%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Gradient%20Checkpointing&sort=relevance>)

72. **Data Parallelism (PyTorch Distributed)** — Li et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Data%20Parallelism%20%28PyTorch%20Distributed%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Data%20Parallelism%20%28PyTorch%20Distributed%29%20Li%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Data%20Parallelism%20%28PyTorch%20Distributed%29&sort=relevance>)

73. **ZeRO: Memory Optimizations Toward Training Trillion Parameter Models** — Rajbhandari et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1910.02054>) · [Direct PDF](<https://arxiv.org/pdf/1910.02054.pdf>)

74. **DeepSpeed** — Rasley et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DeepSpeed&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DeepSpeed%20Rasley%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DeepSpeed&sort=relevance>)

75. **Megatron-LM: Training Multi-Billion Parameter Language Models** — Shoeybi et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1909.08053>) · [Direct PDF](<https://arxiv.org/pdf/1909.08053.pdf>)

## 5. Convolutional Neural Networks & Computer Vision

76. **Neocognitron: A Self-Organizing Neural Network Model** — Fukushima · 1980 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Neocognitron%3A%20A%20Self-Organizing%20Neural%20Network%20Model&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Neocognitron%3A%20A%20Self-Organizing%20Neural%20Network%20Model%20Fukushima%201980>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Neocognitron%3A%20A%20Self-Organizing%20Neural%20Network%20Model&sort=relevance>)

77. **Gradient-Based Learning Applied to Document Recognition (LeNet)** — LeCun et al. · 1998 · **Type:** Research-paper candidate. [Supplied direct source](<https://yann.lecun.com/exdb/publis/pdf/lecun-98.pdf>)

78. **ImageNet Classification with Deep CNNs (AlexNet)** — Krizhevsky, Sutskever, Hinton · 2012 · **Type:** Research-paper candidate. [Supplied direct source](<https://papers.nips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html>)

79. **Visualizing and Understanding Convolutional Networks (ZFNet)** — Zeiler & Fergus · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1311.2901>) · [Direct PDF](<https://arxiv.org/pdf/1311.2901.pdf>)
   - **Audit:** WRONG SUPPLIED MAPPING — The supplied link named a different work and is corrected in this edition. Original target: “Multi-epoch VLBA H2O maser observations toward the massive YSOs AFGL 2591 VLA 2 and VLA 3”.

80. **Very Deep Convolutional Networks (VGGNet)** — Simonyan & Zisserman · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1409.1556>) · [Direct PDF](<https://arxiv.org/pdf/1409.1556.pdf>)

81. **Going Deeper with Convolutions (GoogLeNet/Inception)** — Szegedy et al. · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1409.4842>) · [Direct PDF](<https://arxiv.org/pdf/1409.4842.pdf>)

82. **Deep Residual Learning for Image Recognition (ResNet)** — He et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1512.03385>) · [Direct PDF](<https://arxiv.org/pdf/1512.03385.pdf>)

83. **Densely Connected Convolutional Networks (DenseNet)** — Huang et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1608.06993>) · [Direct PDF](<https://arxiv.org/pdf/1608.06993.pdf>)

84. **Aggregated Residual Transformations (ResNeXt)** — Xie et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1611.05431>) · [Direct PDF](<https://arxiv.org/pdf/1611.05431.pdf>)

85. **Squeeze-and-Excitation Networks (SENet)** — Hu, Shen, Sun · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Squeeze-and-Excitation%20Networks%20%28SENet%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Squeeze-and-Excitation%20Networks%20%28SENet%29%20Hu%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Squeeze-and-Excitation%20Networks%20%28SENet%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #961.

86. **EfficientNet: Rethinking Model Scaling** — Tan & Le · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1905.11946>) · [Direct PDF](<https://arxiv.org/pdf/1905.11946.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #747.

87. **Network In Network** — Lin, Chen, Yan · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1312.4400>) · [Direct PDF](<https://arxiv.org/pdf/1312.4400.pdf>)

88. **Spatial Transformer Networks** — Jaderberg et al. · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1506.02025>) · [Direct PDF](<https://arxiv.org/pdf/1506.02025.pdf>)

89. **Feature Pyramid Networks for Object Detection** — Lin et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1612.03144>) · [Direct PDF](<https://arxiv.org/pdf/1612.03144.pdf>)

90. **Faster R-CNN** — Ren et al. · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1506.01497>) · [Direct PDF](<https://arxiv.org/pdf/1506.01497.pdf>)

91. **You Only Look Once (YOLO)** — Redmon et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1506.02640>) · [Direct PDF](<https://arxiv.org/pdf/1506.02640.pdf>)

92. **SSD: Single Shot MultiBox Detector** — Liu et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1512.02325>) · [Direct PDF](<https://arxiv.org/pdf/1512.02325.pdf>)

93. **Focal Loss for Dense Object Detection (RetinaNet)** — Lin et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1708.02002>) · [Direct PDF](<https://arxiv.org/pdf/1708.02002.pdf>)

94. **Mask R-CNN** — He et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1703.06870>) · [Direct PDF](<https://arxiv.org/pdf/1703.06870.pdf>)

95. **U-Net: CNNs for Biomedical Image Segmentation** — Ronneberger et al. · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1505.04597>) · [Direct PDF](<https://arxiv.org/pdf/1505.04597.pdf>)

96. **Fully Convolutional Networks for Semantic Segmentation** — Long, Shelhamer, Darrell · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1411.4038>) · [Direct PDF](<https://arxiv.org/pdf/1411.4038.pdf>)

97. **DeepLab: Semantic Image Segmentation with Deep CNNs and CRFs** — Chen et al. · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1412.7062>) · [Direct PDF](<https://arxiv.org/pdf/1412.7062.pdf>)

98. **An Image is Worth 16x16 Words (ViT)** — Dosovitskiy et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2010.11929>) · [Direct PDF](<https://arxiv.org/pdf/2010.11929.pdf>)

99. **Swin Transformer** — Liu et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2103.14030>) · [Direct PDF](<https://arxiv.org/pdf/2103.14030.pdf>)

100. **ConvNeXt: A ConvNet for the 2020s** — Liu et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2201.03545>) · [Direct PDF](<https://arxiv.org/pdf/2201.03545.pdf>)

## 6. Recurrent Neural Networks & Sequence Modeling

101. **Finding Structure in Time** — Elman · 1990 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Finding%20Structure%20in%20Time&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Finding%20Structure%20in%20Time%20Elman%201990>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Finding%20Structure%20in%20Time&sort=relevance>)

102. **Long Short-Term Memory (LSTM)** — Hochreiter & Schmidhuber · 1997 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.bioinf.jku.at/publications/older/2604.pdf>)

103. **Learning to Forget: Continual Prediction with LSTM** — Gers, Schmidhuber, Cummins · 2000 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Learning%20to%20Forget%3A%20Continual%20Prediction%20with%20LSTM&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Learning%20to%20Forget%3A%20Continual%20Prediction%20with%20LSTM%20Gers%202000>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Learning%20to%20Forget%3A%20Continual%20Prediction%20with%20LSTM&sort=relevance>)

104. **Learning Phrase Representations using RNN Encoder-Decoder (GRU)** — Cho et al. · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1406.1078>) · [Direct PDF](<https://arxiv.org/pdf/1406.1078.pdf>)

105. **Sequence to Sequence Learning with Neural Networks** — Sutskever, Vinyals, Le · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1409.3215>) · [Direct PDF](<https://arxiv.org/pdf/1409.3215.pdf>)

106. **Neural Machine Translation by Jointly Learning to Align and Translate** — Bahdanau, Cho, Bengio · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1409.0473>) · [Direct PDF](<https://arxiv.org/pdf/1409.0473.pdf>)

107. **Effective Approaches to Attention-based Neural Machine Translation** — Luong, Pham, Manning · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1508.04025>) · [Direct PDF](<https://arxiv.org/pdf/1508.04025.pdf>)

108. **Connectionist Temporal Classification (CTC)** — Graves et al. · 2006 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.cs.toronto.edu/~graves/icml_2006.pdf>)

109. **Speech Recognition with Deep Recurrent Neural Networks** — Graves, Mohamed, Hinton · 2013 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1303.5778>) · [Direct PDF](<https://arxiv.org/pdf/1303.5778.pdf>)

110. **Pointer Networks** — Vinyals, Fortunato, Jaitly · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1506.03134>) · [Direct PDF](<https://arxiv.org/pdf/1506.03134.pdf>)

111. **Generating Sequences with Recurrent Neural Networks** — Graves · 2013 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1308.0850>) · [Direct PDF](<https://arxiv.org/pdf/1308.0850.pdf>)

112. **Bidirectional Recurrent Neural Networks** — Schuster & Paliwal · 1997 · **Type:** Research-paper candidate. [Supplied direct source](<https://ieeexplore.ieee.org/document/650093>)

113. **Show and Tell: A Neural Image Caption Generator** — Vinyals et al. · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1411.4555>) · [Direct PDF](<https://arxiv.org/pdf/1411.4555.pdf>)

114. **Neural Turing Machines** — Graves, Wayne, Danihelka · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1410.5401>) · [Direct PDF](<https://arxiv.org/pdf/1410.5401.pdf>)

115. **Memory Networks** — Weston, Chopra, Bordes · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1410.3916>) · [Direct PDF](<https://arxiv.org/pdf/1410.3916.pdf>)

116. **Temporal Convolutional Networks (TCN)** — Bai, Kolter, Koltun · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1803.01271>) · [Direct PDF](<https://arxiv.org/pdf/1803.01271.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #966.

117. **Quasi-Recurrent Neural Networks** — Bradbury et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1611.01576>) · [Direct PDF](<https://arxiv.org/pdf/1611.01576.pdf>)

118. **Independently Recurrent Neural Network (IndRNN)** — Li et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Independently%20Recurrent%20Neural%20Network%20%28IndRNN%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Independently%20Recurrent%20Neural%20Network%20%28IndRNN%29%20Li%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Independently%20Recurrent%20Neural%20Network%20%28IndRNN%29&sort=relevance>)

119. **Relational Recurrent Neural Networks** — Santoro et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Relational%20Recurrent%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Relational%20Recurrent%20Neural%20Networks%20Santoro%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Relational%20Recurrent%20Neural%20Networks&sort=relevance>)

120. **WaveNet: A Generative Model for Raw Audio** — van den Oord et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1609.03499>) · [Direct PDF](<https://arxiv.org/pdf/1609.03499.pdf>)

## 7. Generative Models

121. **Auto-Encoding Variational Bayes (VAE)** — Kingma & Welling · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1312.6114>) · [Direct PDF](<https://arxiv.org/pdf/1312.6114.pdf>)

122. **Generative Adversarial Networks (GAN)** — Goodfellow et al. · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1406.2661>) · [Direct PDF](<https://arxiv.org/pdf/1406.2661.pdf>)

123. **Conditional Generative Adversarial Nets** — Mirza & Osindero · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1411.1784>) · [Direct PDF](<https://arxiv.org/pdf/1411.1784.pdf>)

124. **Unsupervised Representation Learning with DCGANs** — Radford, Metz, Chintala · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1511.06434>) · [Direct PDF](<https://arxiv.org/pdf/1511.06434.pdf>)

125. **Improved Techniques for Training GANs** — Salimans et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1606.03498>) · [Direct PDF](<https://arxiv.org/pdf/1606.03498.pdf>)

126. **Wasserstein GAN** — Arjovsky, Chintala, Bottou · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1701.07875>) · [Direct PDF](<https://arxiv.org/pdf/1701.07875.pdf>)

127. **Progressive Growing of GANs** — Karras, Aila, Laine, Lehtinen · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1710.10196>) · [Direct PDF](<https://arxiv.org/pdf/1710.10196.pdf>)

128. **A Style-Based Generator Architecture (StyleGAN)** — Karras, Laine, Aila · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1812.04948>) · [Direct PDF](<https://arxiv.org/pdf/1812.04948.pdf>)

129. **Analyzing and Improving StyleGAN (StyleGAN2)** — Karras et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1912.04958>) · [Direct PDF](<https://arxiv.org/pdf/1912.04958.pdf>)

130. **Image-to-Image Translation with Conditional GANs (Pix2Pix)** — Isola et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1611.07004>) · [Direct PDF](<https://arxiv.org/pdf/1611.07004.pdf>)

131. **Unpaired Image-to-Image Translation (CycleGAN)** — Zhu et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1703.10593>) · [Direct PDF](<https://arxiv.org/pdf/1703.10593.pdf>)

132. **Semantic Image Synthesis with Spatially-Adaptive Normalization (SPADE)** — Park et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1903.07291>) · [Direct PDF](<https://arxiv.org/pdf/1903.07291.pdf>)

133. **Neural Discrete Representation Learning (VQ-VAE)** — van den Oord, Vinyals, Kavukcuoglu · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1711.00937>) · [Direct PDF](<https://arxiv.org/pdf/1711.00937.pdf>)

134. **Generating Diverse High-Fidelity Images with VQ-VAE-2** — Razavi, van den Oord, Vinyals · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1906.00446>) · [Direct PDF](<https://arxiv.org/pdf/1906.00446.pdf>)

135. **DALL·E: Zero-Shot Text-to-Image Generation (dVAE component)** — Ramesh et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2102.12092>) · [Direct PDF](<https://arxiv.org/pdf/2102.12092.pdf>)

136. **Glow: Generative Flow with Invertible 1x1 Convolutions** — Kingma & Dhariwal · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1807.03039>) · [Direct PDF](<https://arxiv.org/pdf/1807.03039.pdf>)

137. **NICE: Non-linear Independent Components Estimation** — Dinh, Krueger, Bengio · 2015 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=NICE%3A%20Non-linear%20Independent%20Components%20Estimation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=NICE%3A%20Non-linear%20Independent%20Components%20Estimation%20Dinh%202015>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=NICE%3A%20Non-linear%20Independent%20Components%20Estimation&sort=relevance>)

138. **Density Estimation Using Real-NVP** — Dinh, Sohl-Dickstein, Bengio · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1605.08803>) · [Direct PDF](<https://arxiv.org/pdf/1605.08803.pdf>)

139. **PixelCNN** — van den Oord et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=PixelCNN&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=PixelCNN%20van%20den%20Oord%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=PixelCNN&sort=relevance>)

140. **Energy-Based Models for Atomic-Scale Modeling** — various · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Energy-Based%20Models%20for%20Atomic-Scale%20Modeling&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Energy-Based%20Models%20for%20Atomic-Scale%20Modeling%20various%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Energy-Based%20Models%20for%20Atomic-Scale%20Modeling&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 8. Attention Mechanisms & Transformers

141. **Attention Is All You Need** — Vaswani et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1706.03762>) · [Direct PDF](<https://arxiv.org/pdf/1706.03762.pdf>)

142. **Self-Attention with Relative Position Representations** — Shaw, Uszkoreit, Vaswani · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1803.02155>) · [Direct PDF](<https://arxiv.org/pdf/1803.02155.pdf>)

143. **Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context** — Dai et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1901.02860>) · [Direct PDF](<https://arxiv.org/pdf/1901.02860.pdf>)

144. **Generating Long Sequences with Sparse Transformers** — Child et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1904.10509>) · [Direct PDF](<https://arxiv.org/pdf/1904.10509.pdf>)

145. **Longformer: The Long-Document Transformer** — Beltagy, Peters, Cohan · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2004.05150>) · [Direct PDF](<https://arxiv.org/pdf/2004.05150.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #184.

146. **Big Bird: Transformers for Longer Sequences** — Zaheer et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2007.14062>) · [Direct PDF](<https://arxiv.org/pdf/2007.14062.pdf>)

147. **Linformer: Self-Attention with Linear Complexity** — Wang et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2006.04768>) · [Direct PDF](<https://arxiv.org/pdf/2006.04768.pdf>)

148. **Performers: Rethinking Attention (FAVOR+)** — Choromanski et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2009.14794>) · [Direct PDF](<https://arxiv.org/pdf/2009.14794.pdf>)

149. **FlashAttention: Fast and Memory-Efficient Exact Attention** — Dao et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2205.14135>) · [Direct PDF](<https://arxiv.org/pdf/2205.14135.pdf>)

150. **FlashAttention-2: Faster Attention with Better Parallelism** — Dao · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2307.08691>) · [Direct PDF](<https://arxiv.org/pdf/2307.08691.pdf>)

151. **Multi-Query Attention** — Shazeer · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1911.02150>) · [Direct PDF](<https://arxiv.org/pdf/1911.02150.pdf>)

152. **GQA: Training Generalized Multi-Query Transformer Models** — Ainslie et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.13245>) · [Direct PDF](<https://arxiv.org/pdf/2305.13245.pdf>)

153. **RoFormer: Enhanced Transformer with Rotary Position Embedding (RoPE)** — Su et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2104.09864>) · [Direct PDF](<https://arxiv.org/pdf/2104.09864.pdf>)

154. **ALiBi: Train Short, Test Long (Attention with Linear Biases)** — Press et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2108.12409>) · [Direct PDF](<https://arxiv.org/pdf/2108.12409.pdf>)

155. **Reformer: The Efficient Transformer** — Kitaev, Kaiser, Levskaya · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2001.04451>) · [Direct PDF](<https://arxiv.org/pdf/2001.04451.pdf>)

156. **Set Transformer** — Lee et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Set%20Transformer&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Set%20Transformer%20Lee%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Set%20Transformer&sort=relevance>)

157. **Universal Transformers** — Dehghani et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1807.03819>) · [Direct PDF](<https://arxiv.org/pdf/1807.03819.pdf>)

158. **Scaling Transformers to 1M tokens (Ring Attention)** — Liu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Scaling%20Transformers%20to%201M%20tokens%20%28Ring%20Attention%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scaling%20Transformers%20to%201M%20tokens%20%28Ring%20Attention%29%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scaling%20Transformers%20to%201M%20tokens%20%28Ring%20Attention%29&sort=relevance>)

159. **Extending Context Window of LLMs via Positional Interpolation** — Chen et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2306.15595>) · [Direct PDF](<https://arxiv.org/pdf/2306.15595.pdf>)

160. **YaRN: Efficient Context Window Extension** — Peng et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2309.00071>) · [Direct PDF](<https://arxiv.org/pdf/2309.00071.pdf>)

## 9. Pre-trained Language Models

161. **Distributed Representations of Words (Word2Vec)** — Mikolov et al. · 2013 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1301.3781>) · [Direct PDF](<https://arxiv.org/pdf/1301.3781.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #162.

162. **Efficient Estimation of Word Representations in Vector Space** — Mikolov et al. · 2013 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1301.3781>) · [Direct PDF](<https://arxiv.org/pdf/1301.3781.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #161.

163. **GloVe: Global Vectors for Word Representation** — Pennington, Socher, Manning · 2014 · **Type:** Research-paper candidate. [Supplied direct source](<https://nlp.stanford.edu/pubs/glove.pdf>)

164. **Enriching Word Vectors with Subword Information (FastText)** — Bojanowski et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1607.04606>) · [Direct PDF](<https://arxiv.org/pdf/1607.04606.pdf>)

165. **Deep Contextualized Word Representations (ELMo)** — Peters et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1802.05365>) · [Direct PDF](<https://arxiv.org/pdf/1802.05365.pdf>)

166. **Semi-supervised Sequence Learning** — Dai & Le · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1511.01432>) · [Direct PDF](<https://arxiv.org/pdf/1511.01432.pdf>)

167. **Universal Language Model Fine-tuning (ULMFiT)** — Howard & Ruder · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1801.06146>) · [Direct PDF](<https://arxiv.org/pdf/1801.06146.pdf>)

168. **BERT: Pre-training of Deep Bidirectional Transformers** — Devlin et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1810.04805>) · [Direct PDF](<https://arxiv.org/pdf/1810.04805.pdf>)

169. **Improving Language Understanding by Generative Pre-Training (GPT-1)** — Radford et al. · 2018 · **Type:** Research-paper candidate. [Supplied direct source](<https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf>)

170. **Language Models are Unsupervised Multitask Learners (GPT-2)** — Radford et al. · 2019 · **Type:** Research-paper candidate. [Supplied direct source](<https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf>)

171. **RoBERTa: A Robustly Optimized BERT Pretraining Approach** — Liu et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1907.11692>) · [Direct PDF](<https://arxiv.org/pdf/1907.11692.pdf>)

172. **ALBERT: A Lite BERT** — Lan et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1909.11942>) · [Direct PDF](<https://arxiv.org/pdf/1909.11942.pdf>)

173. **DistilBERT** — Sanh et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1910.01108>) · [Direct PDF](<https://arxiv.org/pdf/1910.01108.pdf>)

174. **XLNet: Generalized Autoregressive Pretraining** — Yang et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1906.08237>) · [Direct PDF](<https://arxiv.org/pdf/1906.08237.pdf>)

175. **ELECTRA: Pre-training Text Encoders as Discriminators** — Clark et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2003.10555>) · [Direct PDF](<https://arxiv.org/pdf/2003.10555.pdf>)

176. **DeBERTa: Decoding-enhanced BERT with Disentangled Attention** — He et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2006.03654>) · [Direct PDF](<https://arxiv.org/pdf/2006.03654.pdf>)

177. **Exploring the Limits of Transfer Learning with a Unified T5** — Raffel et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1910.10683>) · [Direct PDF](<https://arxiv.org/pdf/1910.10683.pdf>)

178. **BART: Denoising Sequence-to-Sequence Pre-training** — Lewis et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1910.13461>) · [Direct PDF](<https://arxiv.org/pdf/1910.13461.pdf>)

179. **mBERT / XLM: Cross-lingual Language Model Pretraining** — Conneau & Lample · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=mBERT%20/%20XLM%3A%20Cross-lingual%20Language%20Model%20Pretraining&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=mBERT%20/%20XLM%3A%20Cross-lingual%20Language%20Model%20Pretraining%20Conneau%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=mBERT%20/%20XLM%3A%20Cross-lingual%20Language%20Model%20Pretraining&sort=relevance>)

180. **XLM-RoBERTa** — Conneau et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=XLM-RoBERTa&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=XLM-RoBERTa%20Conneau%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=XLM-RoBERTa&sort=relevance>)

181. **Sentence-BERT: Sentence Embeddings using Siamese BERT** — Reimers & Gurevych · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1908.10084>) · [Direct PDF](<https://arxiv.org/pdf/1908.10084.pdf>)

182. **ERNIE: Enhanced Language RepresentatioN with Informative Entities** — Zhang et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ERNIE%3A%20Enhanced%20Language%20RepresentatioN%20with%20Informative%20Entities&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ERNIE%3A%20Enhanced%20Language%20RepresentatioN%20with%20Informative%20Entities%20Zhang%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ERNIE%3A%20Enhanced%20Language%20RepresentatioN%20with%20Informative%20Entities&sort=relevance>)

183. **SpanBERT: Improving Pre-training by Representing and Predicting Spans** — Joshi et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SpanBERT%3A%20Improving%20Pre-training%20by%20Representing%20and%20Predicting%20Spans&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SpanBERT%3A%20Improving%20Pre-training%20by%20Representing%20and%20Predicting%20Spans%20Joshi%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SpanBERT%3A%20Improving%20Pre-training%20by%20Representing%20and%20Predicting%20Spans&sort=relevance>)

184. **Longformer for Long Document Understanding** — Beltagy et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Longformer%20for%20Long%20Document%20Understanding&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Longformer%20for%20Long%20Document%20Understanding%20Beltagy%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Longformer%20for%20Long%20Document%20Understanding&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #145.

185. **Unified Pre-training for Program Understanding and Generation (UniXcoder)** — Guo et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Unified%20Pre-training%20for%20Program%20Understanding%20and%20Generation%20%28UniXcoder%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Unified%20Pre-training%20for%20Program%20Understanding%20and%20Generation%20%28UniXcoder%29%20Guo%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Unified%20Pre-training%20for%20Program%20Understanding%20and%20Generation%20%28UniXcoder%29&sort=relevance>)

## 10. Large Language Models (LLMs)

186. **Language Models are Few-Shot Learners (GPT-3)** — Brown et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2005.14165>) · [Direct PDF](<https://arxiv.org/pdf/2005.14165.pdf>)

187. **GPT-4 Technical Report** — OpenAI · 2023 · **Type:** Technical or industry report. [Canonical arXiv record](<https://arxiv.org/abs/2303.08774>) · [Direct PDF](<https://arxiv.org/pdf/2303.08774.pdf>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

188. **PaLM: Scaling Language Modeling with Pathways** — Chowdhery et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2204.02311>) · [Direct PDF](<https://arxiv.org/pdf/2204.02311.pdf>)

189. **PaLM 2 Technical Report** — Google · 2023 · **Type:** Technical or industry report. [Canonical arXiv record](<https://arxiv.org/abs/2305.10403>) · [Direct PDF](<https://arxiv.org/pdf/2305.10403.pdf>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

190. **Gemini: A Family of Highly Capable Multimodal Models** — Google DeepMind · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2312.11805>) · [Direct PDF](<https://arxiv.org/pdf/2312.11805.pdf>)

191. **LLaMA: Open and Efficient Foundation Language Models** — Touvron et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2302.13971>) · [Direct PDF](<https://arxiv.org/pdf/2302.13971.pdf>)

192. **Llama 2: Open Foundation and Fine-Tuned Chat Models** — Touvron et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2307.09288>) · [Direct PDF](<https://arxiv.org/pdf/2307.09288.pdf>)

193. **Llama 3** — Meta · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2407.21783>) · [Direct PDF](<https://arxiv.org/pdf/2407.21783.pdf>)

194. **The Claude Model Card and Evaluations** — Anthropic · 2023 · **Type:** Product documentation or standard. [arXiv search](<https://arxiv.org/search/?query=The%20Claude%20Model%20Card%20and%20Evaluations&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Claude%20Model%20Card%20and%20Evaluations%20Anthropic%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Claude%20Model%20Card%20and%20Evaluations&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as product documentation or standard, not as a research paper.

195. **Mistral 7B** — Jiang et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2310.06825>) · [Direct PDF](<https://arxiv.org/pdf/2310.06825.pdf>)

196. **Mixtral of Experts** — Jiang et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2401.04088>) · [Direct PDF](<https://arxiv.org/pdf/2401.04088.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #556.

197. **Phi-1: Textbooks Are All You Need** — Gunasekar et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2306.11644>) · [Direct PDF](<https://arxiv.org/pdf/2306.11644.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #731, #782.

198. **Phi-2** — Microsoft · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Phi-2&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Phi-2%20Microsoft%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Phi-2&sort=relevance>)

199. **Phi-3 Technical Report** — Abdin et al. · 2024 · **Type:** Technical or industry report. [Canonical arXiv record](<https://arxiv.org/abs/2404.14219>) · [Direct PDF](<https://arxiv.org/pdf/2404.14219.pdf>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

200. **Chinchilla: Training Compute-Optimal Large Language Models** — Hoffmann et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2203.15556>) · [Direct PDF](<https://arxiv.org/pdf/2203.15556.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #722.

201. **Gopher** — Rae et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2112.11446>) · [Direct PDF](<https://arxiv.org/pdf/2112.11446.pdf>)

202. **OPT: Open Pre-trained Transformer** — Zhang et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2205.01068>) · [Direct PDF](<https://arxiv.org/pdf/2205.01068.pdf>)

203. **BLOOM: A 176B-Parameter Open-Access Multilingual LM** — BigScience · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2211.05100>) · [Direct PDF](<https://arxiv.org/pdf/2211.05100.pdf>)

204. **GLM-130B: An Open Bilingual Pre-Trained Model** — Zeng et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=GLM-130B%3A%20An%20Open%20Bilingual%20Pre-Trained%20Model&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=GLM-130B%3A%20An%20Open%20Bilingual%20Pre-Trained%20Model%20Zeng%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=GLM-130B%3A%20An%20Open%20Bilingual%20Pre-Trained%20Model&sort=relevance>)

205. **Falcon LLM** — TII · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Falcon%20LLM&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Falcon%20LLM%20TII%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Falcon%20LLM&sort=relevance>)

206. **Yi: Open Foundation Models** — 01.AI · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Yi%3A%20Open%20Foundation%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Yi%3A%20Open%20Foundation%20Models%2001.AI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Yi%3A%20Open%20Foundation%20Models&sort=relevance>)

207. **Qwen Technical Report** — Alibaba · 2023 · **Type:** Technical or industry report. [arXiv search](<https://arxiv.org/search/?query=Qwen%20Technical%20Report&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Qwen%20Technical%20Report%20Alibaba%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Qwen%20Technical%20Report&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

208. **Qwen2 Technical Report** — Alibaba · 2024 · **Type:** Technical or industry report. [Canonical arXiv record](<https://arxiv.org/abs/2407.10671>) · [Direct PDF](<https://arxiv.org/pdf/2407.10671.pdf>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

209. **DeepSeek LLM** — DeepSeek · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DeepSeek%20LLM&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DeepSeek%20LLM%20DeepSeek%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DeepSeek%20LLM&sort=relevance>)

210. **DeepSeek-V2: A Strong, Economical, and Efficient MoE LLM** — DeepSeek · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2405.04434>) · [Direct PDF](<https://arxiv.org/pdf/2405.04434.pdf>)

211. **DeepSeek-V3 Technical Report** — DeepSeek · 2024 · **Type:** Technical or industry report. [Canonical arXiv record](<https://arxiv.org/abs/2412.19437>) · [Direct PDF](<https://arxiv.org/pdf/2412.19437.pdf>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

212. **Gemma: Open Models Based on Gemini** — Google DeepMind · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2403.08295>) · [Direct PDF](<https://arxiv.org/pdf/2403.08295.pdf>)

213. **Command R+** — Cohere · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Command%20R%2B&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Command%20R%2B%20Cohere%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Command%20R%2B&sort=relevance>)

214. **Jamba: A Hybrid Transformer-Mamba Model** — AI21 Labs · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2403.19887>) · [Direct PDF](<https://arxiv.org/pdf/2403.19887.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #744.

215. **OLMo: Accelerating the Science of Language Models** — Groeneveld et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2402.00838>) · [Direct PDF](<https://arxiv.org/pdf/2402.00838.pdf>)

## 11. LLM Training & Alignment

216. **Training Language Models to Follow Instructions (InstructGPT)** — Ouyang et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2203.02155>) · [Direct PDF](<https://arxiv.org/pdf/2203.02155.pdf>)

217. **Learning to Summarize from Human Feedback** — Stiennon et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2009.01325>) · [Direct PDF](<https://arxiv.org/pdf/2009.01325.pdf>)

218. **Constitutional AI: Harmlessness from AI Feedback (CAI)** — Bai et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2212.08073>) · [Direct PDF](<https://arxiv.org/pdf/2212.08073.pdf>)

219. **Direct Preference Optimization (DPO)** — Rafailov et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.18290>) · [Direct PDF](<https://arxiv.org/pdf/2305.18290.pdf>)

220. **RLHF: A Survey** — Casper et al. · 2023 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=RLHF%3A%20A%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=RLHF%3A%20A%20Survey%20Casper%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=RLHF%3A%20A%20Survey&sort=relevance>)

221. **PPO: Proximal Policy Optimization Algorithms** — Schulman et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1707.06347>) · [Direct PDF](<https://arxiv.org/pdf/1707.06347.pdf>)

222. **Self-Instruct: Aligning LMs with Self-Generated Instructions** — Wang et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2212.10560>) · [Direct PDF](<https://arxiv.org/pdf/2212.10560.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #781.

223. **LIMA: Less Is More for Alignment** — Zhou et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.11206>) · [Direct PDF](<https://arxiv.org/pdf/2305.11206.pdf>)

224. **Scaling Instruction-Finetuned Language Models (Flan-T5/PaLM)** — Chung et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2210.11416>) · [Direct PDF](<https://arxiv.org/pdf/2210.11416.pdf>)

225. **FLAN: Finetuned Language Net** — Wei et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2109.01652>) · [Direct PDF](<https://arxiv.org/pdf/2109.01652.pdf>)

226. **The Flan Collection** — Longpre et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Flan%20Collection&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Flan%20Collection%20Longpre%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Flan%20Collection&sort=relevance>)

227. **Alpaca: A Strong, Replicable Instruction-Following Model** — Taori et al. · 2023 · **Type:** Research-paper candidate. [Supplied direct source](<https://crfm.stanford.edu/2023/03/13/alpaca.html>)

228. **Vicuna: An Open-Source Chatbot** — Chiang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Vicuna%3A%20An%20Open-Source%20Chatbot&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Vicuna%3A%20An%20Open-Source%20Chatbot%20Chiang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Vicuna%3A%20An%20Open-Source%20Chatbot&sort=relevance>)

229. **WizardLM: Empowering Large Language Models to Follow Complex Instructions** — Xu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=WizardLM%3A%20Empowering%20Large%20Language%20Models%20to%20Follow%20Complex%20Instructions&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=WizardLM%3A%20Empowering%20Large%20Language%20Models%20to%20Follow%20Complex%20Instructions%20Xu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=WizardLM%3A%20Empowering%20Large%20Language%20Models%20to%20Follow%20Complex%20Instructions&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #786.

230. **Orca: Progressive Learning from Complex Explanation Traces** — Mukherjee et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Orca%3A%20Progressive%20Learning%20from%20Complex%20Explanation%20Traces&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Orca%3A%20Progressive%20Learning%20from%20Complex%20Explanation%20Traces%20Mukherjee%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Orca%3A%20Progressive%20Learning%20from%20Complex%20Explanation%20Traces&sort=relevance>)

231. **Zephyr: Direct Distillation of LM Alignment** — Tunstall et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Zephyr%3A%20Direct%20Distillation%20of%20LM%20Alignment&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Zephyr%3A%20Direct%20Distillation%20of%20LM%20Alignment%20Tunstall%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Zephyr%3A%20Direct%20Distillation%20of%20LM%20Alignment&sort=relevance>)

232. **UltraFeedback** — Cui et al. · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=UltraFeedback&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=UltraFeedback%20Cui%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=UltraFeedback&sort=relevance>)

233. **Reinforcement Learning from AI Feedback (RLAIF)** — Lee et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Reinforcement%20Learning%20from%20AI%20Feedback%20%28RLAIF%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Reinforcement%20Learning%20from%20AI%20Feedback%20%28RLAIF%29%20Lee%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Reinforcement%20Learning%20from%20AI%20Feedback%20%28RLAIF%29&sort=relevance>)

234. **KTO: Model Alignment as Prospect Theoretic Optimization** — Ethayarajh et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2402.01306>) · [Direct PDF](<https://arxiv.org/pdf/2402.01306.pdf>)

235. **ORPO: Monolithic Preference Optimization without Reference Model** — Hong et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2403.07691>) · [Direct PDF](<https://arxiv.org/pdf/2403.07691.pdf>)

236. **SimPO: Simple Preference Optimization with a Reference-Free Reward** — Meng et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SimPO%3A%20Simple%20Preference%20Optimization%20with%20a%20Reference-Free%20Reward&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SimPO%3A%20Simple%20Preference%20Optimization%20with%20a%20Reference-Free%20Reward%20Meng%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SimPO%3A%20Simple%20Preference%20Optimization%20with%20a%20Reference-Free%20Reward&sort=relevance>)

237. **LoRA: Low-Rank Adaptation of Large Language Models** — Hu et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2106.09685>) · [Direct PDF](<https://arxiv.org/pdf/2106.09685.pdf>)

238. **QLoRA: Efficient Finetuning of Quantized LLMs** — Dettmers et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.14314>) · [Direct PDF](<https://arxiv.org/pdf/2305.14314.pdf>)

239. **Prefix-Tuning** — Li & Liang · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2101.00190>) · [Direct PDF](<https://arxiv.org/pdf/2101.00190.pdf>)

240. **P-Tuning v2** — Liu et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=P-Tuning%20v2&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=P-Tuning%20v2%20Liu%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=P-Tuning%20v2&sort=relevance>)

241. **The Power of Scale for Parameter-Efficient Prompt Tuning** — Lester, Al-Rfou, Constant · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Power%20of%20Scale%20for%20Parameter-Efficient%20Prompt%20Tuning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Power%20of%20Scale%20for%20Parameter-Efficient%20Prompt%20Tuning%20Lester%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Power%20of%20Scale%20for%20Parameter-Efficient%20Prompt%20Tuning&sort=relevance>)

242. **DoRA: Weight-Decomposed Low-Rank Adaptation** — Liu et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2402.09353>) · [Direct PDF](<https://arxiv.org/pdf/2402.09353.pdf>)

243. **Adapters: Parameter-Efficient Transfer Learning** — Houlsby et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Adapters%3A%20Parameter-Efficient%20Transfer%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Adapters%3A%20Parameter-Efficient%20Transfer%20Learning%20Houlsby%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Adapters%3A%20Parameter-Efficient%20Transfer%20Learning&sort=relevance>)

244. **NEFTune: Noisy Embeddings Improve Instruction Finetuning** — Jain et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=NEFTune%3A%20Noisy%20Embeddings%20Improve%20Instruction%20Finetuning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=NEFTune%3A%20Noisy%20Embeddings%20Improve%20Instruction%20Finetuning%20Jain%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=NEFTune%3A%20Noisy%20Embeddings%20Improve%20Instruction%20Finetuning&sort=relevance>)

245. **Self-Play Fine-Tuning (SPIN)** — Chen et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Self-Play%20Fine-Tuning%20%28SPIN%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Self-Play%20Fine-Tuning%20%28SPIN%29%20Chen%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Self-Play%20Fine-Tuning%20%28SPIN%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #788.

## 12. Prompt Engineering & In-Context Learning

246. **A Survey on In-Context Learning** — Dong et al. · 2023 · **Type:** Survey or review. [Canonical arXiv record](<https://arxiv.org/abs/2301.00234>) · [Direct PDF](<https://arxiv.org/pdf/2301.00234.pdf>)

247. **What Makes In-Context Learning Work?** — Min et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=What%20Makes%20In-Context%20Learning%20Work%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=What%20Makes%20In-Context%20Learning%20Work%3F%20Min%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=What%20Makes%20In-Context%20Learning%20Work%3F&sort=relevance>)

248. **Rethinking the Role of Demonstrations** — Min et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Rethinking%20the%20Role%20of%20Demonstrations&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Rethinking%20the%20Role%20of%20Demonstrations%20Min%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Rethinking%20the%20Role%20of%20Demonstrations&sort=relevance>)

249. **Fantastically Ordered Prompts and Where to Find Them** — Lu et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Fantastically%20Ordered%20Prompts%20and%20Where%20to%20Find%20Them&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Fantastically%20Ordered%20Prompts%20and%20Where%20to%20Find%20Them%20Lu%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Fantastically%20Ordered%20Prompts%20and%20Where%20to%20Find%20Them&sort=relevance>)

250. **Calibrate Before Use: Improving Few-Shot Performance of LMs** — Zhao et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Calibrate%20Before%20Use%3A%20Improving%20Few-Shot%20Performance%20of%20LMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Calibrate%20Before%20Use%3A%20Improving%20Few-Shot%20Performance%20of%20LMs%20Zhao%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Calibrate%20Before%20Use%3A%20Improving%20Few-Shot%20Performance%20of%20LMs&sort=relevance>)

251. **Large Language Models are Zero-Shot Reasoners** — Kojima et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2205.11916>) · [Direct PDF](<https://arxiv.org/pdf/2205.11916.pdf>)

252. **Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm** — Reynolds & McDonough · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Prompt%20Programming%20for%20Large%20Language%20Models%3A%20Beyond%20the%20Few-Shot%20Paradigm&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Prompt%20Programming%20for%20Large%20Language%20Models%3A%20Beyond%20the%20Few-Shot%20Paradigm%20Reynolds%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Prompt%20Programming%20for%20Large%20Language%20Models%3A%20Beyond%20the%20Few-Shot%20Paradigm&sort=relevance>)

253. **Structured Prompting: Scaling In-Context Learning to 1000 Examples** — Hao et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Structured%20Prompting%3A%20Scaling%20In-Context%20Learning%20to%201000%20Examples&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Structured%20Prompting%3A%20Scaling%20In-Context%20Learning%20to%201000%20Examples%20Hao%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Structured%20Prompting%3A%20Scaling%20In-Context%20Learning%20to%201000%20Examples&sort=relevance>)

254. **Many-Shot In-Context Learning** — Agarwal et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Many-Shot%20In-Context%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Many-Shot%20In-Context%20Learning%20Agarwal%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Many-Shot%20In-Context%20Learning&sort=relevance>)

255. **Pre-train, Prompt, and Predict: A Systematic Survey** — Liu et al. · 2023 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Pre-train%2C%20Prompt%2C%20and%20Predict%3A%20A%20Systematic%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Pre-train%2C%20Prompt%2C%20and%20Predict%3A%20A%20Systematic%20Survey%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Pre-train%2C%20Prompt%2C%20and%20Predict%3A%20A%20Systematic%20Survey&sort=relevance>)

256. **The Power of Prompting** — Microsoft · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Power%20of%20Prompting&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Power%20of%20Prompting%20Microsoft%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Power%20of%20Prompting&sort=relevance>)

257. **Meta-Prompting: Enhancing LLMs with the Expert Interviewing Technique** — Suzgun & Kalai · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Meta-Prompting%3A%20Enhancing%20LLMs%20with%20the%20Expert%20Interviewing%20Technique&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Meta-Prompting%3A%20Enhancing%20LLMs%20with%20the%20Expert%20Interviewing%20Technique%20Suzgun%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Meta-Prompting%3A%20Enhancing%20LLMs%20with%20the%20Expert%20Interviewing%20Technique&sort=relevance>)

258. **System 2 Attention** — Weston & Sukhbaatar · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=System%202%20Attention&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=System%202%20Attention%20Weston%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=System%202%20Attention&sort=relevance>)

259. **Anthropic's Prompt Engineering Guide** — Anthropic · 2024 · **Type:** Product documentation or standard. [arXiv search](<https://arxiv.org/search/?query=Anthropic%27s%20Prompt%20Engineering%20Guide&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Anthropic%27s%20Prompt%20Engineering%20Guide%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Anthropic%27s%20Prompt%20Engineering%20Guide&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as product documentation or standard, not as a research paper.

260. **OpenAI's Prompt Engineering Guide** — OpenAI · 2023 · **Type:** Product documentation or standard. [arXiv search](<https://arxiv.org/search/?query=OpenAI%27s%20Prompt%20Engineering%20Guide&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=OpenAI%27s%20Prompt%20Engineering%20Guide%20OpenAI%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=OpenAI%27s%20Prompt%20Engineering%20Guide&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as product documentation or standard, not as a research paper.

## 13. Reasoning & Chain-of-Thought

261. **Chain-of-Thought Prompting Elicits Reasoning in LLMs** — Wei et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2201.11903>) · [Direct PDF](<https://arxiv.org/pdf/2201.11903.pdf>)

262. **Self-Consistency Improves Chain of Thought Reasoning** — Wang et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2203.11171>) · [Direct PDF](<https://arxiv.org/pdf/2203.11171.pdf>)

263. **Tree of Thoughts: Deliberate Problem Solving with LLMs** — Yao et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.10601>) · [Direct PDF](<https://arxiv.org/pdf/2305.10601.pdf>)

264. **Graph of Thoughts** — Besta et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Graph%20of%20Thoughts&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Graph%20of%20Thoughts%20Besta%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Graph%20of%20Thoughts&sort=relevance>)

265. **Least-to-Most Prompting Enables Complex Reasoning** — Zhou et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Least-to-Most%20Prompting%20Enables%20Complex%20Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Least-to-Most%20Prompting%20Enables%20Complex%20Reasoning%20Zhou%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Least-to-Most%20Prompting%20Enables%20Complex%20Reasoning&sort=relevance>)

266. **Automatic Chain of Thought Prompting (Auto-CoT)** — Zhang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Automatic%20Chain%20of%20Thought%20Prompting%20%28Auto-CoT%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Automatic%20Chain%20of%20Thought%20Prompting%20%28Auto-CoT%29%20Zhang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Automatic%20Chain%20of%20Thought%20Prompting%20%28Auto-CoT%29&sort=relevance>)

267. **Complexity-Based Prompting for Multi-Step Reasoning** — Fu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Complexity-Based%20Prompting%20for%20Multi-Step%20Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Complexity-Based%20Prompting%20for%20Multi-Step%20Reasoning%20Fu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Complexity-Based%20Prompting%20for%20Multi-Step%20Reasoning&sort=relevance>)

268. **Program of Thoughts Prompting (PoT)** — Chen et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Program%20of%20Thoughts%20Prompting%20%28PoT%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Program%20of%20Thoughts%20Prompting%20%28PoT%29%20Chen%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Program%20of%20Thoughts%20Prompting%20%28PoT%29&sort=relevance>)

269. **PAL: Program-Aided Language Models** — Gao et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=PAL%3A%20Program-Aided%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=PAL%3A%20Program-Aided%20Language%20Models%20Gao%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=PAL%3A%20Program-Aided%20Language%20Models&sort=relevance>)

270. **Faithful Chain-of-Thought Reasoning** — Lyu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Faithful%20Chain-of-Thought%20Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Faithful%20Chain-of-Thought%20Reasoning%20Lyu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Faithful%20Chain-of-Thought%20Reasoning&sort=relevance>)

271. **Measuring Mathematical Problem Solving With the MATH Dataset** — Hendrycks et al. · 2021 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Measuring%20Mathematical%20Problem%20Solving%20With%20the%20MATH%20Dataset&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Measuring%20Mathematical%20Problem%20Solving%20With%20the%20MATH%20Dataset%20Hendrycks%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Measuring%20Mathematical%20Problem%20Solving%20With%20the%20MATH%20Dataset&sort=relevance>)

272. **GSM8K: Training Verifiers to Solve Math Word Problems** — Cobbe et al. · 2021 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=GSM8K%3A%20Training%20Verifiers%20to%20Solve%20Math%20Word%20Problems&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=GSM8K%3A%20Training%20Verifiers%20to%20Solve%20Math%20Word%20Problems%20Cobbe%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=GSM8K%3A%20Training%20Verifiers%20to%20Solve%20Math%20Word%20Problems&sort=relevance>)

273. **Let's Verify Step by Step** — Lightman et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.20050>) · [Direct PDF](<https://arxiv.org/pdf/2305.20050.pdf>)

274. **STaR: Self-Taught Reasoner** — Zelikman et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2203.14465>) · [Direct PDF](<https://arxiv.org/pdf/2203.14465.pdf>)

275. **Quiet-STaR: LMs Can Teach Themselves to Think Before Speaking** — Zelikman et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2403.09629>) · [Direct PDF](<https://arxiv.org/pdf/2403.09629.pdf>)

276. **Orca 2: Teaching Small LMs How to Reason** — Mitra et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Orca%202%3A%20Teaching%20Small%20LMs%20How%20to%20Reason&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Orca%202%3A%20Teaching%20Small%20LMs%20How%20to%20Reason%20Mitra%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Orca%202%3A%20Teaching%20Small%20LMs%20How%20to%20Reason&sort=relevance>)

277. **Reflexion: Language Agents with Verbal Reinforcement** — Shinn et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.11366>) · [Direct PDF](<https://arxiv.org/pdf/2303.11366.pdf>)

278. **o1 / Large Reasoning Models** — OpenAI · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=o1%20/%20Large%20Reasoning%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=o1%20/%20Large%20Reasoning%20Models%20OpenAI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=o1%20/%20Large%20Reasoning%20Models&sort=relevance>)

279. **DeepSeek-R1: Incentivizing Reasoning Capability in LLMs** — DeepSeek · 2025 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2501.12948>) · [Direct PDF](<https://arxiv.org/pdf/2501.12948.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #915.

280. **Scaling LLM Test-Time Compute Optimally** — Snell et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Scaling%20LLM%20Test-Time%20Compute%20Optimally&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scaling%20LLM%20Test-Time%20Compute%20Optimally%20Snell%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scaling%20LLM%20Test-Time%20Compute%20Optimally&sort=relevance>)

281. **Think before you speak: Training LMs with Pause Tokens** — Goyal et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Think%20before%20you%20speak%3A%20Training%20LMs%20with%20Pause%20Tokens&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Think%20before%20you%20speak%3A%20Training%20LMs%20with%20Pause%20Tokens%20Goyal%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Think%20before%20you%20speak%3A%20Training%20LMs%20with%20Pause%20Tokens&sort=relevance>)

282. **Journey Learning: From Superficial to Reasoning** — Wang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Journey%20Learning%3A%20From%20Superficial%20to%20Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Journey%20Learning%3A%20From%20Superficial%20to%20Reasoning%20Wang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Journey%20Learning%3A%20From%20Superficial%20to%20Reasoning&sort=relevance>)

283. **Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers** — Qi et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Mutual%20Reasoning%20Makes%20Smaller%20LLMs%20Stronger%20Problem-Solvers&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Mutual%20Reasoning%20Makes%20Smaller%20LLMs%20Stronger%20Problem-Solvers%20Qi%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Mutual%20Reasoning%20Makes%20Smaller%20LLMs%20Stronger%20Problem-Solvers&sort=relevance>)

284. **Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms** — Various · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Beyond%20Chain-of-Thought%3A%20A%20Survey%20of%20Chain-of-X%20Paradigms&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Beyond%20Chain-of-Thought%3A%20A%20Survey%20of%20Chain-of-X%20Paradigms%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Beyond%20Chain-of-Thought%3A%20A%20Survey%20of%20Chain-of-X%20Paradigms&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

285. **Rethinking LLM Reasoning: Are LLMs Truly Reasoning or Reciting?** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Rethinking%20LLM%20Reasoning%3A%20Are%20LLMs%20Truly%20Reasoning%20or%20Reciting%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Rethinking%20LLM%20Reasoning%3A%20Are%20LLMs%20Truly%20Reasoning%20or%20Reciting%3F%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Rethinking%20LLM%20Reasoning%3A%20Are%20LLMs%20Truly%20Reasoning%20or%20Reciting%3F&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 14. Retrieval-Augmented Generation (RAG)

286. **Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks** — Lewis et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2005.11401>) · [Direct PDF](<https://arxiv.org/pdf/2005.11401.pdf>)

287. **REALM: Retrieval-Augmented Language Model Pre-Training** — Guu et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=REALM%3A%20Retrieval-Augmented%20Language%20Model%20Pre-Training&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=REALM%3A%20Retrieval-Augmented%20Language%20Model%20Pre-Training%20Guu%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=REALM%3A%20Retrieval-Augmented%20Language%20Model%20Pre-Training&sort=relevance>)

288. **Dense Passage Retrieval (DPR)** — Karpukhin et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2004.04906>) · [Direct PDF](<https://arxiv.org/pdf/2004.04906.pdf>)

289. **ColBERT: Efficient and Effective Passage Search** — Khattab & Zaharia · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2004.12832>) · [Direct PDF](<https://arxiv.org/pdf/2004.12832.pdf>)

290. **ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction** — Santhanam et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ColBERTv2%3A%20Effective%20and%20Efficient%20Retrieval%20via%20Lightweight%20Late%20Interaction&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ColBERTv2%3A%20Effective%20and%20Efficient%20Retrieval%20via%20Lightweight%20Late%20Interaction%20Santhanam%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ColBERTv2%3A%20Effective%20and%20Efficient%20Retrieval%20via%20Lightweight%20Late%20Interaction&sort=relevance>)

291. **RETRO: Improving Language Models by Retrieving from Trillions of Tokens** — Borgeaud et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2112.04426>) · [Direct PDF](<https://arxiv.org/pdf/2112.04426.pdf>)

292. **Atlas: Few-shot Learning with Retrieval Augmented LMs** — Izacard et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Atlas%3A%20Few-shot%20Learning%20with%20Retrieval%20Augmented%20LMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Atlas%3A%20Few-shot%20Learning%20with%20Retrieval%20Augmented%20LMs%20Izacard%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Atlas%3A%20Few-shot%20Learning%20with%20Retrieval%20Augmented%20LMs&sort=relevance>)

293. **Self-RAG: Learning to Retrieve, Generate, and Critique** — Asai et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2310.11511>) · [Direct PDF](<https://arxiv.org/pdf/2310.11511.pdf>)

294. **Corrective RAG (CRAG)** — Yan et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Corrective%20RAG%20%28CRAG%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Corrective%20RAG%20%28CRAG%29%20Yan%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Corrective%20RAG%20%28CRAG%29&sort=relevance>)

295. **Adaptive RAG** — Jeong et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Adaptive%20RAG&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Adaptive%20RAG%20Jeong%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Adaptive%20RAG&sort=relevance>)

296. **RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval** — Sarthi et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2401.18059>) · [Direct PDF](<https://arxiv.org/pdf/2401.18059.pdf>)

297. **HyDE: Precise Zero-Shot Dense Retrieval without Relevance Labels** — Gao et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=HyDE%3A%20Precise%20Zero-Shot%20Dense%20Retrieval%20without%20Relevance%20Labels&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=HyDE%3A%20Precise%20Zero-Shot%20Dense%20Retrieval%20without%20Relevance%20Labels%20Gao%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=HyDE%3A%20Precise%20Zero-Shot%20Dense%20Retrieval%20without%20Relevance%20Labels&sort=relevance>)

298. **Lost in the Middle: How LMs Use Long Contexts** — Liu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Lost%20in%20the%20Middle%3A%20How%20LMs%20Use%20Long%20Contexts&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Lost%20in%20the%20Middle%3A%20How%20LMs%20Use%20Long%20Contexts%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Lost%20in%20the%20Middle%3A%20How%20LMs%20Use%20Long%20Contexts&sort=relevance>)

299. **Retrieval-Augmented Generation: A Survey** — Gao et al. · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Retrieval-Augmented%20Generation%3A%20A%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Retrieval-Augmented%20Generation%3A%20A%20Survey%20Gao%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Retrieval-Augmented%20Generation%3A%20A%20Survey&sort=relevance>)

300. **From RAG to Agentic RAG** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=From%20RAG%20to%20Agentic%20RAG&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=From%20RAG%20to%20Agentic%20RAG%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=From%20RAG%20to%20Agentic%20RAG&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

301. **Graph RAG: Unlocking LLM Discovery on Narrative Private Data** — Edge et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2404.16130>) · [Direct PDF](<https://arxiv.org/pdf/2404.16130.pdf>)

302. **Contextual Retrieval** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Contextual%20Retrieval&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Contextual%20Retrieval%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Contextual%20Retrieval&sort=relevance>)

303. **Embedding Models (E5, BGE, GTE)** — Various · 2023-2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Embedding%20Models%20%28E5%2C%20BGE%2C%20GTE%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Embedding%20Models%20%28E5%2C%20BGE%2C%20GTE%29%20Various%202023-2024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Embedding%20Models%20%28E5%2C%20BGE%2C%20GTE%29&sort=relevance>)
   - **Audit:** NON-PAPER OR MULTI-WORK RECORD — A multi-model topic container, not one citable work; the importer requires a four-digit year.

304. **MTEB: Massive Text Embedding Benchmark** — Muennighoff et al. · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=MTEB%3A%20Massive%20Text%20Embedding%20Benchmark&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MTEB%3A%20Massive%20Text%20Embedding%20Benchmark%20Muennighoff%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MTEB%3A%20Massive%20Text%20Embedding%20Benchmark&sort=relevance>)

305. **Matryoshka Representation Learning** — Kusupati et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Matryoshka%20Representation%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Matryoshka%20Representation%20Learning%20Kusupati%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Matryoshka%20Representation%20Learning&sort=relevance>)

## 15. Tool Use & Function Calling

306. **Toolformer: Language Models Can Teach Themselves to Use Tools** — Schick et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2302.04761>) · [Direct PDF](<https://arxiv.org/pdf/2302.04761.pdf>)

307. **WebGPT: Browser-Assisted Question-Answering** — Nakano et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=WebGPT%3A%20Browser-Assisted%20Question-Answering&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=WebGPT%3A%20Browser-Assisted%20Question-Answering%20Nakano%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=WebGPT%3A%20Browser-Assisted%20Question-Answering&sort=relevance>)

308. **ART: Automatic Multi-step Reasoning and Tool-use** — Paranjape et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ART%3A%20Automatic%20Multi-step%20Reasoning%20and%20Tool-use&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ART%3A%20Automatic%20Multi-step%20Reasoning%20and%20Tool-use%20Paranjape%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ART%3A%20Automatic%20Multi-step%20Reasoning%20and%20Tool-use&sort=relevance>)

309. **Gorilla: Large Language Model Connected with Massive APIs** — Patil et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.15334>) · [Direct PDF](<https://arxiv.org/pdf/2305.15334.pdf>)

310. **ToolLLM: Facilitating LLMs to Master 16000+ APIs** — Qin et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ToolLLM%3A%20Facilitating%20LLMs%20to%20Master%2016000%2B%20APIs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ToolLLM%3A%20Facilitating%20LLMs%20to%20Master%2016000%2B%20APIs%20Qin%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ToolLLM%3A%20Facilitating%20LLMs%20to%20Master%2016000%2B%20APIs&sort=relevance>)

311. **API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs** — Li et al. · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=API-Bank%3A%20A%20Comprehensive%20Benchmark%20for%20Tool-Augmented%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=API-Bank%3A%20A%20Comprehensive%20Benchmark%20for%20Tool-Augmented%20LLMs%20Li%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=API-Bank%3A%20A%20Comprehensive%20Benchmark%20for%20Tool-Augmented%20LLMs&sort=relevance>)

312. **TaskMatrix.AI: Completing Tasks by Connecting Foundation Models** — Liang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=TaskMatrix.AI%3A%20Completing%20Tasks%20by%20Connecting%20Foundation%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=TaskMatrix.AI%3A%20Completing%20Tasks%20by%20Connecting%20Foundation%20Models%20Liang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=TaskMatrix.AI%3A%20Completing%20Tasks%20by%20Connecting%20Foundation%20Models&sort=relevance>)

313. **HuggingGPT: Solving AI Tasks with ChatGPT and HuggingFace** — Shen et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.17580>) · [Direct PDF](<https://arxiv.org/pdf/2303.17580.pdf>)

314. **Chameleon: Plug-and-Play Compositional Reasoning** — Lu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Chameleon%3A%20Plug-and-Play%20Compositional%20Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Chameleon%3A%20Plug-and-Play%20Compositional%20Reasoning%20Lu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Chameleon%3A%20Plug-and-Play%20Compositional%20Reasoning&sort=relevance>)

315. **NexusRaven: Function Calling LLM** — Nexusflow · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=NexusRaven%3A%20Function%20Calling%20LLM&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=NexusRaven%3A%20Function%20Calling%20LLM%20Nexusflow%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=NexusRaven%3A%20Function%20Calling%20LLM&sort=relevance>)

316. **Anthropic's Tool Use Documentation** — Anthropic · 2024 · **Type:** Product documentation or standard. [arXiv search](<https://arxiv.org/search/?query=Anthropic%27s%20Tool%20Use%20Documentation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Anthropic%27s%20Tool%20Use%20Documentation%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Anthropic%27s%20Tool%20Use%20Documentation&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as product documentation or standard, not as a research paper.

317. **Model Context Protocol (MCP)** — Anthropic · 2024 · **Type:** Product documentation or standard. [arXiv search](<https://arxiv.org/search/?query=Model%20Context%20Protocol%20%28MCP%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Model%20Context%20Protocol%20%28MCP%29%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Model%20Context%20Protocol%20%28MCP%29&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as product documentation or standard, not as a research paper.

318. **RestGPT: Connecting LLMs with Real-World RESTful APIs** — Song et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=RestGPT%3A%20Connecting%20LLMs%20with%20Real-World%20RESTful%20APIs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=RestGPT%3A%20Connecting%20LLMs%20with%20Real-World%20RESTful%20APIs%20Song%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=RestGPT%3A%20Connecting%20LLMs%20with%20Real-World%20RESTful%20APIs&sort=relevance>)

319. **MRKL Systems** — Karpas et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MRKL%20Systems&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MRKL%20Systems%20Karpas%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MRKL%20Systems&sort=relevance>)

320. **Faithful Reasoning Using Large Language Models** — Creswell & Shanahan · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Faithful%20Reasoning%20Using%20Large%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Faithful%20Reasoning%20Using%20Large%20Language%20Models%20Creswell%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Faithful%20Reasoning%20Using%20Large%20Language%20Models&sort=relevance>)

## 16. AI Agents — Foundations

321. **A Survey on Large Language Model based Autonomous Agents** — Wang et al. · 2023 · **Type:** Survey or review. [Canonical arXiv record](<https://arxiv.org/abs/2308.11432>) · [Direct PDF](<https://arxiv.org/pdf/2308.11432.pdf>)

322. **ReAct: Synergizing Reasoning and Acting in LMs** — Yao et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2210.03629>) · [Direct PDF](<https://arxiv.org/pdf/2210.03629.pdf>)

323. **Voyager: An Open-Ended Embodied Agent with LLMs** — Wang et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.16291>) · [Direct PDF](<https://arxiv.org/pdf/2305.16291.pdf>)

324. **Generative Agents: Interactive Simulacra of Human Behavior** — Park et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2304.03442>) · [Direct PDF](<https://arxiv.org/pdf/2304.03442.pdf>)

325. **AutoGPT (technical overview)** — Richards · 2023 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=AutoGPT%20%28technical%20overview%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AutoGPT%20%28technical%20overview%29%20Richards%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AutoGPT%20%28technical%20overview%29&sort=relevance>)

326. **BabyAGI** — Nakajima · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=BabyAGI&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=BabyAGI%20Nakajima%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=BabyAGI&sort=relevance>)

327. **LangChain Agents (technical paper)** — Chase · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LangChain%20Agents%20%28technical%20paper%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LangChain%20Agents%20%28technical%20paper%29%20Chase%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LangChain%20Agents%20%28technical%20paper%29&sort=relevance>)

328. **AgentBench: Evaluating LLMs as Agents** — Liu et al. · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=AgentBench%3A%20Evaluating%20LLMs%20as%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AgentBench%3A%20Evaluating%20LLMs%20as%20Agents%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AgentBench%3A%20Evaluating%20LLMs%20as%20Agents&sort=relevance>)

329. **BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents** — Liu et al. · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=BOLAA%3A%20Benchmarking%20and%20Orchestrating%20LLM-augmented%20Autonomous%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=BOLAA%3A%20Benchmarking%20and%20Orchestrating%20LLM-augmented%20Autonomous%20Agents%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=BOLAA%3A%20Benchmarking%20and%20Orchestrating%20LLM-augmented%20Autonomous%20Agents&sort=relevance>)

330. **The Rise and Potential of LLM Based Agents: A Survey** — Xi et al. · 2023 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=The%20Rise%20and%20Potential%20of%20LLM%20Based%20Agents%3A%20A%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Rise%20and%20Potential%20of%20LLM%20Based%20Agents%3A%20A%20Survey%20Xi%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Rise%20and%20Potential%20of%20LLM%20Based%20Agents%3A%20A%20Survey&sort=relevance>)

331. **Agent Foundation Models** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Agent%20Foundation%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Agent%20Foundation%20Models%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Agent%20Foundation%20Models&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

332. **OpenAgents: An Open Platform for Language Agents** — Xie et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=OpenAgents%3A%20An%20Open%20Platform%20for%20Language%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=OpenAgents%3A%20An%20Open%20Platform%20for%20Language%20Agents%20Xie%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=OpenAgents%3A%20An%20Open%20Platform%20for%20Language%20Agents&sort=relevance>)

333. **AgentTuning: Enabling Generalized Agent Abilities For LLMs** — Zeng et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AgentTuning%3A%20Enabling%20Generalized%20Agent%20Abilities%20For%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AgentTuning%3A%20Enabling%20Generalized%20Agent%20Abilities%20For%20LLMs%20Zeng%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AgentTuning%3A%20Enabling%20Generalized%20Agent%20Abilities%20For%20LLMs&sort=relevance>)

334. **Cognitive Architectures for Language Agents (CoALA)** — Sumers et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2309.02427>) · [Direct PDF](<https://arxiv.org/pdf/2309.02427.pdf>)

335. **Language Agents: From Next-Token Prediction to Digital Automation** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Language%20Agents%3A%20From%20Next-Token%20Prediction%20to%20Digital%20Automation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Language%20Agents%3A%20From%20Next-Token%20Prediction%20to%20Digital%20Automation%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Language%20Agents%3A%20From%20Next-Token%20Prediction%20to%20Digital%20Automation&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

336. **Agent-as-a-Judge** — Zhuge et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Agent-as-a-Judge&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Agent-as-a-Judge%20Zhuge%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Agent-as-a-Judge&sort=relevance>)

337. **AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agents** — Liu et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AgentLite%3A%20A%20Lightweight%20Library%20for%20Building%20and%20Advancing%20Task-Oriented%20LLM%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AgentLite%3A%20A%20Lightweight%20Library%20for%20Building%20and%20Advancing%20Task-Oriented%20LLM%20Agents%20Liu%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AgentLite%3A%20A%20Lightweight%20Library%20for%20Building%20and%20Advancing%20Task-Oriented%20LLM%20Agents&sort=relevance>)

338. **Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security** — Li et al. · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Personal%20LLM%20Agents%3A%20Insights%20and%20Survey%20about%20the%20Capability%2C%20Efficiency%20and%20Security&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Personal%20LLM%20Agents%3A%20Insights%20and%20Survey%20about%20the%20Capability%2C%20Efficiency%20and%20Security%20Li%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Personal%20LLM%20Agents%3A%20Insights%20and%20Survey%20about%20the%20Capability%2C%20Efficiency%20and%20Security&sort=relevance>)

339. **Agent Smith: Scaling LLM Agents** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Agent%20Smith%3A%20Scaling%20LLM%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Agent%20Smith%3A%20Scaling%20LLM%20Agents%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Agent%20Smith%3A%20Scaling%20LLM%20Agents&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

340. **Internet of Agents** — Chen et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Internet%20of%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Internet%20of%20Agents%20Chen%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Internet%20of%20Agents&sort=relevance>)

## 17. AI Agents — Planning & Reasoning

341. **Planning with Large Language Models for Code Generation** — Zhang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Planning%20with%20Large%20Language%20Models%20for%20Code%20Generation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Planning%20with%20Large%20Language%20Models%20for%20Code%20Generation%20Zhang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Planning%20with%20Large%20Language%20Models%20for%20Code%20Generation&sort=relevance>)

342. **LLM+P: Empowering LLMs with Optimal Planning Proficiency** — Liu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLM%2BP%3A%20Empowering%20LLMs%20with%20Optimal%20Planning%20Proficiency&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLM%2BP%3A%20Empowering%20LLMs%20with%20Optimal%20Planning%20Proficiency%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLM%2BP%3A%20Empowering%20LLMs%20with%20Optimal%20Planning%20Proficiency&sort=relevance>)

343. **Describe, Explain, Plan and Select (DEPS)** — Wang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Describe%2C%20Explain%2C%20Plan%20and%20Select%20%28DEPS%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Describe%2C%20Explain%2C%20Plan%20and%20Select%20%28DEPS%29%20Wang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Describe%2C%20Explain%2C%20Plan%20and%20Select%20%28DEPS%29&sort=relevance>)

344. **Plan-and-Solve Prompting** — Wang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Plan-and-Solve%20Prompting&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Plan-and-Solve%20Prompting%20Wang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Plan-and-Solve%20Prompting&sort=relevance>)

345. **AdaPlanner: Adaptive Planning from Feedback** — Sun et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AdaPlanner%3A%20Adaptive%20Planning%20from%20Feedback&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AdaPlanner%3A%20Adaptive%20Planning%20from%20Feedback%20Sun%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AdaPlanner%3A%20Adaptive%20Planning%20from%20Feedback&sort=relevance>)

346. **LATS: Language Agent Tree Search** — Zhou et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2310.04406>) · [Direct PDF](<https://arxiv.org/pdf/2310.04406.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #758.

347. **Inner Monologue: Embodied Reasoning through Planning with Language Models** — Huang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Inner%20Monologue%3A%20Embodied%20Reasoning%20through%20Planning%20with%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Inner%20Monologue%3A%20Embodied%20Reasoning%20through%20Planning%20with%20Language%20Models%20Huang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Inner%20Monologue%3A%20Embodied%20Reasoning%20through%20Planning%20with%20Language%20Models&sort=relevance>)

348. **SayPlan: Grounding Large Language Models using 3D Scene Graphs** — Rana et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SayPlan%3A%20Grounding%20Large%20Language%20Models%20using%203D%20Scene%20Graphs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SayPlan%3A%20Grounding%20Large%20Language%20Models%20using%203D%20Scene%20Graphs%20Rana%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SayPlan%3A%20Grounding%20Large%20Language%20Models%20using%203D%20Scene%20Graphs&sort=relevance>)

349. **LLM Powered Autonomous Agents (Lilian Weng blog-paper)** — Weng · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLM%20Powered%20Autonomous%20Agents%20%28Lilian%20Weng%20blog-paper%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLM%20Powered%20Autonomous%20Agents%20%28Lilian%20Weng%20blog-paper%29%20Weng%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLM%20Powered%20Autonomous%20Agents%20%28Lilian%20Weng%20blog-paper%29&sort=relevance>)

350. **Reasoning with Language Model is Planning with World Model** — Hao et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Reasoning%20with%20Language%20Model%20is%20Planning%20with%20World%20Model&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Reasoning%20with%20Language%20Model%20is%20Planning%20with%20World%20Model%20Hao%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Reasoning%20with%20Language%20Model%20is%20Planning%20with%20World%20Model&sort=relevance>)

351. **Everything of Thoughts: Defying the Law of Penrose Triangle** — Ding et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Everything%20of%20Thoughts%3A%20Defying%20the%20Law%20of%20Penrose%20Triangle&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Everything%20of%20Thoughts%3A%20Defying%20the%20Law%20of%20Penrose%20Triangle%20Ding%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Everything%20of%20Thoughts%3A%20Defying%20the%20Law%20of%20Penrose%20Triangle&sort=relevance>)

352. **SWE-Agent: Agent-Computer Interfaces Enable Automated Software Engineering** — Yang et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2405.15793>) · [Direct PDF](<https://arxiv.org/pdf/2405.15793.pdf>)

353. **Trial and Error: Exploration-Based Trajectory Optimization** — Wang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Trial%20and%20Error%3A%20Exploration-Based%20Trajectory%20Optimization&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Trial%20and%20Error%3A%20Exploration-Based%20Trajectory%20Optimization%20Wang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Trial%20and%20Error%3A%20Exploration-Based%20Trajectory%20Optimization&sort=relevance>)

354. **AgentCoder: Multi-Agent-based Code Generation** — Huang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AgentCoder%3A%20Multi-Agent-based%20Code%20Generation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AgentCoder%3A%20Multi-Agent-based%20Code%20Generation%20Huang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AgentCoder%3A%20Multi-Agent-based%20Code%20Generation&sort=relevance>)

355. **Lumos: Learning Agents with Unified Data, Modular Design, and Open-Source LLMs** — Yin et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Lumos%3A%20Learning%20Agents%20with%20Unified%20Data%2C%20Modular%20Design%2C%20and%20Open-Source%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Lumos%3A%20Learning%20Agents%20with%20Unified%20Data%2C%20Modular%20Design%2C%20and%20Open-Source%20LLMs%20Yin%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Lumos%3A%20Learning%20Agents%20with%20Unified%20Data%2C%20Modular%20Design%2C%20and%20Open-Source%20LLMs&sort=relevance>)

## 18. Multi-Agent Systems

356. **CAMEL: Communicative Agents for "Mind" Exploration of LLM Society** — Li et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.17760>) · [Direct PDF](<https://arxiv.org/pdf/2303.17760.pdf>)

357. **AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation** — Wu et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2308.08155>) · [Direct PDF](<https://arxiv.org/pdf/2308.08155.pdf>)

358. **MetaGPT: Meta Programming for Multi-Agent Collaborative Framework** — Hong et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2308.00352>) · [Direct PDF](<https://arxiv.org/pdf/2308.00352.pdf>)

359. **ChatDev: Communicative Agents for Software Development** — Qian et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ChatDev%3A%20Communicative%20Agents%20for%20Software%20Development&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ChatDev%3A%20Communicative%20Agents%20for%20Software%20Development%20Qian%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ChatDev%3A%20Communicative%20Agents%20for%20Software%20Development&sort=relevance>)

360. **AgentVerse: Facilitating Multi-Agent Collaboration** — Chen et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AgentVerse%3A%20Facilitating%20Multi-Agent%20Collaboration&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AgentVerse%3A%20Facilitating%20Multi-Agent%20Collaboration%20Chen%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AgentVerse%3A%20Facilitating%20Multi-Agent%20Collaboration&sort=relevance>)

361. **Multi-Agent Debate** — Du et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Multi-Agent%20Debate&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Multi-Agent%20Debate%20Du%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Multi-Agent%20Debate&sort=relevance>)

362. **Improving Factuality and Reasoning via Multi-Agent Debate** — Liang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Improving%20Factuality%20and%20Reasoning%20via%20Multi-Agent%20Debate&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Improving%20Factuality%20and%20Reasoning%20via%20Multi-Agent%20Debate%20Liang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Improving%20Factuality%20and%20Reasoning%20via%20Multi-Agent%20Debate&sort=relevance>)

363. **Dynamic LLM-Agent Network: Mixture of Agents** — Wang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Dynamic%20LLM-Agent%20Network%3A%20Mixture%20of%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Dynamic%20LLM-Agent%20Network%3A%20Mixture%20of%20Agents%20Wang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Dynamic%20LLM-Agent%20Network%3A%20Mixture%20of%20Agents&sort=relevance>)

364. **Mixture of Agents (MoA)** — Together AI · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Mixture%20of%20Agents%20%28MoA%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Mixture%20of%20Agents%20%28MoA%29%20Together%20AI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Mixture%20of%20Agents%20%28MoA%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #906.

365. **CrewAI (technical overview)** — Moura · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=CrewAI%20%28technical%20overview%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=CrewAI%20%28technical%20overview%29%20Moura%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=CrewAI%20%28technical%20overview%29&sort=relevance>)

366. **Society of Mind (SoM) for LLMs** — Zhuge et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Society%20of%20Mind%20%28SoM%29%20for%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Society%20of%20Mind%20%28SoM%29%20for%20LLMs%20Zhuge%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Society%20of%20Mind%20%28SoM%29%20for%20LLMs&sort=relevance>)

367. **Corex: Pushing the Boundaries of Complex Reasoning through Multi-Agent Collaboration** — Sun et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Corex%3A%20Pushing%20the%20Boundaries%20of%20Complex%20Reasoning%20through%20Multi-Agent%20Collaboration&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Corex%3A%20Pushing%20the%20Boundaries%20of%20Complex%20Reasoning%20through%20Multi-Agent%20Collaboration%20Sun%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Corex%3A%20Pushing%20the%20Boundaries%20of%20Complex%20Reasoning%20through%20Multi-Agent%20Collaboration&sort=relevance>)

368. **MAD: Multi-Agent Debate with Large Language Models** — Liang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MAD%3A%20Multi-Agent%20Debate%20with%20Large%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MAD%3A%20Multi-Agent%20Debate%20with%20Large%20Language%20Models%20Liang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MAD%3A%20Multi-Agent%20Debate%20with%20Large%20Language%20Models&sort=relevance>)

369. **AgentScope: A Flexible yet Robust Multi-Agent Platform** — Gao et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AgentScope%3A%20A%20Flexible%20yet%20Robust%20Multi-Agent%20Platform&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AgentScope%3A%20A%20Flexible%20yet%20Robust%20Multi-Agent%20Platform%20Gao%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AgentScope%3A%20A%20Flexible%20yet%20Robust%20Multi-Agent%20Platform&sort=relevance>)

370. **LLM-Blender: Ensembling LLMs with Pairwise Ranking** — Jiang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLM-Blender%3A%20Ensembling%20LLMs%20with%20Pairwise%20Ranking&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLM-Blender%3A%20Ensembling%20LLMs%20with%20Pairwise%20Ranking%20Jiang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLM-Blender%3A%20Ensembling%20LLMs%20with%20Pairwise%20Ranking&sort=relevance>)

## 19. Code Generation & Software Engineering Agents

371. **Evaluating Large Language Models Trained on Code (Codex)** — Chen et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2107.03374>) · [Direct PDF](<https://arxiv.org/pdf/2107.03374.pdf>)

372. **HumanEval** — Chen et al. · 2021 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=HumanEval&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=HumanEval%20Chen%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=HumanEval&sort=relevance>)

373. **CodeGen: An Open Large Language Model for Code** — Nijkamp et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=CodeGen%3A%20An%20Open%20Large%20Language%20Model%20for%20Code&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=CodeGen%3A%20An%20Open%20Large%20Language%20Model%20for%20Code%20Nijkamp%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=CodeGen%3A%20An%20Open%20Large%20Language%20Model%20for%20Code&sort=relevance>)

374. **StarCoder: May the Source Be with You!** — Li et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.06161>) · [Direct PDF](<https://arxiv.org/pdf/2305.06161.pdf>)

375. **StarCoder 2** — Lozhkov et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=StarCoder%202&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=StarCoder%202%20Lozhkov%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=StarCoder%202&sort=relevance>)

376. **Code Llama: Open Foundation Models for Code** — Rozière et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2308.12950>) · [Direct PDF](<https://arxiv.org/pdf/2308.12950.pdf>)

377. **DeepSeek-Coder** — Guo et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DeepSeek-Coder&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DeepSeek-Coder%20Guo%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DeepSeek-Coder&sort=relevance>)

378. **WizardCoder: Empowering Code LLMs with Evol-Instruct** — Luo et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=WizardCoder%3A%20Empowering%20Code%20LLMs%20with%20Evol-Instruct&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=WizardCoder%3A%20Empowering%20Code%20LLMs%20with%20Evol-Instruct%20Luo%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=WizardCoder%3A%20Empowering%20Code%20LLMs%20with%20Evol-Instruct&sort=relevance>)

379. **SWE-bench: Can Language Models Resolve Real-World GitHub Issues?** — Jimenez et al. · 2024 · **Type:** Benchmark or dataset paper. [Canonical arXiv record](<https://arxiv.org/abs/2310.06770>) · [Direct PDF](<https://arxiv.org/pdf/2310.06770.pdf>)

380. **SWE-Agent** — Yang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SWE-Agent&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SWE-Agent%20Yang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SWE-Agent&sort=relevance>)

381. **Devin: AI Software Engineer** — Cognition · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Devin%3A%20AI%20Software%20Engineer&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Devin%3A%20AI%20Software%20Engineer%20Cognition%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Devin%3A%20AI%20Software%20Engineer&sort=relevance>)

382. **OpenDevin / OpenHands** — Wang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=OpenDevin%20/%20OpenHands&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=OpenDevin%20/%20OpenHands%20Wang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=OpenDevin%20/%20OpenHands&sort=relevance>)

383. **Aider: AI Pair Programming in Terminal** — Gauthier · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Aider%3A%20AI%20Pair%20Programming%20in%20Terminal&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Aider%3A%20AI%20Pair%20Programming%20in%20Terminal%20Gauthier%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Aider%3A%20AI%20Pair%20Programming%20in%20Terminal&sort=relevance>)

384. **Claude Code (Anthropic)** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Claude%20Code%20%28Anthropic%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Claude%20Code%20%28Anthropic%29%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Claude%20Code%20%28Anthropic%29&sort=relevance>)

385. **AlphaCode: Competition-Level Code Generation** — Li et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2203.07814>) · [Direct PDF](<https://arxiv.org/pdf/2203.07814.pdf>)

386. **AlphaCode 2** — Google DeepMind · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AlphaCode%202&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AlphaCode%202%20Google%20DeepMind%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AlphaCode%202&sort=relevance>)

387. **CodeT: Code Generation with Generated Tests** — Chen et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=CodeT%3A%20Code%20Generation%20with%20Generated%20Tests&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=CodeT%3A%20Code%20Generation%20with%20Generated%20Tests%20Chen%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=CodeT%3A%20Code%20Generation%20with%20Generated%20Tests&sort=relevance>)

388. **Self-Debugging: Teaching LLMs to Self-Debug** — Chen et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Self-Debugging%3A%20Teaching%20LLMs%20to%20Self-Debug&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Self-Debugging%3A%20Teaching%20LLMs%20to%20Self-Debug%20Chen%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Self-Debugging%3A%20Teaching%20LLMs%20to%20Self-Debug&sort=relevance>)

389. **MBPP: Mostly Basic Python Programming** — Austin et al. · 2021 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=MBPP%3A%20Mostly%20Basic%20Python%20Programming&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MBPP%3A%20Mostly%20Basic%20Python%20Programming%20Austin%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MBPP%3A%20Mostly%20Basic%20Python%20Programming&sort=relevance>)

390. **Repository-Level Code Completion** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Repository-Level%20Code%20Completion&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Repository-Level%20Code%20Completion%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Repository-Level%20Code%20Completion&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 20. Reinforcement Learning Foundations

391. **Reinforcement Learning: An Introduction** — Sutton & Barto · 1998/2018 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=Reinforcement%20Learning%3A%20An%20Introduction&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Reinforcement%20Learning%3A%20An%20Introduction%20Sutton%201998/2018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Reinforcement%20Learning%3A%20An%20Introduction&sort=relevance>)
   - **Audit:** NON-PAPER OR MULTI-WORK RECORD — Two book editions are combined; the importer requires a single four-digit year.

392. **Q-Learning** — Watkins & Dayan · 1992 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Q-Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Q-Learning%20Watkins%201992>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Q-Learning&sort=relevance>)

393. **Policy Gradient Methods for RL with Function Approximation** — Sutton et al. · 2000 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Policy%20Gradient%20Methods%20for%20RL%20with%20Function%20Approximation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Policy%20Gradient%20Methods%20for%20RL%20with%20Function%20Approximation%20Sutton%202000>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Policy%20Gradient%20Methods%20for%20RL%20with%20Function%20Approximation&sort=relevance>)

394. **Actor-Critic Algorithms** — Konda & Tsitsiklis · 2000 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Actor-Critic%20Algorithms&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Actor-Critic%20Algorithms%20Konda%202000>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Actor-Critic%20Algorithms&sort=relevance>)

395. **Simple Statistical Gradient-Following Algorithms (REINFORCE)** — Williams · 1992 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Simple%20Statistical%20Gradient-Following%20Algorithms%20%28REINFORCE%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Simple%20Statistical%20Gradient-Following%20Algorithms%20%28REINFORCE%29%20Williams%201992>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Simple%20Statistical%20Gradient-Following%20Algorithms%20%28REINFORCE%29&sort=relevance>)

396. **Temporal-Difference Learning** — Sutton · 1988 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Temporal-Difference%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Temporal-Difference%20Learning%20Sutton%201988>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Temporal-Difference%20Learning&sort=relevance>)

397. **Between MDPs and semi-MDPs: A Framework for Temporal Abstraction (Options)** — Sutton, Precup, Singh · 1999 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Between%20MDPs%20and%20semi-MDPs%3A%20A%20Framework%20for%20Temporal%20Abstraction%20%28Options%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Between%20MDPs%20and%20semi-MDPs%3A%20A%20Framework%20for%20Temporal%20Abstraction%20%28Options%29%20Sutton%201999>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Between%20MDPs%20and%20semi-MDPs%3A%20A%20Framework%20for%20Temporal%20Abstraction%20%28Options%29&sort=relevance>)

398. **Multi-Agent Reinforcement Learning: A Selective Overview** — Zhang et al. · 2021 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Multi-Agent%20Reinforcement%20Learning%3A%20A%20Selective%20Overview&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Multi-Agent%20Reinforcement%20Learning%3A%20A%20Selective%20Overview%20Zhang%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Multi-Agent%20Reinforcement%20Learning%3A%20A%20Selective%20Overview&sort=relevance>)

399. **Reward Shaping** — Ng, Harada, Russell · 1999 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Reward%20Shaping&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Reward%20Shaping%20Ng%201999>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Reward%20Shaping&sort=relevance>)

400. **Exploration and Exploitation in RL** — various · 2000s · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Exploration%20and%20Exploitation%20in%20RL&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Exploration%20and%20Exploitation%20in%20RL%20various%202000s>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Exploration%20and%20Exploitation%20in%20RL&sort=relevance>)
   - **Audit:** NON-PAPER OR MULTI-WORK RECORD — A topic label rather than a uniquely identifiable publication.

## 21. Deep Reinforcement Learning

401. **Playing Atari with Deep Reinforcement Learning (DQN)** — Mnih et al. · 2013 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1312.5602>) · [Direct PDF](<https://arxiv.org/pdf/1312.5602.pdf>)

402. **Human-Level Control through Deep RL (DQN Nature)** — Mnih et al. · 2015 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/nature14236>)

403. **Deep Reinforcement Learning with Double Q-Learning** — van Hasselt et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1509.06461>) · [Direct PDF](<https://arxiv.org/pdf/1509.06461.pdf>)

404. **Prioritized Experience Replay** — Schaul et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1511.05952>) · [Direct PDF](<https://arxiv.org/pdf/1511.05952.pdf>)

405. **Dueling Network Architectures for Deep RL** — Wang et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Dueling%20Network%20Architectures%20for%20Deep%20RL&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Dueling%20Network%20Architectures%20for%20Deep%20RL%20Wang%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Dueling%20Network%20Architectures%20for%20Deep%20RL&sort=relevance>)

406. **Asynchronous Methods for Deep RL (A3C)** — Mnih et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1602.01783>) · [Direct PDF](<https://arxiv.org/pdf/1602.01783.pdf>)

407. **Continuous Control with Deep RL (DDPG)** — Lillicrap et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1509.02971>) · [Direct PDF](<https://arxiv.org/pdf/1509.02971.pdf>)

408. **Trust Region Policy Optimization (TRPO)** — Schulman et al. · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1502.05477>) · [Direct PDF](<https://arxiv.org/pdf/1502.05477.pdf>)

409. **High-Dimensional Continuous Control Using Generalized Advantage Estimation (GAE)** — Schulman et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=High-Dimensional%20Continuous%20Control%20Using%20Generalized%20Advantage%20Estimation%20%28GAE%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=High-Dimensional%20Continuous%20Control%20Using%20Generalized%20Advantage%20Estimation%20%28GAE%29%20Schulman%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=High-Dimensional%20Continuous%20Control%20Using%20Generalized%20Advantage%20Estimation%20%28GAE%29&sort=relevance>)

410. **Soft Actor-Critic (SAC)** — Haarnoja et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1801.01290>) · [Direct PDF](<https://arxiv.org/pdf/1801.01290.pdf>)

411. **Mastering the Game of Go with Deep Neural Networks (AlphaGo)** — Silver et al. · 2016 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/nature16961>)

412. **Mastering Go without Human Knowledge (AlphaGo Zero)** — Silver et al. · 2017 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/nature24270>)

413. **A General RL Algorithm that Masters Chess, Shogi, and Go (AlphaZero)** — Silver et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1712.01815>) · [Direct PDF](<https://arxiv.org/pdf/1712.01815.pdf>)

414. **MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model** — Schrittwieser et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1911.08265>) · [Direct PDF](<https://arxiv.org/pdf/1911.08265.pdf>)

415. **OpenAI Five** — OpenAI · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=OpenAI%20Five&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=OpenAI%20Five%20OpenAI%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=OpenAI%20Five&sort=relevance>)

416. **AlphaStar: Mastering StarCraft II** — Vinyals et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AlphaStar%3A%20Mastering%20StarCraft%20II&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AlphaStar%3A%20Mastering%20StarCraft%20II%20Vinyals%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AlphaStar%3A%20Mastering%20StarCraft%20II&sort=relevance>)

417. **Curiosity-Driven Exploration (ICM)** — Pathak et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1705.05363>) · [Direct PDF](<https://arxiv.org/pdf/1705.05363.pdf>)

418. **World Models** — Ha & Schmidhuber · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1803.10122>) · [Direct PDF](<https://arxiv.org/pdf/1803.10122.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #691.

419. **Dream to Control (Dreamer)** — Hafner et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Dream%20to%20Control%20%28Dreamer%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Dream%20to%20Control%20%28Dreamer%29%20Hafner%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Dream%20to%20Control%20%28Dreamer%29&sort=relevance>)

420. **DreamerV3: Mastering Diverse Domains through World Models** — Hafner et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2301.04104>) · [Direct PDF](<https://arxiv.org/pdf/2301.04104.pdf>)

## 22. Multimodal AI

421. **Multimodal Learning with Deep Boltzmann Machines** — Srivastava & Salakhutdinov · 2012 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Multimodal%20Learning%20with%20Deep%20Boltzmann%20Machines&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Multimodal%20Learning%20with%20Deep%20Boltzmann%20Machines%20Srivastava%202012>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Multimodal%20Learning%20with%20Deep%20Boltzmann%20Machines&sort=relevance>)

422. **Visual Question Answering (VQA)** — Antol et al. · 2015 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Visual%20Question%20Answering%20%28VQA%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Visual%20Question%20Answering%20%28VQA%29%20Antol%202015>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Visual%20Question%20Answering%20%28VQA%29&sort=relevance>)

423. **ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations** — Lu et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ViLBERT%3A%20Pretraining%20Task-Agnostic%20Visiolinguistic%20Representations&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ViLBERT%3A%20Pretraining%20Task-Agnostic%20Visiolinguistic%20Representations%20Lu%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ViLBERT%3A%20Pretraining%20Task-Agnostic%20Visiolinguistic%20Representations&sort=relevance>)

424. **LXMERT: Learning Cross-Modality Encoder Representations** — Tan & Bansal · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LXMERT%3A%20Learning%20Cross-Modality%20Encoder%20Representations&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LXMERT%3A%20Learning%20Cross-Modality%20Encoder%20Representations%20Tan%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LXMERT%3A%20Learning%20Cross-Modality%20Encoder%20Representations&sort=relevance>)

425. **Oscar: Object-Semantics Aligned Pre-training** — Li et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Oscar%3A%20Object-Semantics%20Aligned%20Pre-training&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Oscar%3A%20Object-Semantics%20Aligned%20Pre-training%20Li%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Oscar%3A%20Object-Semantics%20Aligned%20Pre-training&sort=relevance>)

426. **UNITER: Universal Image-Text Representation Learning** — Chen et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=UNITER%3A%20Universal%20Image-Text%20Representation%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=UNITER%3A%20Universal%20Image-Text%20Representation%20Learning%20Chen%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=UNITER%3A%20Universal%20Image-Text%20Representation%20Learning&sort=relevance>)

427. **VinVL: Revisiting Visual Representations in VL Models** — Zhang et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=VinVL%3A%20Revisiting%20Visual%20Representations%20in%20VL%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=VinVL%3A%20Revisiting%20Visual%20Representations%20in%20VL%20Models%20Zhang%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=VinVL%3A%20Revisiting%20Visual%20Representations%20in%20VL%20Models&sort=relevance>)

428. **Perceiver: General Perception with Iterative Attention** — Jaegle et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2103.03206>) · [Direct PDF](<https://arxiv.org/pdf/2103.03206.pdf>)

429. **Perceiver IO** — Jaegle et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Perceiver%20IO&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Perceiver%20IO%20Jaegle%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Perceiver%20IO&sort=relevance>)

430. **Data2Vec: A General Framework for Self-Supervised Learning** — Baevski et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2202.03555>) · [Direct PDF](<https://arxiv.org/pdf/2202.03555.pdf>)

431. **ImageBind: One Embedding Space to Bind Them All** — Girdhar et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.05665>) · [Direct PDF](<https://arxiv.org/pdf/2305.05665.pdf>)

432. **Meta-Transformer: A Unified Framework for Multimodal Learning** — Zhang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Meta-Transformer%3A%20A%20Unified%20Framework%20for%20Multimodal%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Meta-Transformer%3A%20A%20Unified%20Framework%20for%20Multimodal%20Learning%20Zhang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Meta-Transformer%3A%20A%20Unified%20Framework%20for%20Multimodal%20Learning&sort=relevance>)

433. **Any-to-Any Generation via Composable Diffusion (CoDi)** — Tang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Any-to-Any%20Generation%20via%20Composable%20Diffusion%20%28CoDi%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Any-to-Any%20Generation%20via%20Composable%20Diffusion%20%28CoDi%29%20Tang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Any-to-Any%20Generation%20via%20Composable%20Diffusion%20%28CoDi%29&sort=relevance>)

434. **NExT-GPT: Any-to-Any Multimodal LLM** — Wu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=NExT-GPT%3A%20Any-to-Any%20Multimodal%20LLM&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=NExT-GPT%3A%20Any-to-Any%20Multimodal%20LLM%20Wu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=NExT-GPT%3A%20Any-to-Any%20Multimodal%20LLM&sort=relevance>)

435. **4M: Massively Multimodal Masked Modeling** — Bachmann et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=4M%3A%20Massively%20Multimodal%20Masked%20Modeling&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=4M%3A%20Massively%20Multimodal%20Masked%20Modeling%20Bachmann%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=4M%3A%20Massively%20Multimodal%20Masked%20Modeling&sort=relevance>)

## 23. Vision-Language Models

436. **CLIP: Learning Transferable Visual Models from Natural Language Supervision** — Radford et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2103.00020>) · [Direct PDF](<https://arxiv.org/pdf/2103.00020.pdf>)

437. **ALIGN: Scaling Up Visual and Vision-Language Representation Learning** — Jia et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ALIGN%3A%20Scaling%20Up%20Visual%20and%20Vision-Language%20Representation%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ALIGN%3A%20Scaling%20Up%20Visual%20and%20Vision-Language%20Representation%20Learning%20Jia%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ALIGN%3A%20Scaling%20Up%20Visual%20and%20Vision-Language%20Representation%20Learning&sort=relevance>)

438. **Florence: A New Foundation Model for Computer Vision** — Yuan et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Florence%3A%20A%20New%20Foundation%20Model%20for%20Computer%20Vision&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Florence%3A%20A%20New%20Foundation%20Model%20for%20Computer%20Vision%20Yuan%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Florence%3A%20A%20New%20Foundation%20Model%20for%20Computer%20Vision&sort=relevance>)

439. **Flamingo: A Visual Language Model for Few-Shot Learning** — Alayrac et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2204.14198>) · [Direct PDF](<https://arxiv.org/pdf/2204.14198.pdf>)

440. **BLIP: Bootstrapping Language-Image Pre-training** — Li et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2201.12086>) · [Direct PDF](<https://arxiv.org/pdf/2201.12086.pdf>)

441. **BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and LLMs** — Li et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2301.12597>) · [Direct PDF](<https://arxiv.org/pdf/2301.12597.pdf>)

442. **LLaVA: Visual Instruction Tuning** — Liu et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2304.08485>) · [Direct PDF](<https://arxiv.org/pdf/2304.08485.pdf>)

443. **LLaVA-1.5** — Liu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLaVA-1.5&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLaVA-1.5%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLaVA-1.5&sort=relevance>)

444. **InstructBLIP** — Dai et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=InstructBLIP&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=InstructBLIP%20Dai%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=InstructBLIP&sort=relevance>)

445. **MiniGPT-4: Enhancing Vision-Language Understanding** — Zhu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MiniGPT-4%3A%20Enhancing%20Vision-Language%20Understanding&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MiniGPT-4%3A%20Enhancing%20Vision-Language%20Understanding%20Zhu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MiniGPT-4%3A%20Enhancing%20Vision-Language%20Understanding&sort=relevance>)

446. **Qwen-VL: A Versatile VLM** — Bai et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Qwen-VL%3A%20A%20Versatile%20VLM&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Qwen-VL%3A%20A%20Versatile%20VLM%20Bai%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Qwen-VL%3A%20A%20Versatile%20VLM&sort=relevance>)

447. **CogVLM: Visual Expert for Pretrained Language Models** — Wang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=CogVLM%3A%20Visual%20Expert%20for%20Pretrained%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=CogVLM%3A%20Visual%20Expert%20for%20Pretrained%20Language%20Models%20Wang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=CogVLM%3A%20Visual%20Expert%20for%20Pretrained%20Language%20Models&sort=relevance>)

448. **InternVL: Scaling Up Vision Foundation Models** — Chen et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2312.14238>) · [Direct PDF](<https://arxiv.org/pdf/2312.14238.pdf>)

449. **Cambrian-1: A Fully Open Vision-Centric Exploration of Multimodal LLMs** — Tong et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2406.16860>) · [Direct PDF](<https://arxiv.org/pdf/2406.16860.pdf>)

450. **SigLIP: Sigmoid Loss for Language Image Pre-Training** — Zhai et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.15343>) · [Direct PDF](<https://arxiv.org/pdf/2303.15343.pdf>)

451. **SAM: Segment Anything** — Kirillov et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2304.02643>) · [Direct PDF](<https://arxiv.org/pdf/2304.02643.pdf>)

452. **SAM 2: Segment Anything in Images and Videos** — Ravi et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2408.00714>) · [Direct PDF](<https://arxiv.org/pdf/2408.00714.pdf>)

453. **Grounding DINO: Marrying DINO with Grounded Pre-Training** — Liu et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.05499>) · [Direct PDF](<https://arxiv.org/pdf/2303.05499.pdf>)

454. **OWL-ViT: Open-World Object Detection with Vision Transformers** — Minderer et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2205.06230>) · [Direct PDF](<https://arxiv.org/pdf/2205.06230.pdf>)

455. **PaLI: A Jointly-Scaled Multilingual Language-Image Model** — Chen et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2209.06794>) · [Direct PDF](<https://arxiv.org/pdf/2209.06794.pdf>)

## 24. Speech & Audio Models

456. **Deep Speech: Scaling Up End-to-End Speech Recognition** — Hannun et al. · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1412.5567>) · [Direct PDF](<https://arxiv.org/pdf/1412.5567.pdf>)

457. **Deep Speech 2** — Amodei et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1512.02595>) · [Direct PDF](<https://arxiv.org/pdf/1512.02595.pdf>)

458. **Listen, Attend and Spell (LAS)** — Chan et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Listen%2C%20Attend%20and%20Spell%20%28LAS%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Listen%2C%20Attend%20and%20Spell%20%28LAS%29%20Chan%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Listen%2C%20Attend%20and%20Spell%20%28LAS%29&sort=relevance>)

459. **wav2vec: Unsupervised Pre-training for Speech Recognition** — Schneider et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=wav2vec%3A%20Unsupervised%20Pre-training%20for%20Speech%20Recognition&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=wav2vec%3A%20Unsupervised%20Pre-training%20for%20Speech%20Recognition%20Schneider%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=wav2vec%3A%20Unsupervised%20Pre-training%20for%20Speech%20Recognition&sort=relevance>)

460. **wav2vec 2.0: A Framework for Self-Supervised Learning of Speech** — Baevski et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2006.11477>) · [Direct PDF](<https://arxiv.org/pdf/2006.11477.pdf>)

461. **HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction** — Hsu et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=HuBERT%3A%20Self-Supervised%20Speech%20Representation%20Learning%20by%20Masked%20Prediction&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=HuBERT%3A%20Self-Supervised%20Speech%20Representation%20Learning%20by%20Masked%20Prediction%20Hsu%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=HuBERT%3A%20Self-Supervised%20Speech%20Representation%20Learning%20by%20Masked%20Prediction&sort=relevance>)

462. **Whisper: Robust Speech Recognition via Large-Scale Weak Supervision** — Radford et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2212.04356>) · [Direct PDF](<https://arxiv.org/pdf/2212.04356.pdf>)

463. **Conformer: Convolution-augmented Transformer for Speech** — Gulati et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2005.08100>) · [Direct PDF](<https://arxiv.org/pdf/2005.08100.pdf>)

464. **Tacotron 2: Natural TTS Synthesis** — Shen et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1712.05884>) · [Direct PDF](<https://arxiv.org/pdf/1712.05884.pdf>)

465. **FastSpeech 2: Fast and High-Quality End-to-End TTS** — Ren et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2006.04558>) · [Direct PDF](<https://arxiv.org/pdf/2006.04558.pdf>)

466. **VALL-E: Neural Codec Language Models are Zero-Shot TTS** — Wang et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2301.02111>) · [Direct PDF](<https://arxiv.org/pdf/2301.02111.pdf>)

467. **XTTS: Massively Multilingual TTS** — Casanova et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=XTTS%3A%20Massively%20Multilingual%20TTS&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=XTTS%3A%20Massively%20Multilingual%20TTS%20Casanova%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=XTTS%3A%20Massively%20Multilingual%20TTS&sort=relevance>)

468. **AudioLM: A Language Modeling Approach to Audio Generation** — Borsos et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AudioLM%3A%20A%20Language%20Modeling%20Approach%20to%20Audio%20Generation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AudioLM%3A%20A%20Language%20Modeling%20Approach%20to%20Audio%20Generation%20Borsos%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AudioLM%3A%20A%20Language%20Modeling%20Approach%20to%20Audio%20Generation&sort=relevance>)

469. **MusicLM: Generating Music from Text** — Agostinelli et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2301.11325>) · [Direct PDF](<https://arxiv.org/pdf/2301.11325.pdf>)

470. **Bark: Text-Prompted Generative Audio Model** — Suno · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Bark%3A%20Text-Prompted%20Generative%20Audio%20Model&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Bark%3A%20Text-Prompted%20Generative%20Audio%20Model%20Suno%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Bark%3A%20Text-Prompted%20Generative%20Audio%20Model&sort=relevance>)

471. **AudioPaLM: A Large Language Model that Can Speak and Listen** — Rubenstein et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AudioPaLM%3A%20A%20Large%20Language%20Model%20that%20Can%20Speak%20and%20Listen&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AudioPaLM%3A%20A%20Large%20Language%20Model%20that%20Can%20Speak%20and%20Listen%20Rubenstein%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AudioPaLM%3A%20A%20Large%20Language%20Model%20that%20Can%20Speak%20and%20Listen&sort=relevance>)

472. **SeamlessM4T: Massively Multilingual & Multimodal Machine Translation** — Meta · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2308.11596>) · [Direct PDF](<https://arxiv.org/pdf/2308.11596.pdf>)

473. **Encodec: High Fidelity Neural Audio Compression** — Défossez et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2210.13438>) · [Direct PDF](<https://arxiv.org/pdf/2210.13438.pdf>)

474. **Voicebox: Text-Guided Multilingual Universal Speech Generation** — Le et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Voicebox%3A%20Text-Guided%20Multilingual%20Universal%20Speech%20Generation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Voicebox%3A%20Text-Guided%20Multilingual%20Universal%20Speech%20Generation%20Le%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Voicebox%3A%20Text-Guided%20Multilingual%20Universal%20Speech%20Generation&sort=relevance>)

475. **SoundStorm: Efficient Parallel Audio Generation** — Borsos et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SoundStorm%3A%20Efficient%20Parallel%20Audio%20Generation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SoundStorm%3A%20Efficient%20Parallel%20Audio%20Generation%20Borsos%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SoundStorm%3A%20Efficient%20Parallel%20Audio%20Generation&sort=relevance>)

## 25. Diffusion Models & Image Generation

476. **Deep Unsupervised Learning using Nonequilibrium Thermodynamics** — Sohl-Dickstein et al. · 2015 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Deep%20Unsupervised%20Learning%20using%20Nonequilibrium%20Thermodynamics&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Deep%20Unsupervised%20Learning%20using%20Nonequilibrium%20Thermodynamics%20Sohl-Dickstein%202015>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Deep%20Unsupervised%20Learning%20using%20Nonequilibrium%20Thermodynamics&sort=relevance>)

477. **Denoising Diffusion Probabilistic Models (DDPM)** — Ho, Jain, Abbeel · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2006.11239>) · [Direct PDF](<https://arxiv.org/pdf/2006.11239.pdf>)

478. **Score-Based Generative Modeling through SDE** — Song et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2011.13456>) · [Direct PDF](<https://arxiv.org/pdf/2011.13456.pdf>)

479. **Denoising Diffusion Implicit Models (DDIM)** — Song, Meng, Ermon · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2010.02502>) · [Direct PDF](<https://arxiv.org/pdf/2010.02502.pdf>)

480. **Classifier-Free Diffusion Guidance** — Ho & Salimans · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2207.12598>) · [Direct PDF](<https://arxiv.org/pdf/2207.12598.pdf>)

481. **High-Resolution Image Synthesis with Latent Diffusion Models (LDM)** — Rombach et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2112.10752>) · [Direct PDF](<https://arxiv.org/pdf/2112.10752.pdf>)

482. **Stable Diffusion** — Stability AI / Rombach et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Stable%20Diffusion&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Stable%20Diffusion%20Stability%20AI%20/%20Rombach%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Stable%20Diffusion&sort=relevance>)

483. **DALL·E 2: Hierarchical Text-Conditional Image Generation** — Ramesh et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2204.06125>) · [Direct PDF](<https://arxiv.org/pdf/2204.06125.pdf>)

484. **Imagen: Photorealistic Text-to-Image Diffusion Models** — Saharia et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2205.11487>) · [Direct PDF](<https://arxiv.org/pdf/2205.11487.pdf>)

485. **eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers** — Balaji et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=eDiff-I%3A%20Text-to-Image%20Diffusion%20Models%20with%20an%20Ensemble%20of%20Expert%20Denoisers&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=eDiff-I%3A%20Text-to-Image%20Diffusion%20Models%20with%20an%20Ensemble%20of%20Expert%20Denoisers%20Balaji%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=eDiff-I%3A%20Text-to-Image%20Diffusion%20Models%20with%20an%20Ensemble%20of%20Expert%20Denoisers&sort=relevance>)

486. **SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis** — Podell et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2307.01952>) · [Direct PDF](<https://arxiv.org/pdf/2307.01952.pdf>)

487. **DALL·E 3** — Betker et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DALL%C2%B7E%203&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DALL%C2%B7E%203%20Betker%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DALL%C2%B7E%203&sort=relevance>)

488. **Midjourney (V5/V6 technical analysis)** — Midjourney · 2023-2024 · **Type:** Product or topic analysis. [arXiv search](<https://arxiv.org/search/?query=Midjourney%20%28V5/V6%20technical%20analysis%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Midjourney%20%28V5/V6%20technical%20analysis%29%20Midjourney%202023-2024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Midjourney%20%28V5/V6%20technical%20analysis%29&sort=relevance>)
   - **Audit:** NON-PAPER OR MULTI-WORK RECORD — A product-analysis container rather than a uniquely identifiable research paper.

489. **Consistency Models** — Song et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.01469>) · [Direct PDF](<https://arxiv.org/pdf/2303.01469.pdf>)

490. **Latent Consistency Models** — Luo et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2310.04378>) · [Direct PDF](<https://arxiv.org/pdf/2310.04378.pdf>)

491. **Rectified Flow (InstaFlow)** — Liu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Rectified%20Flow%20%28InstaFlow%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Rectified%20Flow%20%28InstaFlow%29%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Rectified%20Flow%20%28InstaFlow%29&sort=relevance>)

492. **Stable Diffusion 3 (MM-DiT)** — Esser et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2403.03206>) · [Direct PDF](<https://arxiv.org/pdf/2403.03206.pdf>)

493. **FLUX** — Black Forest Labs · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=FLUX&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=FLUX%20Black%20Forest%20Labs%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=FLUX&sort=relevance>)

494. **ControlNet: Adding Conditional Control to Text-to-Image Diffusion** — Zhang & Agrawala · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2302.05543>) · [Direct PDF](<https://arxiv.org/pdf/2302.05543.pdf>)

495. **IP-Adapter: Text Compatible Image Prompt Adapter** — Ye et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2308.06721>) · [Direct PDF](<https://arxiv.org/pdf/2308.06721.pdf>)

496. **DreamBooth: Fine Tuning Text-to-Image Models for Subject-Driven Generation** — Ruiz et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2208.12242>) · [Direct PDF](<https://arxiv.org/pdf/2208.12242.pdf>)

497. **Textual Inversion: An Image is Worth One Word** — Gal et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2208.01618>) · [Direct PDF](<https://arxiv.org/pdf/2208.01618.pdf>)

498. **InstructPix2Pix: Learning to Follow Image Editing Instructions** — Brooks et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2211.09800>) · [Direct PDF](<https://arxiv.org/pdf/2211.09800.pdf>)

499. **Scalable Diffusion Models with Transformers (DiT)** — Peebles & Xie · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2212.09748>) · [Direct PDF](<https://arxiv.org/pdf/2212.09748.pdf>)

500. **Flow Matching for Generative Modeling** — Lipman et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2210.02747>) · [Direct PDF](<https://arxiv.org/pdf/2210.02747.pdf>)

## 26. Video Generation & Understanding

501. **Video Diffusion Models** — Ho et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2204.03458>) · [Direct PDF](<https://arxiv.org/pdf/2204.03458.pdf>)

502. **Imagen Video: High Definition Video Generation** — Ho et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Imagen%20Video%3A%20High%20Definition%20Video%20Generation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Imagen%20Video%3A%20High%20Definition%20Video%20Generation%20Ho%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Imagen%20Video%3A%20High%20Definition%20Video%20Generation&sort=relevance>)

503. **Make-A-Video: Text-to-Video Generation without Text-Video Data** — Singer et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2209.14792>) · [Direct PDF](<https://arxiv.org/pdf/2209.14792.pdf>)

504. **VideoPoet: A Large Language Model for Zero-Shot Video Generation** — Kondratyuk et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2312.14125>) · [Direct PDF](<https://arxiv.org/pdf/2312.14125.pdf>)

505. **Sora: Video Generation Models as World Simulators** — OpenAI · 2024 · **Type:** Research-paper candidate. [Supplied direct source](<https://openai.com/research/video-generation-models-as-world-simulators>)

506. **Stable Video Diffusion** — Blattmann et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2311.15127>) · [Direct PDF](<https://arxiv.org/pdf/2311.15127.pdf>)

507. **AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models** — Guo et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2307.04725>) · [Direct PDF](<https://arxiv.org/pdf/2307.04725.pdf>)

508. **Kling (technical analysis)** — Kuaishou · 2024 · **Type:** Product or topic analysis. [arXiv search](<https://arxiv.org/search/?query=Kling%20%28technical%20analysis%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Kling%20%28technical%20analysis%29%20Kuaishou%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Kling%20%28technical%20analysis%29&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as product or topic analysis, not as a research paper.

509. **Gen-2: Generating Novel Videos with Text, Images, or Clips** — Runway · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Gen-2%3A%20Generating%20Novel%20Videos%20with%20Text%2C%20Images%2C%20or%20Clips&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Gen-2%3A%20Generating%20Novel%20Videos%20with%20Text%2C%20Images%2C%20or%20Clips%20Runway%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Gen-2%3A%20Generating%20Novel%20Videos%20with%20Text%2C%20Images%2C%20or%20Clips&sort=relevance>)

510. **VideoCrafter: A Toolkit for Text-to-Video Generation** — Chen et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=VideoCrafter%3A%20A%20Toolkit%20for%20Text-to-Video%20Generation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=VideoCrafter%3A%20A%20Toolkit%20for%20Text-to-Video%20Generation%20Chen%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=VideoCrafter%3A%20A%20Toolkit%20for%20Text-to-Video%20Generation&sort=relevance>)

511. **VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training** — Tong et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2203.12602>) · [Direct PDF](<https://arxiv.org/pdf/2203.12602.pdf>)

512. **TimeSformer: Is Space-Time Attention All You Need for Video Understanding?** — Bertasius et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2102.05095>) · [Direct PDF](<https://arxiv.org/pdf/2102.05095.pdf>)

513. **ViViT: A Video Vision Transformer** — Arnab et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2103.15691>) · [Direct PDF](<https://arxiv.org/pdf/2103.15691.pdf>)

514. **InternVideo: General Video Foundation Models** — Wang et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=InternVideo%3A%20General%20Video%20Foundation%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=InternVideo%3A%20General%20Video%20Foundation%20Models%20Wang%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=InternVideo%3A%20General%20Video%20Foundation%20Models&sort=relevance>)

515. **Video-LLaVA** — Lin et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Video-LLaVA&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Video-LLaVA%20Lin%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Video-LLaVA&sort=relevance>)

## 27. Graph Neural Networks

516. **Semi-Supervised Classification with Graph Convolutional Networks (GCN)** — Kipf & Welling · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1609.02907>) · [Direct PDF](<https://arxiv.org/pdf/1609.02907.pdf>)

517. **Attention-Based Graph Neural Network (GAT)** — Veličković et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1710.10903>) · [Direct PDF](<https://arxiv.org/pdf/1710.10903.pdf>)

518. **Inductive Representation Learning on Large Graphs (GraphSAGE)** — Hamilton, Ying, Leskovec · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1706.02216>) · [Direct PDF](<https://arxiv.org/pdf/1706.02216.pdf>)

519. **How Powerful are Graph Neural Networks? (GIN)** — Xu et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1810.00826>) · [Direct PDF](<https://arxiv.org/pdf/1810.00826.pdf>)

520. **Message Passing Neural Networks (MPNN)** — Gilmer et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1704.01212>) · [Direct PDF](<https://arxiv.org/pdf/1704.01212.pdf>)

521. **Graph Isomorphism Network (GIN)** — Xu et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Graph%20Isomorphism%20Network%20%28GIN%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Graph%20Isomorphism%20Network%20%28GIN%29%20Xu%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Graph%20Isomorphism%20Network%20%28GIN%29&sort=relevance>)

522. **Relational Graph Convolutional Networks (R-GCN)** — Schlichtkrull et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1703.06103>) · [Direct PDF](<https://arxiv.org/pdf/1703.06103.pdf>)

523. **Graph Transformer Networks** — Yun et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1911.06455>) · [Direct PDF](<https://arxiv.org/pdf/1911.06455.pdf>)

524. **DeepWalk: Online Learning of Social Representations** — Perozzi, Al-Rfou, Skiena · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1403.6652>) · [Direct PDF](<https://arxiv.org/pdf/1403.6652.pdf>)

525. **Node2Vec: Scalable Feature Learning for Networks** — Grover & Leskovec · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1607.00653>) · [Direct PDF](<https://arxiv.org/pdf/1607.00653.pdf>)

526. **A Comprehensive Survey on Graph Neural Networks** — Wu et al. · 2021 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=A%20Comprehensive%20Survey%20on%20Graph%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Comprehensive%20Survey%20on%20Graph%20Neural%20Networks%20Wu%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Comprehensive%20Survey%20on%20Graph%20Neural%20Networks&sort=relevance>)

527. **Equivariant Graph Neural Networks** — Satorras et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Equivariant%20Graph%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Equivariant%20Graph%20Neural%20Networks%20Satorras%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Equivariant%20Graph%20Neural%20Networks&sort=relevance>)

528. **Graph Neural Networks: A Review of Methods and Applications** — Zhou et al. · 2020 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Graph%20Neural%20Networks%3A%20A%20Review%20of%20Methods%20and%20Applications&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Graph%20Neural%20Networks%3A%20A%20Review%20of%20Methods%20and%20Applications%20Zhou%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Graph%20Neural%20Networks%3A%20A%20Review%20of%20Methods%20and%20Applications&sort=relevance>)

529. **Spectral Graph Theory and Deep Learning** — Bruna et al. · 2014 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Spectral%20Graph%20Theory%20and%20Deep%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Spectral%20Graph%20Theory%20and%20Deep%20Learning%20Bruna%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Spectral%20Graph%20Theory%20and%20Deep%20Learning&sort=relevance>)

530. **Geometric Deep Learning** — Bronstein et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2104.13478>) · [Direct PDF](<https://arxiv.org/pdf/2104.13478.pdf>)

## 28. Efficient AI & Model Compression

531. **Deep Compression: Pruning, Quantization, Huffman Coding** — Han et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1510.00149>) · [Direct PDF](<https://arxiv.org/pdf/1510.00149.pdf>)

532. **MobileNets: Efficient CNNs for Mobile Vision Applications** — Howard et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1704.04861>) · [Direct PDF](<https://arxiv.org/pdf/1704.04861.pdf>)

533. **MobileNetV2: Inverted Residuals and Linear Bottlenecks** — Sandler et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1801.04381>) · [Direct PDF](<https://arxiv.org/pdf/1801.04381.pdf>)

534. **ShuffleNet: Extremely Efficient CNN for Mobile Devices** — Zhang et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1707.01083>) · [Direct PDF](<https://arxiv.org/pdf/1707.01083.pdf>)

535. **SqueezeNet: AlexNet-level Accuracy with 50x Fewer Parameters** — Iandola et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1602.07360>) · [Direct PDF](<https://arxiv.org/pdf/1602.07360.pdf>)

536. **LLM.int8(): 8-bit Matrix Multiplication for Transformers** — Dettmers et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2208.07339>) · [Direct PDF](<https://arxiv.org/pdf/2208.07339.pdf>)

537. **GPTQ: Accurate Post-Training Quantization for GPT** — Frantar et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2210.17323>) · [Direct PDF](<https://arxiv.org/pdf/2210.17323.pdf>)

538. **AWQ: Activation-aware Weight Quantization** — Lin et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2306.00978>) · [Direct PDF](<https://arxiv.org/pdf/2306.00978.pdf>)

539. **SqueezeLLM: Dense-and-Sparse Quantization** — Kim et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SqueezeLLM%3A%20Dense-and-Sparse%20Quantization&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SqueezeLLM%3A%20Dense-and-Sparse%20Quantization%20Kim%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SqueezeLLM%3A%20Dense-and-Sparse%20Quantization&sort=relevance>)

540. **The Era of 1-bit LLMs: Training-time 1.58-bit Models (BitNet)** — Ma et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2402.17764>) · [Direct PDF](<https://arxiv.org/pdf/2402.17764.pdf>)

541. **SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot** — Frantar & Alistarh · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2301.00774>) · [Direct PDF](<https://arxiv.org/pdf/2301.00774.pdf>)

542. **Wanda: A Simple and Effective Pruning Approach** — Sun et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Wanda%3A%20A%20Simple%20and%20Effective%20Pruning%20Approach&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Wanda%3A%20A%20Simple%20and%20Effective%20Pruning%20Approach%20Sun%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Wanda%3A%20A%20Simple%20and%20Effective%20Pruning%20Approach&sort=relevance>)

543. **Flash-Decoding for Long Context Inference** — Dao et al. · 2023 · **Type:** Research-paper candidate. [Supplied direct source](<https://crfm.stanford.edu/2023/10/12/flashdecoding.html>)

544. **Speculative Decoding** — Leviathan, Kalman, Matias · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2211.17192>) · [Direct PDF](<https://arxiv.org/pdf/2211.17192.pdf>)

545. **Medusa: Simple LLM Inference Acceleration** — Cai et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2401.10774>) · [Direct PDF](<https://arxiv.org/pdf/2401.10774.pdf>)

546. **vLLM: Efficient Memory Management for LLM Serving (PagedAttention)** — Kwon et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2309.06180>) · [Direct PDF](<https://arxiv.org/pdf/2309.06180.pdf>)

547. **TensorRT-LLM** — NVIDIA · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=TensorRT-LLM&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=TensorRT-LLM%20NVIDIA%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=TensorRT-LLM&sort=relevance>)

548. **llama.cpp** — Gerganov · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=llama.cpp&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=llama.cpp%20Gerganov%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=llama.cpp&sort=relevance>)

549. **Ollama (technical overview)** — Ollama · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Ollama%20%28technical%20overview%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Ollama%20%28technical%20overview%29%20Ollama%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Ollama%20%28technical%20overview%29&sort=relevance>)

550. **GGUF: GPT-Generated Unified Format** — Gerganov · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=GGUF%3A%20GPT-Generated%20Unified%20Format&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=GGUF%3A%20GPT-Generated%20Unified%20Format%20Gerganov%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=GGUF%3A%20GPT-Generated%20Unified%20Format&sort=relevance>)

## 29. Mixture of Experts

551. **Adaptive Mixtures of Local Experts** — Jacobs et al. · 1991 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.cs.toronto.edu/~hinton/absps/jjnh91.pdf>)

552. **Outrageously Large Neural Networks: The Sparsely-Gated MoE Layer** — Shazeer et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1701.06538>) · [Direct PDF](<https://arxiv.org/pdf/1701.06538.pdf>)

553. **Switch Transformers: Scaling to Trillion Parameter Models** — Fedus, Zoph, Shazeer · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2101.03961>) · [Direct PDF](<https://arxiv.org/pdf/2101.03961.pdf>)

554. **GLaM: Efficient Scaling of Language Models with MoE** — Du et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=GLaM%3A%20Efficient%20Scaling%20of%20Language%20Models%20with%20MoE&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=GLaM%3A%20Efficient%20Scaling%20of%20Language%20Models%20with%20MoE%20Du%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=GLaM%3A%20Efficient%20Scaling%20of%20Language%20Models%20with%20MoE&sort=relevance>)

555. **ST-MoE: Designing Stable and Transferable Sparse Expert Models** — Zoph et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ST-MoE%3A%20Designing%20Stable%20and%20Transferable%20Sparse%20Expert%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ST-MoE%3A%20Designing%20Stable%20and%20Transferable%20Sparse%20Expert%20Models%20Zoph%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ST-MoE%3A%20Designing%20Stable%20and%20Transferable%20Sparse%20Expert%20Models&sort=relevance>)

556. **Mixtral of Experts** — Jiang et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2401.04088>) · [Direct PDF](<https://arxiv.org/pdf/2401.04088.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #196.

557. **DeepSeekMoE: Towards Ultimate Expert Specialization** — Dai et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2401.06066>) · [Direct PDF](<https://arxiv.org/pdf/2401.06066.pdf>)

558. **Unified Scaling Laws for Routed Language Models** — Clark et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Unified%20Scaling%20Laws%20for%20Routed%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Unified%20Scaling%20Laws%20for%20Routed%20Language%20Models%20Clark%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Unified%20Scaling%20Laws%20for%20Routed%20Language%20Models&sort=relevance>)

559. **MoE-Mamba** — Pioro et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2401.04081>) · [Direct PDF](<https://arxiv.org/pdf/2401.04081.pdf>)

560. **Skywork-MoE: A Deep Dive into Training Techniques for MoE LLMs** — Wei et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Skywork-MoE%3A%20A%20Deep%20Dive%20into%20Training%20Techniques%20for%20MoE%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Skywork-MoE%3A%20A%20Deep%20Dive%20into%20Training%20Techniques%20for%20MoE%20LLMs%20Wei%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Skywork-MoE%3A%20A%20Deep%20Dive%20into%20Training%20Techniques%20for%20MoE%20LLMs&sort=relevance>)

## 30. Knowledge Distillation & Transfer Learning

561. **Distilling the Knowledge in a Neural Network** — Hinton, Vinyals, Dean · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1503.02531>) · [Direct PDF](<https://arxiv.org/pdf/1503.02531.pdf>)

562. **Born Again Neural Networks** — Furlanello et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1805.04770>) · [Direct PDF](<https://arxiv.org/pdf/1805.04770.pdf>)

563. **Be Your Own Teacher: Improve the Performance of CNNs via Self-Distillation** — Zhang et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Be%20Your%20Own%20Teacher%3A%20Improve%20the%20Performance%20of%20CNNs%20via%20Self-Distillation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Be%20Your%20Own%20Teacher%3A%20Improve%20the%20Performance%20of%20CNNs%20via%20Self-Distillation%20Zhang%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Be%20Your%20Own%20Teacher%3A%20Improve%20the%20Performance%20of%20CNNs%20via%20Self-Distillation&sort=relevance>)

564. **TinyBERT: Distilling BERT for Natural Language Understanding** — Jiao et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1909.10351>) · [Direct PDF](<https://arxiv.org/pdf/1909.10351.pdf>)

565. **MiniLM: Deep Self-Attention Distillation** — Wang et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MiniLM%3A%20Deep%20Self-Attention%20Distillation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MiniLM%3A%20Deep%20Self-Attention%20Distillation%20Wang%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MiniLM%3A%20Deep%20Self-Attention%20Distillation&sort=relevance>)

566. **How Good Are You at Transferring? A Survey on Transfer Learning** — Zhuang et al. · 2020 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=How%20Good%20Are%20You%20at%20Transferring%3F%20A%20Survey%20on%20Transfer%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=How%20Good%20Are%20You%20at%20Transferring%3F%20A%20Survey%20on%20Transfer%20Learning%20Zhuang%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=How%20Good%20Are%20You%20at%20Transferring%3F%20A%20Survey%20on%20Transfer%20Learning&sort=relevance>)

567. **Domain Adaptation for Object Recognition** — Saenko et al. · 2010 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Domain%20Adaptation%20for%20Object%20Recognition&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Domain%20Adaptation%20for%20Object%20Recognition%20Saenko%202010>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Domain%20Adaptation%20for%20Object%20Recognition&sort=relevance>)

568. **DeCAF: A Deep Convolutional Activation Feature** — Donahue et al. · 2014 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DeCAF%3A%20A%20Deep%20Convolutional%20Activation%20Feature&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DeCAF%3A%20A%20Deep%20Convolutional%20Activation%20Feature%20Donahue%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DeCAF%3A%20A%20Deep%20Convolutional%20Activation%20Feature&sort=relevance>)

569. **How Transferable are Features in Deep Neural Networks?** — Yosinski et al. · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1411.1792>) · [Direct PDF](<https://arxiv.org/pdf/1411.1792.pdf>)

570. **LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement** — Lee et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLM2LLM%3A%20Boosting%20LLMs%20with%20Novel%20Iterative%20Data%20Enhancement&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLM2LLM%3A%20Boosting%20LLMs%20with%20Novel%20Iterative%20Data%20Enhancement%20Lee%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLM2LLM%3A%20Boosting%20LLMs%20with%20Novel%20Iterative%20Data%20Enhancement&sort=relevance>)

## 31. Continual & Lifelong Learning

571. **Overcoming Catastrophic Forgetting in Neural Networks (EWC)** — Kirkpatrick et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1612.00796>) · [Direct PDF](<https://arxiv.org/pdf/1612.00796.pdf>)

572. **Progressive Neural Networks** — Rusu et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1606.04671>) · [Direct PDF](<https://arxiv.org/pdf/1606.04671.pdf>)

573. **Continual Lifelong Learning with Neural Networks: A Review** — Parisi et al. · 2019 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Continual%20Lifelong%20Learning%20with%20Neural%20Networks%3A%20A%20Review&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Continual%20Lifelong%20Learning%20with%20Neural%20Networks%3A%20A%20Review%20Parisi%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Continual%20Lifelong%20Learning%20with%20Neural%20Networks%3A%20A%20Review&sort=relevance>)

574. **Learning without Forgetting** — Li & Hoiem · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1606.09282>) · [Direct PDF](<https://arxiv.org/pdf/1606.09282.pdf>)

575. **PackNet: Adding Multiple Tasks to a Single Network** — Mallya & Lazebnik · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=PackNet%3A%20Adding%20Multiple%20Tasks%20to%20a%20Single%20Network&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=PackNet%3A%20Adding%20Multiple%20Tasks%20to%20a%20Single%20Network%20Mallya%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=PackNet%3A%20Adding%20Multiple%20Tasks%20to%20a%20Single%20Network&sort=relevance>)

576. **Experience Replay for Continual Learning** — Rolnick et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Experience%20Replay%20for%20Continual%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Experience%20Replay%20for%20Continual%20Learning%20Rolnick%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Experience%20Replay%20for%20Continual%20Learning&sort=relevance>)

577. **A Comprehensive Survey of Continual Learning** — De Lange et al. · 2022 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=A%20Comprehensive%20Survey%20of%20Continual%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Comprehensive%20Survey%20of%20Continual%20Learning%20De%20Lange%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Comprehensive%20Survey%20of%20Continual%20Learning&sort=relevance>)

578. **Continual Pre-training of Language Models** — Gururangan et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Continual%20Pre-training%20of%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Continual%20Pre-training%20of%20Language%20Models%20Gururangan%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Continual%20Pre-training%20of%20Language%20Models&sort=relevance>)

579. **TRACE: A Comprehensive Benchmark for Continual Learning in LLMs** — Wang et al. · 2024 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=TRACE%3A%20A%20Comprehensive%20Benchmark%20for%20Continual%20Learning%20in%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=TRACE%3A%20A%20Comprehensive%20Benchmark%20for%20Continual%20Learning%20in%20LLMs%20Wang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=TRACE%3A%20A%20Comprehensive%20Benchmark%20for%20Continual%20Learning%20in%20LLMs&sort=relevance>)

580. **Online Continual Learning for LLMs** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Online%20Continual%20Learning%20for%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Online%20Continual%20Learning%20for%20LLMs%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Online%20Continual%20Learning%20for%20LLMs&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 32. Self-Supervised & Contrastive Learning

581. **Momentum Contrast for Unsupervised Visual Representation Learning (MoCo)** — He et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1911.05722>) · [Direct PDF](<https://arxiv.org/pdf/1911.05722.pdf>)

582. **MoCo v2** — Chen et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MoCo%20v2&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MoCo%20v2%20Chen%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MoCo%20v2&sort=relevance>)

583. **A Simple Framework for Contrastive Learning (SimCLR)** — Chen et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2002.05709>) · [Direct PDF](<https://arxiv.org/pdf/2002.05709.pdf>)

584. **Bootstrap Your Own Latent (BYOL)** — Grill et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2006.07733>) · [Direct PDF](<https://arxiv.org/pdf/2006.07733.pdf>)

585. **Barlow Twins: Self-Supervised Learning via Redundancy Reduction** — Zbontar et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2103.03230>) · [Direct PDF](<https://arxiv.org/pdf/2103.03230.pdf>)

586. **VICReg: Variance-Invariance-Covariance Regularization** — Bardes, Ponce, LeCun · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2105.04906>) · [Direct PDF](<https://arxiv.org/pdf/2105.04906.pdf>)

587. **DINO: Emerging Properties in Self-Supervised Vision Transformers** — Caron et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2104.14294>) · [Direct PDF](<https://arxiv.org/pdf/2104.14294.pdf>)

588. **DINOv2: Learning Robust Visual Features** — Oquab et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2304.07193>) · [Direct PDF](<https://arxiv.org/pdf/2304.07193.pdf>)

589. **Masked Autoencoders Are Scalable Vision Learners (MAE)** — He et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2111.06377>) · [Direct PDF](<https://arxiv.org/pdf/2111.06377.pdf>)

590. **SimMIM: A Simple Framework for Masked Image Modeling** — Xie et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SimMIM%3A%20A%20Simple%20Framework%20for%20Masked%20Image%20Modeling&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SimMIM%3A%20A%20Simple%20Framework%20for%20Masked%20Image%20Modeling%20Xie%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SimMIM%3A%20A%20Simple%20Framework%20for%20Masked%20Image%20Modeling&sort=relevance>)

591. **BEiT: BERT Pre-Training of Image Transformers** — Bao, Dong, Wei · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2106.08254>) · [Direct PDF](<https://arxiv.org/pdf/2106.08254.pdf>)

592. **Supervised Contrastive Learning** — Khosla et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2004.11362>) · [Direct PDF](<https://arxiv.org/pdf/2004.11362.pdf>)

593. **Understanding Contrastive Representation Learning** — Arora et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Understanding%20Contrastive%20Representation%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Understanding%20Contrastive%20Representation%20Learning%20Arora%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Understanding%20Contrastive%20Representation%20Learning&sort=relevance>)

594. **Self-Supervised Learning: Generative or Contrastive (Survey)** — Liu et al. · 2021 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Self-Supervised%20Learning%3A%20Generative%20or%20Contrastive%20%28Survey%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Self-Supervised%20Learning%3A%20Generative%20or%20Contrastive%20%28Survey%29%20Liu%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Self-Supervised%20Learning%3A%20Generative%20or%20Contrastive%20%28Survey%29&sort=relevance>)

595. **I-JEPA: Image Joint-Embedding Predictive Architecture** — Assran et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2301.08243>) · [Direct PDF](<https://arxiv.org/pdf/2301.08243.pdf>)

## 33. Federated & Privacy-Preserving AI

596. **Communication-Efficient Learning of Deep Networks from Decentralized Data (FedAvg)** — McMahan et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1602.05629>) · [Direct PDF](<https://arxiv.org/pdf/1602.05629.pdf>)

597. **Federated Learning: Challenges, Methods, and Future Directions** — Li et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Federated%20Learning%3A%20Challenges%2C%20Methods%2C%20and%20Future%20Directions&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Federated%20Learning%3A%20Challenges%2C%20Methods%2C%20and%20Future%20Directions%20Li%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Federated%20Learning%3A%20Challenges%2C%20Methods%2C%20and%20Future%20Directions&sort=relevance>)

598. **Advances and Open Problems in Federated Learning** — Kairouz et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Advances%20and%20Open%20Problems%20in%20Federated%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Advances%20and%20Open%20Problems%20in%20Federated%20Learning%20Kairouz%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Advances%20and%20Open%20Problems%20in%20Federated%20Learning&sort=relevance>)

599. **Deep Learning with Differential Privacy** — Abadi et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1607.00133>) · [Direct PDF](<https://arxiv.org/pdf/1607.00133.pdf>)

600. **The Algorithmic Foundations of Differential Privacy** — Dwork & Roth · 2014 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Algorithmic%20Foundations%20of%20Differential%20Privacy&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Algorithmic%20Foundations%20of%20Differential%20Privacy%20Dwork%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Algorithmic%20Foundations%20of%20Differential%20Privacy&sort=relevance>)

601. **SecureML: A System for Scalable Privacy-Preserving ML** — Mohassel & Zhang · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SecureML%3A%20A%20System%20for%20Scalable%20Privacy-Preserving%20ML&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SecureML%3A%20A%20System%20for%20Scalable%20Privacy-Preserving%20ML%20Mohassel%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SecureML%3A%20A%20System%20for%20Scalable%20Privacy-Preserving%20ML&sort=relevance>)

602. **Membership Inference Attacks Against ML Models** — Shokri et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1610.05820>) · [Direct PDF](<https://arxiv.org/pdf/1610.05820.pdf>)

603. **Machine Unlearning** — Bourtoule et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2103.03279>) · [Direct PDF](<https://arxiv.org/pdf/2103.03279.pdf>)

604. **FedProx: Federated Optimization in Heterogeneous Networks** — Li et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=FedProx%3A%20Federated%20Optimization%20in%20Heterogeneous%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=FedProx%3A%20Federated%20Optimization%20in%20Heterogeneous%20Networks%20Li%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=FedProx%3A%20Federated%20Optimization%20in%20Heterogeneous%20Networks&sort=relevance>)

605. **Scaffold: Stochastic Controlled Averaging for FL** — Karimireddy et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Scaffold%3A%20Stochastic%20Controlled%20Averaging%20for%20FL&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scaffold%3A%20Stochastic%20Controlled%20Averaging%20for%20FL%20Karimireddy%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scaffold%3A%20Stochastic%20Controlled%20Averaging%20for%20FL&sort=relevance>)

## 34. Evaluation & Benchmarks

606. **ImageNet Large Scale Visual Recognition Challenge** — Russakovsky et al. · 2015 · **Type:** Benchmark or dataset paper. [Canonical arXiv record](<https://arxiv.org/abs/1409.0575>) · [Direct PDF](<https://arxiv.org/pdf/1409.0575.pdf>)

607. **SQuAD: 100,000+ Questions for Machine Comprehension** — Rajpurkar et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1606.05250>) · [Direct PDF](<https://arxiv.org/pdf/1606.05250.pdf>)

608. **GLUE: A Multi-Task Benchmark for NLU** — Wang et al. · 2019 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=GLUE%3A%20A%20Multi-Task%20Benchmark%20for%20NLU&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=GLUE%3A%20A%20Multi-Task%20Benchmark%20for%20NLU%20Wang%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=GLUE%3A%20A%20Multi-Task%20Benchmark%20for%20NLU&sort=relevance>)

609. **SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding** — Wang et al. · 2019 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=SuperGLUE%3A%20A%20Stickier%20Benchmark%20for%20General-Purpose%20Language%20Understanding&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SuperGLUE%3A%20A%20Stickier%20Benchmark%20for%20General-Purpose%20Language%20Understanding%20Wang%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SuperGLUE%3A%20A%20Stickier%20Benchmark%20for%20General-Purpose%20Language%20Understanding&sort=relevance>)

610. **MMLU: Measuring Massive Multitask Language Understanding** — Hendrycks et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2009.03300>) · [Direct PDF](<https://arxiv.org/pdf/2009.03300.pdf>)

611. **MMLU-Pro** — Wang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MMLU-Pro&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MMLU-Pro%20Wang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MMLU-Pro&sort=relevance>)

612. **BIG-Bench: Beyond the Imitation Game** — Srivastava et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2206.04615>) · [Direct PDF](<https://arxiv.org/pdf/2206.04615.pdf>)

613. **HellaSwag: Can a Machine Really Finish Your Sentence?** — Zellers et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=HellaSwag%3A%20Can%20a%20Machine%20Really%20Finish%20Your%20Sentence%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=HellaSwag%3A%20Can%20a%20Machine%20Really%20Finish%20Your%20Sentence%3F%20Zellers%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=HellaSwag%3A%20Can%20a%20Machine%20Really%20Finish%20Your%20Sentence%3F&sort=relevance>)

614. **ARC: Think You Have Solved Question Answering?** — Clark et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ARC%3A%20Think%20You%20Have%20Solved%20Question%20Answering%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ARC%3A%20Think%20You%20Have%20Solved%20Question%20Answering%3F%20Clark%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ARC%3A%20Think%20You%20Have%20Solved%20Question%20Answering%3F&sort=relevance>)

615. **WinoGrande: An Adversarial Winograd Schema Challenge** — Sakaguchi et al. · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=WinoGrande%3A%20An%20Adversarial%20Winograd%20Schema%20Challenge&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=WinoGrande%3A%20An%20Adversarial%20Winograd%20Schema%20Challenge%20Sakaguchi%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=WinoGrande%3A%20An%20Adversarial%20Winograd%20Schema%20Challenge&sort=relevance>)

616. **TruthfulQA: Measuring How Models Mimic Human Falsehoods** — Lin, Hilton, Evans · 2022 · **Type:** Benchmark or dataset paper. [Canonical arXiv record](<https://arxiv.org/abs/2109.07958>) · [Direct PDF](<https://arxiv.org/pdf/2109.07958.pdf>)

617. **Chatbot Arena / LMSYS Leaderboard** — Zheng et al. · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Chatbot%20Arena%20/%20LMSYS%20Leaderboard&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Chatbot%20Arena%20/%20LMSYS%20Leaderboard%20Zheng%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Chatbot%20Arena%20/%20LMSYS%20Leaderboard&sort=relevance>)

618. **MT-Bench** — Zheng et al. · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=MT-Bench&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MT-Bench%20Zheng%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MT-Bench&sort=relevance>)

619. **AlpacaEval** — Li et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AlpacaEval&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AlpacaEval%20Li%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AlpacaEval&sort=relevance>)

620. **GPQA: A Graduate-Level Google-Proof Q&A Benchmark** — Rein et al. · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=GPQA%3A%20A%20Graduate-Level%20Google-Proof%20Q%26A%20Benchmark&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=GPQA%3A%20A%20Graduate-Level%20Google-Proof%20Q%26A%20Benchmark%20Rein%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=GPQA%3A%20A%20Graduate-Level%20Google-Proof%20Q%26A%20Benchmark&sort=relevance>)

621. **LiveBench: A Challenging, Contamination-Free LLM Benchmark** — White et al. · 2024 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=LiveBench%3A%20A%20Challenging%2C%20Contamination-Free%20LLM%20Benchmark&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LiveBench%3A%20A%20Challenging%2C%20Contamination-Free%20LLM%20Benchmark%20White%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LiveBench%3A%20A%20Challenging%2C%20Contamination-Free%20LLM%20Benchmark&sort=relevance>)

622. **MATH and GSM8K (mentioned earlier)** — Hendrycks; Cobbe · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MATH%20and%20GSM8K%20%28mentioned%20earlier%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MATH%20and%20GSM8K%20%28mentioned%20earlier%29%20Hendrycks%3B%20Cobbe%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MATH%20and%20GSM8K%20%28mentioned%20earlier%29&sort=relevance>)

623. **HumanEval and MBPP (mentioned earlier)** — Chen; Austin · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=HumanEval%20and%20MBPP%20%28mentioned%20earlier%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=HumanEval%20and%20MBPP%20%28mentioned%20earlier%29%20Chen%3B%20Austin%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=HumanEval%20and%20MBPP%20%28mentioned%20earlier%29&sort=relevance>)

624. **MGSM: Multilingual Grade School Math** — Shi et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MGSM%3A%20Multilingual%20Grade%20School%20Math&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MGSM%3A%20Multilingual%20Grade%20School%20Math%20Shi%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MGSM%3A%20Multilingual%20Grade%20School%20Math&sort=relevance>)

625. **SimpleQA** — OpenAI · 2024 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=SimpleQA&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SimpleQA%20OpenAI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SimpleQA&sort=relevance>)

626. **IFEval: Instruction-Following Evaluation** — Zhou et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=IFEval%3A%20Instruction-Following%20Evaluation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=IFEval%3A%20Instruction-Following%20Evaluation%20Zhou%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=IFEval%3A%20Instruction-Following%20Evaluation&sort=relevance>)

627. **MuSR: Multi-Step Soft Reasoning** — Sprague et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MuSR%3A%20Multi-Step%20Soft%20Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MuSR%3A%20Multi-Step%20Soft%20Reasoning%20Sprague%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MuSR%3A%20Multi-Step%20Soft%20Reasoning&sort=relevance>)

628. **Agentic Benchmarks: WebArena, OSWorld, etc.** — Zhou; Xie et al. · 2024 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Agentic%20Benchmarks%3A%20WebArena%2C%20OSWorld%2C%20etc.&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Agentic%20Benchmarks%3A%20WebArena%2C%20OSWorld%2C%20etc.%20Zhou%3B%20Xie%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Agentic%20Benchmarks%3A%20WebArena%2C%20OSWorld%2C%20etc.&sort=relevance>)

629. **COCO: Common Objects in Context** — Lin et al. · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1405.0312>) · [Direct PDF](<https://arxiv.org/pdf/1405.0312.pdf>)

630. **Holistic Evaluation of Language Models (HELM)** — Liang et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2211.09110>) · [Direct PDF](<https://arxiv.org/pdf/2211.09110.pdf>)

## 35. AI Safety, Ethics & Alignment

631. **Concrete Problems in AI Safety** — Amodei et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1606.06565>) · [Direct PDF](<https://arxiv.org/pdf/1606.06565.pdf>)

632. **AI Alignment: A Comprehensive Survey** — Ji et al. · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=AI%20Alignment%3A%20A%20Comprehensive%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AI%20Alignment%3A%20A%20Comprehensive%20Survey%20Ji%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AI%20Alignment%3A%20A%20Comprehensive%20Survey&sort=relevance>)

633. **The Alignment Problem (overview)** — Christian · 2020 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=The%20Alignment%20Problem%20%28overview%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Alignment%20Problem%20%28overview%29%20Christian%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Alignment%20Problem%20%28overview%29&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper.

634. **Reward Hacking in Reinforcement Learning** — Skalse et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Reward%20Hacking%20in%20Reinforcement%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Reward%20Hacking%20in%20Reinforcement%20Learning%20Skalse%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Reward%20Hacking%20in%20Reinforcement%20Learning&sort=relevance>)

635. **Scalable Agent Alignment via Reward Modeling (Anthropic)** — Leike et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Scalable%20Agent%20Alignment%20via%20Reward%20Modeling%20%28Anthropic%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scalable%20Agent%20Alignment%20via%20Reward%20Modeling%20%28Anthropic%29%20Leike%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scalable%20Agent%20Alignment%20via%20Reward%20Modeling%20%28Anthropic%29&sort=relevance>)

636. **Language Models Don't Always Say What They Think: Unfaithful Explanations** — Turpin et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Language%20Models%20Don%27t%20Always%20Say%20What%20They%20Think%3A%20Unfaithful%20Explanations&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Language%20Models%20Don%27t%20Always%20Say%20What%20They%20Think%3A%20Unfaithful%20Explanations%20Turpin%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Language%20Models%20Don%27t%20Always%20Say%20What%20They%20Think%3A%20Unfaithful%20Explanations&sort=relevance>)

637. **Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training** — Hubinger et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2401.05566>) · [Direct PDF](<https://arxiv.org/pdf/2401.05566.pdf>)

638. **Risks from Learned Optimization in Advanced ML Systems (Mesa-Optimizers)** — Hubinger et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Risks%20from%20Learned%20Optimization%20in%20Advanced%20ML%20Systems%20%28Mesa-Optimizers%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Risks%20from%20Learned%20Optimization%20in%20Advanced%20ML%20Systems%20%28Mesa-Optimizers%29%20Hubinger%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Risks%20from%20Learned%20Optimization%20in%20Advanced%20ML%20Systems%20%28Mesa-Optimizers%29&sort=relevance>)

639. **Goal Misgeneralization: Why Correct Specifications Aren't Enough** — Shah et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Goal%20Misgeneralization%3A%20Why%20Correct%20Specifications%20Aren%27t%20Enough&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Goal%20Misgeneralization%3A%20Why%20Correct%20Specifications%20Aren%27t%20Enough%20Shah%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Goal%20Misgeneralization%3A%20Why%20Correct%20Specifications%20Aren%27t%20Enough&sort=relevance>)

640. **Red Teaming Language Models to Reduce Harms** — Ganguli et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2202.03286>) · [Direct PDF](<https://arxiv.org/pdf/2202.03286.pdf>)

641. **Red Teaming Language Models with Language Models** — Perez et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Red%20Teaming%20Language%20Models%20with%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Red%20Teaming%20Language%20Models%20with%20Language%20Models%20Perez%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Red%20Teaming%20Language%20Models%20with%20Language%20Models&sort=relevance>)

642. **Universal and Transferable Adversarial Attacks on Aligned LMs** — Zou et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2307.15043>) · [Direct PDF](<https://arxiv.org/pdf/2307.15043.pdf>)

643. **Jailbroken: How Does LLM Safety Training Fail?** — Wei et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Jailbroken%3A%20How%20Does%20LLM%20Safety%20Training%20Fail%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Jailbroken%3A%20How%20Does%20LLM%20Safety%20Training%20Fail%3F%20Wei%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Jailbroken%3A%20How%20Does%20LLM%20Safety%20Training%20Fail%3F&sort=relevance>)

644. **Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior** — Pan et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Do%20the%20Rewards%20Justify%20the%20Means%3F%20Measuring%20Trade-Offs%20Between%20Rewards%20and%20Ethical%20Behavior&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Do%20the%20Rewards%20Justify%20the%20Means%3F%20Measuring%20Trade-Offs%20Between%20Rewards%20and%20Ethical%20Behavior%20Pan%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Do%20the%20Rewards%20Justify%20the%20Means%3F%20Measuring%20Trade-Offs%20Between%20Rewards%20and%20Ethical%20Behavior&sort=relevance>)

645. **The Model Spec (OpenAI)** — OpenAI · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Model%20Spec%20%28OpenAI%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Model%20Spec%20%28OpenAI%29%20OpenAI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Model%20Spec%20%28OpenAI%29&sort=relevance>)

646. **Claude's Character (Anthropic)** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Claude%27s%20Character%20%28Anthropic%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Claude%27s%20Character%20%28Anthropic%29%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Claude%27s%20Character%20%28Anthropic%29&sort=relevance>)

647. **Responsible Scaling Policies (RSPs)** — Anthropic · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Responsible%20Scaling%20Policies%20%28RSPs%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Responsible%20Scaling%20Policies%20%28RSPs%29%20Anthropic%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Responsible%20Scaling%20Policies%20%28RSPs%29&sort=relevance>)

648. **Managing AI Risks in an Era of Rapid Progress** — Bengio et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Managing%20AI%20Risks%20in%20an%20Era%20of%20Rapid%20Progress&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Managing%20AI%20Risks%20in%20an%20Era%20of%20Rapid%20Progress%20Bengio%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Managing%20AI%20Risks%20in%20an%20Era%20of%20Rapid%20Progress&sort=relevance>)

649. **Pause Giant AI Experiments: An Open Letter** — Future of Life · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Pause%20Giant%20AI%20Experiments%3A%20An%20Open%20Letter&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Pause%20Giant%20AI%20Experiments%3A%20An%20Open%20Letter%20Future%20of%20Life%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Pause%20Giant%20AI%20Experiments%3A%20An%20Open%20Letter&sort=relevance>)

650. **On the Dangers of Stochastic Parrots** — Bender et al. · 2021 · **Type:** Research-paper candidate. [Supplied direct source](<https://dl.acm.org/doi/10.1145/3442188.3445922>)

651. **Fairness and Machine Learning** — Barocas, Hardt, Narayanan · 2019 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=Fairness%20and%20Machine%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Fairness%20and%20Machine%20Learning%20Barocas%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Fairness%20and%20Machine%20Learning&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper.

652. **Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification** — Buolamwini & Gebru · 2018 · **Type:** Research-paper candidate. [Supplied direct source](<https://proceedings.mlr.press/v81/buolamwini18a.html>)

653. **Datasheets for Datasets** — Gebru et al. · 2021 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Datasheets%20for%20Datasets&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Datasheets%20for%20Datasets%20Gebru%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Datasheets%20for%20Datasets&sort=relevance>)

654. **Model Cards for Model Reporting** — Mitchell et al. · 2019 · **Type:** Technical or industry report. [arXiv search](<https://arxiv.org/search/?query=Model%20Cards%20for%20Model%20Reporting&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Model%20Cards%20for%20Model%20Reporting%20Mitchell%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Model%20Cards%20for%20Model%20Reporting&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

655. **Blueprint for an AI Bill of Rights** — White House · 2022 · **Type:** Policy or official document. [arXiv search](<https://arxiv.org/search/?query=Blueprint%20for%20an%20AI%20Bill%20of%20Rights&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Blueprint%20for%20an%20AI%20Bill%20of%20Rights%20White%20House%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Blueprint%20for%20an%20AI%20Bill%20of%20Rights&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as policy or official document, not as a research paper.

## 36. Interpretability & Explainability

656. **"Why Should I Trust You?" Explaining the Predictions of Any Classifier (LIME)** — Ribeiro et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1602.04938>) · [Direct PDF](<https://arxiv.org/pdf/1602.04938.pdf>)

657. **A Unified Approach to Interpreting Model Predictions (SHAP)** — Lundberg & Lee · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1705.07874>) · [Direct PDF](<https://arxiv.org/pdf/1705.07874.pdf>)

658. **Grad-CAM: Visual Explanations from Deep Networks** — Selvaraju et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1610.02391>) · [Direct PDF](<https://arxiv.org/pdf/1610.02391.pdf>)

659. **Attention is not Explanation** — Jain & Wallace · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1902.10186>) · [Direct PDF](<https://arxiv.org/pdf/1902.10186.pdf>)

660. **Attention is not not Explanation** — Wiegreffe & Pinter · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Attention%20is%20not%20not%20Explanation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Attention%20is%20not%20not%20Explanation%20Wiegreffe%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Attention%20is%20not%20not%20Explanation&sort=relevance>)

661. **Network Dissection: Quantifying Interpretability of Deep Visual Representations** — Bau et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1704.05796>) · [Direct PDF](<https://arxiv.org/pdf/1704.05796.pdf>)

662. **Zoom In: An Introduction to Circuits** — Olah et al. · 2020 · **Type:** Research-paper candidate. [Supplied direct source](<https://distill.pub/2020/circuits/zoom-in/>)

663. **A Mathematical Framework for Transformer Circuits** — Elhage et al. · 2021 · **Type:** Research-paper candidate. [Supplied direct source](<https://transformer-circuits.pub/2021/framework/index.html>)

664. **Toy Models of Superposition** — Elhage et al. · 2022 · **Type:** Research-paper candidate. [Supplied direct source](<https://transformer-circuits.pub/2022/toy_model/index.html>)

665. **Towards Monosemanticity: Decomposing Language Models with Dictionary Learning** — Bricken et al. · 2023 · **Type:** Research-paper candidate. [Supplied direct source](<https://transformer-circuits.pub/2023/monosemantic-features>)

666. **Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet** — Templeton et al. · 2024 · **Type:** Research-paper candidate. [Supplied direct source](<https://transformer-circuits.pub/2024/scaling-monosemanticity/>)

667. **Representation Engineering: A Top-Down Approach to AI Transparency** — Zou et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Representation%20Engineering%3A%20A%20Top-Down%20Approach%20to%20AI%20Transparency&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Representation%20Engineering%3A%20A%20Top-Down%20Approach%20to%20AI%20Transparency%20Zou%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Representation%20Engineering%3A%20A%20Top-Down%20Approach%20to%20AI%20Transparency&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #987.

668. **Inference-Time Intervention: Eliciting Truthful Answers from a Language Model** — Li et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2306.03341>) · [Direct PDF](<https://arxiv.org/pdf/2306.03341.pdf>)

669. **Sparse Probing for LLM Representations** — Gurnee et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Sparse%20Probing%20for%20LLM%20Representations&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Sparse%20Probing%20for%20LLM%20Representations%20Gurnee%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Sparse%20Probing%20for%20LLM%20Representations&sort=relevance>)

670. **Circuit Discovery in LLMs** — Conmy et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Circuit%20Discovery%20in%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Circuit%20Discovery%20in%20LLMs%20Conmy%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Circuit%20Discovery%20in%20LLMs&sort=relevance>)

671. **Interpretability in the Wild** — Bills et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Interpretability%20in%20the%20Wild&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Interpretability%20in%20the%20Wild%20Bills%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Interpretability%20in%20the%20Wild&sort=relevance>)

672. **The Geometry of Truth: Emergent Linear Structure in LLM Representations** — Marks & Tegmark · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2310.06824>) · [Direct PDF](<https://arxiv.org/pdf/2310.06824.pdf>)

673. **Polysemanticity and Capacity in Neural Networks** — Anthropic · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Polysemanticity%20and%20Capacity%20in%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Polysemanticity%20and%20Capacity%20in%20Neural%20Networks%20Anthropic%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Polysemanticity%20and%20Capacity%20in%20Neural%20Networks&sort=relevance>)

674. **Probing Classifiers: Promises, Shortcomings, and Advances** — Belinkov · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Probing%20Classifiers%3A%20Promises%2C%20Shortcomings%2C%20and%20Advances&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Probing%20Classifiers%3A%20Promises%2C%20Shortcomings%2C%20and%20Advances%20Belinkov%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Probing%20Classifiers%3A%20Promises%2C%20Shortcomings%2C%20and%20Advances&sort=relevance>)

675. **Causal Abstraction for Faithful Model Interpretation** — Geiger et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Causal%20Abstraction%20for%20Faithful%20Model%20Interpretation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Causal%20Abstraction%20for%20Faithful%20Model%20Interpretation%20Geiger%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Causal%20Abstraction%20for%20Faithful%20Model%20Interpretation&sort=relevance>)

## 37. Robotics & Embodied AI

676. **Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan)** — Ahn et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2204.01691>) · [Direct PDF](<https://arxiv.org/pdf/2204.01691.pdf>)

677. **RT-1: Robotics Transformer for Real-World Control** — Brohan et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2212.06817>) · [Direct PDF](<https://arxiv.org/pdf/2212.06817.pdf>)

678. **RT-2: Vision-Language-Action Models** — Brohan et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2307.15818>) · [Direct PDF](<https://arxiv.org/pdf/2307.15818.pdf>)

679. **Open X-Embodiment: Robotic Learning Datasets and RT-X Models** — Open X-Embodiment · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Open%20X-Embodiment%3A%20Robotic%20Learning%20Datasets%20and%20RT-X%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Open%20X-Embodiment%3A%20Robotic%20Learning%20Datasets%20and%20RT-X%20Models%20Open%20X-Embodiment%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Open%20X-Embodiment%3A%20Robotic%20Learning%20Datasets%20and%20RT-X%20Models&sort=relevance>)

680. **PaLM-E: An Embodied Multimodal Language Model** — Driess et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.03378>) · [Direct PDF](<https://arxiv.org/pdf/2303.03378.pdf>)

681. **Code as Policies: Language Model Programs for Embodied Control** — Liang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Code%20as%20Policies%3A%20Language%20Model%20Programs%20for%20Embodied%20Control&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Code%20as%20Policies%3A%20Language%20Model%20Programs%20for%20Embodied%20Control%20Liang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Code%20as%20Policies%3A%20Language%20Model%20Programs%20for%20Embodied%20Control&sort=relevance>)

682. **Language Models as Zero-Shot Planners** — Huang et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Language%20Models%20as%20Zero-Shot%20Planners&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Language%20Models%20as%20Zero-Shot%20Planners%20Huang%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Language%20Models%20as%20Zero-Shot%20Planners&sort=relevance>)

683. **TidyBot: Personalized Robot Assistance with LLMs** — Wu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=TidyBot%3A%20Personalized%20Robot%20Assistance%20with%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=TidyBot%3A%20Personalized%20Robot%20Assistance%20with%20LLMs%20Wu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=TidyBot%3A%20Personalized%20Robot%20Assistance%20with%20LLMs&sort=relevance>)

684. **RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation** — Bousmalis et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=RoboCat%3A%20A%20Self-Improving%20Generalist%20Agent%20for%20Robotic%20Manipulation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=RoboCat%3A%20A%20Self-Improving%20Generalist%20Agent%20for%20Robotic%20Manipulation%20Bousmalis%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=RoboCat%3A%20A%20Self-Improving%20Generalist%20Agent%20for%20Robotic%20Manipulation&sort=relevance>)

685. **Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ALOHA)** — Zhao et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2304.13705>) · [Direct PDF](<https://arxiv.org/pdf/2304.13705.pdf>)

686. **Octo: An Open-Source Generalist Robot Policy** — Team et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Octo%3A%20An%20Open-Source%20Generalist%20Robot%20Policy&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Octo%3A%20An%20Open-Source%20Generalist%20Robot%20Policy%20Team%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Octo%3A%20An%20Open-Source%20Generalist%20Robot%20Policy&sort=relevance>)

687. **π0: A Vision-Language-Action Flow Model for General Robot Control** — Physical Intelligence · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=%CF%800%3A%20A%20Vision-Language-Action%20Flow%20Model%20for%20General%20Robot%20Control&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=%CF%800%3A%20A%20Vision-Language-Action%20Flow%20Model%20for%20General%20Robot%20Control%20Physical%20Intelligence%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=%CF%800%3A%20A%20Vision-Language-Action%20Flow%20Model%20for%20General%20Robot%20Control&sort=relevance>)

688. **NVIDIA Isaac and Omniverse for Robotics** — NVIDIA · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=NVIDIA%20Isaac%20and%20Omniverse%20for%20Robotics&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=NVIDIA%20Isaac%20and%20Omniverse%20for%20Robotics%20NVIDIA%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=NVIDIA%20Isaac%20and%20Omniverse%20for%20Robotics&sort=relevance>)

689. **Dexterous Manipulation with Reinforcement Learning** — OpenAI (Rubik's cube) · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Dexterous%20Manipulation%20with%20Reinforcement%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Dexterous%20Manipulation%20with%20Reinforcement%20Learning%20OpenAI%20%28Rubik%27s%20cube%29%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Dexterous%20Manipulation%20with%20Reinforcement%20Learning&sort=relevance>)

690. **Mobile ALOHA: Learning Bimanual Mobile Manipulation** — Fu et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2401.02117>) · [Direct PDF](<https://arxiv.org/pdf/2401.02117.pdf>)

## 38. World Models & Simulation

691. **World Models (Ha & Schmidhuber)** — Ha & Schmidhuber · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=World%20Models%20%28Ha%20%26%20Schmidhuber%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=World%20Models%20%28Ha%20%26%20Schmidhuber%29%20Ha%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=World%20Models%20%28Ha%20%26%20Schmidhuber%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #418.

692. **A Path Towards Autonomous Machine Intelligence** — LeCun · 2022 · **Type:** Research-paper candidate. [Supplied direct source](<https://openreview.net/pdf?id=BZ5a1r-kVsf>)

693. **Genie: Generative Interactive Environments** — Bruce et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2402.15391>) · [Direct PDF](<https://arxiv.org/pdf/2402.15391.pdf>)

694. **DIAMOND: Diffusion for World Modeling** — Alonso et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DIAMOND%3A%20Diffusion%20for%20World%20Modeling&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DIAMOND%3A%20Diffusion%20for%20World%20Modeling%20Alonso%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DIAMOND%3A%20Diffusion%20for%20World%20Modeling&sort=relevance>)

695. **UniSim: Learning Interactive Real-World Simulators** — Yang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=UniSim%3A%20Learning%20Interactive%20Real-World%20Simulators&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=UniSim%3A%20Learning%20Interactive%20Real-World%20Simulators%20Yang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=UniSim%3A%20Learning%20Interactive%20Real-World%20Simulators&sort=relevance>)

696. **Genesis: Generative and Universal Physics Engine** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Genesis%3A%20Generative%20and%20Universal%20Physics%20Engine&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Genesis%3A%20Generative%20and%20Universal%20Physics%20Engine%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Genesis%3A%20Generative%20and%20Universal%20Physics%20Engine&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

697. **Learning General World Models in a Handful of Reward-Free Deployments** — Seo et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Learning%20General%20World%20Models%20in%20a%20Handful%20of%20Reward-Free%20Deployments&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Learning%20General%20World%20Models%20in%20a%20Handful%20of%20Reward-Free%20Deployments%20Seo%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Learning%20General%20World%20Models%20in%20a%20Handful%20of%20Reward-Free%20Deployments&sort=relevance>)

698. **GAIA-1: A Generative World Model for Autonomous Driving** — Hu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=GAIA-1%3A%20A%20Generative%20World%20Model%20for%20Autonomous%20Driving&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=GAIA-1%3A%20A%20Generative%20World%20Model%20for%20Autonomous%20Driving%20Hu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=GAIA-1%3A%20A%20Generative%20World%20Model%20for%20Autonomous%20Driving&sort=relevance>)

699. **Language Models Meet World Models** — Hao et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Language%20Models%20Meet%20World%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Language%20Models%20Meet%20World%20Models%20Hao%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Language%20Models%20Meet%20World%20Models&sort=relevance>)

700. **Video as World Model** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Video%20as%20World%20Model&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Video%20as%20World%20Model%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Video%20as%20World%20Model&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 39. Scientific AI & Domain Applications

701. **AlphaFold: Highly Accurate Protein Structure Prediction** — Jumper et al. · 2021 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/s41586-021-03819-2>)

702. **AlphaFold 2 (full paper)** — Jumper et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AlphaFold%202%20%28full%20paper%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AlphaFold%202%20%28full%20paper%29%20Jumper%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AlphaFold%202%20%28full%20paper%29&sort=relevance>)

703. **AlphaFold 3** — Abramson et al. · 2024 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/s41586-024-07487-w>)

704. **RoseTTAFold** — Baek et al. · 2021 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.science.org/doi/10.1126/science.abj8754>)

705. **ESM-2: Language Models of Protein Sequences at the Scale of Evolution** — Lin et al. · 2023 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.science.org/doi/10.1126/science.ade2574>)

706. **ProteinMPNN: Robust Protein Sequence Design** — Dauparas et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ProteinMPNN%3A%20Robust%20Protein%20Sequence%20Design&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ProteinMPNN%3A%20Robust%20Protein%20Sequence%20Design%20Dauparas%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ProteinMPNN%3A%20Robust%20Protein%20Sequence%20Design&sort=relevance>)

707. **GNoME: Scaling Deep Learning for Materials Discovery** — Merchant et al. · 2023 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/s41586-023-06735-9>)

708. **MatterGen: A Generative Model for Inorganic Materials Design** — Zeni et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MatterGen%3A%20A%20Generative%20Model%20for%20Inorganic%20Materials%20Design&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MatterGen%3A%20A%20Generative%20Model%20for%20Inorganic%20Materials%20Design%20Zeni%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MatterGen%3A%20A%20Generative%20Model%20for%20Inorganic%20Materials%20Design&sort=relevance>)

709. **FunSearch: Mathematical Discoveries from Program Search** — Romera-Paredes et al. · 2024 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/s41586-023-06924-6>)

710. **AlphaGeometry: Solving Olympiad Geometry without Human Demonstrations** — Trinh et al. · 2024 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.nature.com/articles/s41586-023-06747-5>)

711. **AlphaProof and AlphaGeometry 2** — Google DeepMind · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AlphaProof%20and%20AlphaGeometry%202&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AlphaProof%20and%20AlphaGeometry%202%20Google%20DeepMind%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AlphaProof%20and%20AlphaGeometry%202&sort=relevance>)

712. **Med-PaLM: Large Language Models Encode Clinical Knowledge** — Singhal et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2212.13138>) · [Direct PDF](<https://arxiv.org/pdf/2212.13138.pdf>)

713. **Med-PaLM 2** — Singhal et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Med-PaLM%202&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Med-PaLM%202%20Singhal%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Med-PaLM%202&sort=relevance>)

714. **PMC-LLaMA: Towards Building Open-source Language Models for Medicine** — Wu et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=PMC-LLaMA%3A%20Towards%20Building%20Open-source%20Language%20Models%20for%20Medicine&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=PMC-LLaMA%3A%20Towards%20Building%20Open-source%20Language%20Models%20for%20Medicine%20Wu%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=PMC-LLaMA%3A%20Towards%20Building%20Open-source%20Language%20Models%20for%20Medicine&sort=relevance>)

715. **BioGPT: Generative Pre-trained Transformer for Biomedical Text** — Luo et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=BioGPT%3A%20Generative%20Pre-trained%20Transformer%20for%20Biomedical%20Text&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=BioGPT%3A%20Generative%20Pre-trained%20Transformer%20for%20Biomedical%20Text%20Luo%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=BioGPT%3A%20Generative%20Pre-trained%20Transformer%20for%20Biomedical%20Text&sort=relevance>)

716. **Galactica: A Large Language Model for Science** — Taylor et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2211.09085>) · [Direct PDF](<https://arxiv.org/pdf/2211.09085.pdf>)

717. **ScienceQA: Science Question Answering** — Lu et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ScienceQA%3A%20Science%20Question%20Answering&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ScienceQA%3A%20Science%20Question%20Answering%20Lu%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ScienceQA%3A%20Science%20Question%20Answering&sort=relevance>)

718. **ChemCrow: Augmenting LLMs with Chemistry Tools** — Bran et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ChemCrow%3A%20Augmenting%20LLMs%20with%20Chemistry%20Tools&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ChemCrow%3A%20Augmenting%20LLMs%20with%20Chemistry%20Tools%20Bran%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ChemCrow%3A%20Augmenting%20LLMs%20with%20Chemistry%20Tools&sort=relevance>)

719. **ClimateBERT** — Webersinke et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ClimateBERT&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ClimateBERT%20Webersinke%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ClimateBERT&sort=relevance>)

720. **GraphCast: Learning Skillful Medium-Range Global Weather Forecasting** — Lam et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2212.12794>) · [Direct PDF](<https://arxiv.org/pdf/2212.12794.pdf>)

## 40. Scaling Laws & Emergent Abilities

721. **Scaling Laws for Neural Language Models** — Kaplan et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2001.08361>) · [Direct PDF](<https://arxiv.org/pdf/2001.08361.pdf>)

722. **Training Compute-Optimal LLMs (Chinchilla)** — Hoffmann et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Training%20Compute-Optimal%20LLMs%20%28Chinchilla%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Training%20Compute-Optimal%20LLMs%20%28Chinchilla%29%20Hoffmann%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Training%20Compute-Optimal%20LLMs%20%28Chinchilla%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #200.

723. **Emergent Abilities of Large Language Models** — Wei et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2206.07682>) · [Direct PDF](<https://arxiv.org/pdf/2206.07682.pdf>)

724. **Are Emergent Abilities of LLMs a Mirage?** — Schaeffer et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2304.15004>) · [Direct PDF](<https://arxiv.org/pdf/2304.15004.pdf>)

725. **Scaling Data-Constrained Language Models** — Muennighoff et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.16264>) · [Direct PDF](<https://arxiv.org/pdf/2305.16264.pdf>)

726. **Scaling Laws for Autoregressive Generative Modeling** — Henighan et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2010.14701>) · [Direct PDF](<https://arxiv.org/pdf/2010.14701.pdf>)

727. **Beyond Neural Scaling Laws** — Caballero et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Beyond%20Neural%20Scaling%20Laws&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Beyond%20Neural%20Scaling%20Laws%20Caballero%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Beyond%20Neural%20Scaling%20Laws&sort=relevance>)

728. **Scaling Laws for Reward Model Overoptimization** — Gao et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Scaling%20Laws%20for%20Reward%20Model%20Overoptimization&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scaling%20Laws%20for%20Reward%20Model%20Overoptimization%20Gao%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scaling%20Laws%20for%20Reward%20Model%20Overoptimization&sort=relevance>)

729. **Scaling Vision Transformers** — Zhai et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Scaling%20Vision%20Transformers&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scaling%20Vision%20Transformers%20Zhai%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scaling%20Vision%20Transformers&sort=relevance>)

730. **Observational Scaling Laws** — Ruan et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Observational%20Scaling%20Laws&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Observational%20Scaling%20Laws%20Ruan%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Observational%20Scaling%20Laws&sort=relevance>)

731. **Textbooks Are All You Need (Phi-1 scaling)** — Gunasekar et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Textbooks%20Are%20All%20You%20Need%20%28Phi-1%20scaling%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Textbooks%20Are%20All%20You%20Need%20%28Phi-1%20scaling%29%20Gunasekar%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Textbooks%20Are%20All%20You%20Need%20%28Phi-1%20scaling%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #197, #782.

732. **The Pile: An 800GB Dataset for Diverse Text Modeling** — Gao et al. · 2020 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=The%20Pile%3A%20An%20800GB%20Dataset%20for%20Diverse%20Text%20Modeling&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Pile%3A%20An%20800GB%20Dataset%20for%20Diverse%20Text%20Modeling%20Gao%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Pile%3A%20An%20800GB%20Dataset%20for%20Diverse%20Text%20Modeling&sort=relevance>)

733. **RedPajama: An Open Dataset for Training LLMs** — Together AI · 2023 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=RedPajama%3A%20An%20Open%20Dataset%20for%20Training%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=RedPajama%3A%20An%20Open%20Dataset%20for%20Training%20LLMs%20Together%20AI%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=RedPajama%3A%20An%20Open%20Dataset%20for%20Training%20LLMs&sort=relevance>)

734. **FineWeb: Decanting the Web for the Finest Text Data** — Penedo et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2406.17557>) · [Direct PDF](<https://arxiv.org/pdf/2406.17557.pdf>)

735. **DataComp: In Search of the Next Generation of Multimodal Datasets** — Gadre et al. · 2023 · **Type:** Benchmark or dataset paper. [Canonical arXiv record](<https://arxiv.org/abs/2304.14108>) · [Direct PDF](<https://arxiv.org/pdf/2304.14108.pdf>)

## 41. Architecture Innovations & Beyond Transformers

736. **Mamba: Linear-Time Sequence Modeling with Selective State Spaces** — Gu & Dao · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2312.00752>) · [Direct PDF](<https://arxiv.org/pdf/2312.00752.pdf>)

737. **Mamba-2: Structured State Space Duality** — Dao & Gu · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2405.21060>) · [Direct PDF](<https://arxiv.org/pdf/2405.21060.pdf>)

738. **Efficiently Modeling Long Sequences with Structured State Spaces (S4)** — Gu et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2111.00396>) · [Direct PDF](<https://arxiv.org/pdf/2111.00396.pdf>)

739. **Hyena Hierarchy: Towards Larger Convolutional Language Models** — Poli et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2302.10866>) · [Direct PDF](<https://arxiv.org/pdf/2302.10866.pdf>)

740. **RWKV: Reinventing RNNs for the Transformer Era** — Peng et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.13048>) · [Direct PDF](<https://arxiv.org/pdf/2305.13048.pdf>)

741. **Griffin: Mixing Gated Linear Recurrences with Local Attention** — De et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2402.19427>) · [Direct PDF](<https://arxiv.org/pdf/2402.19427.pdf>)

742. **RetNet: Retentive Network: A Successor to Transformer** — Sun et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2307.08621>) · [Direct PDF](<https://arxiv.org/pdf/2307.08621.pdf>)

743. **xLSTM: Extended Long Short-Term Memory** — Beck et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2405.04517>) · [Direct PDF](<https://arxiv.org/pdf/2405.04517.pdf>)

744. **Jamba: Hybrid Transformer-Mamba** — AI21 Labs · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Jamba%3A%20Hybrid%20Transformer-Mamba&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Jamba%3A%20Hybrid%20Transformer-Mamba%20AI21%20Labs%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Jamba%3A%20Hybrid%20Transformer-Mamba&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #214.

745. **Mixture of Depths: Dynamically Allocating Compute in Transformer-Based Models** — Raposo et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2404.02258>) · [Direct PDF](<https://arxiv.org/pdf/2404.02258.pdf>)

746. **Neural Architecture Search (NAS)** — Zoph & Le · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1611.01578>) · [Direct PDF](<https://arxiv.org/pdf/1611.01578.pdf>)

747. **EfficientNet (NAS-derived)** — Tan & Le · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=EfficientNet%20%28NAS-derived%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=EfficientNet%20%28NAS-derived%29%20Tan%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=EfficientNet%20%28NAS-derived%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #86.

748. **Differential Transformer** — Ye et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2410.05258>) · [Direct PDF](<https://arxiv.org/pdf/2410.05258.pdf>)

749. **Kolmogorov-Arnold Networks (KAN)** — Liu et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2404.19756>) · [Direct PDF](<https://arxiv.org/pdf/2404.19756.pdf>)

750. **TTT: Learning to (Learn at Test Time)** — Sun et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2407.04620>) · [Direct PDF](<https://arxiv.org/pdf/2407.04620.pdf>)

## 42. AGI Theory & Frontier Research

751. **Sparks of AGI: Early Experiments with GPT-4** — Bubeck et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.12712>) · [Direct PDF](<https://arxiv.org/pdf/2303.12712.pdf>)

752. **Levels of AGI** — Morris et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Levels%20of%20AGI&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Levels%20of%20AGI%20Morris%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Levels%20of%20AGI&sort=relevance>)

753. **The Bitter Lesson** — Sutton · 2019 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.incompleteideas.net/IncIdeas/BitterLesson.html>)

754. **Reward is Enough** — Silver et al. · 2021 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.sciencedirect.com/science/article/pii/S0004370221000862>)

755. **Artificial General Intelligence: Concept, State of the Art, and Future Prospects** — Goertzel · 2014 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Artificial%20General%20Intelligence%3A%20Concept%2C%20State%20of%20the%20Art%2C%20and%20Future%20Prospects&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Artificial%20General%20Intelligence%3A%20Concept%2C%20State%20of%20the%20Art%2C%20and%20Future%20Prospects%20Goertzel%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Artificial%20General%20Intelligence%3A%20Concept%2C%20State%20of%20the%20Art%2C%20and%20Future%20Prospects&sort=relevance>)

756. **On the Measure of Intelligence** — Chollet · 2019 · **Type:** Benchmark or dataset paper. [Canonical arXiv record](<https://arxiv.org/abs/1911.01547>) · [Direct PDF](<https://arxiv.org/pdf/1911.01547.pdf>)

757. **ARC Prize and ARC-AGI** — Chollet et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ARC%20Prize%20and%20ARC-AGI&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ARC%20Prize%20and%20ARC-AGI%20Chollet%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ARC%20Prize%20and%20ARC-AGI&sort=relevance>)

758. **Language Agent Tree Search Unifies Reasoning, Acting, and Planning** — Zhou et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Language%20Agent%20Tree%20Search%20Unifies%20Reasoning%2C%20Acting%2C%20and%20Planning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Language%20Agent%20Tree%20Search%20Unifies%20Reasoning%2C%20Acting%2C%20and%20Planning%20Zhou%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Language%20Agent%20Tree%20Search%20Unifies%20Reasoning%2C%20Acting%2C%20and%20Planning&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #346.

759. **Superintelligence: Paths, Dangers, Strategies (summary)** — Bostrom · 2014 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Superintelligence%3A%20Paths%2C%20Dangers%2C%20Strategies%20%28summary%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Superintelligence%3A%20Paths%2C%20Dangers%2C%20Strategies%20%28summary%29%20Bostrom%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Superintelligence%3A%20Paths%2C%20Dangers%2C%20Strategies%20%28summary%29&sort=relevance>)

760. **The Case for AI Safety** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Case%20for%20AI%20Safety&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Case%20for%20AI%20Safety%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Case%20for%20AI%20Safety&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 43. Knowledge Graphs & Structured Knowledge

761. **Translating Embeddings for Modeling Multi-relational Data (TransE)** — Bordes et al. · 2013 · **Type:** Research-paper candidate. [Supplied direct source](<https://papers.nips.cc/paper/2013/hash/1cecc7a77928ca8133fa24680a88d2f9-Abstract.html>)

762. **Knowledge Graph Embedding by Translating on Hyperplanes (TransH)** — Wang et al. · 2014 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Knowledge%20Graph%20Embedding%20by%20Translating%20on%20Hyperplanes%20%28TransH%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Knowledge%20Graph%20Embedding%20by%20Translating%20on%20Hyperplanes%20%28TransH%29%20Wang%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Knowledge%20Graph%20Embedding%20by%20Translating%20on%20Hyperplanes%20%28TransH%29&sort=relevance>)

763. **RotatE: Knowledge Graph Embedding by Relational Rotation** — Sun et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=RotatE%3A%20Knowledge%20Graph%20Embedding%20by%20Relational%20Rotation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=RotatE%3A%20Knowledge%20Graph%20Embedding%20by%20Relational%20Rotation%20Sun%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=RotatE%3A%20Knowledge%20Graph%20Embedding%20by%20Relational%20Rotation&sort=relevance>)

764. **KGQA: Knowledge Graph Question Answering** — Various · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=KGQA%3A%20Knowledge%20Graph%20Question%20Answering&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=KGQA%3A%20Knowledge%20Graph%20Question%20Answering%20Various%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=KGQA%3A%20Knowledge%20Graph%20Question%20Answering&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

765. **Unifying LLMs and Knowledge Graphs: A Roadmap** — Pan et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Unifying%20LLMs%20and%20Knowledge%20Graphs%3A%20A%20Roadmap&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Unifying%20LLMs%20and%20Knowledge%20Graphs%3A%20A%20Roadmap%20Pan%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Unifying%20LLMs%20and%20Knowledge%20Graphs%3A%20A%20Roadmap&sort=relevance>)

766. **Think-on-Graph: Deep and Responsible Reasoning of LLMs on KGs** — Sun et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Think-on-Graph%3A%20Deep%20and%20Responsible%20Reasoning%20of%20LLMs%20on%20KGs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Think-on-Graph%3A%20Deep%20and%20Responsible%20Reasoning%20of%20LLMs%20on%20KGs%20Sun%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Think-on-Graph%3A%20Deep%20and%20Responsible%20Reasoning%20of%20LLMs%20on%20KGs&sort=relevance>)

767. **KnowPrompt: Knowledge-aware Prompt-tuning** — Chen et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=KnowPrompt%3A%20Knowledge-aware%20Prompt-tuning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=KnowPrompt%3A%20Knowledge-aware%20Prompt-tuning%20Chen%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=KnowPrompt%3A%20Knowledge-aware%20Prompt-tuning&sort=relevance>)

768. **Wikidata and Knowledge Graphs at Scale** — Vrandečić & Krötzsch · 2014 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Wikidata%20and%20Knowledge%20Graphs%20at%20Scale&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Wikidata%20and%20Knowledge%20Graphs%20at%20Scale%20Vrande%C4%8Di%C4%87%202014>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Wikidata%20and%20Knowledge%20Graphs%20at%20Scale&sort=relevance>)

769. **QA-GNN: Reasoning with Language Models and KGs** — Yasunaga et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=QA-GNN%3A%20Reasoning%20with%20Language%20Models%20and%20KGs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=QA-GNN%3A%20Reasoning%20with%20Language%20Models%20and%20KGs%20Yasunaga%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=QA-GNN%3A%20Reasoning%20with%20Language%20Models%20and%20KGs&sort=relevance>)

770. **Graph-based Deep Learning for NLP** — Various · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Graph-based%20Deep%20Learning%20for%20NLP&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Graph-based%20Deep%20Learning%20for%20NLP%20Various%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Graph-based%20Deep%20Learning%20for%20NLP&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 44. Natural Language Understanding & Generation

771. **Named Entity Recognition with Bidirectional LSTM-CNNs** — Chiu & Nichols · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Named%20Entity%20Recognition%20with%20Bidirectional%20LSTM-CNNs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Named%20Entity%20Recognition%20with%20Bidirectional%20LSTM-CNNs%20Chiu%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Named%20Entity%20Recognition%20with%20Bidirectional%20LSTM-CNNs&sort=relevance>)

772. **End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF** — Ma & Hovy · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=End-to-end%20Sequence%20Labeling%20via%20Bi-directional%20LSTM-CNNs-CRF&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=End-to-end%20Sequence%20Labeling%20via%20Bi-directional%20LSTM-CNNs-CRF%20Ma%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=End-to-end%20Sequence%20Labeling%20via%20Bi-directional%20LSTM-CNNs-CRF&sort=relevance>)

773. **Neural Machine Translation of Rare Words with Subword Units (BPE)** — Sennrich et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1508.07909>) · [Direct PDF](<https://arxiv.org/pdf/1508.07909.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #921.

774. **SentencePiece** — Kudo & Richardson · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SentencePiece&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SentencePiece%20Kudo%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SentencePiece&sort=relevance>)

775. **A Call for Clarity in Reporting BLEU Scores** — Post · 2018 · **Type:** Technical or industry report. [arXiv search](<https://arxiv.org/search/?query=A%20Call%20for%20Clarity%20in%20Reporting%20BLEU%20Scores&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Call%20for%20Clarity%20in%20Reporting%20BLEU%20Scores%20Post%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Call%20for%20Clarity%20in%20Reporting%20BLEU%20Scores&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

776. **BERTScore: Evaluating Text Generation with BERT** — Zhang et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1904.09675>) · [Direct PDF](<https://arxiv.org/pdf/1904.09675.pdf>)

777. **ROUGE: A Package for Automatic Evaluation of Summaries** — Lin · 2004 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ROUGE%3A%20A%20Package%20for%20Automatic%20Evaluation%20of%20Summaries&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ROUGE%3A%20A%20Package%20for%20Automatic%20Evaluation%20of%20Summaries%20Lin%202004>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ROUGE%3A%20A%20Package%20for%20Automatic%20Evaluation%20of%20Summaries&sort=relevance>)

778. **Get To The Point: Summarization with Pointer-Generator Networks** — See et al. · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Get%20To%20The%20Point%3A%20Summarization%20with%20Pointer-Generator%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Get%20To%20The%20Point%3A%20Summarization%20with%20Pointer-Generator%20Networks%20See%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Get%20To%20The%20Point%3A%20Summarization%20with%20Pointer-Generator%20Networks&sort=relevance>)

779. **Abstractive Text Summarization using Seq-to-Seq RNNs** — Nallapati et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Abstractive%20Text%20Summarization%20using%20Seq-to-Seq%20RNNs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Abstractive%20Text%20Summarization%20using%20Seq-to-Seq%20RNNs%20Nallapati%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Abstractive%20Text%20Summarization%20using%20Seq-to-Seq%20RNNs&sort=relevance>)

780. **A Structured Self-Attentive Sentence Embedding** — Lin et al. · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=A%20Structured%20Self-Attentive%20Sentence%20Embedding&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Structured%20Self-Attentive%20Sentence%20Embedding%20Lin%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Structured%20Self-Attentive%20Sentence%20Embedding&sort=relevance>)

## 45. Synthetic Data & Data Generation

781. **Self-Instruct** — Wang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Self-Instruct&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Self-Instruct%20Wang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Self-Instruct&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #222.

782. **Textbooks Are All You Need** — Gunasekar et al. · 2023 · **Type:** Book or textbook. [arXiv search](<https://arxiv.org/search/?query=Textbooks%20Are%20All%20You%20Need&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Textbooks%20Are%20All%20You%20Need%20Gunasekar%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Textbooks%20Are%20All%20You%20Need&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as book or textbook, not as a research paper; DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #197, #731.

783. **Cosmopedia: Creating Large-Scale Synthetic Data** — HuggingFace · 2024 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Cosmopedia%3A%20Creating%20Large-Scale%20Synthetic%20Data&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Cosmopedia%3A%20Creating%20Large-Scale%20Synthetic%20Data%20HuggingFace%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Cosmopedia%3A%20Creating%20Large-Scale%20Synthetic%20Data&sort=relevance>)

784. **Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling** — Maini et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Rephrasing%20the%20Web%3A%20A%20Recipe%20for%20Compute%20and%20Data-Efficient%20Language%20Modeling&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Rephrasing%20the%20Web%3A%20A%20Recipe%20for%20Compute%20and%20Data-Efficient%20Language%20Modeling%20Maini%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Rephrasing%20the%20Web%3A%20A%20Recipe%20for%20Compute%20and%20Data-Efficient%20Language%20Modeling&sort=relevance>)

785. **Magpie: Alignment Data Synthesis from Scratch** — Xu et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Magpie%3A%20Alignment%20Data%20Synthesis%20from%20Scratch&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Magpie%3A%20Alignment%20Data%20Synthesis%20from%20Scratch%20Xu%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Magpie%3A%20Alignment%20Data%20Synthesis%20from%20Scratch&sort=relevance>)

786. **WizardLM: Empowering LLMs to Follow Complex Instructions** — Xu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=WizardLM%3A%20Empowering%20LLMs%20to%20Follow%20Complex%20Instructions&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=WizardLM%3A%20Empowering%20LLMs%20to%20Follow%20Complex%20Instructions%20Xu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=WizardLM%3A%20Empowering%20LLMs%20to%20Follow%20Complex%20Instructions&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #229.

787. **AgentInstruct: Toward Generative Teaching with Agentic Flows** — Mitra et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AgentInstruct%3A%20Toward%20Generative%20Teaching%20with%20Agentic%20Flows&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AgentInstruct%3A%20Toward%20Generative%20Teaching%20with%20Agentic%20Flows%20Mitra%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AgentInstruct%3A%20Toward%20Generative%20Teaching%20with%20Agentic%20Flows&sort=relevance>)

788. **SPIN: Self-Play Fine-Tuning Converts Weak LMs to Strong** — Chen et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SPIN%3A%20Self-Play%20Fine-Tuning%20Converts%20Weak%20LMs%20to%20Strong&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SPIN%3A%20Self-Play%20Fine-Tuning%20Converts%20Weak%20LMs%20to%20Strong%20Chen%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SPIN%3A%20Self-Play%20Fine-Tuning%20Converts%20Weak%20LMs%20to%20Strong&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #245.

789. **DataDreamer: A Tool for Synthetically Generating, Transforming, and Analyzing NLP Data** — Patel et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DataDreamer%3A%20A%20Tool%20for%20Synthetically%20Generating%2C%20Transforming%2C%20and%20Analyzing%20NLP%20Data&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DataDreamer%3A%20A%20Tool%20for%20Synthetically%20Generating%2C%20Transforming%2C%20and%20Analyzing%20NLP%20Data%20Patel%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DataDreamer%3A%20A%20Tool%20for%20Synthetically%20Generating%2C%20Transforming%2C%20and%20Analyzing%20NLP%20Data&sort=relevance>)

790. **Scaling Synthetic Data Creation with 1B Personas** — Chan et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Scaling%20Synthetic%20Data%20Creation%20with%201B%20Personas&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scaling%20Synthetic%20Data%20Creation%20with%201B%20Personas%20Chan%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scaling%20Synthetic%20Data%20Creation%20with%201B%20Personas&sort=relevance>)

## 46. Long Context & Memory Systems

791. **Memorizing Transformers** — Wu et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Memorizing%20Transformers&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Memorizing%20Transformers%20Wu%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Memorizing%20Transformers&sort=relevance>)

792. **∞-former: Infinite Memory Transformer** — Martins et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=%E2%88%9E-former%3A%20Infinite%20Memory%20Transformer&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=%E2%88%9E-former%3A%20Infinite%20Memory%20Transformer%20Martins%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=%E2%88%9E-former%3A%20Infinite%20Memory%20Transformer&sort=relevance>)

793. **LongNet: Scaling Transformers to 1B Tokens** — Ding et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LongNet%3A%20Scaling%20Transformers%20to%201B%20Tokens&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LongNet%3A%20Scaling%20Transformers%20to%201B%20Tokens%20Ding%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LongNet%3A%20Scaling%20Transformers%20to%201B%20Tokens&sort=relevance>)

794. **Leave No Context Behind: Efficient Infinite-Context Transformers with Infini-attention** — Munkhdalai et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Leave%20No%20Context%20Behind%3A%20Efficient%20Infinite-Context%20Transformers%20with%20Infini-attention&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Leave%20No%20Context%20Behind%3A%20Efficient%20Infinite-Context%20Transformers%20with%20Infini-attention%20Munkhdalai%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Leave%20No%20Context%20Behind%3A%20Efficient%20Infinite-Context%20Transformers%20with%20Infini-attention&sort=relevance>)

795. **Needle in a Haystack Evaluation** — Kamradt · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Needle%20in%20a%20Haystack%20Evaluation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Needle%20in%20a%20Haystack%20Evaluation%20Kamradt%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Needle%20in%20a%20Haystack%20Evaluation&sort=relevance>)

796. **Ruler: What's the Real Context Size of Your LLM?** — Hsieh et al. · 2024 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Ruler%3A%20What%27s%20the%20Real%20Context%20Size%20of%20Your%20LLM%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Ruler%3A%20What%27s%20the%20Real%20Context%20Size%20of%20Your%20LLM%3F%20Hsieh%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Ruler%3A%20What%27s%20the%20Real%20Context%20Size%20of%20Your%20LLM%3F&sort=relevance>)

797. **MemGPT: Towards LLMs as Operating Systems** — Packer et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2310.08560>) · [Direct PDF](<https://arxiv.org/pdf/2310.08560.pdf>)

798. **LongRoPE: Extending LLM Context Window Beyond 2M Tokens** — Ding et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LongRoPE%3A%20Extending%20LLM%20Context%20Window%20Beyond%202M%20Tokens&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LongRoPE%3A%20Extending%20LLM%20Context%20Window%20Beyond%202M%20Tokens%20Ding%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LongRoPE%3A%20Extending%20LLM%20Context%20Window%20Beyond%202M%20Tokens&sort=relevance>)

799. **Claude's 200K Context (Anthropic)** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Claude%27s%20200K%20Context%20%28Anthropic%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Claude%27s%20200K%20Context%20%28Anthropic%29%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Claude%27s%20200K%20Context%20%28Anthropic%29&sort=relevance>)

800. **Gemini 1.5 Pro: Long Context (1M tokens)** — Google DeepMind · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2403.05530>) · [Direct PDF](<https://arxiv.org/pdf/2403.05530.pdf>)

## 47. Structured & Constrained Generation

801. **Outlines: Structured Text Generation** — Willard & Louf · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2307.09702>) · [Direct PDF](<https://arxiv.org/pdf/2307.09702.pdf>)

802. **Guidance: Constrained Generation** — Microsoft · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Guidance%3A%20Constrained%20Generation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Guidance%3A%20Constrained%20Generation%20Microsoft%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Guidance%3A%20Constrained%20Generation&sort=relevance>)

803. **Grammar of Thought: LLM Structured Reasoning** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Grammar%20of%20Thought%3A%20LLM%20Structured%20Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Grammar%20of%20Thought%3A%20LLM%20Structured%20Reasoning%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Grammar%20of%20Thought%3A%20LLM%20Structured%20Reasoning&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

804. **LMQL: Language Model Query Language** — Beurer-Kellner et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2212.06094>) · [Direct PDF](<https://arxiv.org/pdf/2212.06094.pdf>)

805. **DSPy: Compiling Declarative Language Model Calls into Pipelines** — Khattab et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2310.03714>) · [Direct PDF](<https://arxiv.org/pdf/2310.03714.pdf>)

806. **SGLang: Efficient Execution of Structured LM Programs** — Zheng et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2312.07104>) · [Direct PDF](<https://arxiv.org/pdf/2312.07104.pdf>)

807. **Instructor: Structured Outputs from LLMs** — Liu · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Instructor%3A%20Structured%20Outputs%20from%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Instructor%3A%20Structured%20Outputs%20from%20LLMs%20Liu%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Instructor%3A%20Structured%20Outputs%20from%20LLMs&sort=relevance>)

808. **TypeChat** — Microsoft · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=TypeChat&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=TypeChat%20Microsoft%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=TypeChat&sort=relevance>)

809. **JSON Mode / Structured Outputs (OpenAI, Anthropic)** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=JSON%20Mode%20/%20Structured%20Outputs%20%28OpenAI%2C%20Anthropic%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=JSON%20Mode%20/%20Structured%20Outputs%20%28OpenAI%2C%20Anthropic%29%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=JSON%20Mode%20/%20Structured%20Outputs%20%28OpenAI%2C%20Anthropic%29&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

810. **Semantic Kernel** — Microsoft · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Semantic%20Kernel&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Semantic%20Kernel%20Microsoft%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Semantic%20Kernel&sort=relevance>)

## 48. AI for Finance (BFSI-relevant)

811. **BloombergGPT: A Large Language Model for Finance** — Wu et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2303.17564>) · [Direct PDF](<https://arxiv.org/pdf/2303.17564.pdf>)

812. **FinGPT: Open-Source Financial Large Language Models** — Yang et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2306.06031>) · [Direct PDF](<https://arxiv.org/pdf/2306.06031.pdf>)

813. **FinBERT: Financial Sentiment Analysis** — Araci · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1908.10063>) · [Direct PDF](<https://arxiv.org/pdf/1908.10063.pdf>)

814. **Deep Learning for Stock Market Prediction** — Various · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Deep%20Learning%20for%20Stock%20Market%20Prediction&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Deep%20Learning%20for%20Stock%20Market%20Prediction%20Various%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Deep%20Learning%20for%20Stock%20Market%20Prediction&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

815. **LLMs for Financial NLP: Opportunities and Challenges** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLMs%20for%20Financial%20NLP%3A%20Opportunities%20and%20Challenges&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLMs%20for%20Financial%20NLP%3A%20Opportunities%20and%20Challenges%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLMs%20for%20Financial%20NLP%3A%20Opportunities%20and%20Challenges&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

816. **Fraud Detection using Machine Learning** — Various · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Fraud%20Detection%20using%20Machine%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Fraud%20Detection%20using%20Machine%20Learning%20Various%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Fraud%20Detection%20using%20Machine%20Learning&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

817. **Credit Scoring with Machine Learning** — Lessmann et al. · 2015 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Credit%20Scoring%20with%20Machine%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Credit%20Scoring%20with%20Machine%20Learning%20Lessmann%202015>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Credit%20Scoring%20with%20Machine%20Learning&sort=relevance>)

818. **KYC/AML with NLP** — Various · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=KYC/AML%20with%20NLP&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=KYC/AML%20with%20NLP%20Various%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=KYC/AML%20with%20NLP&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

819. **Explainable AI for Banking** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Explainable%20AI%20for%20Banking&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Explainable%20AI%20for%20Banking%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Explainable%20AI%20for%20Banking&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

820. **Document AI for Financial Services** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Document%20AI%20for%20Financial%20Services&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Document%20AI%20for%20Financial%20Services%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Document%20AI%20for%20Financial%20Services&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 49. Document AI & Information Extraction

821. **LayoutLM: Pre-training of Text and Layout for Document AI** — Xu et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1912.13318>) · [Direct PDF](<https://arxiv.org/pdf/1912.13318.pdf>)

822. **LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking** — Huang et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2204.08387>) · [Direct PDF](<https://arxiv.org/pdf/2204.08387.pdf>)

823. **Donut: Document Understanding Transformer** — Kim et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2111.15664>) · [Direct PDF](<https://arxiv.org/pdf/2111.15664.pdf>)

824. **PaddleOCR** — PaddlePaddle · 2020 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=PaddleOCR&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=PaddleOCR%20PaddlePaddle%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=PaddleOCR&sort=relevance>)

825. **TrOCR: Transformer-based Optical Character Recognition** — Li et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=TrOCR%3A%20Transformer-based%20Optical%20Character%20Recognition&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=TrOCR%3A%20Transformer-based%20Optical%20Character%20Recognition%20Li%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=TrOCR%3A%20Transformer-based%20Optical%20Character%20Recognition&sort=relevance>)

826. **Table Transformer (TATR)** — Smock et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Table%20Transformer%20%28TATR%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Table%20Transformer%20%28TATR%29%20Smock%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Table%20Transformer%20%28TATR%29&sort=relevance>)

827. **DocPrompting: Generating Code by Retrieving the Docs** — Zhou et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DocPrompting%3A%20Generating%20Code%20by%20Retrieving%20the%20Docs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DocPrompting%3A%20Generating%20Code%20by%20Retrieving%20the%20Docs%20Zhou%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DocPrompting%3A%20Generating%20Code%20by%20Retrieving%20the%20Docs&sort=relevance>)

828. **Nougat: Neural Optical Understanding for Academic Documents** — Blecher et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2308.13418>) · [Direct PDF](<https://arxiv.org/pdf/2308.13418.pdf>)

829. **ColPali: Efficient Document Retrieval with Vision Language Models** — Faysse et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2407.01449>) · [Direct PDF](<https://arxiv.org/pdf/2407.01449.pdf>)

830. **Marker: PDF to Markdown Conversion** — Datalab · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Marker%3A%20PDF%20to%20Markdown%20Conversion&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Marker%3A%20PDF%20to%20Markdown%20Conversion%20Datalab%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Marker%3A%20PDF%20to%20Markdown%20Conversion&sort=relevance>)

## 50. Autonomous Driving & Navigation

831. **End to End Learning for Self-Driving Cars** — Bojarski et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1604.07316>) · [Direct PDF](<https://arxiv.org/pdf/1604.07316.pdf>)

832. **PointNet: Deep Learning on Point Sets** — Qi et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1612.00593>) · [Direct PDF](<https://arxiv.org/pdf/1612.00593.pdf>)

833. **PointNet++: Deep Hierarchical Feature Learning on Point Sets** — Qi et al. · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=PointNet%2B%2B%3A%20Deep%20Hierarchical%20Feature%20Learning%20on%20Point%20Sets&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=PointNet%2B%2B%3A%20Deep%20Hierarchical%20Feature%20Learning%20on%20Point%20Sets%20Qi%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=PointNet%2B%2B%3A%20Deep%20Hierarchical%20Feature%20Learning%20on%20Point%20Sets&sort=relevance>)

834. **VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection** — Zhou & Tuzel · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=VoxelNet%3A%20End-to-End%20Learning%20for%20Point%20Cloud%20Based%203D%20Object%20Detection&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=VoxelNet%3A%20End-to-End%20Learning%20for%20Point%20Cloud%20Based%203D%20Object%20Detection%20Zhou%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=VoxelNet%3A%20End-to-End%20Learning%20for%20Point%20Cloud%20Based%203D%20Object%20Detection&sort=relevance>)

835. **BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images** — Li et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2203.17270>) · [Direct PDF](<https://arxiv.org/pdf/2203.17270.pdf>)

836. **UniAD: Planning-Oriented Autonomous Driving** — Hu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=UniAD%3A%20Planning-Oriented%20Autonomous%20Driving&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=UniAD%3A%20Planning-Oriented%20Autonomous%20Driving%20Hu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=UniAD%3A%20Planning-Oriented%20Autonomous%20Driving&sort=relevance>)

837. **LLM-based Driving Agents** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLM-based%20Driving%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLM-based%20Driving%20Agents%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLM-based%20Driving%20Agents&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

838. **DriveLM: Driving with Graph Visual Question Answering** — Sima et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DriveLM%3A%20Driving%20with%20Graph%20Visual%20Question%20Answering&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DriveLM%3A%20Driving%20with%20Graph%20Visual%20Question%20Answering%20Sima%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DriveLM%3A%20Driving%20with%20Graph%20Visual%20Question%20Answering&sort=relevance>)

839. **DriveGPT4: Interpretable End-to-end Autonomous Driving** — Xu et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DriveGPT4%3A%20Interpretable%20End-to-end%20Autonomous%20Driving&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DriveGPT4%3A%20Interpretable%20End-to-end%20Autonomous%20Driving%20Xu%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DriveGPT4%3A%20Interpretable%20End-to-end%20Autonomous%20Driving&sort=relevance>)

840. **Waymo Open Dataset** — Sun et al. · 2020 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=Waymo%20Open%20Dataset&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Waymo%20Open%20Dataset%20Sun%202020>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Waymo%20Open%20Dataset&sort=relevance>)

## 51. AI Infrastructure & MLOps

841. **TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems** — Abadi et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=TensorFlow%3A%20Large-Scale%20Machine%20Learning%20on%20Heterogeneous%20Systems&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=TensorFlow%3A%20Large-Scale%20Machine%20Learning%20on%20Heterogeneous%20Systems%20Abadi%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=TensorFlow%3A%20Large-Scale%20Machine%20Learning%20on%20Heterogeneous%20Systems&sort=relevance>)

842. **PyTorch: An Imperative Style, High-Performance Deep Learning Library** — Paszke et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=PyTorch%3A%20An%20Imperative%20Style%2C%20High-Performance%20Deep%20Learning%20Library&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=PyTorch%3A%20An%20Imperative%20Style%2C%20High-Performance%20Deep%20Learning%20Library%20Paszke%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=PyTorch%3A%20An%20Imperative%20Style%2C%20High-Performance%20Deep%20Learning%20Library&sort=relevance>)

843. **JAX: Composable Transformations of Python+NumPy** — Bradbury et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=JAX%3A%20Composable%20Transformations%20of%20Python%2BNumPy&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=JAX%3A%20Composable%20Transformations%20of%20Python%2BNumPy%20Bradbury%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=JAX%3A%20Composable%20Transformations%20of%20Python%2BNumPy&sort=relevance>)

844. **MLflow: A Platform for the ML Lifecycle** — Zaharia et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=MLflow%3A%20A%20Platform%20for%20the%20ML%20Lifecycle&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=MLflow%3A%20A%20Platform%20for%20the%20ML%20Lifecycle%20Zaharia%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=MLflow%3A%20A%20Platform%20for%20the%20ML%20Lifecycle&sort=relevance>)

845. **Hidden Technical Debt in ML Systems** — Sculley et al. · 2015 · **Type:** Research-paper candidate. [Supplied direct source](<https://papers.nips.cc/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html>)

846. **Ray: A Distributed Framework for Emerging AI Applications** — Moritz et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Ray%3A%20A%20Distributed%20Framework%20for%20Emerging%20AI%20Applications&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Ray%3A%20A%20Distributed%20Framework%20for%20Emerging%20AI%20Applications%20Moritz%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Ray%3A%20A%20Distributed%20Framework%20for%20Emerging%20AI%20Applications&sort=relevance>)

847. **Kubernetes for ML Workloads (Kubeflow)** — Various · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Kubernetes%20for%20ML%20Workloads%20%28Kubeflow%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Kubernetes%20for%20ML%20Workloads%20%28Kubeflow%29%20Various%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Kubernetes%20for%20ML%20Workloads%20%28Kubeflow%29&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

848. **Feature Stores for ML** — Baylor et al. · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Feature%20Stores%20for%20ML&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Feature%20Stores%20for%20ML%20Baylor%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Feature%20Stores%20for%20ML&sort=relevance>)

849. **Continuous Delivery for ML (CD4ML)** — Sato et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Continuous%20Delivery%20for%20ML%20%28CD4ML%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Continuous%20Delivery%20for%20ML%20%28CD4ML%29%20Sato%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Continuous%20Delivery%20for%20ML%20%28CD4ML%29&sort=relevance>)

850. **NVIDIA GPU Architecture: From Pascal to Hopper** — NVIDIA · 2016-2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=NVIDIA%20GPU%20Architecture%3A%20From%20Pascal%20to%20Hopper&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=NVIDIA%20GPU%20Architecture%3A%20From%20Pascal%20to%20Hopper%20NVIDIA%202016-2024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=NVIDIA%20GPU%20Architecture%3A%20From%20Pascal%20to%20Hopper&sort=relevance>)
   - **Audit:** NON-PAPER OR MULTI-WORK RECORD — A multi-generation document family rather than one work.

## 52. Agentic Web & Computer Use

851. **WebArena: A Realistic Web Environment for Building Autonomous Agents** — Zhou et al. · 2024 · **Type:** Benchmark or dataset paper. [Canonical arXiv record](<https://arxiv.org/abs/2307.13854>) · [Direct PDF](<https://arxiv.org/pdf/2307.13854.pdf>)

852. **Mind2Web: Towards a Generalist Agent for the Web** — Deng et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Mind2Web%3A%20Towards%20a%20Generalist%20Agent%20for%20the%20Web&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Mind2Web%3A%20Towards%20a%20Generalist%20Agent%20for%20the%20Web%20Deng%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Mind2Web%3A%20Towards%20a%20Generalist%20Agent%20for%20the%20Web&sort=relevance>)

853. **OS-World: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments** — Xie et al. · 2024 · **Type:** Benchmark or dataset paper. [arXiv search](<https://arxiv.org/search/?query=OS-World%3A%20Benchmarking%20Multimodal%20Agents%20for%20Open-Ended%20Tasks%20in%20Real%20Computer%20Environments&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=OS-World%3A%20Benchmarking%20Multimodal%20Agents%20for%20Open-Ended%20Tasks%20in%20Real%20Computer%20Environments%20Xie%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=OS-World%3A%20Benchmarking%20Multimodal%20Agents%20for%20Open-Ended%20Tasks%20in%20Real%20Computer%20Environments&sort=relevance>)

854. **Computer Use (Anthropic Claude)** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Computer%20Use%20%28Anthropic%20Claude%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Computer%20Use%20%28Anthropic%20Claude%29%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Computer%20Use%20%28Anthropic%20Claude%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #913.

855. **CogAgent: A Visual Language Model for GUI Agents** — Hong et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2312.08914>) · [Direct PDF](<https://arxiv.org/pdf/2312.08914.pdf>)

856. **AppAgent: Multimodal Agents as Smartphone Users** — Zhang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AppAgent%3A%20Multimodal%20Agents%20as%20Smartphone%20Users&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AppAgent%3A%20Multimodal%20Agents%20as%20Smartphone%20Users%20Zhang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AppAgent%3A%20Multimodal%20Agents%20as%20Smartphone%20Users&sort=relevance>)

857. **ScreenAgent: A Vision Language Model-Driven Computer Control Agent** — Niu et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=ScreenAgent%3A%20A%20Vision%20Language%20Model-Driven%20Computer%20Control%20Agent&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=ScreenAgent%3A%20A%20Vision%20Language%20Model-Driven%20Computer%20Control%20Agent%20Niu%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=ScreenAgent%3A%20A%20Vision%20Language%20Model-Driven%20Computer%20Control%20Agent&sort=relevance>)

858. **UFO: A UI-Focused Agent for Windows OS Interaction** — Zhang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=UFO%3A%20A%20UI-Focused%20Agent%20for%20Windows%20OS%20Interaction&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=UFO%3A%20A%20UI-Focused%20Agent%20for%20Windows%20OS%20Interaction%20Zhang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=UFO%3A%20A%20UI-Focused%20Agent%20for%20Windows%20OS%20Interaction&sort=relevance>)

859. **WebVoyager: Building an End-to-End Web Agent** — He et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=WebVoyager%3A%20Building%20an%20End-to-End%20Web%20Agent&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=WebVoyager%3A%20Building%20an%20End-to-End%20Web%20Agent%20He%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=WebVoyager%3A%20Building%20an%20End-to-End%20Web%20Agent&sort=relevance>)

860. **Anthropic MCP for Agents** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Anthropic%20MCP%20for%20Agents&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Anthropic%20MCP%20for%20Agents%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Anthropic%20MCP%20for%20Agents&sort=relevance>)

## 53. 3D Generation & Neural Rendering

861. **NeRF: Representing Scenes as Neural Radiance Fields** — Mildenhall et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2003.08934>) · [Direct PDF](<https://arxiv.org/pdf/2003.08934.pdf>)

862. **Instant NGP** — Müller et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2201.05989>) · [Direct PDF](<https://arxiv.org/pdf/2201.05989.pdf>)

863. **3D Gaussian Splatting for Real-Time Radiance Field Rendering** — Kerbl et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2308.04079>) · [Direct PDF](<https://arxiv.org/pdf/2308.04079.pdf>)
   - **Audit:** WRONG SUPPLIED MAPPING — The supplied link named a different work and is corrected in this edition. Original target: “Flexible Techniques for Differentiable Rendering with 3D Gaussians”.

864. **DreamFusion: Text-to-3D using 2D Diffusion** — Poole et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2209.14988>) · [Direct PDF](<https://arxiv.org/pdf/2209.14988.pdf>)

865. **Magic3D: High-Resolution Text-to-3D Content Creation** — Lin et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Magic3D%3A%20High-Resolution%20Text-to-3D%20Content%20Creation&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Magic3D%3A%20High-Resolution%20Text-to-3D%20Content%20Creation%20Lin%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Magic3D%3A%20High-Resolution%20Text-to-3D%20Content%20Creation&sort=relevance>)

866. **Zero-1-to-3: Zero-shot One Image to 3D Object** — Liu et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Zero-1-to-3%3A%20Zero-shot%20One%20Image%20to%203D%20Object&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Zero-1-to-3%3A%20Zero-shot%20One%20Image%20to%203D%20Object%20Liu%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Zero-1-to-3%3A%20Zero-shot%20One%20Image%20to%203D%20Object&sort=relevance>)

867. **Point-E: A System for Generating 3D Point Clouds from Complex Prompts** — Nichol et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Point-E%3A%20A%20System%20for%20Generating%203D%20Point%20Clouds%20from%20Complex%20Prompts&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Point-E%3A%20A%20System%20for%20Generating%203D%20Point%20Clouds%20from%20Complex%20Prompts%20Nichol%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Point-E%3A%20A%20System%20for%20Generating%203D%20Point%20Clouds%20from%20Complex%20Prompts&sort=relevance>)

868. **Shap-E: Generating Conditional 3D Implicit Functions** — Jun & Nichol · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Shap-E%3A%20Generating%20Conditional%203D%20Implicit%20Functions&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Shap-E%3A%20Generating%20Conditional%203D%20Implicit%20Functions%20Jun%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Shap-E%3A%20Generating%20Conditional%203D%20Implicit%20Functions&sort=relevance>)

869. **Meshy / TripoSR / LRM** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Meshy%20/%20TripoSR%20/%20LRM&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Meshy%20/%20TripoSR%20/%20LRM%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Meshy%20/%20TripoSR%20/%20LRM&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

870. **4D Generation: Recent Progress and Opportunities** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=4D%20Generation%3A%20Recent%20Progress%20and%20Opportunities&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=4D%20Generation%3A%20Recent%20Progress%20and%20Opportunities%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=4D%20Generation%3A%20Recent%20Progress%20and%20Opportunities&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 54. Adversarial ML & Robustness

871. **Intriguing Properties of Neural Networks** — Szegedy et al. · 2014 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1312.6199>) · [Direct PDF](<https://arxiv.org/pdf/1312.6199.pdf>)

872. **Explaining and Harnessing Adversarial Examples (FGSM)** — Goodfellow, Shlens, Szegedy · 2015 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1412.6572>) · [Direct PDF](<https://arxiv.org/pdf/1412.6572.pdf>)

873. **Towards Deep Learning Models Resistant to Adversarial Attacks (PGD)** — Madry et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1706.06083>) · [Direct PDF](<https://arxiv.org/pdf/1706.06083.pdf>)

874. **Certified Adversarial Robustness via Randomized Smoothing** — Cohen, Rosenfeld, Kolter · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Certified%20Adversarial%20Robustness%20via%20Randomized%20Smoothing&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Certified%20Adversarial%20Robustness%20via%20Randomized%20Smoothing%20Cohen%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Certified%20Adversarial%20Robustness%20via%20Randomized%20Smoothing&sort=relevance>)

875. **DeepFool: A Simple and Accurate Method to Fool Deep Neural Networks** — Moosavi-Dezfooli et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DeepFool%3A%20A%20Simple%20and%20Accurate%20Method%20to%20Fool%20Deep%20Neural%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DeepFool%3A%20A%20Simple%20and%20Accurate%20Method%20to%20Fool%20Deep%20Neural%20Networks%20Moosavi-Dezfooli%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DeepFool%3A%20A%20Simple%20and%20Accurate%20Method%20to%20Fool%20Deep%20Neural%20Networks&sort=relevance>)

876. **Adversarial Examples in the Physical World** — Kurakin et al. · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Adversarial%20Examples%20in%20the%20Physical%20World&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Adversarial%20Examples%20in%20the%20Physical%20World%20Kurakin%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Adversarial%20Examples%20in%20the%20Physical%20World&sort=relevance>)

877. **Robust and Accurate Object Detection via Adversarial Learning** — Chen et al. · 2021 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Robust%20and%20Accurate%20Object%20Detection%20via%20Adversarial%20Learning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Robust%20and%20Accurate%20Object%20Detection%20via%20Adversarial%20Learning%20Chen%202021>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Robust%20and%20Accurate%20Object%20Detection%20via%20Adversarial%20Learning&sort=relevance>)

878. **Adversarial Attacks on LLMs** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Adversarial%20Attacks%20on%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Adversarial%20Attacks%20on%20LLMs%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Adversarial%20Attacks%20on%20LLMs&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

879. **Prompt Injection Attacks** — Various · 2023-2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Prompt%20Injection%20Attacks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Prompt%20Injection%20Attacks%20Various%202023-2024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Prompt%20Injection%20Attacks&sort=relevance>)
   - **Audit:** NON-PAPER OR MULTI-WORK RECORD — A research topic rather than a uniquely identifiable publication.

880. **Backdoor Attacks on LLMs** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Backdoor%20Attacks%20on%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Backdoor%20Attacks%20on%20LLMs%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Backdoor%20Attacks%20on%20LLMs&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 55. Meta-Learning & Few-Shot Learning

881. **Model-Agnostic Meta-Learning (MAML)** — Finn, Abbeel, Levine · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1703.03400>) · [Direct PDF](<https://arxiv.org/pdf/1703.03400.pdf>)

882. **Matching Networks for One Shot Learning** — Vinyals et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1606.04080>) · [Direct PDF](<https://arxiv.org/pdf/1606.04080.pdf>)

883. **Prototypical Networks for Few-Shot Learning** — Snell, Swersky, Zemel · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1703.05175>) · [Direct PDF](<https://arxiv.org/pdf/1703.05175.pdf>)

884. **Learning to Learn with Compound HD Models** — Santoro et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Learning%20to%20Learn%20with%20Compound%20HD%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Learning%20to%20Learn%20with%20Compound%20HD%20Models%20Santoro%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Learning%20to%20Learn%20with%20Compound%20HD%20Models&sort=relevance>)

885. **Meta-Learning: A Survey** — Hospedales et al. · 2022 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Meta-Learning%3A%20A%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Meta-Learning%3A%20A%20Survey%20Hospedales%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Meta-Learning%3A%20A%20Survey&sort=relevance>)

886. **Reptile: A Scalable Metalearning Algorithm** — Nichol, Achiam, Schulman · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Reptile%3A%20A%20Scalable%20Metalearning%20Algorithm&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Reptile%3A%20A%20Scalable%20Metalearning%20Algorithm%20Nichol%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Reptile%3A%20A%20Scalable%20Metalearning%20Algorithm&sort=relevance>)

887. **LEO: Latent Embedding Optimization** — Rusu et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LEO%3A%20Latent%20Embedding%20Optimization&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LEO%3A%20Latent%20Embedding%20Optimization%20Rusu%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LEO%3A%20Latent%20Embedding%20Optimization&sort=relevance>)

888. **How to Train Your MAML** — Antoniou et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=How%20to%20Train%20Your%20MAML&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=How%20to%20Train%20Your%20MAML%20Antoniou%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=How%20to%20Train%20Your%20MAML&sort=relevance>)

889. **Few-Shot Learning via Saliency-Guided Hallucination** — Zhang et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Few-Shot%20Learning%20via%20Saliency-Guided%20Hallucination&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Few-Shot%20Learning%20via%20Saliency-Guided%20Hallucination%20Zhang%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Few-Shot%20Learning%20via%20Saliency-Guided%20Hallucination&sort=relevance>)

890. **P>M>F: Pre-training, Meta-training, Fine-tuning** — Hu et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=P%3EM%3EF%3A%20Pre-training%2C%20Meta-training%2C%20Fine-tuning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=P%3EM%3EF%3A%20Pre-training%2C%20Meta-training%2C%20Fine-tuning%20Hu%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=P%3EM%3EF%3A%20Pre-training%2C%20Meta-training%2C%20Fine-tuning&sort=relevance>)

## 56. Neuro-Symbolic AI

891. **Neural-Symbolic Computing: An Effective Methodology for Principled Integration** — Garcez et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Neural-Symbolic%20Computing%3A%20An%20Effective%20Methodology%20for%20Principled%20Integration&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Neural-Symbolic%20Computing%3A%20An%20Effective%20Methodology%20for%20Principled%20Integration%20Garcez%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Neural-Symbolic%20Computing%3A%20An%20Effective%20Methodology%20for%20Principled%20Integration&sort=relevance>)

892. **Neural Module Networks** — Andreas et al. · 2016 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Neural%20Module%20Networks&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Neural%20Module%20Networks%20Andreas%202016>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Neural%20Module%20Networks&sort=relevance>)

893. **End-to-End Differentiable Proving (NTP)** — Rocktäschel & Riedel · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=End-to-End%20Differentiable%20Proving%20%28NTP%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=End-to-End%20Differentiable%20Proving%20%28NTP%29%20Rockt%C3%A4schel%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=End-to-End%20Differentiable%20Proving%20%28NTP%29&sort=relevance>)

894. **DeepProbLog: Neural Probabilistic Logic Programming** — Manhaeve et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DeepProbLog%3A%20Neural%20Probabilistic%20Logic%20Programming&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DeepProbLog%3A%20Neural%20Probabilistic%20Logic%20Programming%20Manhaeve%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DeepProbLog%3A%20Neural%20Probabilistic%20Logic%20Programming&sort=relevance>)

895. **Scallop: A Language for Neurosymbolic Programming** — Li et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Scallop%3A%20A%20Language%20for%20Neurosymbolic%20Programming&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scallop%3A%20A%20Language%20for%20Neurosymbolic%20Programming%20Li%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scallop%3A%20A%20Language%20for%20Neurosymbolic%20Programming&sort=relevance>)

896. **LLMs and Symbolic Reasoning** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLMs%20and%20Symbolic%20Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLMs%20and%20Symbolic%20Reasoning%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLMs%20and%20Symbolic%20Reasoning&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

897. **Program Synthesis with LLMs** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Program%20Synthesis%20with%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Program%20Synthesis%20with%20LLMs%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Program%20Synthesis%20with%20LLMs&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

898. **Logic-LM: Empowering LLMs with Symbolic Solvers** — Pan et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Logic-LM%3A%20Empowering%20LLMs%20with%20Symbolic%20Solvers&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Logic-LM%3A%20Empowering%20LLMs%20with%20Symbolic%20Solvers%20Pan%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Logic-LM%3A%20Empowering%20LLMs%20with%20Symbolic%20Solvers&sort=relevance>)

899. **Code-as-Reasoning** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Code-as-Reasoning&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Code-as-Reasoning%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Code-as-Reasoning&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

900. **Binding Language Models in Symbolic Languages** — Cheng et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Binding%20Language%20Models%20in%20Symbolic%20Languages&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Binding%20Language%20Models%20in%20Symbolic%20Languages%20Cheng%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Binding%20Language%20Models%20in%20Symbolic%20Languages&sort=relevance>)

## 57. AI Governance & Policy

901. **EU AI Act (technical analysis)** — EU · 2024 · **Type:** Policy or official document. [arXiv search](<https://arxiv.org/search/?query=EU%20AI%20Act%20%28technical%20analysis%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=EU%20AI%20Act%20%28technical%20analysis%29%20EU%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=EU%20AI%20Act%20%28technical%20analysis%29&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as policy or official document, not as a research paper.

902. **Executive Order on AI Safety (US)** — White House · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Executive%20Order%20on%20AI%20Safety%20%28US%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Executive%20Order%20on%20AI%20Safety%20%28US%29%20White%20House%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Executive%20Order%20on%20AI%20Safety%20%28US%29&sort=relevance>)

903. **Governing AI: A Blueprint for the Future** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Governing%20AI%3A%20A%20Blueprint%20for%20the%20Future&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Governing%20AI%3A%20A%20Blueprint%20for%20the%20Future%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Governing%20AI%3A%20A%20Blueprint%20for%20the%20Future&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

904. **International AI Safety Report** — AI Safety Summit · 2024 · **Type:** Technical or industry report. [arXiv search](<https://arxiv.org/search/?query=International%20AI%20Safety%20Report&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=International%20AI%20Safety%20Report%20AI%20Safety%20Summit%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=International%20AI%20Safety%20Report&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

905. **Foundation Model Transparency Index** — Bommasani et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Foundation%20Model%20Transparency%20Index&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Foundation%20Model%20Transparency%20Index%20Bommasani%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Foundation%20Model%20Transparency%20Index&sort=relevance>)

## 58. Emerging Frontiers (2024–2025)

906. **Mixture of Agents (MoA)** — Wang et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Mixture%20of%20Agents%20%28MoA%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Mixture%20of%20Agents%20%28MoA%29%20Wang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Mixture%20of%20Agents%20%28MoA%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #364.

907. **Inference Scaling Laws** — Sardana & Frankle · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Inference%20Scaling%20Laws&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Inference%20Scaling%20Laws%20Sardana%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Inference%20Scaling%20Laws&sort=relevance>)

908. **LLM Operating Systems** — Packer et al.; Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=LLM%20Operating%20Systems&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLM%20Operating%20Systems%20Packer%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLM%20Operating%20Systems&sort=relevance>)

909. **SWE-bench Verified** — OpenAI · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SWE-bench%20Verified&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SWE-bench%20Verified%20OpenAI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SWE-bench%20Verified&sort=relevance>)

910. **Anthropic's Responsible Scaling Policy (RSP) Update** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Anthropic%27s%20Responsible%20Scaling%20Policy%20%28RSP%29%20Update&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Anthropic%27s%20Responsible%20Scaling%20Policy%20%28RSP%29%20Update%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Anthropic%27s%20Responsible%20Scaling%20Policy%20%28RSP%29%20Update&sort=relevance>)

911. **Q* / Process Reward Models** — Various (speculated/inferred) · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Q%2A%20/%20Process%20Reward%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Q%2A%20/%20Process%20Reward%20Models%20Various%20%28speculated/inferred%29%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Q%2A%20/%20Process%20Reward%20Models&sort=relevance>)
   - **Audit:** UNVERIFIED — The record itself attributes the work to 'Various (speculated/inferred)' and does not identify a publication. Not verified/citable until repaired.

912. **OpenAI o1 and o1-pro** — OpenAI · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=OpenAI%20o1%20and%20o1-pro&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=OpenAI%20o1%20and%20o1-pro%20OpenAI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=OpenAI%20o1%20and%20o1-pro&sort=relevance>)

913. **Claude 3.5 Sonnet and Computer Use** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Claude%203.5%20Sonnet%20and%20Computer%20Use&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Claude%203.5%20Sonnet%20and%20Computer%20Use%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Claude%203.5%20Sonnet%20and%20Computer%20Use&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #854.

914. **Gemini 2.0 Flash** — Google DeepMind · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Gemini%202.0%20Flash&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Gemini%202.0%20Flash%20Google%20DeepMind%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Gemini%202.0%20Flash&sort=relevance>)

915. **DeepSeek-R1** — DeepSeek · 2025 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=DeepSeek-R1&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=DeepSeek-R1%20DeepSeek%202025>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=DeepSeek-R1&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #279.

916. **Grok-2** — xAI · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Grok-2&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Grok-2%20xAI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Grok-2&sort=relevance>)

917. **Nemotron-4** — NVIDIA · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Nemotron-4&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Nemotron-4%20NVIDIA%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Nemotron-4&sort=relevance>)

918. **Apple Intelligence (technical overview)** — Apple · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Apple%20Intelligence%20%28technical%20overview%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Apple%20Intelligence%20%28technical%20overview%29%20Apple%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Apple%20Intelligence%20%28technical%20overview%29&sort=relevance>)

919. **Llama 3.1 405B** — Meta · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Llama%203.1%20405B&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Llama%203.1%20405B%20Meta%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Llama%203.1%20405B&sort=relevance>)

920. **Scaling Test-Time Compute (Survey)** — Various · 2025 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Scaling%20Test-Time%20Compute%20%28Survey%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Scaling%20Test-Time%20Compute%20%28Survey%29%20Various%202025>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Scaling%20Test-Time%20Compute%20%28Survey%29&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

## 59. Tokenization & Data Processing

921. **Byte Pair Encoding (BPE) for NMT** — Sennrich et al. · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1508.07909>) · [Direct PDF](<https://arxiv.org/pdf/1508.07909.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #773.

922. **SentencePiece: Unsupervised Text Tokenizer** — Kudo & Richardson · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=SentencePiece%3A%20Unsupervised%20Text%20Tokenizer&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=SentencePiece%3A%20Unsupervised%20Text%20Tokenizer%20Kudo%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=SentencePiece%3A%20Unsupervised%20Text%20Tokenizer&sort=relevance>)

923. **Tokenizer Choice Matters** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Tokenizer%20Choice%20Matters&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Tokenizer%20Choice%20Matters%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Tokenizer%20Choice%20Matters&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

924. **MegaByte: Predicting Million-Byte Sequences** — Yu et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2305.07185>) · [Direct PDF](<https://arxiv.org/pdf/2305.07185.pdf>)

925. **BLT: Byte Latent Transformer** — Meta · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=BLT%3A%20Byte%20Latent%20Transformer&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=BLT%3A%20Byte%20Latent%20Transformer%20Meta%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=BLT%3A%20Byte%20Latent%20Transformer&sort=relevance>)

926. **The Tokenizer Landscape for LLMs** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Tokenizer%20Landscape%20for%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Tokenizer%20Landscape%20for%20LLMs%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Tokenizer%20Landscape%20for%20LLMs&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

927. **Tiktoken (OpenAI)** — OpenAI · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Tiktoken%20%28OpenAI%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Tiktoken%20%28OpenAI%29%20OpenAI%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Tiktoken%20%28OpenAI%29&sort=relevance>)

928. **Data Deduplication for LLM Training** — Lee et al. · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Data%20Deduplication%20for%20LLM%20Training&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Data%20Deduplication%20for%20LLM%20Training%20Lee%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Data%20Deduplication%20for%20LLM%20Training&sort=relevance>)

929. **Quality Filtering for LLM Training Data** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Quality%20Filtering%20for%20LLM%20Training%20Data&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Quality%20Filtering%20for%20LLM%20Training%20Data%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Quality%20Filtering%20for%20LLM%20Training%20Data&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

930. **Dolma: An Open Corpus of Trillion Tokens** — Soldaini et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Dolma%3A%20An%20Open%20Corpus%20of%20Trillion%20Tokens&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Dolma%3A%20An%20Open%20Corpus%20of%20Trillion%20Tokens%20Soldaini%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Dolma%3A%20An%20Open%20Corpus%20of%20Trillion%20Tokens&sort=relevance>)

## 60. Surveys, Tutorials & Meta-Papers

931. **A Survey of Large Language Models** — Zhao et al. · 2023 · **Type:** Survey or review. [Canonical arXiv record](<https://arxiv.org/abs/2303.18223>) · [Direct PDF](<https://arxiv.org/pdf/2303.18223.pdf>)

932. **Harnessing the Power of LLMs in Practice** — Yang et al. · 2023 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Harnessing%20the%20Power%20of%20LLMs%20in%20Practice&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Harnessing%20the%20Power%20of%20LLMs%20in%20Practice%20Yang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Harnessing%20the%20Power%20of%20LLMs%20in%20Practice&sort=relevance>)

933. **A Survey on Hallucination in LLMs** — Huang et al. · 2023 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=A%20Survey%20on%20Hallucination%20in%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Survey%20on%20Hallucination%20in%20LLMs%20Huang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Survey%20on%20Hallucination%20in%20LLMs&sort=relevance>)

934. **Siren's Song in the AI Ocean: A Survey on Hallucination** — Zhang et al. · 2023 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Siren%27s%20Song%20in%20the%20AI%20Ocean%3A%20A%20Survey%20on%20Hallucination&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Siren%27s%20Song%20in%20the%20AI%20Ocean%3A%20A%20Survey%20on%20Hallucination%20Zhang%202023>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Siren%27s%20Song%20in%20the%20AI%20Ocean%3A%20A%20Survey%20on%20Hallucination&sort=relevance>)

935. **A Survey on Multimodal Large Language Models** — Yin et al. · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=A%20Survey%20on%20Multimodal%20Large%20Language%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Survey%20on%20Multimodal%20Large%20Language%20Models%20Yin%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Survey%20on%20Multimodal%20Large%20Language%20Models&sort=relevance>)

936. **A Survey on Evaluation of LLMs** — Chang et al. · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=A%20Survey%20on%20Evaluation%20of%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Survey%20on%20Evaluation%20of%20LLMs%20Chang%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Survey%20on%20Evaluation%20of%20LLMs&sort=relevance>)

937. **A Comprehensive Survey on Vector Database** — Han et al. · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=A%20Comprehensive%20Survey%20on%20Vector%20Database&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Comprehensive%20Survey%20on%20Vector%20Database%20Han%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Comprehensive%20Survey%20on%20Vector%20Database&sort=relevance>)

938. **Retrieval-Augmented Generation for AI-Generated Content: A Survey** — Zhao et al. · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Retrieval-Augmented%20Generation%20for%20AI-Generated%20Content%3A%20A%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Retrieval-Augmented%20Generation%20for%20AI-Generated%20Content%3A%20A%20Survey%20Zhao%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Retrieval-Augmented%20Generation%20for%20AI-Generated%20Content%3A%20A%20Survey&sort=relevance>)

939. **Tool Learning with Foundation Models** — Qin et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Tool%20Learning%20with%20Foundation%20Models&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Tool%20Learning%20with%20Foundation%20Models%20Qin%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Tool%20Learning%20with%20Foundation%20Models&sort=relevance>)

940. **LLM Agents: A Survey of Applications** — Various · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=LLM%20Agents%3A%20A%20Survey%20of%20Applications&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=LLM%20Agents%3A%20A%20Survey%20of%20Applications%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=LLM%20Agents%3A%20A%20Survey%20of%20Applications&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

941. **The Landscape of Emerging AI Agent Architectures** — Masterman et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Landscape%20of%20Emerging%20AI%20Agent%20Architectures&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Landscape%20of%20Emerging%20AI%20Agent%20Architectures%20Masterman%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Landscape%20of%20Emerging%20AI%20Agent%20Architectures&sort=relevance>)

942. **A Survey on Self-Evolution of LLMs** — Tao et al. · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=A%20Survey%20on%20Self-Evolution%20of%20LLMs&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=A%20Survey%20on%20Self-Evolution%20of%20LLMs%20Tao%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=A%20Survey%20on%20Self-Evolution%20of%20LLMs&sort=relevance>)

943. **From RAG to Rich: Retrieval-Augmented Generation Survey** — Various · 2024 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=From%20RAG%20to%20Rich%3A%20Retrieval-Augmented%20Generation%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=From%20RAG%20to%20Rich%3A%20Retrieval-Augmented%20Generation%20Survey%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=From%20RAG%20to%20Rich%3A%20Retrieval-Augmented%20Generation%20Survey&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

944. **AI Agents That Matter** — Kapoor et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=AI%20Agents%20That%20Matter&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=AI%20Agents%20That%20Matter%20Kapoor%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=AI%20Agents%20That%20Matter&sort=relevance>)

945. **Position: What Can LLMs Do for ME?** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Position%3A%20What%20Can%20LLMs%20Do%20for%20ME%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Position%3A%20What%20Can%20LLMs%20Do%20for%20ME%3F%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Position%3A%20What%20Can%20LLMs%20Do%20for%20ME%3F&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

946. **The State of AI Report 2024** — Benaich & Hogarth · 2024 · **Type:** Technical or industry report. [arXiv search](<https://arxiv.org/search/?query=The%20State%20of%20AI%20Report%202024&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20State%20of%20AI%20Report%202024%20Benaich%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20State%20of%20AI%20Report%202024&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

947. **Stanford AI Index Report 2024** — Stanford · 2024 · **Type:** Technical or industry report. [arXiv search](<https://arxiv.org/search/?query=Stanford%20AI%20Index%20Report%202024&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Stanford%20AI%20Index%20Report%202024%20Stanford%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Stanford%20AI%20Index%20Report%202024&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as technical or industry report, not as a research paper.

948. **Attention Mechanisms in Computer Vision: A Survey** — Guo et al. · 2022 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Attention%20Mechanisms%20in%20Computer%20Vision%3A%20A%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Attention%20Mechanisms%20in%20Computer%20Vision%3A%20A%20Survey%20Guo%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Attention%20Mechanisms%20in%20Computer%20Vision%3A%20A%20Survey&sort=relevance>)

949. **Vision Transformers: A Survey** — Khan et al. · 2022 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Vision%20Transformers%3A%20A%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Vision%20Transformers%3A%20A%20Survey%20Khan%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Vision%20Transformers%3A%20A%20Survey&sort=relevance>)

950. **Efficient Transformers: A Survey** — Tay et al. · 2022 · **Type:** Survey or review. [arXiv search](<https://arxiv.org/search/?query=Efficient%20Transformers%3A%20A%20Survey&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Efficient%20Transformers%3A%20A%20Survey%20Tay%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Efficient%20Transformers%3A%20A%20Survey&sort=relevance>)

## 61. Additional Foundational & Influential Papers

951. **The Unreasonable Effectiveness of Data** — Halevy, Norvig, Pereira · 2009 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Unreasonable%20Effectiveness%20of%20Data&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Unreasonable%20Effectiveness%20of%20Data%20Halevy%202009>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Unreasonable%20Effectiveness%20of%20Data&sort=relevance>)

952. **Batch Normalization (revisited for theory)** — Santurkar et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Batch%20Normalization%20%28revisited%20for%20theory%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Batch%20Normalization%20%28revisited%20for%20theory%29%20Santurkar%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Batch%20Normalization%20%28revisited%20for%20theory%29&sort=relevance>)

953. **Mixup: Beyond Empirical Risk Minimization** — Zhang et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1710.09412>) · [Direct PDF](<https://arxiv.org/pdf/1710.09412.pdf>)

954. **CutMix: Regularization Strategy to Train Strong Classifiers** — Yun et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1905.04899>) · [Direct PDF](<https://arxiv.org/pdf/1905.04899.pdf>)

955. **RandAugment: Practical Automated Data Augmentation** — Cubuk et al. · 2020 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1909.13719>) · [Direct PDF](<https://arxiv.org/pdf/1909.13719.pdf>)

956. **AutoAugment: Learning Augmentation Strategies from Data** — Cubuk et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1805.09501>) · [Direct PDF](<https://arxiv.org/pdf/1805.09501.pdf>)

957. **Label Smoothing Revisited** — Müller et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Label%20Smoothing%20Revisited&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Label%20Smoothing%20Revisited%20M%C3%BCller%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Label%20Smoothing%20Revisited&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #958.

958. **When Does Label Smoothing Help?** — Müller et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=When%20Does%20Label%20Smoothing%20Help%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=When%20Does%20Label%20Smoothing%20Help%3F%20M%C3%BCller%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=When%20Does%20Label%20Smoothing%20Help%3F&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #957.

959. **Swish: A Self-Gated Activation Function** — Ramachandran et al. · 2017 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Swish%3A%20A%20Self-Gated%20Activation%20Function&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Swish%3A%20A%20Self-Gated%20Activation%20Function%20Ramachandran%202017>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Swish%3A%20A%20Self-Gated%20Activation%20Function&sort=relevance>)

960. **GELU: Gaussian Error Linear Units** — Hendrycks & Gimpel · 2016 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1606.08415>) · [Direct PDF](<https://arxiv.org/pdf/1606.08415.pdf>)

961. **Squeeze-and-Excitation Networks** — Hu et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1709.01507>) · [Direct PDF](<https://arxiv.org/pdf/1709.01507.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #85.

962. **CBAM: Convolutional Block Attention Module** — Woo et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1807.06521>) · [Direct PDF](<https://arxiv.org/pdf/1807.06521.pdf>)

963. **Non-local Neural Networks** — Wang et al. · 2018 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1711.07971>) · [Direct PDF](<https://arxiv.org/pdf/1711.07971.pdf>)

964. **Deformable Convolutional Networks** — Dai et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1703.06211>) · [Direct PDF](<https://arxiv.org/pdf/1703.06211.pdf>)

965. **Dynamic Routing Between Capsules** — Sabour, Frosst, Hinton · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1710.09829>) · [Direct PDF](<https://arxiv.org/pdf/1710.09829.pdf>)

966. **An Empirical Evaluation of Generic Convolutional and Recurrent Networks (TCN paper)** — Bai et al. · 2018 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=An%20Empirical%20Evaluation%20of%20Generic%20Convolutional%20and%20Recurrent%20Networks%20%28TCN%20paper%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=An%20Empirical%20Evaluation%20of%20Generic%20Convolutional%20and%20Recurrent%20Networks%20%28TCN%20paper%29%20Bai%202018>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=An%20Empirical%20Evaluation%20of%20Generic%20Convolutional%20and%20Recurrent%20Networks%20%28TCN%20paper%29&sort=relevance>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #116.

967. **Weight Standardization** — Qiao et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Weight%20Standardization&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Weight%20Standardization%20Qiao%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Weight%20Standardization&sort=relevance>)

968. **Fixup Initialization: Residual Learning Without Normalization** — Zhang et al. · 2019 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Fixup%20Initialization%3A%20Residual%20Learning%20Without%20Normalization&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Fixup%20Initialization%3A%20Residual%20Learning%20Without%20Normalization%20Zhang%202019>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Fixup%20Initialization%3A%20Residual%20Learning%20Without%20Normalization&sort=relevance>)

969. **Deep Sets** — Zaheer et al. · 2017 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1703.06114>) · [Direct PDF](<https://arxiv.org/pdf/1703.06114.pdf>)

970. **Attention Augmented Convolutional Networks** — Bello et al. · 2019 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/1904.09925>) · [Direct PDF](<https://arxiv.org/pdf/1904.09925.pdf>)

971. **Do Vision Transformers See Like Convolutional Neural Networks?** — Raghu et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2108.08810>) · [Direct PDF](<https://arxiv.org/pdf/2108.08810.pdf>)

972. **What Do Vision Transformers Learn?** — Park & Kim · 2022 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=What%20Do%20Vision%20Transformers%20Learn%3F&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=What%20Do%20Vision%20Transformers%20Learn%3F%20Park%202022>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=What%20Do%20Vision%20Transformers%20Learn%3F&sort=relevance>)

973. **Scaling Vision with Sparse MoE** — Riquelme et al. · 2021 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2106.05974>) · [Direct PDF](<https://arxiv.org/pdf/2106.05974.pdf>)

974. **Token Merging: Your ViT but Faster** — Bolya et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2210.09461>) · [Direct PDF](<https://arxiv.org/pdf/2210.09461.pdf>)

975. **FNet: Mixing Tokens with Fourier Transforms** — Lee-Thorp et al. · 2022 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2105.03824>) · [Direct PDF](<https://arxiv.org/pdf/2105.03824.pdf>)

## 62. Final Frontier Papers (976–1000)

976. **OpenAI o3 (analysis)** — OpenAI · 2024 · **Type:** Product or topic analysis. [arXiv search](<https://arxiv.org/search/?query=OpenAI%20o3%20%28analysis%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=OpenAI%20o3%20%28analysis%29%20OpenAI%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=OpenAI%20o3%20%28analysis%29&sort=relevance>)
   - **Audit:** NON-PAPER TYPE — Classified for discovery as product or topic analysis, not as a research paper; UNVERIFIED — A product-analysis label without a canonical publication title. Not verified/citable until repaired.

977. **Claude 4 / Opus 4 (model card)** — Anthropic · 2025 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Claude%204%20/%20Opus%204%20%28model%20card%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Claude%204%20/%20Opus%204%20%28model%20card%29%20Anthropic%202025>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Claude%204%20/%20Opus%204%20%28model%20card%29&sort=relevance>)

978. **Gemini 2.0 Pro** — Google DeepMind · 2025 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Gemini%202.0%20Pro&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Gemini%202.0%20Pro%20Google%20DeepMind%202025>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Gemini%202.0%20Pro&sort=relevance>)

979. **Llama 4** — Meta · 2025 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Llama%204&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Llama%204%20Meta%202025>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Llama%204&sort=relevance>)

980. **Multi-Token Prediction** — Gloeckle et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2404.19737>) · [Direct PDF](<https://arxiv.org/pdf/2404.19737.pdf>)

981. **Native Multi-Modality vs Late Fusion** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Native%20Multi-Modality%20vs%20Late%20Fusion&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Native%20Multi-Modality%20vs%20Late%20Fusion%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Native%20Multi-Modality%20vs%20Late%20Fusion&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved; UNVERIFIED — Exact-title search did not identify a matching publication; authors are listed as Various. Not verified/citable until repaired.

982. **RLVR: Reinforcement Learning with Verifiable Rewards** — Various · 2025 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=RLVR%3A%20Reinforcement%20Learning%20with%20Verifiable%20Rewards&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=RLVR%3A%20Reinforcement%20Learning%20with%20Verifiable%20Rewards%20Various%202025>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=RLVR%3A%20Reinforcement%20Learning%20with%20Verifiable%20Rewards&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved; UNVERIFIED — A method-family label with authors listed as Various rather than a unique work. Not verified/citable until repaired.

983. **Constitutional AI 2.0 Concepts** — Anthropic · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Constitutional%20AI%202.0%20Concepts&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Constitutional%20AI%202.0%20Concepts%20Anthropic%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Constitutional%20AI%202.0%20Concepts&sort=relevance>)
   - **Audit:** UNVERIFIED — Exact-title search did not identify an Anthropic publication with this title. Not verified/citable until repaired.

984. **Post-Training Scaling** — Various · 2025 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Post-Training%20Scaling&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Post-Training%20Scaling%20Various%202025>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Post-Training%20Scaling&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved; UNVERIFIED — A topic label with authors listed as Various rather than a unique work. Not verified/citable until repaired.

985. **Weak-to-Strong Generalization** — Burns et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2312.09390>) · [Direct PDF](<https://arxiv.org/pdf/2312.09390.pdf>)

986. **The Platonic Representation Hypothesis** — Huh et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2405.07987>) · [Direct PDF](<https://arxiv.org/pdf/2405.07987.pdf>)

987. **Representation Engineering** — Zou et al. · 2023 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2310.01405>) · [Direct PDF](<https://arxiv.org/pdf/2310.01405.pdf>)
   - **Audit:** DUPLICATE — The same or a materially identical work appears elsewhere in the curriculum. Related rows: #667.

988. **LLMs as Optimizers (OPRO)** — Yang et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2309.03409>) · [Direct PDF](<https://arxiv.org/pdf/2309.03409.pdf>)

989. **Can LLMs Generate Novel Research Ideas?** — Si et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2409.04109>) · [Direct PDF](<https://arxiv.org/pdf/2409.04109.pdf>)

990. **The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery** — Lu et al. · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2408.06292>) · [Direct PDF](<https://arxiv.org/pdf/2408.06292.pdf>)

991. **Absolute Zero: Reinforced Self-Play Reasoning with Zero Data** — Zhao et al. · 2025 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2505.03335>) · [Direct PDF](<https://arxiv.org/pdf/2505.03335.pdf>)

992. **Agentic Workflows and Compound AI Systems** — Zaharia et al. · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Agentic%20Workflows%20and%20Compound%20AI%20Systems&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Agentic%20Workflows%20and%20Compound%20AI%20Systems%20Zaharia%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Agentic%20Workflows%20and%20Compound%20AI%20Systems&sort=relevance>)

993. **From Models to Compound AI Systems** — Matei Zaharia · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=From%20Models%20to%20Compound%20AI%20Systems&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=From%20Models%20to%20Compound%20AI%20Systems%20Matei%20Zaharia%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=From%20Models%20to%20Compound%20AI%20Systems&sort=relevance>)

994. **LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions** — Chuanneng Sun, Songjun Huang, Dario Pompili · 2024 · **Type:** Research-paper candidate. [Canonical arXiv record](<https://arxiv.org/abs/2405.11106>) · [Direct PDF](<https://arxiv.org/pdf/2405.11106.pdf>)
   - **Audit:** VAGUE AUTHOR CORRECTED — The source used a vague attribution; named authorship is supplied by the audit; METADATA CORRECTED — Canonical title, authorship and arXiv identity were supplied by the audit.

995. **Self-Improving AI Systems** — Various · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Self-Improving%20AI%20Systems&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Self-Improving%20AI%20Systems%20Various%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Self-Improving%20AI%20Systems&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved; UNVERIFIED — A generic topic label with authors listed as Various. Not verified/citable until repaired.

996. **Frontier AI Safety Commitments (Seoul Declaration)** — Various govts · 2024 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=Frontier%20AI%20Safety%20Commitments%20%28Seoul%20Declaration%29&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=Frontier%20AI%20Safety%20Commitments%20%28Seoul%20Declaration%29%20Various%20govts%202024>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=Frontier%20AI%20Safety%20Commitments%20%28Seoul%20Declaration%29&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved.

997. **Situational Awareness** — Aschenbrenner · 2024 · **Type:** Research-paper candidate. [Supplied direct source](<https://situational-awareness.ai/>)

998. **What We Learned from a Year of Building with LLMs** — Yan et al. · 2024 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms/>)

999. **Building Effective Agents** — Anthropic · 2024 · **Type:** Research-paper candidate. [Supplied direct source](<https://www.anthropic.com/research/building-effective-agents>)

1000. **The Next Wave: AI Systems that Reason, Act, and Learn** — Various · 2025 · **Type:** Research-paper candidate. [arXiv search](<https://arxiv.org/search/?query=The%20Next%20Wave%3A%20AI%20Systems%20that%20Reason%2C%20Act%2C%20and%20Learn&searchtype=all>) · [Google Scholar](<https://scholar.google.com/scholar?q=The%20Next%20Wave%3A%20AI%20Systems%20that%20Reason%2C%20Act%2C%20and%20Learn%20Various%202025>) · [Semantic Scholar](<https://www.semanticscholar.org/search?q=The%20Next%20Wave%3A%20AI%20Systems%20that%20Reason%2C%20Act%2C%20and%20Learn&sort=relevance>)
   - **Audit:** VAGUE AUTHOR — Named authorship is unresolved; UNVERIFIED — Exact-title search did not identify a matching publication; authors are listed as Various. Not verified/citable until repaired.
