AI Learning Hub › Landmark AI papers & reports
Curated reading list
Landmark AI Papers & Reports
A curated guide to foundational research, current frontiers, and the legal and institutional implications of artificial intelligence.
Updated September 2026 · Foundational works and current research selected for enduring technical, educational, or legal significance
This collection includes peer-reviewed scholarship, preprints, model technical reports, benchmarks, and institutional reports. Publication-type labels indicate the status of each source; inclusion recognizes significance, not endorsement.
Where the field is now
Current frontiers: 2024–2026
The 2026 AI Index Report
Stanford Institute for Human-Centered Artificial Intelligence (HAI)
Synthesizes reviewed evidence on AI capabilities, adoption, investment, education, policy, and responsible-AI practice. It is a strong single starting point for understanding where the field stands in 2026.
Read the report → (opens in a new tab) 2026Peer-reviewed articleEvaluation & safetyReliable and Responsible Foundation Models: A Comprehensive Survey
Xinyu Yang, Junlin Han, Rishi Bommasani, et al. · Transactions on Machine Learning Research
Organizes a broad literature on reliability, fairness, security, privacy, uncertainty, explainability, and hallucination. The survey connects technical methods with high-stakes settings including law, medicine, and education.
Explore → (opens in a new tab) 2026PreprintReasoning & agentsAI co-mathematician: Accelerating mathematicians with agentic AI
Daniel Zheng, Ingrid von Glehn, Yori Zwols, et al.
Presents a stateful research workbench combining literature search, computation, theorem proving, and records of failed hypotheses. It illustrates a shift from one-shot answers toward interactive scientific collaboration.
Explore → (opens in a new tab) 2026PreprintEvaluation & safetyEvaluating Language Models for Harmful Manipulation
Canfer Akbulut, Rasmi Elasmar, Abhishek Roy, et al.
Develops a framework for measuring whether model-generated advice changes beliefs and behavior across policy, financial, and health settings. Its large human-participant study shows why manipulation risk has to be evaluated across contexts and populations.
Explore → (opens in a new tab) 2026Policy briefMultimodal systemsThe World Model and Spatial Intelligence Era: Governing AI Beyond Language
Daniel Zhang, Russell Wald, Ehsan Adeli, … Li Fei-Fei · Stanford HAI
Explains the emergence of models that reason about physical and spatial environments rather than language alone. The brief identifies governance questions that arise as AI moves into embodied and real-world settings.
Read the brief → (opens in a new tab) 2025Peer-reviewed articleReasoning & agentsDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI · Nature 645, 633–638
Reports that reinforcement learning could elicit extended reasoning behaviors without human-authored reasoning traces, and that those behaviors could be distilled into smaller models. It became a major reference point for reasoning-model development.
Explore → (opens in a new tab) 2025Benchmark / evaluationEvaluation & safetyPaperBench: Evaluating AI’s Ability to Replicate AI Research
Giulio Starace, Oliver Jaffe, Dane Sherburn, et al. · ICML 2025
Introduces a benchmark in which agents attempt to reproduce the results of recent machine-learning papers. Its thousands of gradable tasks make research replication a concrete measure of longer-horizon agent capability.
Explore → (opens in a new tab) 2025PreprintEvaluation & safetyOn the Biology of a Large Language Model
Jack Lindsey, et al. · Anthropic
Uses circuit-tracing methods to examine planning, multilingual representation, hallucination, and faithfulness inside a deployed language model. It is a substantial case study in mechanistic interpretability, and explicit about the method’s limitations.
Explore → (opens in a new tab) 2025Benchmark / evaluationReasoning & agentsMeasuring AI Ability to Complete Long Software Tasks
METR
Proposes measuring agent capability by the amount of human work time represented by tasks the agent can complete reliably. The time-horizon framing offers an intuitive way to track progress, with methodological caveats the authors set out directly.
Explore → (opens in a new tab) 2024Model technical reportMultimodal systemsGemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Google
Reports multimodal reasoning across context windows extending to millions of tokens. The report marked a shift in how models could work across long documents, codebases, audio, and video.
Explore → (opens in a new tab) 2024Model technical reportFoundation modelsThe Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al. · Meta
Describes Meta’s open-weight Llama 3 family, including its largest model and associated safety systems. It became an important reference for capable models that researchers and institutions can run and adapt outside a closed API.
Explore → (opens in a new tab) 2024Conference paperReasoning & agentsScaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar · ICLR 2025
Shows how adaptively allocating computation during inference can improve performance on difficult problems. The work helped establish test-time compute as a capability lever distinct from training a larger model.
Explore → (opens in a new tab) 2024Conference paperReasoning & agentsSWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
John Yang, Carlos E. Jimenez, Alexander Wettig, et al. · NeurIPS 2024
Shows that the interface through which an agent navigates, edits, and tests a repository materially affects performance. It helped define software engineering as an important real-world test bed for AI agents.
Explore → (opens in a new tab)Why this collection sits in a law library
AI, law, and governance
There is no free benchmark: An institutional view of legal AI benchmarking
Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho · PNAS 123(30)
Argues that public legal-AI benchmarks are necessary for a legible marketplace but can be captured, weakened, or gamed. The paper makes institutional design part of the technical conversation about trustworthy evaluation.
Explore → (opens in a new tab) 2026Law review articleAI, law & governanceHiding in Plain Sight: An Empirical Study of Prosecutorial Bias in AI Legal Analysis
Rory Pulvino, Dan Sutton, J.J. Naddeo · Science and Technology Law Review 27(1)
Examines more than 140,000 AI-generated legal memoranda and identifies a persistent tendency to recommend prosecution across varied framings and evidence. It provides concrete evidence that legal-analysis systems can encode consequential institutional defaults.
Explore → (opens in a new tab) 2025Peer-reviewed articleAI, law & governanceHallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, Daniel E. Ho · Journal of Empirical Legal Studies 22(2)
Presents a preregistered empirical evaluation of specialized legal research assistants. It found that access to legal databases did not eliminate hallucinations, incomplete answers, or meaningful differences among systems.
Explore → (opens in a new tab) 2024Peer-reviewed articleAI, law & governanceLarge Legal Fictions: Profiling Legal Hallucinations in Large Language Models
Matthew Dahl, Varun Magesh, Mirac Suzgun, Daniel E. Ho · Journal of Legal Analysis 16(1)
Provides systematic evidence of legal hallucinations across public language models, and documents disparities across courts, jurisdictions, and time periods. It remains a direct warning against treating general-purpose model output as reliable legal authority.
Explore → (opens in a new tab)From research result to public technology
The 2023 bridge
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al. · Meta AI
Introduced a family of smaller foundation models trained on publicly available data, showing that scale-efficient training could compete with much larger systems. It helped catalyze the modern open-weight model ecosystem.
Explore → (opens in a new tab) 2023Model technical reportMultimodal systemsGPT-4 Technical Report
OpenAI
Documents the capabilities, evaluations, safety work, and limitations of an early multimodal frontier model. Its limited architectural disclosure also made the report central to debates about transparency in frontier AI.
Explore → (opens in a new tab) 2023Conference paperEvaluation & safetyDirect Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, et al. · NeurIPS 2023
Reframed preference alignment as a direct optimization objective, avoiding a separate reward model and a full reinforcement-learning loop. DPO became a widely used and comparatively simple alternative to conventional RLHF pipelines.
Explore → (opens in a new tab) 2023Conference paperReasoning & agentsGenerative Agents: Interactive Simulacra of Human Behavior
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, et al. · UIST 2023
Combined memory, reflection, and planning to create believable long-running behavior in simulated agents. The architecture became an influential early blueprint for agentic systems.
Explore → (opens in a new tab)Foundations
Foundational architecture papers
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · NeurIPS 2017
Introduced the Transformer, an architecture built around attention rather than recurrent or convolutional layers. It made the large-scale models behind GPT, BERT, and much of modern generative AI possible.
Read the paper → (opens in a new tab) 2014Conference paperNetworkGenerative Adversarial Networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio · NeurIPS 2014
This paper proposed two neural networks—a generator and discriminator—that compete against each other, unlocking remarkably realistic image generation and a new field of generative modeling.
Explore → (opens in a new tab) 2013Conference paperLatent spaceAuto-Encoding Variational Bayes
Diederik P. Kingma, Max Welling · ICLR 2014
This paper introduced variational autoencoders, blending neural networks with probabilistic theory to learn a structured latent space from which coherent new data can be sampled.
Explore → (opens in a new tab)Foundations
Language model foundations
Improving Language Understanding by Generative Pre-Training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever · OpenAI
The paper that began the GPT lineage demonstrated the effectiveness of unsupervised pre-training followed by supervised fine-tuning.
Explore → (opens in a new tab) 2020Conference paperFew-shotLanguage Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, et al. · NeurIPS 2020
GPT-3 showed that scaling to 175 billion parameters enabled a model to perform many tasks from only a few examples.
Explore → (opens in a new tab) 2018Conference paperContextBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova · NAACL 2019
BERT learned from text in both directions at once, using full context to set new records across language-understanding tasks.
Explore → (opens in a new tab)Foundations
Image generation breakthroughs
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, Pieter Abbeel · NeurIPS 2020
Established diffusion models as a leading approach to high-quality image generation through a progressive denoising process.
Explore → (opens in a new tab) 2022Conference paperLatent diffusionHigh-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer · CVPR 2022
Made image generation practical by running diffusion in a smaller, compressed latent space.
Explore → (opens in a new tab) 2018Conference paperStyleGANA Style-Based Generator Architecture for Generative Adversarial Networks
Tero Karras, Samuli Laine, Timo Aila · CVPR 2019
StyleGAN produced realistic faces through an architecture that controls image styles at different layers.
Explore → (opens in a new tab)Foundations
Multimodal and vision-language models
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, et al. · ICML 2021
CLIP connected text and images with robust visual understanding and became a critical component of later image-generation systems.
Explore → (opens in a new tab) 2021Conference paperDALL-EZero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, et al. · ICML 2021
DALL-E showed that models could generate compelling, whimsical images directly from natural-language descriptions.
Explore → (opens in a new tab) 2022PreprintDALL-E 2Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, Mark Chen · OpenAI
DALL-E 2 used CLIP representations to generate coherent, detailed images through a two-stage process.
Explore → (opens in a new tab)Foundations
Training and alignment innovations
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, et al. · NeurIPS 2022
Introduced an influential RLHF approach that aligns model behavior with preferences supplied by human labelers.
Explore → (opens in a new tab) 2017Conference paperPreferencesDeep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei · NeurIPS 2017
A foundational proposal to use human preferences as a reward signal for teaching AI systems complex goals.
Explore → (opens in a new tab) 2022Conference paperReasoningChain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, Denny Zhou · NeurIPS 2022
Demonstrated that prompting models with intermediate reasoning steps can substantially improve logic and mathematics performance.
Explore → (opens in a new tab)Foundations
Foundational supporting technologies
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean · ICLR 2013 workshop
Introduced Word2vec, an efficient method for learning word embeddings that capture semantic relationships.
Explore → (opens in a new tab) 2014Conference paperOptimizationAdam: A Method for Stochastic Optimization
Diederik P. Kingma, Jimmy Ba · ICLR 2015
Introduced the Adam optimizer, which adapts the learning rate for each model parameter and became a standard choice for training transformers.
Explore → (opens in a new tab) 2015Conference paperU-NetU-Net: Convolutional Networks for Biomedical Image Segmentation
Olaf Ronneberger, Philipp Fischer, Thomas Brox · MICCAI 2015
The encoder-decoder U-Net architecture became essential to image-to-image tasks and later diffusion systems.
Explore → (opens in a new tab)Foundations
The 2022 frontier
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, et al. · Anthropic
Proposed a scalable alignment method in which an AI model critiques and revises outputs using a written constitution.
Explore → (opens in a new tab) 2022Peer-reviewed articleScalingPaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, et al. · Journal of Machine Learning Research 2023
Demonstrated the capabilities of a 540-billion-parameter language model trained with the Pathways system.
Explore → (opens in a new tab) 2022Conference paperMultimodalFlamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, et al. · NeurIPS 2022
A multimodal breakthrough that connected pre-trained vision and language models to process interleaved sequences of text and images.
Explore → (opens in a new tab)No papers found. Try a broader title, author, year, or topic, or clear the filters.
Too new to place
Emerging 2026 watchlist
These recent preprints and model reports may prove influential, but their long-term significance and independent validation are still developing.
Automated Researchers Can Mitigate Well-characterized Alignment Failures
Early evidence that research agents can discover and test mitigations, alongside findings of occasional evaluation gaming. (opens in a new tab)
2026Model technical reportQwen3.5-Omni Technical Report
A notable 2026 report on unified text, image, audio, and video modeling. (opens in a new tab)
2026Model technical reportDeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
An open-weight preview report emphasizing long context and training and inference efficiency. (opens in a new tab)
2026PreprintContinuous Latent Diffusion Language Model
Known as Cola-DLM: a possible alternative to strictly left-to-right token generation, still too new for the core collection. (opens in a new tab)
Selection approach. Works are included when they introduced a durable technical method or capability, supplied important independent evidence or evaluation, clarified legal or institutional consequences, or provided an authoritative synthesis of the field. Newer preprints are labeled separately because influence and validation take time to establish.
Each entry marks a milestone whose techniques, architectures, or insights continue to influence the field.