AI Learning Hub › Landmark AI papers & reports

Curated reading list

Landmark AI Papers & Reports

A curated guide to foundational research, current frontiers, and the legal and institutional implications of artificial intelligence.

Updated September 2026 · Foundational works and current research selected for enduring technical, educational, or legal significance

This collection includes peer-reviewed scholarship, preprints, model technical reports, benchmarks, and institutional reports. Publication-type labels indicate the status of each source; inclusion recognizes significance, not endorsement.

Where the field is now

Current frontiers: 2024–2026

2026Institutional reportEvaluation & safety

The 2026 AI Index Report

Stanford Institute for Human-Centered Artificial Intelligence (HAI)

Synthesizes reviewed evidence on AI capabilities, adoption, investment, education, policy, and responsible-AI practice. It is a strong single starting point for understanding where the field stands in 2026.

Read the report → (opens in a new tab)
2026Peer-reviewed articleEvaluation & safety

Reliable and Responsible Foundation Models: A Comprehensive Survey

Xinyu Yang, Junlin Han, Rishi Bommasani, et al. · Transactions on Machine Learning Research

Organizes a broad literature on reliability, fairness, security, privacy, uncertainty, explainability, and hallucination. The survey connects technical methods with high-stakes settings including law, medicine, and education.

Explore → (opens in a new tab)
2026PreprintReasoning & agents

AI co-mathematician: Accelerating mathematicians with agentic AI

Daniel Zheng, Ingrid von Glehn, Yori Zwols, et al.

Presents a stateful research workbench combining literature search, computation, theorem proving, and records of failed hypotheses. It illustrates a shift from one-shot answers toward interactive scientific collaboration.

Explore → (opens in a new tab)
2026PreprintEvaluation & safety

Evaluating Language Models for Harmful Manipulation

Canfer Akbulut, Rasmi Elasmar, Abhishek Roy, et al.

Develops a framework for measuring whether model-generated advice changes beliefs and behavior across policy, financial, and health settings. Its large human-participant study shows why manipulation risk has to be evaluated across contexts and populations.

Explore → (opens in a new tab)
2026Policy briefMultimodal systems

The World Model and Spatial Intelligence Era: Governing AI Beyond Language

Daniel Zhang, Russell Wald, Ehsan Adeli, … Li Fei-Fei · Stanford HAI

Explains the emergence of models that reason about physical and spatial environments rather than language alone. The brief identifies governance questions that arise as AI moves into embodied and real-world settings.

Read the brief → (opens in a new tab)
2025Peer-reviewed articleReasoning & agents

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI · Nature 645, 633–638

Reports that reinforcement learning could elicit extended reasoning behaviors without human-authored reasoning traces, and that those behaviors could be distilled into smaller models. It became a major reference point for reasoning-model development.

Explore → (opens in a new tab)
2025Benchmark / evaluationEvaluation & safety

PaperBench: Evaluating AI’s Ability to Replicate AI Research

Giulio Starace, Oliver Jaffe, Dane Sherburn, et al. · ICML 2025

Introduces a benchmark in which agents attempt to reproduce the results of recent machine-learning papers. Its thousands of gradable tasks make research replication a concrete measure of longer-horizon agent capability.

Explore → (opens in a new tab)
2025PreprintEvaluation & safety

On the Biology of a Large Language Model

Jack Lindsey, et al. · Anthropic

Uses circuit-tracing methods to examine planning, multilingual representation, hallucination, and faithfulness inside a deployed language model. It is a substantial case study in mechanistic interpretability, and explicit about the method’s limitations.

Explore → (opens in a new tab)
2025Benchmark / evaluationReasoning & agents

Measuring AI Ability to Complete Long Software Tasks

METR

Proposes measuring agent capability by the amount of human work time represented by tasks the agent can complete reliably. The time-horizon framing offers an intuitive way to track progress, with methodological caveats the authors set out directly.

Explore → (opens in a new tab)
2024Model technical reportMultimodal systems

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Google

Reports multimodal reasoning across context windows extending to millions of tokens. The report marked a shift in how models could work across long documents, codebases, audio, and video.

Explore → (opens in a new tab)
2024Model technical reportFoundation models

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al. · Meta

Describes Meta’s open-weight Llama 3 family, including its largest model and associated safety systems. It became an important reference for capable models that researchers and institutions can run and adapt outside a closed API.

Explore → (opens in a new tab)
2024Conference paperReasoning & agents

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar · ICLR 2025

Shows how adaptively allocating computation during inference can improve performance on difficult problems. The work helped establish test-time compute as a capability lever distinct from training a larger model.

Explore → (opens in a new tab)
2024Conference paperReasoning & agents

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

John Yang, Carlos E. Jimenez, Alexander Wettig, et al. · NeurIPS 2024

Shows that the interface through which an agent navigates, edits, and tests a repository materially affects performance. It helped define software engineering as an important real-world test bed for AI agents.

Explore → (opens in a new tab)

Why this collection sits in a law library

AI, law, and governance

2026Peer-reviewed articleAI, law & governance

There is no free benchmark: An institutional view of legal AI benchmarking

Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho · PNAS 123(30)

Argues that public legal-AI benchmarks are necessary for a legible marketplace but can be captured, weakened, or gamed. The paper makes institutional design part of the technical conversation about trustworthy evaluation.

Explore → (opens in a new tab)
2026Law review articleAI, law & governance

Hiding in Plain Sight: An Empirical Study of Prosecutorial Bias in AI Legal Analysis

Rory Pulvino, Dan Sutton, J.J. Naddeo · Science and Technology Law Review 27(1)

Examines more than 140,000 AI-generated legal memoranda and identifies a persistent tendency to recommend prosecution across varied framings and evidence. It provides concrete evidence that legal-analysis systems can encode consequential institutional defaults.

Explore → (opens in a new tab)
2025Peer-reviewed articleAI, law & governance

Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools

Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, Daniel E. Ho · Journal of Empirical Legal Studies 22(2)

Presents a preregistered empirical evaluation of specialized legal research assistants. It found that access to legal databases did not eliminate hallucinations, incomplete answers, or meaningful differences among systems.

Explore → (opens in a new tab)
2024Peer-reviewed articleAI, law & governance

Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models

Matthew Dahl, Varun Magesh, Mirac Suzgun, Daniel E. Ho · Journal of Legal Analysis 16(1)

Provides systematic evidence of legal hallucinations across public language models, and documents disparities across courts, jurisdictions, and time periods. It remains a direct warning against treating general-purpose model output as reliable legal authority.

Explore → (opens in a new tab)

From research result to public technology

The 2023 bridge

2023PreprintFoundation models

LLaMA: Open and Efficient Foundation Language Models

Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al. · Meta AI

Introduced a family of smaller foundation models trained on publicly available data, showing that scale-efficient training could compete with much larger systems. It helped catalyze the modern open-weight model ecosystem.

Explore → (opens in a new tab)
2023Model technical reportMultimodal systems

GPT-4 Technical Report

OpenAI

Documents the capabilities, evaluations, safety work, and limitations of an early multimodal frontier model. Its limited architectural disclosure also made the report central to debates about transparency in frontier AI.

Explore → (opens in a new tab)
2023Conference paperEvaluation & safety

Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Rafael Rafailov, Archit Sharma, Eric Mitchell, et al. · NeurIPS 2023

Reframed preference alignment as a direct optimization objective, avoiding a separate reward model and a full reinforcement-learning loop. DPO became a widely used and comparatively simple alternative to conventional RLHF pipelines.

Explore → (opens in a new tab)
2023Conference paperReasoning & agents

Generative Agents: Interactive Simulacra of Human Behavior

Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, et al. · UIST 2023

Combined memory, reflection, and planning to create believable long-running behavior in simulated agents. The architecture became an influential early blueprint for agentic systems.

Explore → (opens in a new tab)

Foundations

Foundational architecture papers

Foundations

Language model foundations

Foundations

Image generation breakthroughs

Foundations

Multimodal and vision-language models

Foundations

Training and alignment innovations

Foundations

Foundational supporting technologies

Foundations

The 2022 frontier

Too new to place

Emerging 2026 watchlist

These recent preprints and model reports may prove influential, but their long-term significance and independent validation are still developing.

Selection approach. Works are included when they introduced a durable technical method or capability, supplied important independent evidence or evaluation, clarified legal or institutional consequences, or provided an authoritative synthesis of the field. Newer preprints are labeled separately because influence and validation take time to establish.

Each entry marks a milestone whose techniques, architectures, or insights continue to influence the field.