Research
I’m interested in the place where formal and informal systems meet: in particular, where the rigor and clarity of computer science collide with the fundamental, human-centered questions of philosophy and psychology.
My research pursues this junction through different perspectives (and fields). I’m interested in computational cognitive science—formal models of human cognition—and in AI alignment—formal models of human norms, desires, goals, and values. Recently, that has meant examining the implicit models of human decision making that are inverted by reinforcement learning from human feedback (RLHF), and exploring reward models as rich but flawed artifacts for expressing human preferences, subject to both computational and representational constraints. I am interested in how these models both succeed and fail.
My work has explored what chatbots do and don’t capture about conversation at its best (and worst); charted the computational structure of everyday decision-making dilemmas and shown that humans embody a resource-rational approximation of optimal solutions; and interrogated how reward models operationalize human preferences, norms, and values by turning normative and non-normative dialogues into scalar reward numbers. Through it all I am most engaged with pursuing rigorous notions of optimality, structure, and (mis)specification as they pertain to the core questions of the human experience. This is not my only interest, but after 20 years it is clear that it is inexhaustible.
There has never been a richer or more fascinating—or more urgent—time to pursue this particular set of questions. It is more than a life’s work, and it demands of us—I have always believed—that we be willing to trespass over traditional disciplinary lines to pursue these questions wherever they lead.
Publications
Post-training makes large language models less human-like
AI epistemic risks: Emerging mechanisms & evidence
AI assistance reduces persistence and hurts independent performance
Resolving Feynman’s restaurant problem reveals optimal solutions and human strategies
Reward models inherit value biases from pretraining
Using adaptive intrinsic motivation in RL to model learning across development
Personhood credentials: Artificial intelligence and the need for privacy-preserving ways to distinguish who is real online
Can reinforcement learning model learning across development? Online lifelong learning through adaptive intrinsic motivation
AI Objectives Institute whitepaper: A research agenda for the production of a flourishing civilization
Comment of the AI Policy and Governance Working Group on the NTIA AI Accountability Policy Request for Comment Docket NTIA-230407-0093
How do humans overcome individual computational limitations by working together?
For a fuller list, see my Google Scholar profile.
Academic affiliations
-
University of California, Berkeley
- Research Fellow, Center for Human-Compatible Artificial Intelligence (2026–Present)
- Affiliate, Center for Human-Compatible Artificial Intelligence (2020–Present)
- Visiting Scholar, Simons Institute for the Theory of Computing (2020)
- Visiting Scholar, CITRIS (2018–2019)
- Visiting Scholar, Institute of Cognitive and Brain Sciences (2012–2015)
- Research Assistant, Computational Cognitive Science Lab (2007)
-
University of Oxford
- Doctoral Researcher, Human Information Processing Lab (2023–2026)
- Clarendon Scholarship (2023–2026)
- Thesis: What Humans Want: Models of Preference, Choice, and Reward in Minds and Machines
- Letter of Commendation, Medical Sciences Division (2026)
- Trappes Exhibition (thesis examination prize), Lincoln College (2026)
-
Institute for Advanced Study
- Member, AI Policy and Governance Working Group (2023–2025)