<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Brian Christian — publications</title><description>New papers and preprints.</description><link>https://brianchristian.org/</link><item><title>Self-blinding and counterfactual self-simulation mitigate biases and sycophancy in large language models</title><link>https://arxiv.org/abs/2601.14553</link><guid isPermaLink="true">https://arxiv.org/abs/2601.14553</guid><description>Brian Christian, Matan Mazor</description><pubDate>Thu, 01 Jan 2026 00:00:00 GMT</pubDate></item><item><title>Post-training makes large language models less human-like</title><link>https://arxiv.org/abs/2605.07632</link><guid isPermaLink="true">https://arxiv.org/abs/2605.07632</guid><description>Marcel Binz, Elif Akata, Abdullah Almaatouq, Brian Christian, Eric Schulz</description><pubDate>Thu, 01 Jan 2026 00:00:00 GMT</pubDate></item><item><title>AI epistemic risks: Emerging mechanisms &amp; evidence</title><link>https://doi.org/10.2139/ssrn.6873005</link><guid isPermaLink="true">https://doi.org/10.2139/ssrn.6873005</guid><description>Mick Yang, Stephen Casper, Jonathan Stray, Jasmine Li, Cameron Jones, Anna Gausen, Natasha Jaques, Brian Christian</description><pubDate>Thu, 01 Jan 2026 00:00:00 GMT</pubDate></item><item><title>AI assistance reduces persistence and hurts independent performance</title><link>https://arxiv.org/abs/2604.04721</link><guid isPermaLink="true">https://arxiv.org/abs/2604.04721</guid><description>Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel Bakker, Rachit Dubey — Proceedings of the Conference on Language Modeling (COLM 2026)</description><pubDate>Thu, 01 Jan 2026 00:00:00 GMT</pubDate></item><item><title>Resolving Feynman’s restaurant problem reveals optimal solutions and human strategies</title><link>https://doi.org/10.1073/pnas.2509612123</link><guid isPermaLink="true">https://doi.org/10.1073/pnas.2509612123</guid><description>Brian Christian, Evan M. Russek, Thomas L. Griffiths — Proceedings of the National Academy of Sciences (PNAS)</description><pubDate>Thu, 01 Jan 2026 00:00:00 GMT</pubDate></item><item><title>Reward models inherit value biases from pretraining</title><link>https://openreview.net/forum?id=dT399j1Azv</link><guid isPermaLink="true">https://openreview.net/forum?id=dT399j1Azv</guid><description>Brian Christian, Jessica A.F. Thompson, Elle Michelle Yang, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, Tsvetomira Dumbalska — Proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026)</description><pubDate>Thu, 01 Jan 2026 00:00:00 GMT</pubDate></item><item><title>Reward model interpretability via optimal and pessimal tokens</title><link>https://brianchristian.org/research/papers/2025_Christian_etal_Reward_Model_Interpretability.pdf</link><guid isPermaLink="true">https://brianchristian.org/research/papers/2025_Christian_etal_Reward_Model_Interpretability.pdf</guid><description>Brian Christian, Hannah Rose Kirk, Jessica A.F. Thompson, Christopher Summerfield, Tsvetomira Dumbalska — Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT 2025)</description><pubDate>Wed, 01 Jan 2025 00:00:00 GMT</pubDate></item><item><title>Computational frameworks for human care</title><link>https://brianchristian.org/research/papers/2025_Christian_Computational_Frameworks_for_Human_Care.pdf</link><guid isPermaLink="true">https://brianchristian.org/research/papers/2025_Christian_Computational_Frameworks_for_Human_Care.pdf</guid><description>Brian Christian — Dædalus: The Social Science of Caregiving</description><pubDate>Wed, 01 Jan 2025 00:00:00 GMT</pubDate></item><item><title>Using adaptive intrinsic motivation in RL to model learning across development</title><link>https://brianchristian.org/research/papers/2024_Sandbrink_etal_Using_Adaptive_Intrinsic_Motivation.pdf</link><guid isPermaLink="true">https://brianchristian.org/research/papers/2024_Sandbrink_etal_Using_Adaptive_Intrinsic_Motivation.pdf</guid><description>Kai Sandbrink, Brian Christian, Linas Nasvytis, Christian Schroeder de Witt, Patrick Butlin — NeurIPS Workshop on Intrinsically-Motivated and Open-Ended Learning (IMOL)</description><pubDate>Mon, 01 Jan 2024 00:00:00 GMT</pubDate></item><item><title>Personhood credentials: Artificial intelligence and the need for privacy-preserving ways to distinguish who is real online</title><link>https://arxiv.org/abs/2408.07892</link><guid isPermaLink="true">https://arxiv.org/abs/2408.07892</guid><description>Steven Adler, Zoë Hitzig, Shrey Jain, Catherine Brewer, Wayne Chang, Renée DiResta, Brian Christian</description><pubDate>Mon, 01 Jan 2024 00:00:00 GMT</pubDate></item><item><title>Can reinforcement learning model learning across development? Online lifelong learning through adaptive intrinsic motivation</title><link>https://brianchristian.org/research/papers/2024_Sandbrink_etal_RL_Model_Learning.pdf</link><guid isPermaLink="true">https://brianchristian.org/research/papers/2024_Sandbrink_etal_RL_Model_Learning.pdf</guid><description>Kai Sandbrink, Brian Christian, Linas Nasvytis, Christian Schroeder de Witt, Patrick Butlin — Proceedings of the 46th Annual Conference of the Cognitive Science Society</description><pubDate>Mon, 01 Jan 2024 00:00:00 GMT</pubDate></item><item><title>AI Objectives Institute whitepaper: A research agenda for the production of a flourishing civilization</title><link>https://ai.objectives.institute/s/AOI-Whitepaper.pdf</link><guid isPermaLink="true">https://ai.objectives.institute/s/AOI-Whitepaper.pdf</guid><description>Peter Eckersley, Max Shron, Tushant Jha, Divya Siddarth, Brittney Gallagher, Carroll Wainwright, Joel Lehman, Brian Christian, Deger Turan</description><pubDate>Sun, 01 Jan 2023 00:00:00 GMT</pubDate></item><item><title>Comment of the AI Policy and Governance Working Group on the NTIA AI Accountability Policy Request for Comment Docket NTIA-230407-0093</title><link>https://www.ias.edu/sites/default/files/AI%20Policy%20and%20Governance%20Working%20Group%20NTIA%20Comment.pdf</link><guid isPermaLink="true">https://www.ias.edu/sites/default/files/AI%20Policy%20and%20Governance%20Working%20Group%20NTIA%20Comment.pdf</guid><description>Aaron Maniam, Alondra Nelson, Ben Garfinkel, Brian Christian, Daniel E. Ho, Dorothy Chou, Helen Toner, Inioluwa Deborah Raji</description><pubDate>Sun, 01 Jan 2023 00:00:00 GMT</pubDate></item><item><title>How do humans overcome individual computational limitations by working together?</title><link>https://doi.org/10.1111/cogs.13232</link><guid isPermaLink="true">https://doi.org/10.1111/cogs.13232</guid><description>Natalia Vélez, Brian Christian, Mathew Hardy, Bill D. Thompson, Thomas L. Griffiths — Cognitive Science</description><pubDate>Sun, 01 Jan 2023 00:00:00 GMT</pubDate></item><item><title>Caching algorithms and rational models of memory</title><link>https://brianchristian.org/research/papers/2014_Press_etal_Caching_Algorithms.pdf</link><guid isPermaLink="true">https://brianchristian.org/research/papers/2014_Press_etal_Caching_Algorithms.pdf</guid><description>Avi Press, M Pacer, Thomas L. Griffiths, Brian Christian — Proceedings of the Annual Meeting of the Cognitive Science Society</description><pubDate>Wed, 01 Jan 2014 00:00:00 GMT</pubDate></item><item><title>Using category structures to test iterated learning as a method for identifying inductive biases</title><link>https://doi.org/10.1080/03640210701801974</link><guid isPermaLink="true">https://doi.org/10.1080/03640210701801974</guid><description>Thomas L. Griffiths, Brian Christian, Michael L. Kalish — Cognitive Science</description><pubDate>Tue, 01 Jan 2008 00:00:00 GMT</pubDate></item><item><title>Revealing priors on category structures through iterated learning</title><link>https://brianchristian.org/research/papers/2006_Griffiths_etal_Revealing_Priors.pdf</link><guid isPermaLink="true">https://brianchristian.org/research/papers/2006_Griffiths_etal_Revealing_Priors.pdf</guid><description>Thomas L. Griffiths, Brian Christian, Michael L. Kalish — Proceedings of the 28th Annual Conference of the Cognitive Science Society</description><pubDate>Sun, 01 Jan 2006 00:00:00 GMT</pubDate></item></channel></rss>