☑ represents archival peer-reviewed publications; * denotes equal contribution.
| Book | AI
as Normal Technology
Arvind Narayanan, Sayash Kapoor Princeton University Press (in preparation) |
| Preprint | Births are difficult to predict even with
rich survey and full-population register data
Elizaveta Sivak, ..., Sayash Kapoor et al. Preprint (2026) |
| Preprint | Can AI agents conduct open-ended AI research?
Early evidence from two case studies · Project page
Peter Kirgis, Sayash Kapoor et al. Preprint (2026) |
| Preprint | Bridging Predictions and Interventions: An
Integrated Framework for Automated Decision-Systems
Inioluwa Deborah Raji*, ..., Sayash Kapoor et al. Preprint (2026) |
| Conference | ☑ Holistic Agent Leaderboard: The Missing
Infrastructure for AI Agent Evaluation · Project page
Sayash Kapoor* et al. International Conference on Learning Representations (ICLR 2026) |
| Conference | ☑ The Limits of Inference Scaling Through
Resampling
Benedikt Stroebl, Sayash Kapoor, Arvind Narayanan International Conference on Learning Representations (ICLR 2026) |
| Conference | ☑ Towards a Science of AI Agent
Reliability · Dashboard
· Blog post
Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, Arvind Narayanan International Conference on Machine Learning (ICML 2026) Also presented at: Workshop on Failure Modes of Agentic AI at ICML 2026 |
| Conference | ☑ FLARE-AI: Flaw Reporting for AI
· OpenReview · ICML poster
Shayne Longpre, ..., Sayash Kapoor et al. International Conference on Machine Learning (ICML 2026) |
| Workshop | Life After Benchmark Saturation: A Case Study
of CORE-Bench · OpenReview
Nitya Nadgir, Sayash Kapoor et al. Second Workshop on Agents in the Wild: Safety, Security, and Beyond at ICML 2026 |
| Workshop | Log analysis is necessary for credible
evaluation of AI agents · OpenReview
Peter Kirgis, Sayash Kapoor et al. Workshop on Failure Modes of Agentic AI at ICML 2026 |
| Preprint | Seven simple steps for log analysis in AI
systems · Inspect
Scout
Magda Dubois, ..., Sayash Kapoor et al. Preprint (2026) |
| Workshop | Open-World Evaluations for Measuring Frontier
AI Capabilities · PDF
· CRUX · Blog post
· OpenReview
Sayash Kapoor et al. Second Workshop on Agents in the Wild: Safety, Security, and Beyond at ICML 2026 |
| Report | International AI Safety Report 2026
Yoshua Bengio, ..., Sayash Kapoor et al. International AI Safety Report (2026) A report on the state of advanced AI capabilities and risks written by 100 AI experts |
| Journal | ☑ The 2025 Foundation Model Transparency
Index · Project page
Alexander Wan, ..., Sayash Kapoor et al. Transactions on Machine Learning Research (TMLR 2026) |
| Online Essay | AI Won't
Automatically Make Legal Services Cheaper
Justin Curl, Sayash Kapoor, Arvind Narayanan Lawfare (2026) |
| Preprint | Bridging Prediction and Intervention
Problems in Social Systems
Lydia T. Liu, ..., Sayash Kapoor et al. Preprint (2025) |
| Report | International AI Safety Report
Yoshua Bengio, ..., Sayash Kapoor et al. International AI Safety Report (2025) A report on the state of advanced AI capabilities and risks written by 100 AI experts |
| Preprint | A Different Approach to AI Safety:
Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI
Safety
Camille François, ..., Sayash Kapoor et al. Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety (2025) |
| Conference | ☑ Position: Build Agent Advocates, Not
Platform Agents
Sayash Kapoor*, Noam Kolt*, Seth Lazar* International Conference on Machine Learning (ICML 2025 Position Paper Track) |
| Journal | ☑ AI
Agents That Matter · Blog post
Sayash Kapoor*, Benedikt Stroebl*, Zachary S. Siegel, Nitya Nadgir, Arvind Narayanan Transactions on Machine Learning Research (TMLR 2025) |
| Conference | ☑ The Leaderboard Illusion
Shivalika Singh, ..., Sayash Kapoor et al. NeurIPS D&B (2025) |
| Conference |
☑ Establishing Best Practices in Building
Rigorous Agentic Benchmarks Yuxuan Zhu, ..., Sayash Kapoor et al. NeurIPS D&B (2025) |
| Journal | Why an overreliance on
AI-driven modelling is bad for science
Arvind Narayanan, Sayash Kapoor Nature (2025) |
| Conference | ☑ Position: In-House
Evaluation Is Not Enough. Towards Robust Third-Party Evaluation and Flaw Disclosure
for General-Purpose AI · arXiv
Shayne Longpre, ..., Sayash Kapoor et al. International Conference on Machine Learning (ICML 2025 Position Paper Track, Spotlight) |
| Conference | ☑ The Reality of AI and Biorisk
Aidan Peppin, ..., Sayash Kapoor et al. ACM Conference on Fairness, Accountability, and Transparency (FAccT 2025) |
| Journal | ☑ The 2024 Foundation Model Transparency
Index Rishi Bommasani, ..., Sayash Kapoor et al. Transactions on Machine Learning Research (TMLR 2025) |
| Journal | ☑ The 2023 Foundation Model Transparency
Index Rishi Bommasani, ..., Sayash Kapoor et al. Transactions on Machine Learning Research (TMLR 2025 Featured certification) |
| Journal | ☑ CORE-Bench: Fostering the Credibility of
Published Research Through a Computational Reproducibility Agent Benchmark Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, Arvind Narayanan Transactions on Machine Learning Research (TMLR 2024) |
| Book | ☑ AI
Snake Oil: What Artificial Intelligence Can Do, What It Can't, and How to Tell the
Difference
Arvind Narayanan, Sayash Kapoor Princeton University Press (2024) Named one of Nature’s 10 best books of 2024, Bloomberg’s 49 best books of 2024, and Forbes’s 10 must-read tech books of 2024. |
| Journal | ☑ Considerations for
governing open foundation models
Rishi Bommasani, Sayash Kapoor et al. Science (2024) |
| Journal | ☑ REFORMS: Consensus-based
Recommendations for Machine-learning-based Science · Blog post
Sayash Kapoor et al. Science Advances (2024) |
| Conference | ☑ On the Societal Impact of Open
Foundation
Models · Blog
post
Sayash Kapoor* et al. International Conference on Machine Learning (ICML 2024 Oral) |
| Conference | ☑ A Safe Harbor for AI Evaluation and
Red
Teaming · Blog
post Shayne Longpre, Sayash Kapoor et al. International Conference on Machine Learning (ICML 2024 Oral) Our open letter to AI companies calling for a safe harbor was signed by over 350 academics, researchers, and civil society members. |
| Journal | ☑ How large language
models can reshape collective intelligence Jason W. Burton, ..., Sayash Kapoor et al. Nature Human Behaviour (2024) |
| Journal | ☑ The Responsible Foundation Model Development
Cheatsheet: A Review of Tools & Resources Shayne Longpre, ..., Sayash Kapoor et al. Transactions on Machine Learning Research (TMLR 2024 Survey certification) |
| Preprint | Towards a Framework for Openness in
Foundation Models: Proceedings from the Columbia Convening on Openness in Artificial
Intelligence Adrien Basdevant, ..., Sayash Kapoor et al. Preprint (2024) |
| Journal | ☑ Promises and pitfalls of artificial
intelligence for legal applications · Blog post
Sayash Kapoor, Peter Henderson, Arvind Narayanan Journal of Cross-disciplinary Research in Computational Law (CRCL 2024) |
| Journal | ☑ Against Predictive
Optimization: On the Legitimacy of Decision-Making Algorithms that Optimize
Predictive Accuracy · Blog
post Angelina Wang*, Sayash Kapoor*, Solon Barocas, Arvind Narayanan ACM Journal on Responsible Computing (JCR 2024) Also presented at: Philosophy, AI, and Society (2023); Data (Re)Makes the World (2023), ACM FAccT (2023) |
| Conference | ☑ Foundation Model Transparency
Reports · Blog post
Rishi Bommasani, ..., Sayash Kapoor et al. AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society (AIES 2024) |
| Journal | ☑ Leakage and the reproducibility
crisis in ML-based science Sayash Kapoor, Arvind Narayanan Patterns (2023) |
| Policy brief |
Considerations for Governing Open Foundation Models
·
Blog post
Rishi Bommasani, Sayash Kapoor et al. Stanford HAI Issue Brief (2023) |
| Comment | The limitations of machine
learning models for predicting scientific replicability M. J. Crockett, Xuechunzi Bai, Sayash Kapoor, Lisa Messeri, and Arvind Narayanan Proceedings of the National Academy of Sciences Comment (PNAS 2023) |
| Online essay | How
to Prepare for the Deluge of Generative AI on Social Media Sayash Kapoor, Arvind Narayanan Knight First Amendment Institute (2023) |
| Conference | ☑ Weaving Privacy and Power: On the Privacy
Practices of Labor Organizers in the U.S. Technology Industry Sayash Kapoor*, Matthew Sun*, Mona Wang*, Klaudia Jaźwińska*, Elizabeth Anne Watkins* ACM Conference on Computer-Supported Cooperative Work and Social Computing (CSCW 2022) 🏆 Impact Recognition Award |
| Conference | ☑ The worst of both worlds: A comparative
analysis of errors in learning from data in psychology and machine learning Jessica Hullman, Sayash Kapoor, Priyanka Nanayakkara, Andrew Gelman, Arvind Narayanan ACM Conference on AI, Ethics, and Society (AIES 2022) |
| Conference | ☑ Controlling polarization in
personalization: an algorithmic framework L. Elisa Celis, Sayash Kapoor, Farnood Salehi, and Nisheeth K. Vishnoi ACM Conference on Fairness, Accountability, and Transparency (FAccT) 2019 🏆 Best Paper Award |
| Journal | ☑ Corruption-tolerant
bandit learning Sayash Kapoor, Kumar Kshitij Patel, and Purushottam Kar Machine Learning (2019) |
| Journal | ☑ A
dashboard for controlling polarization in personalization L. Elisa Celis, Sayash Kapoor, Farnood Salehi, Vijay Keswani, and Nisheeth K. Vishnoi AI Communications (2019) |
| Conference | ☑ Balanced news using constrained
bandit-based personalization Sayash Kapoor, Vijay Keswani, Nisheeth K. Vishnoi, and L. Elisa Celis IJCAI Demos Track (2018) |