Portrait
ॐ

Hello, my name is Sonia.

I care about understanding humans as thinking, feeling, computational machines and using these insights to build artificial intelligence that better serves a diversity of human intelligence. I study how machines act as social partners, and use the tools of computational cognitive science and human-AI interaction to evaluate and improve their behavior. I view scientific research as a creative endeavor, and one vehicle for the thought-forms I hope to project into the world.

I am currently on leave from my PhD in Computer Science at Harvard University. I am grateful to be advised by Tomer Ullman, and to be supported by an NSF Graduate Research Fellowship and Kempner Institute Graduate Fellowship. Previously, I spent time collaborating with researchers at the UK AI Security Institute and Google DeepMind on human-AI complementarity for scalable oversight, was a Predoctoral Young Investigator at the Allen Institute for Artificial Intelligence (AI2) working on personalized descriptions of scientific concepts, and completed my undergraduate studies at Princeton University, where I got my start in computational cognitive science building models of word-color associations in humans to understand synesthesia.

Selected Publications

sparkle spotlight talk

Directing large language models to follow the letter or spirit of the law

Peng Qian, Andrew Li, Sam Chen, Sonia K. Murthy, Yonatan Belinkov, Tomer D. Ullman

arXiv (2026)

Beyond Anthropomorphism: a Spectrum of Interface Metaphors for LLMs

Jianna So, Connie Cheng, Sonia K. Murthy

CHI Extended Abstracts (2026)

Cognitive models can reveal interpretable value trade-offs in language models

Sonia K. Murthy, Rosie Zhao, Jennifer Hu, Sham Kakade, Markus Wulfmeier, Peng Qian, Tomer Ullman

ICLR (2026)

An earlier iteration of this work appeared as spotlight talks at thesparklePragmatic Reasoning in Language Models workshop @ COLM 2025 andsparkleInterpreting Cognition in Deep Learning Models workshop @ NeurIPS 2025

Priors in Time: Missing Inductive Biases for Language Model Interpretability

Ekdeep Singh Lubana*, Can Rager*, Sai Sumedh R. Hindupur*, Valerie Costa, Greta Tuckute, Oam Patel, Sonia K. Murthy, Thomas Fel, Daniel Wurgaft, Eric J. Bigelow, Johnny Lin, Demba Ba, Martin Wattenberg, Fernanda Viegas, Melanie Weber, Aaron Mueller

ICLR (2026)

An earlier iteration of this work appeared at the Interpreting Cognition in Deep Learning Models workshop @ NeurIPS 2025

One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity

Sonia K. Murthy, Tomer Ullman, Jennifer Hu

NAACL (2025)

Please see Google Scholar for a full list of publications.

Invited talks

February 2026 Brown University, ANCOR seminar series

November 2025 Google DeepMind, VOICES team

October 2025 Annual Meeting of the Society for Neuroeconomics, "AI in Neuroeconomics" panel

If you would like to chat about research, or are a minority student considering graduate school in Psychology or Computer Science, please feel free to reach out to me at and I will do my best to respond!