Our paper accepted at NeurIPS 2026!
Delighted to share that our paper, “Tracing Moral Foundations in Large Language Models”, has been accepted at NeurIPS 2026! A huge thank-you to my wonderful coauthors. I’m looking forward to sharing our work in Atlanta!
For psychology, our findings provide computational evidence for moral pluralism over single-axis models of harm, showing that multidimensional moral representations can emerge purely from the statistical regularities of human language.
For computer science, we show that this moral geometry is established during pretraining and selectively rewired by alignment. Using sparse autoencoders (SAEs) and causal steering, we identify interpretable mechanisms that directly shape downstream model judgments.