Julie Zhu
Program Information
PhD
Research Category
Research Interests
AI Safety
Research Project
My research focuses on trustworthy machine learning, drawing on cognitive science and interpretability methods to understand how foundation models represent information, reason, and fail. This approach treats model failures, such as breakdowns under distribution shift and unwarranted confidence, as cognitive science treats human error, using them as evidence about underlying mechanisms. A central question is when and why these failures occur and what they reveal about the reliability and limitations of large language models.