Julie Zhu

Julie Zhu

P
Program Information

PhD

👤
◆
Research Category
Artificial Intelligence
★
Research Interests
AI Safety

Research Project

My research focuses on trustworthy machine learning, drawing on cognitive science and interpretability methods to understand how foundation models represent information, reason, and fail. This approach treats model failures, such as breakdowns under distribution shift and unwarranted confidence, as cognitive science treats human error, using them as evidence about underlying mechanisms. A central question is when and why these failures occur and what they reveal about the reliability and limitations of large language models.