I'm a researcher working on language model interpretability, control, and equitable AI systems.
Currently, I'm a Non-Trivial Fellow advised by Clinton Morimoto and an ML Research Intern at the ECAI Lab at Stevens Institute of Technology advised by Haohang Li and Zining Zhu, where I work on mechanistic interpretability and control. Previously, I worked with Alan Sun on circuit stability as a measure for grokking and forgetting, and with Utkarsh Sharma at Algoverse on MURMUR, a cross-lingual multimodal benchmark for low-resource languages. I also interned at Pocket FM working on long-form expressive narrative TTS. I also founded the AI Equity Project, a national nonprofit building technology with and for communities historically excluded from its design.
I'm broadly interested in understanding and improving language models, with a focus on who they fail and why. My work has approached this through three axes: (a) interpretability: what internal mechanisms drive model behavior? (b) control: how do we intervene on those mechanisms reliably? (c) equity: where do current systems systematically fail underrepresented populations, and how do we measure that?
Most recently I've become interested in (1) whether mechanistic signals can guide data curation to resolve tradeoffs in model training that appear algorithmic but are actually artifacts of data composition; (2) whether the generalization properties RL develops can be replicated through targeted SFT; and (3) whether AI systems used in group decision-making faithfully represent the populations they're supposed to serve, and how to measure when they don't.
(Ordered chronologically)