Research
Bias, Alignment, and Social Meaning in Language Models
I study how subjective and often implicit social values embedded in language, and how large language models pick up, reproduce, or distort those signals. My work explores what this means both for understanding society through text and for building AI systems that are safe and trustworthy.
Bias in LLMs
Studying how large language models encode and express bias, including alignment behaviors that can undermine trustworthy AI.
Measuring Sycophancy of LLMs in Moral Dilemmas
Evaluating sycophantic bias across seven LLMs by measuring agreement and disagreement shifts and refusal rates under varying response restrictions and realistic framings.
Measuring Social Dynamics with LLMs
Using LLMs to detect and quantify social phenomena in real-world text.
Moderating Effects of Community Values on Media Bias and Incivility
Fine-tuned language models on 4M+ Reddit comments to detect incivility and seven community values, then used mixed-effects models to quantify how community norms moderate the relationship between media bias and incivility.
Computational Social Science
Applying computational text methods to study broader social phenomena in science and scholarship.
Measuring Linguistic Mimesis in Global Science
Developing text-based computational metrics to quantify international influence and field disruption in scientific research.
Quantifying Intellectual Humility in Scientific Literature
Built an annotated dataset of human values and humility, and fine-tuned language models to detect subtle, implicit expressions of intellectual humility in research papers.