Yeaeun Kwon

Research

Bias, Alignment, and Social Meaning in Language Models

I study how subjective and often implicit social values embedded in language, and how large language models pick up, reproduce, or distort those signals. My work explores what this means both for understanding society through text and for building AI systems that are safe and trustworthy.

Bias in LLMs

Studying how large language models encode and express bias, including alignment behaviors that can undermine trustworthy AI.

Measuring Sycophancy of LLMs in Moral Dilemmas

Evaluating sycophantic bias across seven LLMs by measuring agreement and disagreement shifts and refusal rates under varying response restrictions and realistic framings.

Measuring Social Dynamics with LLMs

Using LLMs to detect and quantify social phenomena in real-world text.

Moderating Effects of Community Values on Media Bias and Incivility

Fine-tuned language models on 4M+ Reddit comments to detect incivility and seven community values, then used mixed-effects models to quantify how community norms moderate the relationship between media bias and incivility.

Computational Social Science

Applying computational text methods to study broader social phenomena in science and scholarship.

Measuring Linguistic Mimesis in Global Science

Developing text-based computational metrics to quantify international influence and field disruption in scientific research.

Quantifying Intellectual Humility in Scientific Literature

Built an annotated dataset of human values and humility, and fine-tuned language models to detect subtle, implicit expressions of intellectual humility in research papers.