Posts by Collection

publications

SHARD: Safe and Helpful Alignment via Self-Reframing Distillation

Published in Findings of EMNLP 2026, 2026

We introduce SHARD, a self-reframing distillation method that rewrites sensitive prompts to surface benign intent, reframes responses into safe and helpful ones, and fine-tunes the model on its self-reframed outputs — improving safe-helpfulness without sacrificing safety.

talks

teaching

INLS 509-001: Information Retrieval

Teaching Assistant, University of North Carolina at Chapel Hill, 2026

Teaching Assistant for INLS 509-001: Information Retrieval (Fall 2026), taught by Yue “Ray” Wang.