
Diagnose
What does safe-looking behavior leave hidden?
Probe fragile reasoning and internal vulnerabilities that output-based evaluations can miss.
I study AI safety and alignment through models’ internal representations, developing ways to detect hidden risks and intervene while preserving useful capabilities.
I am a final-year Computer Science Ph.D. candidate at UIUC, advised by Prof. Sanmi Koyejo and Prof. Nancy Amato, and a visiting researcher at Stanford’s Trustworthy AI Research (STAIR) Lab. My broader interests include robustness under distribution shift and applications in climate science and healthcare.
Interested in discussing AI safety, alignment, or a potential research collaboration? Book a time below to connect with me.
Internal representations as evidence for evaluation,
signals for alignment, and targets for intervention.

What does safe-looking behavior leave hidden?
Probe fragile reasoning and internal vulnerabilities that output-based evaluations can miss.

How can representations guide better alignment?
Use latent-space feedback during learning to complement token-level preference objectives.

When should a safety mechanism step in?
Apply safety mechanisms selectively, balancing protection with useful model capabilities.

Identify the features and pathways that causally contribute to safe behavior in open-weight models.
Representation-level safety evaluation ↗
Detect emerging risks across reasoning and action, then redirect agents while preserving legitimate task progress.
CLEAR: selective intervention ↗If latent risk can be detected, can we intervene only when necessary?
Chengxiao Wang*, Enyi Jiang*, Xiaojing Liao, Sanmi Koyejo (* equal contribution)
arXiv preprint · 2026
CLEAR uses a hidden-state gate to adjust a safety adapter’s strength, reducing attack success while preserving useful capabilities.
Does safe-looking behavior necessarily imply a safe internal model state?
Enyi Jiang*, Anders Gjølbye*, Yibo Jacky Zhang, Sanmi Koyejo (* equal contribution)
arXiv preprint · 2026
Models can pass static refusal tests yet remain vulnerable to small internal perturbations, revealing a gap in behavioral safety evaluation.
Can alignment objectives operate directly over representations?
Enyi Jiang, Yibo Jacky Zhang, Yinglun Xu, Andreas Haupt, Nancy Amato, Sanmi Koyejo
arXiv preprint · 2026
GANPO regularizes preference optimization in latent space to provide more robust feedback under distribution shift and noise.
The past doesn’t define us. ❤️
May we, by God's grace, create beauty, discover essence, and perceive connections.