Portrait of Enyi (Olivia) Jiang
Portrait
Enyi (Olivia) Jiang
PhD @ UIUC
Visiting PhD @ Stanford

About Me

I study AI safety and alignment through models’ internal representations, developing ways to detect hidden risks and intervene while preserving useful capabilities.

I am a final-year Computer Science Ph.D. candidate at UIUC, advised by Prof. Sanmi Koyejo and Prof. Nancy Amato, and a visiting researcher at Stanford’s Trustworthy AI Research (STAIR) Lab. My broader interests include robustness under distribution shift and applications in climate science and healthcare.

Interested in discussing AI safety, alignment, or a potential research collaboration? Book a time below to connect with me.

Research · Enyi (Olivia) Jiang

Understanding AI safety
from the inside out.

Internal representations as evidence for evaluation,
signals for alignment, and targets for intervention.

Current & ongoing research
Looking ahead · future agenda
An emerging risk is detected and an agent trajectory is redirected.

Agent monitoring & steering

Detect emerging risks across reasoning and action, then redirect agents while preserving legitimate task progress.

Builds on CLEAR: selective intervention ↗

If latent risk can be detected, can we intervene only when necessary?

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

Chengxiao Wang*, Enyi Jiang*, Xiaojing Liao, Sanmi Koyejo (* equal contribution)

arXiv preprint · 2026

CLEAR uses a hidden-state gate to adjust a safety adapter’s strength, reducing attack success while preserving useful capabilities.

Does safe-looking behavior necessarily imply a safe internal model state?

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Enyi Jiang*, Anders Gjølbye*, Yibo Jacky Zhang, Sanmi Koyejo (* equal contribution)

arXiv preprint · 2026

Models can pass static refusal tests yet remain vulnerable to small internal perturbations, revealing a gap in behavioral safety evaluation.

Can alignment objectives operate directly over representations?

Latent Adversarial Regularization for Offline Preference Optimization

Enyi Jiang, Yibo Jacky Zhang, Yinglun Xu, Andreas Haupt, Nancy Amato, Sanmi Koyejo

arXiv preprint · 2026

GANPO regularizes preference optimization in latent space to provide more robust feedback under distribution shift and noise.

Education

  • University of Illinois at Urbana-Champaign
    Computer Science
    Ph.D. Student
    2023 - present
  • University of Illinois at Urbana-Champaign
    Computer Science
    MS Student (thesis-track)
    2020 - 2023
  • University of Illinois at Urbana-Champaign
    Electrical and Computer Engineering
    Undergraduate Student
    2016 - 2020
  • Zhejiang University
    ZJU-UIUC Institute
    Undergraduate Student
    2016 - 2020

Experience

  • Stanford University
    Department of Computer Science
    Visiting Student Researcher
    2026 - present
  • Meta
    Research Intern
    Summer 2025/2026

The past doesn’t define us. ❤️

May we, by God's grace, create beauty, discover essence, and perceive connections.