I am a second-year Computer Science PhD student at the University of Illinois Chicago, advised by Prof. Philip S. Yu. I work on LLM and agent evaluation, model behavior, personalization, and trustworthy AI.
My current research focuses on:
- reliable evaluation of multi-turn and personalized AI agents;
- behavioral probes for understanding LLM traits, robustness, and alignment;
- long-term user modeling, memory, and intention-aware personalization.
I am open to full-time AI research internships year-round, during the academic year as well as Summer 2027.
- Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update — EMNLP 2026 Findings · code and data
- Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits — preprint, 2026 · code
- Scaling LLM Agent Learning with Data Synthesis: A Comprehensive Survey — preprint, 2026
- EX-FEVER: A Dataset for Multi-hop Explainable Fact Verification — ACL 2024 Findings · code and data
- Interpretable Multimodal Out-of-Context Detection with Soft Logic Regularization — ICASSP 2024 Oral
- Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data — arXiv preprint · data
- CSI: behavioral evaluation toolkit for probing LLM traits beyond self-report.
- Awesome-LLM-based-Evaluators: curated research on LLM-as-a-Judge and behavioral evaluation.
- EX-FEVER: benchmark, dataset, and code for multi-hop explainable fact verification.
- ClaudeUsageWidget: a macOS desktop widget for monitoring Claude AI usage limits and reset times.
