Skip to content
View dependentsign's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report dependentsign

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
dependentsign/README.md

Hi, I'm Huanhuan Ma

I am a second-year Computer Science PhD student at the University of Illinois Chicago, advised by Prof. Philip S. Yu. I work on LLM and agent evaluation, model behavior, personalization, and trustworthy AI.

My current research focuses on:

  • reliable evaluation of multi-turn and personalized AI agents;
  • behavioral probes for understanding LLM traits, robustness, and alignment;
  • long-term user modeling, memory, and intention-aware personalization.

I am open to full-time AI research internships year-round, during the academic year as well as Summer 2027.

Selected work

Open source

  • CSI: behavioral evaluation toolkit for probing LLM traits beyond self-report.
  • Awesome-LLM-based-Evaluators: curated research on LLM-as-a-Judge and behavioral evaluation.
  • EX-FEVER: benchmark, dataset, and code for multi-hop explainable fact verification.

Side project

  • ClaudeUsageWidget: a macOS desktop widget for monitoring Claude AI usage limits and reset times.

Homepage · Google Scholar · LinkedIn

Pinned Loading

  1. sycophancy-rational-updating sycophancy-rational-updating Public

    Code and data for "Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update" (Findings of EMNLP 2026)

    Python

  2. CSI CSI Public

    Core Sentiment Inventory (CSI): reliable behavioral evaluation of LLM personality traits beyond self-report.

    Python 3

  3. Awesome-LLM-based-Evaluators Awesome-LLM-based-Evaluators Public

    ✨✨Latest Papers about LLM-based Evaluators

    33 2

  4. EX-FEVER EX-FEVER Public

    A Dataset for Multi-hop Explainable Fact Verification (ACL 2024 findings)

    Python 12

  5. ClaudeUsageWidget ClaudeUsageWidget Public

    macOS desktop widget (WidgetKit) for monitoring Claude AI usage limits and reset times

    Swift 53 12