I hold an MSc in Computer Science from Nanjing University, and my research focuses on the intersection of Software Engineering (SE) and Artificial Intelligence.
🎓 Status: I have completed my Master's degree and am actively seeking PhD opportunities in AI evaluation, model auditing, and trustworthy systems.
My work focuses on empirical evaluation, benchmark validity, and auditing automated systems:
- Model Auditing & Bias Probing: Developing black-box testing methodologies (e.g., Combinatorial Interaction Testing) to stress-test failure boundaries, behavioral inconsistencies, and systemic bias in foundation models.
- Empirical Benchmarks: Designing robust, counterfactual, and task-driven benchmark suites to evaluate LLMs and agentic pipelines beyond surface-level metrics.
- SE for AI / Reliable Tooling: Structuring agent workflows, Model Context Protocol (MCP) integrations, and verification mechanisms for reliable execution.
- LLM Evaluation: Studied empirical LLM benchmarks, failure modes, and RAG architectures as a Visiting Scholar at York University.
- AI Testing: Investigated combinatorial black-box testing (CIT) to uncover multi-attribute failure modes in deep learning models.
- Agent Frameworks: Contributed to open-source agent orchestration (ISEK) and authored a book chapter on intelligent agent development (From DeepSeek to Manus, Tsinghua University Press).
- Email: peixuanxia@gmail.com
- Looking for: PhD opportunities / Research collaborations in LLM Auditing & Evaluation, and empirical SE.
- Pronouns: She/Her
- Beyond Code: Passionate about light hiking, crochet & sewing, intersectional feminism, and animal welfare.


