← First-author publications

Co-authored publications

Publications and ongoing work on evaluation, agents, reasoning, and alignment.

  1. COLLABORATIONICML 2026

    ProbeLLM: Automating Principled Diagnosis of LLM Failures

    Y. Huang, Z. Jiang, Y. Ma, Y. Jiang, Xiangqi Wang, et al.

  2. COLLABORATIONNeurIPS 2025

    DyFlow: Dynamic Workflow Framework for Agentic Reasoning

    Y. Wang, Z. Xu, Y. Huang, Xiangqi Wang, et al.

    Dynamic workflow construction for tool-using and multi-agent reasoning, supported by theoretical analysis.

  3. COLLABORATIONACL 2025

    CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP

    T. Yang, L. Dai, Xiangqi Wang, M. Cheng, Y. Tian, X. Zhang

    Contributed to implementation and research visualization.

  4. COLLABORATIONFindings of EMNLP 2025

    Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study

    Y. Zhou, J. Ye, Z. Ling, …, Xiangqi Wang, X. Zhang

  5. COLLABORATIONICLR 2026

    On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

    Y. Huang, C. Gao, S. Wu, H. Wang, Xiangqi Wang, et al.

    Contributed to the assessment framework, methodology, and large-scale experiments.

  6. COLLABORATIONCIKM 2025

    Think it Image by Image: Multi-Image Moral Reasoning of Large Vision-Language Models

    C. Gao, Y. Huang, Xiangqi Wang, S. Wu, N. Chawla, X. Zhang

    Contributed to manuscript development and experimental analysis.

  7. COLLABORATIONCOLM 2025

    Exposing and Patching the Flaws of Large Language Models in Social Character Simulation

    Y. Huang, Z. Yuan, Y. Zhou, K. Guo, Xiangqi Wang, et al.

    Related preprint: Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations? (2024; arXiv:2410.23426). Contributed simulation scenarios, reliability metrics, and analysis.

  8. COLLABORATIONPreprint, 2026Preprint

    Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents

    Y. Zhou, K. Guo, H. Zhuang, Xiangqi Wang, Y. Huang, et al.

  9. COLLABORATIONPreprint, 2026Preprint

    SenseMath: Do LLMs Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation

    H. Zhuang, Xiangqi Wang, Y. Shen, Y. Cheng, X. Zhang

  10. COLLABORATIONOngoing workOngoing

    ValueLence: A Dashboard for In-Depth Value Probing of LLMs

    Y. Huang, J. Ye, Z. Liu, Y. Li, Xiangqi Wang, et al.

    Contributed to the evaluation and analysis pipeline.