Xiangqi Wang

Xiangqi (Shawn) Wang

I am a PhD student in Computer Science and Engineering at the University of Notre Dame, advised by Prof. Xiangliang Zhang. My research focuses on continual learning for language models.

Continual learning for language models

Language models must adapt to new experience without losing reliable reasoning or aligned behavior. My research studies this problem through test-time training, reinforcement learning, self-distillation, and preference optimization: how to select reasoning actions, enable adaptation, assess reasoning quality, limit behavioral drift, and balance competing objectives in human–model interaction.

A sequential learning perspective

The schematic below relates action selection and learning gains to future learning value, subject to a constraint on protected utility.

Pt(θt)= maxatAdaReasoner E[ gtPrimeTTT: reachableCE-PO: valid +γΔt Pt+1(θt+1)LGSD: preserved] s.t.UH(θt)UH(θ0)Dignified Peers: explicit protected utilityC.
Figure 1. A schematic view connecting distinct research contributions. Each method has its own learning objective, assumptions, and evaluation. The equation can be scrolled horizontally.
Notation and interpretation

In this schematic, θt denotes model parameters; at, a selected action; gt, an immediate learning gain; and Pt, a long-term learning value. The factor γΔt discounts future value. UH denotes protected utility, θ0 the reference model, and C an allowable utility decrease. These are conceptual definitions for the diagram; their operational meaning depends on the individual research setting.

Research questions and corresponding first-author work
ComponentResearch question and approachRelated work
Action selection
at
How should a model select reasoning actions for different problems?Reinforcement learning for an adaptive reasoning controller.AdaReasonerNeurIPS 2025 · Spotlight
Test-time adaptation
gt: reachability
How can initialization prepare a model to learn effectively at test time?Meta-learned initialization for test-time training.PrimeTTTEMNLP 2026 · Main
Reasoning coherence
gt: validity
How can policy optimization account for reasoning coherence alongside answer correctness?Coherence-aware reinforcement learning.CE-POEMNLP 2026 · Main
Behavioral preservation
Pt+1
How can self-distillation improve reasoning while limiting behavioral drift?Locality-constrained self-distillation.LGSDICLR 2027 · Submission
Preference alignment
UH
How can preference optimization balance truthfulness, anti-sycophancy, empathy, and creativity?Dual preference optimization with multiple behavioral objectives.Dignified PeersFindings of EMNLP 2026

First-author publications

Continual learning, reasoning, online learning, and AI systems.

  1. FIRST AUTHOREMNLP 2026 Main

    PrimeTTT: Priming Test-Time Training for LLMs via Meta-Learned Initialization

    Xiangqi Wang, D. Aishan, Y. Zhou, Y. Huang, X. Zhang

  2. FIRST AUTHOREMNLP 2026 Main

    Right Answers for the Right Reasons: Coherence-Enhanced Policy Optimization for Large Language Models

    Xiangqi Wang, Y. Huang, Y. Zhou, X. Luo, K. Guo, Y. Ma, X. Zhang

    Preprint title: Causally-Enhanced Reinforcement Policy Optimization (arXiv:2509.23095).

  3. FIRST AUTHORFindings of EMNLP 2026

    From Evasive Servants to Dignified Peers via Dual Preference Optimization

    Xiangqi Wang, Y. Huang, H. Zhuang, K. Guo, X. Zhang

    Preprint title: Dual Optimal: Make Your LLM Peer-like with Dignity (arXiv:2604.00979).

  4. FIRST AUTHORNeurIPS 2025 Spotlight

    AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking

    Xiangqi Wang, Y. Huang, Y. Wang, X. Luo, K. Guo, Y. Zhou, X. Zhang

  5. FIRST AUTHORTMLR 2025

    Fair Online Influence Maximization

    Xiangqi Wang, S. Zhang, J. E. A. Escamilla, Q. Wu, X. Zhang, J. Kang, H. Wang

  6. FIRST AUTHORWSDM 2025

    WildlifeLookup: A Chatbot Facilitating Wildlife Management with Accessible Data and Insights

    Xiangqi Wang, T. Yang, J. Rohr, B. Scheffers, N. Chawla, X. Zhang

  7. FIRST AUTHORICLR 2027 SubmissionUnder review

    Locality-Guided Self-Distillation

    Xiangqi Wang, D. Aishan, K. Guo, Y. Huang, X. Zhang

  8. FIRST AUTHORSIGMOD SubmissionUnder review

    SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL

    Xiangqi Wang, N. H. Pham, O. Hassanzadeh, D. Subramanian, X. Zhang

Co-authored publications

Evaluation, agents, reasoning, and trustworthy AI.

Experience and education

View full CV
FALL 2026
INDUSTRY

SAP

Part-time iXp Research Intern

Remote, USA

MAY — AUG 2026
INDUSTRY

IBM Research

Research Intern · Yorktown Heights, NY

Developed SAGE, a unified algebra and adaptive execution framework for AI functions in SQL, optimizing model selection and query execution for quality, cost, and latency.

2024 — PRESENT
EDUCATION

University of Notre Dame

PhD Student · Computer Science and Engineering

Advisor: Prof. Xiangliang Zhang
South Bend, Indiana

2020 — 2024
EDUCATION

University of Science and Technology of China

B.Sc. in Computer Science · First-Class Honors

School of the Gifted Young · Hefei, China

Prior mentorshipQi Liu ↗Huazheng Wang ↗Jian Kang ↗Zhuoran Yang ↗

Research software

Implementations and research systems.

Tools

PyTorch · Transformers · VERL · LLaMA-Factory · SQL · Git · Docker

Teaching, service, and research support

Mentorship & teaching

I guide student-led projects from problem selection and method design through experiments and writing.

Teaching assistant for Graduate Machine Learning at Notre Dame, and Computer Programming A and Data Analysis and Practice at USTC.

Academic service

Reviewer for NeurIPS, ICML, ICLR, KDD, and ACL Rolling Review (2025–2026), WWW (2025), and KDD (2024).

Journal reviewing: TMLR (2025) and IEEE Transactions on Computational Social Systems. Additional service: IEEE NASC Competition.

Research support

NSF DISCOVER ACCESS computing allocation and OpenAI Researcher Access API credits (2025), plus Anthropic research computing support.

Contributions to data and knowledge integration through NSF Proto-OKN, with GPU research infrastructure through NSF ACCESS.

Honors & community participation

Honors

  • Chinese Mathematics Competition · Provincial First Prize, 2022
  • ICPC China Regional · Competition prize, 2021
  • USTC Outstanding Student · Annual award
  • SGY 87–00 Scholarship · Nominee

Talks & participation

NSF Proto-OKN Annual Meeting; SASC Symposium; ACM WSDM Tutorial; EMNLP 2024; Third Reinforcement Learning Conference (remote).

Contact