SAP
Part-time iXp Research Intern
Remote, USA

I am a PhD student in Computer Science and Engineering at the University of Notre Dame, advised by Prof. Xiangliang Zhang. My research focuses on continual learning for language models.
Language models must adapt to new experience without losing reliable reasoning or aligned behavior. My research studies this problem through test-time training, reinforcement learning, self-distillation, and preference optimization: how to select reasoning actions, enable adaptation, assess reasoning quality, limit behavioral drift, and balance competing objectives in human–model interaction.
The schematic below relates action selection and learning gains to future learning value, subject to a constraint on protected utility.
In this schematic, θt denotes model parameters; at, a selected action; gt, an immediate learning gain; and Pt, a long-term learning value. The factor γΔt discounts future value. UH denotes protected utility, θ0 the reference model, and C an allowable utility decrease. These are conceptual definitions for the diagram; their operational meaning depends on the individual research setting.
| Component | Research question and approach | Related work |
|---|---|---|
| Action selection at | How should a model select reasoning actions for different problems?Reinforcement learning for an adaptive reasoning controller. | AdaReasonerNeurIPS 2025 · Spotlight |
| Test-time adaptation gt: reachability | How can initialization prepare a model to learn effectively at test time?Meta-learned initialization for test-time training. | PrimeTTTEMNLP 2026 · Main |
| Reasoning coherence gt: validity | How can policy optimization account for reasoning coherence alongside answer correctness?Coherence-aware reinforcement learning. | CE-POEMNLP 2026 · Main |
| Behavioral preservation Pt+1 | How can self-distillation improve reasoning while limiting behavioral drift?Locality-constrained self-distillation. | LGSDICLR 2027 · Submission |
| Preference alignment UH | How can preference optimization balance truthfulness, anti-sycophancy, empathy, and creativity?Dual preference optimization with multiple behavioral objectives. | Dignified PeersFindings of EMNLP 2026 |
Continual learning, reasoning, online learning, and AI systems.
Preprint title: Causally-Enhanced Reinforcement Policy Optimization (arXiv:2509.23095).
Preprint title: Dual Optimal: Make Your LLM Peer-like with Dignity (arXiv:2604.00979).
No papers match that search. Try a different topic or .
Evaluation, agents, reasoning, and trustworthy AI.
Part-time iXp Research Intern
Remote, USA
Research Intern · Yorktown Heights, NY
Developed SAGE, a unified algebra and adaptive execution framework for AI functions in SQL, optimizing model selection and query execution for quality, cost, and latency.
PhD Student · Computer Science and Engineering
Advisor: Prof. Xiangliang Zhang
South Bend, Indiana
B.Sc. in Computer Science · First-Class Honors
School of the Gifted Young · Hefei, China
Prior mentorshipQi Liu ↗Huazheng Wang ↗Jian Kang ↗Zhuoran Yang ↗
Implementations and research systems.
VERL extended with Jacobian-based causal scores and a parallel reward backend for policy optimization.
A modular reasoning controller with reproducible training and inference pipelines. The official implementation.
Fairness-aware combinatorial bandits with scalable optimization and online evaluation on public social graphs.
A searchable wildlife knowledge interface and graph-RAG chatbot for species, habitat, and policy queries.
PyTorch · Transformers · VERL · LLaMA-Factory · SQL · Git · Docker
I guide student-led projects from problem selection and method design through experiments and writing.
Teaching assistant for Graduate Machine Learning at Notre Dame, and Computer Programming A and Data Analysis and Practice at USTC.
Reviewer for NeurIPS, ICML, ICLR, KDD, and ACL Rolling Review (2025–2026), WWW (2025), and KDD (2024).
Journal reviewing: TMLR (2025) and IEEE Transactions on Computational Social Systems. Additional service: IEEE NASC Competition.
NSF DISCOVER ACCESS computing allocation and OpenAI Researcher Access API credits (2025), plus Anthropic research computing support.
Contributions to data and knowledge integration through NSF Proto-OKN, with GPU research infrastructure through NSF ACCESS.
NSF Proto-OKN Annual Meeting; SASC Symposium; ACM WSDM Tutorial; EMNLP 2024; Third Reinforcement Learning Conference (remote).