Undergraduate Student · Tsinghua University

Building language systems that reason, plan, and act with people.

I'm Zixuan (Alex) Wang, a Mathematics and Physics undergraduate at Tsinghua University, with a minor in Artificial Intelligence. I work on agentic reinforcement learning, agent evaluation, and LLM personalization.

I am applying for 2027 Fall graduate program in the U.S.

I am Zixuan (Alex) Wang, a Mathematics and Physics undergraduate at Tsinghua University (entered in 2023), with a minor in Artificial Intelligence. I study agentic reinforcement learning and LLM personalization, with an interest in how agents learn from interaction and adapt to new tasks and users.

I have worked as a research assistant at Carnegie Mellon University, advised by Prof. Andrea Zanette. During my Fall 2025 exchange at UC San Diego, I was an undergraduate researcher at the UCSD MixLab (HDSI) with Dr. Zhen Wang. Earlier, I was a research intern at MiroMind AI, mentored by Dr. Yuntao Chen.

Recent Updates
  • Mind2Dialogue was accepted to the NeurIPS 2026 UserSim Workshop.
  • Harness Learning was accepted to the COLM 2026 LLA Workshop as a Spotlight Oral.
Research Interests

My recent work explores how agents can adapt their harnesses from execution feedback, how user simulation can provide supervision for human-aware models, and how to evaluate the user information retained in long-term agent memory.

Agent Learning

Reinforcement learning for harness adaptation, tool use, and sustained interaction.

Human-Aware Models

User simulation and structured supervision for personalization and reasoning about user intent.

Agent Evaluation

Evaluating adaptation across tasks and the user information agents retain in memory.

2023 — present
B.S. in Math&Physics
Tsinghua University
Minor in Artificial Intelligence · GPA: 3.9/4.0
Sep 2025 — Jan 2026
Visiting Student in Computer Science
UC San Diego
Marshall College · UCEAP · GPA: 3.9/4.0
Intern · Internship
Looki
Apr 2026 — Present · 1 mo
Agentic AIHuman-Centered AILarge Language Models
Research Assistant
Apr 2026 — Sep 2026 · 6 mos
Agentic AILarge Language Models
Research Intern · Internship
Jun 2025 — Nov 2025 · 6 mos
Agentic AILarge Language Models
Research Assistant
Halıcıoğlu Data Science Institute — UC San Diego (HDSI) — MixLab, advised by Dr. Zhen Wang
Nov 2025 — Apr 2026 · 6 mos
Interactive AILarge Language Models

Papers & Preprints

Teaser figure for Harness Learning Enables Generalizable Test-Time Adaptation LLA Workshop

Harness Learning Enables Generalizable Test-Time Adaptation

Alvin Zhang*, Xuecheng Liu*, Zixuan Wang*, Fahim Tajwar, Daman Arora, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette
COLM 2026 LLA Workshop — Accepted, Spotlight Oral
Training a proposer to adapt a solver's executable harness from execution feedback, with transfer to unseen tasks.
Teaser figure for Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States UserSim Workshop

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

Zixuan Wang*, Yufan Zhou*, Jinzhou Tang*, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang
NeurIPS 2026 UserSim Workshop — Accepted
Workshop version: Mind2Dialogue: Training Human-Aware Language Models through Shared-State User Simulation
Shared-state user simulation provides privileged supervision for personalization and theory-of-mind reasoning.
Teaser figure for Drift Calibration in Geometric Eye Tracking Systems Preprint

Drift Calibration in Geometric Eye Tracking Systems

Jiaqi Liu*, Zixuan Wang*, Yuhong Zhang, Dingkang Liang, Jane Hanqi Li, Tzyy-Ping Jung, Gert Cauwenberghs†
Preprint, 2026
A calibration-focused eye-tracking dataset and lightweight neural refinement for residual gaze error.
Teaser figure for MemAudit: Auditing Long-Term Agent Memory via Hidden User-State Recovery Preprint

MemAudit: Auditing Long-Term Agent Memory via Hidden User-State Recovery

Enze Ma, Yufan Zhou, Wei-Chieh Huang, Jie Yang, Huanhuan Ma, Zixuan Wang, Chengze Li, Chunyu Miao, Philip S. Yu, Zhen Wang
Preprint, 2026
The MEMPROBE benchmark measures how much hidden user state can be recovered from an agent's long-term memory.

Selected Writing

Technical Blogs

2/2026
Teaser figure for Multi-hop QA Synthesis for Large-Scale Deep Research Agent

Multi-hop QA Synthesis for Large-Scale Deep Research Agent

Multi-hop QAData SynthesisDeep Research
1/2026
Teaser figure for Stabilizing Large-scale MoE Agentic Reinforcement Learning Training

Course Notes

5/2025
Teaser figure for 3D Visual Computing Course Notes

3D Visual Computing Course Notes

Notes on 3D visual computing, including geometry processing, rendering, and 3D reconstruction techniques.

Computer Graphics3D VisionCourse Notes
1/2025
Teaser figure for Machine Learning Course Notes - Learning Theory

Machine Learning Course Notes - Learning Theory

Comprehensive notes on learning theory, covering PAC learning, VC dimension, and statistical learning foundations.

Machine LearningLearning TheoryCourse Notes
Teaser figure for Ego-embodied Reasoner: Egocentric Embodied Reasoning and Planning with MLLM via Reinforcement Learning Course

Ego-embodied Reasoner: Egocentric Embodied Reasoning and Planning with MLLM via Reinforcement Learning

Deep Reinforcement Learning Course Project, Instructed by Professor Huazhe Xu
Teaser figure for Unconditional and Image-conditioned 3D Generation Course

Unconditional and Image-conditioned 3D Generation

3D Visual Computing Course Project, instructed by Professor Li Yi
Teaser figure for Human Skeleton and Skin Generation Course

Human Skeleton and Skin Generation

Fundamentals of Computer Graphics Course Project, the 5-th Jittor AI competition, track 2
Teaser figure for Project Reading Report Course

Project Reading Report

Object Oriented Programming Course Project

I build with agents, and agents run on tokens. Every chart below is live telemetry from my own machines — every token my coding agents have read, cached, and written.

0
tokens, all time
0
cache hit rate
0
api-equivalent value
0
tokens written back
~/.claude/*.jsonl ccusage tokens.json this page ✳