About me.
I study how multimodal decision systems can learn reliably and improve continuously in the real world.
Hi, I am a second-year master's student in the Department of Automation at Tsinghua University. My work follows a closed learning loop: constructing long-tail data, learning grounded reward signals, and post-training agents through failure feedback and interaction.
I am particularly interested in autonomous driving and embodied intelligence. I also explore quantitative research as an independent application of the same discipline: forming hypotheses from data and making decisions under uncertainty.
I believe complex systems become understandable when their patterns are observed rigorously. Good data, explicit feedback, and careful statistical analysis turn those observations into better decisions.
Feel free to reach out for discussion or collaboration.
RESEARCH THREADS
From data and feedback to better decisions.
Across my work, the same question recurs: how can a system encounter failures, receive grounded feedback, and become more reliable after each iteration?
-
01
Long-tail learning
Constructing challenging data and scenarios for robust decision-making.
-
02
Multimodal evaluation
Learning richer reward signals from visual context, rules, and intent.
-
03
Post-training agents
Making VLA systems improve through failure feedback and interaction.
SELECTED RESEARCH
Representative work, ordered with first-author and co-first-author contributions first.
Publications and Preprints
DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving
Stackelberg Autonomous-Background Vehicle Modeling for Continual Policy Improvement
World-in-Loop: Online Correction via Event-Triggered World Models for Robust VLA Policies
Educations
Master's Student
Department of Automation, Tsinghua University, Beijing, China.
Undergraduate
Department of Automation, Tsinghua University, Beijing, China.