M.S. student in Intelligent Technology
Macau University of Science and Technology
I work on GUI / browser agents — how to rigorously evaluate them on real websites, and how to train them with reinforcement learning and distillation. I am the first author of CAP-Bench (COLM 2026), a benchmark of 420 cross-site tasks over 108 real-world websites with a verifiable agent-as-a-judge evaluation framework.
Previously, I was a research assistant at TsinghuaNLP & OpenBMB (advised by Prof. Zhiyuan Liu) working on MiniCPM fine-tuning, AgentCPM, and RL for agents, and an algorithm & benchmark engineer at Fellou AI. I also build and open-source agent infrastructure. I am seeking Ph.D. opportunities (Fall 2027) in agents, VLMs, and RL.