Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
Entropy-curve control for stable long-horizon RLVR training.
I am a PhD candidate in the Department of Computer Science at Purdue University, advised by Ruqi Zhang. I completed my BE in computer science and technology at Tianjin University, advised by Changqing Zhang.
My research develops statistical foundations for stable and efficient post-training frameworks. I currently focus on improving the performance boundary of LLM RL, including RL entropy control, and agentic MoE RL. More broadly, I work on test-time scaling, preference alignment, multimodal LLM safety, and Bayesian/statistical methods for robust ML.
Entropy-curve control for stable long-horizon RLVR training.
Speculative decoding reframed as efficient reward-guided alignment.
Segment-level rejection sampling for faster aligned generation.
Decision-theoretic utilities for reliable long-tailed predictions.
A flatness-aware sampling method for Bayesian deep learning.
Uncertainty-aware expert engagement for long-tailed recognition.
Community-aware contrastive learning for graph representation.