Picture for Xuewei Yang

Xuewei Yang

Simple-OPD: Demystifying Warm-up for On-policy Distillation

Add code
Aug 07, 2026
Viaarxiv icon

Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning

Add code
May 30, 2026
Viaarxiv icon

S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models

Add code
Sep 26, 2025
Figure 1 for S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
Figure 2 for S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
Figure 3 for S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
Figure 4 for S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
Viaarxiv icon