Verifier-Induced Support Reshaping in On-Policy Optimization Paper • 2608.00220 • Published 23 days ago • 5
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs Paper • 2603.22446 • Published Mar 23 • 11
Verifier-Induced Support Reshaping in On-Policy Optimization Paper • 2608.00220 • Published 23 days ago • 5
Verifier-Induced Support Reshaping in On-Policy Optimization Paper • 2608.00220 • Published 23 days ago • 5
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published 17 days ago • 56
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published 24 days ago • 73
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published 20 days ago • 142
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining Paper • 2605.14747 • Published May 14 • 147
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models Paper • 2604.10866 • Published Apr 13 • 69