Submitted by JeonghyeKim 37 Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization Microsoft 3
Submitted by Sayan Deb Sarkar 32 CoPE-VideoLM: Codec Primitives For Efficient Video Language Models Microsoft 2
Submitted by Zijie Chen 44 Improving Data and Reward Design for Scientific Reasoning in Large Language Models Microsoft 2
Submitted by Mingqian Feng 6 SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks Microsoft 11 2
Submitted by Jialiang Zhu 20 RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents Microsoft 80 1
Submitted by Xiao Liu 34 MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration Microsoft 3
Submitted by Minh-Quan Le 24 PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards Microsoft 2
Submitted by Mingqian Feng 22 Statistical Estimation of Adversarial Risk in Large Language Models under Best-of-N Sampling Microsoft 3