Submitted by Yinheng Li 13 LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them Microsoft 2
Submitted by Sungho Park 82 ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization Microsoft 2
Submitted by Wenbo Pan 164 The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Microsoft 33 6
Submitted by Kaixiang Zhao 114 When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Microsoft 7 5
Submitted by Sungho Park 66 AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Microsoft 237 2