Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR Paper • 2609.08650 • Published 7 days ago • 12
1010happy/BALANCED_Teacher_r14_train_gptmini-gemma-3-4b-it-seed1010 Image-Text-to-Text • 4B • Updated Aug 13 • 9 • 1
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Paper • 2607.25659 • Published Jul 28 • 85
timm/mobilenetv3_small_100.lamb_in1k Image Classification • 2.55M • Updated Oct 19, 2025 • 17.2M • 121