Reasoning Models Struggle to Control their Chains of Thought Paper • 2603.05706 • Published Mar 5 • 41
STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving Paper • 2502.00212 • Published Jan 31, 2025 • 4
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published Jul 30 • 39
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published Jul 30 • 34
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published Jul 30 • 52
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 62