LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows Paper • 2609.15863 • Published 6 days ago • 52
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 12 days ago • 162
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Paper • 2606.19338 • Published Jun 17 • 52
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion Paper • 2605.30265 • Published May 28 • 22
DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes Paper • 2605.28421 • Published May 27 • 46
ACC: Compiling Agent Trajectories for Long-Context Training Paper • 2605.21850 • Published May 21 • 62
Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning Paper • 2604.05404 • Published Apr 7 • 45
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model Paper • 2603.21986 • Published Mar 23 • 125