Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering Paper • 2608.30468 • Published 6 days ago • 34
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Paper • 2608.25529 • Published 11 days ago • 17
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 5 days ago • 94
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published 10 days ago • 91
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 13 days ago • 205
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published 17 days ago • 12