NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 28 days ago • 23
Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts Paper • 2608.00574 • Published Aug 1 • 8
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published Jul 13 • 80
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models Paper • 2606.19297 • Published Jun 17 • 80
AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models Paper • 2607.02269 • Published Jul 2 • 9
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding Paper • 2605.29707 • Published May 28 • 152
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 253
stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_surplexity_t1_g5_run1_metrics Viewer • Updated Jun 2 • 164 • 14 • 1