IntelligenceLab/Long-Horizon-Terminal-Bench Benchmark • Updated 14 minutes ago • 154 • 23k • 139
Self-Supervised Scaling of Terminal Environments for Scientific Domains Paper • 2610.02710 • Published 10 days ago • 13
Self-Supervised Scaling of Terminal Environments for Scientific Domains Paper • 2610.02710 • Published 10 days ago • 13
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite Paper • 2610.02826 • Published 10 days ago • 103
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 17 days ago • 145
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 17 days ago • 145
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design Paper • 2609.22086 • Published 24 days ago • 34
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges Paper • 2501.02189 • Published Jan 4, 2025 • 1