view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 2 days ago • 47
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 2 days ago • 45
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion Paper • 2608.26794 • Published 9 days ago • 16
Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models Paper • 2608.23478 • Published 12 days ago • 30
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 9 days ago • 36
BVD: Big Video Dataset Collection A 10-Million-Hour Open Video Dataset for Multimodal Pre-training • 11 items • Updated 10 days ago • 7
EgoSuite-Open100K Collection The largest fully-annotated open egocentric human dataset. 100,000 hours across 15,000+ tasks and scenes. • 3 items • Updated 17 days ago • 51
view article Article Skill Evolving: What Happens When Agents Learn From Each Other AmberLJC • Feb 13 • 1
TPIPS Collection This is the collection of models and data for our paper "The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric" • 7 items • Updated Aug 1 • 1
NaFlexCLAP Collection Experiments in OpenCLIP (main branch) with `timm` NaFlexViT audio encoder + modern text encoder for variable-time , variable-length text CLAP models • 4 items • Updated 26 days ago • 3
LTX-2.5 Collection LTX-2.5 base models, quantized models and accompanying LoRAs and IC-LoRAs • 5 items • Updated 4 days ago • 56
You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences Paper • 2606.15956 • Published Jun 14 • 13
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • Jul 27 • 486