UXO UXO Family: Open-source unified customization model bytedance-research/UNO Image-to-Image • Updated Aug 22, 2025 • 182 bytedance-research/USO Text-to-Image • Updated Sep 1, 2025 • 228 • 191 bytedance-research/UMO Text-to-Image • Updated Sep 12, 2025 • 406 • 61
Vidi Vidi model collection for multimodal video understanding and creation bytedance-research/Vidi-7B 9B • Updated Dec 15, 2025 • 19 • 18 bytedance-research/Vidi1.5-9B 10B • Updated Jan 22 • 24 • 12
Valley Valley Family: Exploring Scalable Vision-Language Design for Multimodal Understanding and Reasoning bytedance-research/Valley3-8B-Instruct 10B • Updated May 25 • 23 • 4 bytedance-research/Valley3-32B-Instruct 34B • Updated May 25 • 18 • 4 bytedance-research/Valley3-8B-Think 10B • Updated May 25 • 15 • 8 bytedance-research/Valley3-32B-Think 34B • Updated May 25 • 13 • 2
UXO UXO Family: Open-source unified customization model bytedance-research/UNO Image-to-Image • Updated Aug 22, 2025 • 182 bytedance-research/USO Text-to-Image • Updated Sep 1, 2025 • 228 • 191 bytedance-research/UMO Text-to-Image • Updated Sep 12, 2025 • 406 • 61
Valley Valley Family: Exploring Scalable Vision-Language Design for Multimodal Understanding and Reasoning bytedance-research/Valley3-8B-Instruct 10B • Updated May 25 • 23 • 4 bytedance-research/Valley3-32B-Instruct 34B • Updated May 25 • 18 • 4 bytedance-research/Valley3-8B-Think 10B • Updated May 25 • 15 • 8 bytedance-research/Valley3-32B-Think 34B • Updated May 25 • 13 • 2
Vidi Vidi model collection for multimodal video understanding and creation bytedance-research/Vidi-7B 9B • Updated Dec 15, 2025 • 19 • 18 bytedance-research/Vidi1.5-9B 10B • Updated Jan 22 • 24 • 12