Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
Abstract
We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trained on connected multi-prompt videos to supervise a block-level causal student through distribution-matching distillation; and (3) Low-cost real-time interaction, combining four-step generation with context-preserving streaming to support 832 x 480 inference at 24 FPS at an estimated server rental cost of approximately USD 0.009 per stream-minute. Zing-0.5 achieves an overall score of 81.0 and a consistency score of 88.5 across 158 WBench Navigation cases. A joint-control demonstration shows a text-directed event change during continued navigation without restarting generation. We release the model weights, inference code, and Zing-SGLang serving implementation to support further work on playable generated worlds.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models (2026)
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU (2026)
- EchoWM: Open and Enterable Omnimodal World Models (2026)
- Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds (2026)
- Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation (2026)
- AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (2026)
- Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.17909 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 2
Collections including this paper 0
No Collection including this paper