ByteDance prepares a real-time spatial video model under Zhang Yiming's watch
ByteDance is preparing a real-time spatial video model with founder Zhang Yiming watching it directly. The model aims to place the viewer inside a three-dimensional scene, and a company spokesperson did not answer Bloomberg's request for comment. Sources caution that the timing is not settled and plans may change.
Artificial Intelligence··Evening
Zhang Yiming is personally overseeing the real-time spatial video model
ByteDance is preparing an AI model for real-time spatial video generation. Founder Zhang Yiming is personally overseeing the work. People familiar with the matter said the timing of a launch is not certain and plans may change. The preparation sits on the company's video generation line.[1], [2]
The model is built on Seedance and tied to Pico headsets
The model is built on Seedance, ByteDance's video generation system. Seedance is the company's existing system for cinematic video. It is expected to produce scenes that respond to the voice and movement of people using ByteDance's Pico headsets. The reported aim is interactive virtual worlds for live streams, short-form dramas and games. Pico headsets are the hardware side of those scenes.[1], [2]
The reported target is 20 frames per second and about 0.05 seconds of latency
The reported targets are 20 frames per second and latency of roughly 0.05 seconds, with the work done on cloud servers rather than on the headset. A company spokesperson did not answer Bloomberg's request for comment. Internal measurements are reported to put ByteDance about 10 percent behind the global state of the art on world models. The company allocated an AI data budget three to four times the size of its rivals' for this work. The model aims to place the viewer inside a three-dimensional scene.[1]