Reka puts text, video and robot actions inside one Rho-1 network
Reka has released a research preview of Rho-1, a model that handles text, images, video and robot actions through one shared context. Demonstrations move from creating and editing a scene to generating control signals in a robot simulation. The company describes a proof of concept, with long-video scene consistency and targeted editing still limited. The preview does not establish reliable deployment on physical robots.
Artificial Intelligence··Evening
Rho-1 shares context across different data types
AI developer Reka released the Rho-1 research preview on 5 October. It combines text, images, video and robot-control actions in one neural network and one shared context window. The architecture shares attention and state across understanding and generation. Text and symbolic commands use discrete tokens, the units a model processes; images, video frames and physical actions use continuous representations.[1], [2]
A scene moves from image creation into video editing
In one company demonstration, Rho-1 creates a lighthouse image, marks the object with a bounding box, turns the image into video and changes the scene into a snowstorm. Other demonstrations insert directional commands into an ongoing video stream. These examples present operations within the shared model context. They are company demonstrations, without independent measurements of reliability across other scenes or user requests.[1]
Robot actions remain a simulation demonstration
The robotics example uses LIBERO, a simulation environment for robot tasks. Rho-1 receives observations and wrist-camera images and emits seven action channels. Reka also describes an inverse dynamics model that infers control signals from internet video. The company presents this stage as a proof of concept. It acknowledges drifting room layouts in long video streams, limited object grounding across video and brittle targeted edits across prompts.[1]