Eigen RadarAI
Analysis

Reka puts text, video and robot actions inside one Rho-1 network

Reka has released a research preview of Rho-1, a model that handles text, images, video and robot actions through one shared context. Demonstrations move from creating and editing a scene to generating control signals in a robot simulation. The company describes a proof of concept, with long-video scene consistency and targeted editing still limited. The preview does not establish reliable deployment on physical robots.

Artificial Intelligence··Evening
A monitor shows a robot arm reaching toward a white cube, with a small lighthouse model on the table beside it.

Rho-1 shares context across different data types

AI developer Reka released the Rho-1 research preview on 5 October. It combines text, images, video and robot-control actions in one neural network and one shared context window. The architecture shares attention and state across understanding and generation. Text and symbolic commands use discrete tokens, the units a model processes; images, video frames and physical actions use continuous representations.[1], [2]

A scene moves from image creation into video editing

In one company demonstration, Rho-1 creates a lighthouse image, marks the object with a bounding box, turns the image into video and changes the scene into a snowstorm. Other demonstrations insert directional commands into an ongoing video stream. These examples present operations within the shared model context. They are company demonstrations, without independent measurements of reliability across other scenes or user requests.[1]

Robot actions remain a simulation demonstration

The robotics example uses LIBERO, a simulation environment for robot tasks. Rho-1 receives observations and wrist-camera images and emits seven action channels. Reka also describes an inverse dynamics model that infers control signals from internet video. The company presents this stage as a proof of concept. It acknowledges drifting room layouts in long video streams, limited object grounding across video and brittle targeted edits across prompts.[1]

References

  1. News sourceRekaReka previews Rho-1’s unified text, vision and robot-action model↩1↩2↩3
  2. News sourceThe DecoderRho-1 processes text, images, video and robot control in one network↩