TwelveLabs adds the worker’s viewpoint to Pegasus video analysis
TwelveLabs has released Pegasus 1.6 with support for footage captured from a person’s or robot’s viewpoint. The model converts video into structured descriptions and time-based labels, including material prepared for robotics work. The update also adds still-image analysis through the existing interface and changes entity recognition. Developers can use the same model and prompts for both photographs and video.
Artificial Intelligence··Night
Pegasus processes footage from the person doing the task
TwelveLabs, a company developing video-understanding models, released Pegasus 1.6 on October 6 with support for first-person footage. That means a camera captures the viewpoint of the person or machine carrying out an activity. Pegasus turns video into structured text and time-based labels. The announced robotics use concerns preparing and reviewing that footage for development teams.[1], [2]
Body cameras bring a different viewpoint
Earlier Pegasus uses largely involved a camera looking at a scene from outside, such as security footage or film recordings. TwelveLabs names body cameras, cameras mounted on robots and remotely operated systems among the new examples. The company also describes improved identification of people and objects. These capabilities are presented for labeling action recordings, cataloguing media and reviewing security footage, with the improvements attributed to the developer.[1]
Still images use the existing video interface
Developers can now submit still images through the same interface and model used for video, according to TwelveLabs. Its examples put product photographs beside product videos, or screenshots beside screen recordings. Existing users retain the same development tools and prompts. The update expands the inputs available to applications and the viewpoint of the footage they process. The announcement provides no independent study establishing that those labels improve a robot’s ability to complete a task.[1]