TwelveLabs debuts Pegasus 1.6 to enhance robotics training data

Summary

TwelveLabs has launched Pegasus 1.6, a new version of its video AI software designed to enhance robotics training data derived from first-person video. This update targets robotics developers and data suppliers, enabling them to convert raw video footage into labeled, timestamped materials, which is crucial for training robots without needing to manually perform every task. The new version builds upon the capabilities of Pegasus 1.5, which already offered video-description and segmentation features, by adding specialized support for egocentric footage. This approach aims to streamline the data preparation process for enterprises by allowing them to document and categorize workflows, while competing tools from companies like Nvidia and Google also cater to similar needs in video annotation and robotics development.

Analysis

Jae Lee: Jae Lee is co-founder and CEO of TwelveLabs. He provided details on the company's shift toward robotics applications and the technical improvements in Pegasus 1.6 during an interview ahead of the release. Lee emphasized the challenges of scaling teleoperation for robot training data and the value of first-person recordings as an alternative source. TwelveLabs: TwelveLabs is a video AI company that develops models for transforming raw video into structured, time-based metadata through description and segmentation. The company initially targeted media, entertainment, sports and advertising use cases before expanding into robotics. In this announcement, TwelveLabs released Pegasus 1.6 to address first-person footage specifically for robotics developers and data suppliers seeking to convert human demonstrations into usable training data. Pegasus 1.6: Pegasus 1.6 is a video-to-text generation and segmentation model from TwelveLabs engineered for native temporal reasoning and end-to-end video understanding. This version adds specialized support for egocentric, first-person footage that earlier models handled poorly. It enables robotics teams to break recordings into labeled actions, generate detailed captions, screen for quality and identify events while also supporting native image analysis. Enterprise Workflows: Companies can use the model to document and organize selected workflows, automate parts of annotation and quality review, and define custom categories for labeling actions in assembly or similar tasks. Competitive Landscape: Pegasus 1.6 operates in a space that overlaps with tools from Nvidia, Encord and Google for video annotation, curation and task interpretation in robotics development. Robotics Data Preparation: TwelveLabs positions its models as training infrastructure that helps convert recordings of physical work into structured data without requiring robots to perform every demonstration during collection.

Categories

techmachine_learningaiai_agents
View Original Tweet