TwelveLabs debuts Pegasus 1.6 to turn first-person video into robot training data

1 hour ago 1



Teaching a robot to fold laundry starts with someone watching hours of footage of humans folding laundry. Then that person writes down every move. TwelveLabs wants to take that second job off the table. On October 6, 2026, the company released Pegasus 1.6, a video intelligence model with native support for egocentric video. That means footage shot from a first-person point of view, built into the model’s core design. The target audience is clear. Robotics and physical AI teams have plenty of raw footage, and they increasingly need ways to turn it into usable training data faster. What Pegasus 1.6 actually does Egocentric video is the kind captured by wearable cameras, teleoperated systems and similar sources. Picture a GoPro strapped to a worker’s head, or the camera feed from a human remotely steering a robotic arm. Pegasus 1.6 is built to handle it. The model automates several workflows that robotics teams typically grind through by hand: Action segmentation and labeling. The model breaks a video into discrete steps and tags what is happening in each one. Dense captioning of spatial interactions. Rather than a one-line summary, the model generates detailed descriptions of how han...

Read Entire Article