Pegasus 1.6, TwelveLabs' first egocentric video model, reads first person footage from a worker mounted camera to extract the grip, slip recovery, and tool interaction cues today's robots still miss.
Robots are getting smarter at planning, language, and even basic manipulation. The footage meant to teach them physical work is still a bottleneck, and a growing set of vendors is now competing to turn first-person video into structured training data.
The latest entry: TwelveLabs, a Seoul-based video AI company, on Tuesday released Pegasus 1.6, its first model built for egocentric video, meaning clips shot from a worker's point of view where the camera sees what the hands see. The company positions the model as an answer to a specific gap in robotics, drone, and autonomous-vehicle data pipelines: most human physical skill, from grip changes to recovery after a slip, has never existed in a form a machine can learn from.
TwelveLabs CEO Jae Lee, in a company statement reported by VentureBeat, frames video understanding as a precondition for useful robotics training. The model's targets include extracting what action is happening, when it happens, what the person is interacting with, and how behavior changes over time.
The release is one vendor's claim, not a deployed robotics outcome. No independent benchmarks or named deployment partners are visible in the launch coverage, and egocentric footage is a narrow input: a worker-facing camera only sees what the worker sees. The release signals that the data layer underneath the robotics boom is becoming its own market.