A worker assembling a component demonstrates more than the finished task: which hand holds the part, when the other reaches for a tool, and how to recover when something slips. TwelveLabs wants to make those details usable for companies teaching robots to do similar work.
The video AI company today released Pegasus 1.6, adding specialized support for footage recorded from the perspective of the person or machine performing a task.
TwelveLabs is targeting robotics developers and data suppliers that need to convert those recordings into labeled, timestamped material for training, building on video-description and segmentation capabilities available in Pegasus 1.5, its prior generation model.
Both are video-to-text generation and video segmentation models designed specifically to transform raw video into structured, time-based metadata.
Unlike general-purpose multimodal LLMs that stitch together still image frames or perform basic clip-based Q&A, the Pegasus line is engineered for native temporal reasoning — that is, cause and effect in sequence, with changes noted over time — and end-to-end video understanding.
For enterprises, the proposition addresses an unglamorous obstacle to automation: turning recordings of physical work into data a development team can actually use. Companies exploring robotics could begin documenting and organizing selected workflows; teams already training robots could automate parts of their annotation and quality-review processes.
The distinction matters: Pegasus supplies descriptions and structure. Robotics developers still must connect those observations to the movements, sensors and control systems that make a machine perform reliably.
“These models are right now super data hungry,” TwelveLabs co-founder and CEO Jae Lee told VentureBeat in an interview ahead of the announcement.
From video archives to human demonstrations
Lee said TwelveLabs initially focused on media, entertainment, sports and advertising, where customers needed to search large video collections and describe their contents. Its move into robotics applies those capabilities to a different problem: identifying useful examples of physical behavior.
Teleoperation, in which a person remotely controls a robot, produces valuable demonstrations but requires both equipment and trained operators. Scaling that collection process consumes capital and time, Lee said.
“It's like one dude and one robot,” he said.
First-person recordings offer another source of demonstrations. A camera worn by a worker can capture tasks without requiring a robot to perform each one during collection. But that footage still needs interpretation: what action happens, which object is involved, which hand performs it, and when the action begins and ends.
Pegasus 1.5 already divided video into timed segments and returned structured descriptions. TwelveLabs itself identified those capabilities as a foundation for robotics annotations before this release. Pegasus 1.6 adds training and support specifically for first-person, or egocentric, footage; the company says earlier versions handled that category poorly. The novelty is that specialization and the claimed improvements, rather than the basic ability to turn video into structured text. Customers define the categories and fields they need, and the model organizes video around those requirements, according to Lee and the company's technical briefing.
For example, an assembly team could ask for separate segments covering component pickup, positioning and fastening, with descriptions of each hand's role. That is an illustrative workflow, rather than a disclosed customer deployment.
TwelveLabs describes five principal workflows: breaking recordings into labeled actions; generating detailed captions; screening footage for quality; finding unusual events and duplicates; and flagging potentially sensitive material before downstream use. Quality checks can cover obstructed views, unstable cameras and unclear actions.
The release also adds native image analysis and improved identification of people and objects, allowing customers to process still images through the same API as video. Those additions extend its usefulness beyond robotics to existing video-analysis applications.
Understanding an action is only part of teaching it
Lee described the immediate robotics opportunity as training infrastructure.
“Right now, Carl, I think we're still in the go go go train mode,” he said.
He said current deployments tend to tackle limited tasks or collect field data that helps developers improve their models. His longer-term hypothesis is that human video could supply a broad base of behavioral knowledge, with teleoperated demonstrations helping adapt that knowledge to actual robots.
That remains a development strategy, rather than a demonstrated guarantee that more video produces a broadly capable robot.
Pegasus can describe limb movements and provide a rough understanding of trajectories, Lee said, but does not supply all the information required to execute them. Pressure, touch and precise control remain separate problems. Robotics teams must combine video-derived information with other data and develop the systems that translate it into action.
The company's technical attachment describes internal evaluations involving action identification, recognition of the hand performing an action, and understanding actions over time. It supplies no numerical scores or reproducible comparison against competitors. The interview and advance materials also do not identify a robotics customer whose operational results can be independently checked.
Lee offered an indication of processing scale, recalling a run that handled approximately 17 years' worth of first-person footage in roughly 18 hours without a failed video. He did not specify the computing resources or evaluation conditions, making that an anecdotal throughput claim rather than a reproducible benchmark or accuracy result.
Where TwelveLabs competes—and where it could supply other developers
The competitive question spans several parts of the robotics development process. Some companies organize training data; others reason about a robot's surroundings or produce its actions.
Nvidia's Cosmos Curator directly overlaps with the data-preparation problem through filtering, annotation and duplicate removal. Its open tooling gives engineering teams an alternative to purchasing a managed video-understanding service, although they still need to assemble and operate their chosen pipeline.
Encord's integration of Nvidia Cosmos Reason 2 and Embed is another close comparison. Encord hosts the models, generates preliminary video labels for human review, and supports searches based on actions. That places it in competition for the annotation and curation work TwelveLabs is targeting.
Google's Gemini Robotics ER 2 overlaps in interpreting video and tracking task progress, but emphasizes planning and coordinating execution. It hands motor execution to lower-level action models or robotics interfaces.
Meanwhile, Black Forest Labs and mimic's FLUX-mimic uses a video-model foundation to produce robot actions. The companies report testing and deployment at Audi, including industrial manipulation work. It illustrates the kind of downstream system that could consume better training data, although no integration with TwelveLabs is established here.
Offering | Main role | Relevant capabilities | Enterprise buying distinction | Published pricing / cost basis (USD) |
TwelveLabs Pegasus 1.6 | Prepare and describe video for training | First-person action segmentation, detailed captions and configurable outputs | API service; customer retains responsibility for robot training and control | Published rates: $1.75 per video hour; $3 per million image-input tokens; $7.50 per million output tokens. Enterprise: custom. |
Nvidia Cosmos Curator | Build data-processing pipelines | Filtering, annotation and deduplication | Open tooling for teams developing their own infrastructure | Open-source software, with no Curator software license fee; GPU/cloud, storage and engineering costs remain. Model licenses vary. |
Encord with Cosmos Reason 2 and Embed | Manage annotation and curation | Automated preliminary labels, reviewer workflows and behavior-based search | Hosted platform combining models with human review | No public dollar rates; contact Encord for a quote covering the hosted platform and model usage. |
Google Gemini Robotics ER 2 | Reason and coordinate robot tasks | Video progress assessment, planning and tool use | API-accessible reasoning; separate systems execute movements | Standard API rates: $1 per million input tokens and $5 per million output tokens, including thinking, through Dec. 31, 2026; $2/$10 from Jan. 1, 2027. Free tier available. |
BFL/mimic FLUX-mimic | Generate actions for manipulation | Video-derived representations translated into robot behavior | Robotics deployment system, addressing execution beyond data preparation | No public FLUX-mimic price found in the product announcement; deployment terms require inquiry. Not advertised as a free open model. |
Pricing checked October 6, 2026. The billing units and product scope differ, so these rates do not establish the cheapest option for a given workload. Pegasus bills video by duration and text output by tokens; Google bills input and output by tokens. TwelveLabs also multiplies billed video duration by the number of segment definitions in a Segment request. Open-source software still requires paid infrastructure or existing computing resources.
These are functional comparisons based on vendor descriptions, not a shared performance ranking. No evidence reviewed establishes that Pegasus 1.6 beats all of these approaches on labeling quality, cost or eventual robot success.
Lee sees developers of action-producing models as potential beneficiaries: “We kind of sit very nicely in different stacks,” he said.
The commercial test is whether TwelveLabs can make video preparation sufficiently accurate, economical and easy to integrate that those teams prefer buying it to developing or sourcing an alternative.
Pricing—and what enterprises should test first
TwelveLabs has retained roughly the same output-token price for Pegasus 1.6 as its predecessor: $1.75 per hour of input video, $3 per million input-image tokens, and $7.50 per million output tokens.
Enterprise contracts use custom pricing; buyers seeking private or self-managed deployments must contact sales.
At the listed input rate, processing 1,000 hours once would cost $1,750 before output and other applicable charges. Dense annotations add output costs. TwelveLabs confirmed the updated pricing for Pegasus 1.6 with VentureBeat after it initially listed the values as twice as expensive on its public website. The article has been updated with this information and is now accurate.
The pricing FAQ also says each segment definition in a Segment request is billed separately, multiplying the submitted duration—a material consideration for teams requesting several categories of annotation.
Lee said the company may also enter a small number of longer-term licensing and custom research arrangements with larger partners.
For an enterprise early in robotics, a useful starting point would be one bounded workflow, such as packing a particular product or assembling a specific part. A pilot could establish whether the recordings contain enough detail, whether generated labels match expert judgment, and how much correction they require before a robotics partner can use them.
That also means settling permissions for recording workers and handling proprietary processes. Automated detection of sensitive imagery helps review; it does not establish permission to collect or train on the footage.
For teams already experimenting with robots, the stronger test is downstream: does adding Pegasus-labeled footage improve task completion on held-out examples, reduce collection requirements or shorten development cycles? Annotation accuracy, missed action boundaries and human correction time should be measured alongside processing speed.
The relevant economic metric is cost per accepted, useful training example, followed by its contribution to better robot behavior. TwelveLabs offers a way to accelerate the work between filming an activity and learning from it. Enterprises will need their own evidence that the resulting data improves the task they actually want to automate.
