Dyna Robotics has introduced DYNA-2, a robotics foundation model trained on more than 1 million hours of egocentric human video, as the company seeks to reduce the amount of robot-specific training data required to automate physical tasks. The Redwood City, California-based company said DYNA-2 uses a World-Action Model architecture that combines next-frame and next-action prediction. Unlike approaches that rely heavily on robot teleoperation data, the model is pre-trained on human video before being adapted to individual robotic systems.
Dyna Robotics said research underlying the model shows robot performance improving as the amount of human video used in pre-training increases. The company described the results as evidence of a human-to-robot scaling relationship, in which knowledge learned from human motion can transfer across different robotic platforms.
DYNA-2 has been tested on stationary robot arms, humanoid prototypes and dexterous robotic hands. According to the company, the model can be adapted to new tasks using several hours of additional task-specific training data. In one test, a pair of five-fingered robotic hands learned to twist open a bottle cap using 13 minutes of local training data.
Across 15 benchmark tasks, Dyna Robotics said models pre-trained on larger quantities of human video consistently produced better results. In high-precision manufacturing tests, task success rates rose from about 20% to between 80% and 90% as pre-training data increased, while the post-training dataset remained unchanged.
The company also compared DYNA-2 with DYNA-1, its earlier Vision-Language-Action model, using matched training steps and datasets. Both models recorded a 99.9% task success rate in one set of real-world evaluations, while DYNA-2 achieved a customer-quality pass rate 1.55 times that of DYNA-1. In a separate zero-shot customer deployment, DYNA-2 recorded an 87% quality pass rate compared with 46% for DYNA-1, according to the company.
Dyna Robotics said DYNA-2 was also able to recover from physical disturbances during manipulation tasks such as food preparation and workspace clearing without human intervention. Its video co-training method increased scores by 133% on instruction-following tasks compared with the company’s baseline, it said. “Action data is scarce, but video is everywhere,” Dyna Robotics co-founder Jason Ma said. He said the company is seeking to use human video to teach robots spatial reasoning and physical interactions without requiring equivalent amounts of robot-collected data.
Dyna Robotics’ robots using the DYNA-1 model are already deployed in hotels, restaurants and laundromats. The company plans to expand the amount of video used for training toward 10 million hours as it develops subsequent versions of its robotics models.
