Skild AI Says Its Robot Model Handles Unfamiliar Tasks From a Single Video, No Retraining
The company reports a 66% per-step success rate on tasks it says fell outside its training data, measured in an internal benchmark that used human help to recover from failures.

Skild AI said its robotics foundation model S1 performed physical tasks it had not been trained on after being shown one video demonstration, with no fine-tuning and no change to the model's weights.
The results were announced Aug. 25 and cover four tasks the company says were absent from S1's pre-training data: potting a plant, cooking pancakes, making pour-over coffee, and assembling a kit. That absence rests on Skild's own search of its own dataset. The tasks run up to 10 minutes and span dozens of manipulation steps.
Most robot learning still requires hours of teleoperation data and a separate training run for each new job. Skild claims S1 instead takes the video as context and must infer the demonstrator's intent and task progress in order to map the work onto its own body and scene. The model is built on NVIDIA computing infrastructure, according to the company.
For the plant potting task, Skild published a timeline running 11 minutes from the start of recording to the robot executing on its own. By that same timeline the demonstration itself was finished five minutes before execution began. All results are company-reported, and no independent evaluation was identified.
On an internal suite of unseen tasks, Skild reported a 66% average per-step success rate from one demonstration, against 9% for a language-prompted vision-language-action model trained on the same 100,000 hours of data. Skild says human intervention was used to recover from failures during rollouts so that every step could be graded.
The figure therefore measures per-step performance under supervision, not how often S1 completes a task unaided.
On tasks drawn from the pre-training distribution, at the smallest training scale the company tested, Skild reported the language-prompted policy ahead of its own in-context model, 53% to 43%. The gap reversed as training data was scaled.
Skild reported that post-training the comparison policy on 2,000 demonstrations of a new task reached 86%, above the 66% S1 reached from a single example. The company estimates one demonstration in context is worth roughly 380 post-training examples, a figure it says it arrived at by interpolating between measured points.
The company also published where its own method degrades. In a scored study varying how far the working scene diverged from the demonstration it had been given, S1's success rate fell from 96% to 46% across five levels of difficulty, with the steepest drop when the scene required a different execution plan than the one demonstrated. In a companion study measuring distance from training conditions rather than from the prompt, Skild reported the language-prompted policy degrading up to three times as much as S1.
Skild raised close to $1.4 billion in January in a round led by SoftBank, at a valuation the company put above $14 billion, more than triple its valuation seven months earlier.
Skild says S1 is already at work with commercial partners. It has not named them, described the work or said whether the deployments are paid.
NEWS BRIEF
Skild AI's Robot Learned to Make Coffee From a Single Video; No Retraining Needed
Skild AI said Aug. 25 that its S1 foundation model performed tasks it says were outside its pre-training data after a single video demonstration, without fine-tuning. The company listed plant potting, pancake cooking, pour-over coffee and kit assembly, tasks running up to 10 minutes.
On an internal benchmark of unseen tasks, Skild reported a 66% average per-step success rate from one demonstration, against 9% for a language-prompted model trained on the same data volume. Human intervention was used to recover from failures during the evaluation, so the figure does not measure unaided task completion.
The company's own results carry three qualifications. On familiar tasks at the smallest training scale tested, the language-prompted policy led 53% to 43%. Conventional post-training on 2,000 demonstrations of a new task reached 86%. And S1's success rate fell to 46% when the working scene diverged furthest from the demonstration it was given. Skild says S1 is at work with unnamed commercial partners.
More articles on this topic
11 articles
ReportYou Could Soon Spot Robots On The Streets In Japan
ReportThis Robot Can Knock Ice Off From Power Lines
ReportRenesas Opens Beijing Lab to Test Humanoid Robots’ “Nervous Systems”
ReportBeyondMimic Lets a Humanoid Reuse Human Motions for New Tasks
ReportNew actuator technology could make humanoids lighter
ReportOpen-source robot AI models gain momentum
ReportSoft robotics breakthrough for industrial applications
ReportMIT shows tactile skin for robot fingertips
ReportNew battery pack doubles humanoid runtime
ReportVision-language models cut robot training time
ReportSelf-learning robot cells move closer to productionBUSINESS NEWS WEEKLY LETTER
The Weekly Letter for Robotics Professionals, Summarizing the Most Important Industry Moves, Launches, Deals and Signals.