ShengShu Technology Unveils Motus2, a Self-Evolving General World Model for Dexterous Manipulation

PR Newswire
Today at 10:00am UTC

ShengShu Technology Unveils Motus2, a Self-Evolving General World Model for Dexterous Manipulation

PR Newswire

SINGAPORE, Sept. 14, 2026 /PRNewswire/ -- ShengShu Technology Co-founder and CEO Yihang Luo unveiled Motus2, a self-evolving general world model for robotic dexterous manipulation, at the 2026 Inclusion Conference on the Bund on September 10.

Motus2 unifies action generation, consequence prediction, and outcome evaluation within a single video-action model with shared parameters. By incorporating model-based reinforcement learning, it turns predictions and evaluations of action outcomes into signals for policy improvement, exploring a path toward continually improving robotic manipulation.

According to the team's research paper, Motus2 achieved an average success rate of 84% across five primary real-robot tasks: placing a ball, multi-finger manipulation, attaching an eraser, screwing in a light bulb, and placing a phone. In a separate policy-optimization study involving phone placement and multi-finger manipulation, combining model-based reinforcement learning with inference-time planning increased the average success rate from 65% to 75%, a gain of 10 percentage points. These real-robot evaluations used 20 trials per task for each method.

Dexterous robotic manipulation requires coordinated handling of spatial relationships, finger movements, sustained contact, and historical information. To address these challenges, the Motus2 team has released real-robot demonstrations of tasks including screwing in a light bulb, tearing paper with both hands, turning pages, opening a beverage can, and finding a hidden object. The model has been evaluated on dual-arm robot platforms equipped with WUJI, WUJI Hand 2, and Sharpa Wave dexterous hands.

Motus2's central advance is bringing the predictive capabilities of a world model into robotic decision-making and learning. Within the unified model, a policy interface generates candidate actions based on language instructions, robot states, and visual history. A simulator interface predicts the future visual states those actions may produce, while an evaluator interface assesses whether the predicted outcomes would advance the task. Together, these capabilities form a closed loop: generate actions, predict consequences, evaluate outcomes, and update the policy.

This loop operates differently during execution and training. During execution, Motus2 uses Best-of-N planning to generate multiple candidate actions, predict and compare their outcomes, and execute the higher-scoring option. After each action chunk, the robot incorporates fresh observations from the real world and plans its next step.

During training, the model uses value scores assigned to candidate actions to update its policy, making it more likely to generate actions that support task completion. This process updates only action-related parameters, while the prediction and evaluation components remain frozen. Successful demonstrations provide targets for action learning, while state changes and outcomes from failed and suboptimal interactions help the model learn action consequences and evaluate task progress. In this context, "self-evolution" refers to a concrete mechanism for improving the policy through model predictions and value feedback.

These capabilities are supported by a combination of large-scale human manipulation experience and robot data. Motus2 draws on an egocentric human dataset comprising approximately 130,000 hours of raw recordings from monocular and stereo sources. The model first learns broad patterns of object changes and manipulation from monocular video, then uses stereo video and human action data to learn spatial relationships and hand-object interactions.

Robot-domain mid-training uses more than 100 hours of robot trajectories and supplementary human-robot alignment data, followed by fine-tuning for target tasks. Under the same target-task fine-tuning procedure, the model pretrained only on egocentric human data achieved an average success rate of 51% across the five primary tasks. Adding robot-domain mid-training increased that figure to 84%. The results highlight the complementary roles of human interaction experience and robot-domain adaptation: the former provides broad manipulation priors, while the latter adapts that experience to the robot's observation and control spaces.

To address task states that cannot be fully understood from a single visual frame, Motus2 also investigates mechanisms for retaining historical information and incorporates a lightweight tactile expert. Historical context helps the model use earlier task cues when dealing with occlusion and multi-step manipulation. The tactile expert reads the latest contact feedback immediately before a short action segment is executed and refines the action accordingly.

Across two real-robot tasks—pulling out a paper cup and tearing paper—adding the tactile expert increased the average success rate from 60% to 72.5%, a gain of 12.5 percentage points. The tactile module reuses intermediate computations from the main model, avoiding the need to rerun the full video backbone for every action refinement and providing contact-sensitive manipulation with feedback beyond vision.

Within ShengShu Technology's five-level roadmap for general world models, Motus2 represents a concrete implementation of L3, "Acting in the World." By incorporating evaluation, feedback, and policy updates, it also explores a technical path toward L4, "Autonomous World Agents." The current research demonstrates policy improvement in specific tasks. Reliable decision-making over longer tasks, retention of important historical information, and long-term autonomous learning in open environments remain areas for further research.

Motus2's model architecture, research paper, and real-robot demonstrations are publicly available. For details, visit the Motus2 project website.

Founded in 2023, ShengShu Technology focuses on the research and development of general world models. The company is committed to building general intelligence systems capable of understanding, predicting, and acting in both digital and physical worlds. For more information, visit the ShengShu Technology website.

Cision View original content to download multimedia:https://www.prnewswire.com/news-releases/shengshu-technology-unveils-motus2-a-self-evolving-general-world-model-for-dexterous-manipulation-302877367.html

SOURCE ShengShu Technology