# ShengShu Technology Proposes a Five-Level Roadmap for General World Models

- Link: https://www.thailand-business-news.com/pr-news/shengshu-technology-proposes-a-five-level-roadmap-for-general-world-models
- Published: 2026-08-24T20:00:00+07:00
- Author: PR Newswire

SINGAPORE, Aug. 24, 2026 /PRNewswire/ — At the 2026 World Robot Conference (WRC),
ShengShu Technology unveiled its latest research on General World Models (GWMs),
proposing a five-level development roadmap charting their evolution from world generation
and real-time interaction to physical action, autonomous agents, and world orchestration.

 [⌊Five-Level Roadmap for General World Models⌉](https://mmx.prnasia.com/media/MS1974256/20260824040808EDT_image_1.jpg?id=OA2906026&p=medium600)

Five-Level Roadmap for General World Models

Jun Zhu, Founder and Chief Scientist of ShengShu Technology and an ACM, IEEE and
AAAI Fellow, presented the research during a keynote address at the forum "The Evolution
of Foundation Models for Embodied Intelligence: From Technology Levels to Industrial
Deployment."

"From the perspective of foundation-model development, our goal is not to build 
another specialized model for a particular task or setting, but to create a general
foundation model that can understand the world, imagine possible futures, and take
action," Zhu said.

**Defining the General World Model from First Principles**

World-model research today spans video generation, environment simulation, robot
decision-making and action control. Yet the field still lacks a shared definition
of what makes such a model truly general.

When people learn to ride a bicycle or drive a car, their movements gradually become
stable and precise. One important reason is that the brain develops an "internal
model" through continuous interaction with the external world, allowing it to anticipate
the consequences of an action.

A General World Model similarly requires three interdependent capabilities:

 1. **Understanding the world: **integrating different sources of information to infer
    the current state;
 2. **Imagining possible futures: **predicting what may happen next, including the 
    consequences of different actions;
 3. **Taking action:** influencing a digital or physical environment to realize goals,
    while using real-world feedback to refine subsequent predictions and decisions.

"A General World Model is not merely a generator, simulator, robot action model 
or policy model, nor is it a collection of isolated capabilities or a simple linear
pipeline," Zhu said. "It is a closed-loop feedback system in which **Understanding,
Imagination, and Action** are tightly connected."

Action is not merely the output of the model. It changes the environment, produces
new information, and feeds into the next cycle of understanding, imagination and
decision-making.

**Data, Architecture and Compute: Three Pillars of General World Models**

Building a General World Model requires three fundamental elements of foundation-
model development: **data, architecture and compute**.

On the data side, GWMs require a multi-layer data pyramid that moves progressively
from observation toward action: **web-scale video, curated domain/instructional 
video, egocentric human video, action-recording human demonstrations, and real-robot
interaction**.

Lower layers provide greater scale and broader coverage, while higher layers are
scarcer and more costly but more directly connect tasks, actions and physical outcomes.
Synthetic data can augment every layer, while "imperfect" data—including failed 
attempts, corrective actions and recovery trajectories—provides valuable learning
signals for recovering from failure.

On architecture, a GWM must process images, video, language and robot actions within
a unified framework. **Mixture-of-Transformers (MoT)** assigns modality-specific
parameters while retaining shared attention for cross-modal interaction, enabling
environment understanding, world-state prediction and action generation to work 
together within the same model.

ShengShu Technology’s World Action Model **Motubrain** is built on this architecture,
jointly modeling understanding, prediction and action to improve the utilization
of heterogeneous data and strengthen transfer across tasks.

Compute affects both large-scale pre-training and real-time deployment. While offline
video generation can tolerate some latency, interactive generation and robot control
must predict and decide before the environment changes. Training infrastructure,
inference acceleration, model distillation and efficient attention are therefore
critical to real-world deployment.

**How Do General World Models Evolve? A Five-Level Roadmap**

GWMs can be divided into five progressively evolving levels. Rather than separate
product directions, they represent a continuous progression toward deeper world 
understanding, stronger interaction and greater autonomy. In the past several years,
ShengShu Technology’s research and deployment efforts have already covered the first
three levels.

**L1: World Generation**

The first step is generating coherent world trajectories. Video helps models learn
objects, motion, spatial relationships and temporal dynamics, laying the foundation
for deeper understanding and imagination.

In 2024, ShengShu Technology launched the video generation foundation model **Vidu**.
Through continued improvements in video quality, temporal consistency and world-
dynamics modeling, Vidu represents the company’s initial implementation of L1.

**L2: Interactive World**

Building on generation, the model receives real-time input and continuously changes
what happens next in response to language, speech or control signals, moving from
one-off generation toward continuous interaction.

Released in July 2026, **Vidu S1** advances this capability into real-time interaction.
Users can intervene through speech and continuously influence what happens next,
allowing the generated world to evolve dynamically in response to external input.

**L3: Actionable World**

At L3, the model moves from the digital into the physical world. Environment understanding,
future prediction and action generation work together, enabling the model to produce
executable robot actions and refine its decisions based on real-world feedback.

**Motus and Motubrain** mark ShengShu Technology’s progression from video generation
toward physical action. Motus was released and fully open-sourced in December 2025.**
Motubrain**, released in April 2026, further unifies environment understanding, 
world-state prediction and action-trajectory generation within a single model, advancing
the GWM from visual simulation toward physical decision-making.

Compared with Motus, Motubrain delivers approximately **10× faster inference** and
can adapt to new robot embodiments with **50–100 human demonstrations**. On the **
RoboTwin 2.0** benchmark, it achieved a score of **96.1, ranking first on the leaderboard**.

Motubrain has also been validated across multiple robot embodiments, including robots
from Galaxea AI, demonstrating generalization across long-horizon tasks, multi-task
scenarios and different robot embodiments.

**L4: Autonomous World Agent**

At L4, the model moves toward autonomous decision-making, proactively perceiving
its environment, decomposing tasks, exploring unknown states and continuously planning
actions around long-term goals.

**L5: World Orchestrator**

At L5, autonomy extends beyond a single agent. The model coordinates robots, digital
agents, humans, tools and other resources, enabling complex task allocation and 
multi-agent collaboration.

From L1 to L5, capabilities build progressively. Without learning and predicting
how the world evolves, stable interaction is difficult; without feedback from real-
world action, it is difficult to develop autonomous learning, long-term planning
and complex coordination.

**Toward L4 and L5: Six Key Challenges Ahead**

While L1 through L3 have already seen concrete technical implementations, significant
gaps remain before higher levels of world intelligence can be achieved. Video models
must continue to improve generation quality, controllability and spatiotemporal 
consistency, while World Action Models need greater stability, generalization and
execution efficiency in complex, open-ended environments.

Advancing toward L4 and L5 will require progress in six key areas:

 ◦ joint evaluation across understanding, imagination, action and transfer;
 ◦ learning physical dynamics such as contact, friction, force and the consequences
   of intervention;
 ◦ persistent and revisable memory;
 ◦ online learning and self-improvement through real-world interaction;
 ◦ efficient closed-loop deployment under latency and resource constraints; and
 ◦ stronger safety and controllability.

Moving forward, ShengShu Technology will continue to strengthen World Generation,
Interactive World and Actionable World, while developing goal formation, active 
exploration, continual learning and long-term memory to lay the foundation for higher-
level Autonomous World Agents.

The longer-term goal is to build a general foundation model capable of continuously
understanding the world, imagining the consequences of different actions, acting
autonomously under explicit constraints, and learning and evolving from real-world
feedback.

When Understanding, Imagination and Action form a true closed loop, world models
can evolve beyond content generation or task-specific robot policies to become a
foundation connecting the digital and physical worlds—and an important step toward
general embodied intelligence.

For more information on the General World Model research, visit: [https://www.shengshu.com/en/general-world-model/](https://www.shengshu.com/en/general-world-model/)

---

  |  This article was produced by Cision PR Newswire, our trusted news partner. The views expressed and the content presented here are solely those of the author and may not fully reflect the opinions of Thailand Business News. |

---

 
**Read the original article :** [ShengShu Technology Proposes a Five-Level Roadmap for General World Models ](http://www.prnasia.com/story/archive/5032067_CN32067_0?rand=184833)
