Artificial intelligence (AI) continues to advance at a rapid pace, but we still face significant challenges indeveloping systems with human-like capabilities.

One of the most promising concepts to overcome these limitations are world simulators (or “world models“, as they are known in English).

Next, we will explore what these world models or world simulators are, why they are so complex to develop and how they could transform technology and our lives.

What are world simulators?

A world simulator is an AI model designed to create an internal representation of your physical and social environment.

Unlike large language models (LLMs), which work by predicting the next word in a sequence, world simulators try to imitate how humans perceive, plan and act in the real world.

These models are capable of simulating environments, building a virtual representation of the world and predicting how events will unfold according to a series of physical rules.

In addition, they can plan actions, generating strategies based on their understanding of the environment and specific objectives. This allows them to adapt to uncertainty, calculating multiple possible outcomes and choosing the most appropriate course of action.

In essence, a world simulator works like us

To better understand what a world model is, a practical example. Imagine an artificial intelligence system in charge of managing a product warehouse. Its objective is to optimize the efficiency of the warehouse and keep it organized.

Instead of simply reacting to every event that occurs, this system uses a world simulator to predict and plan its actions.

The world simulator has an internal representation of the warehouse, including the layout of the shelves, the stock of products and the routes of the robots that transport goods.

One day, the system receives notification that a large batch of products is going to arrive. Instead of waiting for products to arrive to decide where to place them, the world simulator evaluates the situation beforehand.

Plan creation

The simulator “imagines” different scenarios of how products could be distributed to optimize space and facilitate future access. Consider factors such as the frequency of product output, their size and weight, and the robot routes.

Simulated execution

Before the products arrive, the simulator performs a series of internal tests, simulating how the robots would move to organize the products. Adjust routes and locations in your virtual model to avoid congestion and ensure efficient layout.

Planned actions

Based on its simulation, the system generates a sequence of actions: designates specific spaces for incoming products, adjusts the robots’ routes and organizes the tasks of human workers to receive the batch in an orderly manner.

Adaptation and optimization

Once products arrive, the system monitors plan execution in real time and makes adjustments as necessary. If something unexpected arises, such as a change in the number of products or their size, the world simulator adapts the plan quickly to maintain efficiency.

This approach allows not only more effective warehouse management, but also a greater ability to anticipate problems and optimize resources, demonstrating the power of world simulators in advanced planning and forecasting.

The challenges of creating World Models

Although the idea of world simulators sounds like something of the present, the reality is that we are still far from achieving fully functional systems. Creating a world simulator is extremely complex due to several factors:

Need for three-dimensional perception

Current systems, such as large language models (LLMs), work with one-dimensional (text) or two-dimensional (images) data.

Simulating the physical world requires modeling spatial, temporal and causal relationships in three dimensions, which is a significant leap in terms of computational power and algorithms.

Understanding cause and effect

To predict the impact of its actions, a world simulator must deduce how objects and agents interact with each other. This involves learning rules that are not always present in the training data.

Ability to reason and plan

These models must be able to make decisions in both the short and long term. For example, planning how to move objects to clean a room or how to design a safe bridge.

Computational efficiency

Building and operating world simulators requires much greater processing power than LLMs, plus enormous amounts of data and time to train them.

Current developments in world simulators

We are still in the early stages of the long road to machines understanding our world the way we humans do.

Currently, developments focus on providing them with spatial awareness, which has led to the advancement of multimodal models and advanced video generation capabilities. Some of the most notable developments include:

Sora from OpenAI

Sora is a video generation model that has stood out for its ability to create realistic and detailed visual representations. Although long promised, OpenAI has not yet released it publicly, although it has revealed some of its capabilities.

Sora uses advanced artificial intelligence techniques to simulate complex scenarios and generate high-quality videos.

This has potential applications in the creation of content for film and video games, as well as in the simulation of training situations for professionals in various industries.

DeepMind Genie 2

DeepMind has developed Genie 2, a model that combines natural language processing and computer vision capabilities to create interactive simulations.

Genie 2 can understand and generate visual scene descriptions, making it useful for applications in robotics, autonomous navigation, and architectural design. Thus you are able to navigate through your “worlds” as if it were a video game.

For example, a robot equipped with Genie 2 can navigate an unfamiliar environment, identify obstacles, and plan routes efficiently.

Oasis of Etched and Decart

Oasis is an interactive, real-time world simulation model that has been developed by the AI laboratory that emerged from the collaboration of Etched and Decart.

Unlike other models that generate video from text, Oasis generates frame-by-frame video from keyboard and mouse input.

This allows users to interact with the virtual world in real time, building structures, breaking blocks and exploring the environment. It is currently used tocreate worlds similar to the game Minecraft.

A paradigm shift

Researchers such as Yann LeCun and Fei-Fei Li are leading research in this field, developing architectures and paradigms that could make world simulators a reality in the next decade.

World simulators represent a paradigm shift in artificial intelligence, with the potential to overcome current limitations and bring us closer to more intelligent and autonomous systems.

Although we are still far from reaching its full potential, interest and investment in this technology ensures that we will see significant advances in the coming years.

This post is also available in: Español Français Русский Italiano