Valerant: Auto-Generating Navigable 3D Game Maps with World Models

Valerant combines video prediction with SLAM to build navigable 3D game maps from a single image — no retraining needed.
Valerant addresses a gap in world action model research for games: existing methods are largely limited to 2D pixel prediction and cannot produce truly navigable 3D spaces. The training-free framework combines three modules — action-conditioned video rollouts for frame prediction, SLAM to convert predicted frames into 3D geometry, and an exploration-driven strategy to continuously expand the map — turning a single image into a persistent 3D game level. Methodologically, it upgrades the world model's role from "predicting pixels" to "constructing usable space." As a preprint, its practical viability regarding accumulated prediction errors and geometric accuracy remains to be verified.
From Video Prediction to Navigable 3D Worlds
World Action Models (WAMs) have rapidly gained traction in the embodied AI field. Their core idea is to couple predictive world modeling with action generation — the model not only predicts how the environment will change, but also uses that prediction to guide an agent's next move. This "imagine first, act second" paradigm allows robots and autonomous driving systems to mentally rehearse the future, enabling more informed decision-making.
However, a newly published arXiv paper (arXiv:2609.09418) points out that while world action models are advancing rapidly in embodied AI, their generalized application to games remains virtually unexplored. Most existing game-oriented approaches patch action-conditioned world models together with external policies and reward functions to approximate WAM-like decision-making — but nearly all of them remain confined to 2D visual observation spaces, with no ability to construct persistent three-dimensional geometry.

The Valerant framework proposed by the research team targets exactly this gap: it aims to transform a pretrained action-conditioned world model into a complete system capable of exploring and constructing 3D game maps.
The Fundamental Difference Between Game Worlds and the Real World
The paper astutely identifies a critical issue that has largely been overlooked: game worlds have no "external substrate."
In autonomous driving and robotics, the physical environment exists independently of any model. No matter what the model predicts, the real world always provides a persistent three-dimensional space in which chosen actions can actually be executed. In other words, the model only needs to "understand" the world — the world itself is already there.
Virtual Worlds Must Be "Instantiated"
Games are fundamentally different. The virtual world itself must be created. Most playable games require a persistent and navigable space, and 3D games additionally demand explicit geometric structure — only then can characters truly move, interact, and explore within it.
This is where a technical gap emerges: action-conditioned video rollouts can generate continuous visual frames, but they provide only image sequences that "look like" the world — not a spatial representation that can actually be navigated. You can have a model "imagine" a stretch of gameplay footage, but you can't have a character genuinely walk through it, collide with objects, or explore it. The core problem Valerant sets out to solve is precisely how to "distill" usable three-dimensional geometry from visual rollouts.
Valerant's Core Mechanism: Three Modules Working Together
One of Valerant's most notable features is that it is a training-free framework. Rather than retraining a large model from scratch, it cleverly combines three existing capability modules and gets them to work in concert.
Predictive Visual Rollouts
Using a pretrained action-conditioned world model, Valerant predicts the visual frames that would follow a given action. This is essentially having the model "envision" what the scene would look like if the character moved in a certain direction.
SLAM-Based 3D Spatial Reconstruction
Visual prediction alone is insufficient for navigation, so Valerant incorporates SLAM (Simultaneous Localization and Mapping) technology. SLAM is well-established in robotics, capable of inferring 3D geometric structure and camera poses from sequential visual observations. Through this step, the originally "virtual" predicted frames are converted into a persistent three-dimensional spatial representation.
Exploration-Driven Action Selection
Finally, Valerant employs an exploration-driven strategy to decide where to move next. It actively selects actions that expand unknown regions and refine the map, causing the 3D map to continuously "grow."
Through the coupling of these three modules, Valerant can start from a single image and progressively expand it into a complete, persistent, navigable 3D game map.
The Significance and Potential Impact of Valerant
Valerant's value manifests on two levels.
At the methodological level, it advances the interaction paradigm of world action models from 2D visual simulation to 3D spatial construction. Previous world models were largely about "predicting pixels," while Valerant is the first to have a world model genuinely take on the responsibility of "constructing usable space." This represents a substantive expansion of WAM's application boundaries.
At the game development level, it offers a new pathway for 3D game map creation that reduces manual labor. Traditional game map design and modeling is an extraordinarily labor-intensive process, requiring artists and level designers to iterate extensively. If a concept image or reference photo could serve as a starting point for a model to automatically explore and generate a navigable three-dimensional scene, it would dramatically improve development efficiency — especially appealing for indie developers and rapid prototyping.
Limitations Worth Scrutinizing
As a newly released preprint, Valerant still has a number of questions that await validation:
- How accurately does SLAM reconstruction perform on purely predicted frames (as opposed to real observations)?
- Will the accumulated prediction errors from the world model cause map "drift" or collapse?
- Can the generated maps meet practical standards in terms of gameplay viability and geometric soundness?
These questions will require more experiments and follow-up work to answer.
Nevertheless, the idea Valerant proposes — upgrading world models from "imagining scenes" to "building worlds" — represents a research direction worth watching. It reminds us that the imaginative potential of generative AI in gaming may extend far beyond producing a playable sequence of frames, reaching toward the automated construction of entire interactive virtual spaces.
Related articles

Invalid Source Material: Unable to Generate a Valid AI/Tech Article
This Twitter source material is an irrelevant marketing tweet with no AI or tech content, making it impossible to generate a valid professional article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material was limited to a single broken tweet with no usable content, making it impossible to produce a complete, high-quality article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material provided was a single vacuous social media tweet with a broken link — insufficient to support writing a complete, factual article.