WorldClaw: Turn One Prompt into an Editable 3D World

Discover how Tencent Hunyuan3D WorldClaw turns text into large, explorable 3D worlds with coherent terrain, editable assets, and agent-driven scene refinement.

01Open-ended prompt
02Structured plan
03Global terrain
04Editable assets
Isometric summer 3D world generated by Tencent Hunyuan3D WorldClaw

Short answer

WorldClaw turns a free-form text prompt into a large, explorable 3D world. It plans the complete scene, builds coherent regional terrain, generates detailed textured objects where they are needed, and uses render-aware agents to improve layout, scale, pose, materials, and ground contact.

Key takeaways

01.

Input: One open-ended text prompt describing the world, regions, objects, style, and spatial relationships.

02.

Terrain: A semantic layout controls continuous height fields, regional materials, vegetation, rocks, and other environmental assets.

03.

Output: Explicit terrain plus independently editable textured object meshes with position, scale, and orientation.

04.

Reported Setup: Claude Opus 4.8, GPT-Image-2, SAM3, SAM3D, Hunyuan3D, Blender 5.1.1, and four NVIDIA H20 GPUs.

WorldClaw pipeline at a glance

TopicInputCore workOutput
01 · PlanAn open-ended text description of the target world.Extract intent, resolve ambiguity, and define regions, terrain, objects, appearance, density, and spatial relationships.A structured scene specification shared by downstream agents.
02 · TerrainThe scene specification and regional constraints.Build a semantic layout, continuous height field, surface materials, environmental scatter, and render-guided corrections.A coherent global terrain with explicit regional semantics.
03 · PopulateThe world plan plus the generated terrain.Compose selected regions, reconstruct object meshes, recover placement, and refine quality, pose, scale, and ground contact.Independently editable textured objects placed on the global terrain.

What WorldClaw Produces

WorldClaw takes one open-ended text prompt and produces a large 3D environment with connected terrain, distinct regions, and populated local areas. The terrain and objects remain separate scene elements rather than being baked into a single image or camera path.

Each object can retain its own textured mesh and placement transform. This supports free-view exploration, object replacement, reuse, and export into conventional 3D workflows.

From One Prompt to a Structured World

The first agent extracts only the requirements stated in the prompt, such as theme, major regions, objects, style, and relative placement. A second planning agent resolves ambiguity and fills in the technical attributes required to construct the scene without discarding the original intent.

The output is a structured specification shared by the terrain and object pipelines. This intermediate plan is what lets the system preserve long-range relationships while different agents work on different parts of the world.

WorldClaw pipeline from intent planning through global terrain and regional object generation to an editable 3D scene
Official WorldClaw pipeline: planning establishes the scene structure before terrain and regional assets are generated and refined.

Building Coherent Terrain Before Adding Detail

WorldClaw turns the plan into a semantic layout map, then builds a continuous height field from region-specific elevation, noise, and landform operators. Materials and reusable environmental assets are assigned using the same region masks, so mountains, plains, water, vegetation, and other terrain features remain spatially connected.

A render-and-review loop inside Blender checks transitions, texture scale, asset scattering, lighting, and visible artifacts. The terrain agent edits only the failing parameters and renders again, protecting the broader regional layout while improving local quality.

WorldClaw terrain generation stages showing height fields, asset scattering, materials, and render-guided refinement
Official terrain workflow showing semantic layout, region-aware construction, environmental scattering, and iterative refinement.

Turning Selected Regions into Editable 3D Assets

WorldClaw does not populate every square meter at maximum detail. A regional planner selects areas whose function and terrain require richer content. It renders the local ground, creates a terrain-aware composition image, separates object instances, reconstructs them as textured meshes, and recovers their scale, orientation, and placement on the terrain.

Render-guided agents then inspect object quality, pose, size, and surface contact. Weak assets can be refined with Hunyuan3D, while terrain and object transforms are adjusted to reduce floating, penetration, or unstable support.

WorldClaw render-guided loops for refining object quality, placement, scale, and terrain contact
Official refinement diagram showing repeated render, inspection, correction, and verification for objects and terrain contact.

What the Final World Contains

The output is an explicit 3D scene rather than a camera-bound video. The terrain remains a manageable surface and the regional objects remain separate textured mesh instances. That representation supports arbitrary viewpoints, object-level replacement, asset reuse, and hand-off to familiar 3D content pipelines.

The official examples cover islands, canyon settlements, snow valleys, desert battlefields, villages, mines, and volcanic environments. Their RGB, instance-mask, normal, and depth views help reveal both visible appearance and underlying scene structure.

Isometric layout of the Frontier Mosaic 3D world generated by WorldClaw
Frontier Mosaic, one of the official WorldClaw examples, combines multiple terrain regions and separately placed scene assets.

A World-Scale 3D Generation Stack

WorldClaw combines Claude Opus 4.8 for agent reasoning, GPT-Image-2 for image generation, SAM3 for segmentation, SAM3D for initial object reconstruction, Hunyuan3D for asset refinement, and Blender 5.1.1 as the scene execution environment. The reported showcase worlds were generated on a server with four NVIDIA H20 GPUs.

The current pipeline focuses on scene construction and visual refinement. Navigation, physics, animation, richer object hierarchies, and deeper runtime engine integration are listed as future work.

Frequently Asked Questions

No. WorldClaw is an agentic framework that coordinates planning, image generation, segmentation, 3D reconstruction, Hunyuan3D refinement, and Blender tools to construct a complete scene.

It targets an explicit 3D scene with a global terrain and separately manageable textured object meshes. Videos and stills are renderings of that underlying scene, not the primary output.

Image-to-3D usually reconstructs one asset. WorldClaw plans a large environment, constructs terrain, decides which regions need detail, and then uses reconstruction and Hunyuan3D tools as parts of a larger scene-building pipeline.

The paper, project page, and GitHub repository are public. As of August 12, 2026, the official repository does not expose runnable inference code, model weights, or a public generation API.

It generates editable 3D scene content, not a finished game. Navigation, physics, animation, interaction logic, and deeper production-engine integration are future development areas.

The reported experiments ran on a server with four NVIDIA H20 GPUs and used Blender 5.1.1 for terrain construction, object placement, refinement, and rendering.

Create assets for your own 3D scenes

Turn a reference image into a detailed 3D character, prop, product, or scene element.

Try Image to 3D