Constructing
3D Scenes with Code

Benchmarking coding agents that turn open-ended scene descriptions into engine-native 3D scenes in Unreal Engine. Measuring spatial reasoning, task fulfillment and physical validity.

Task Format

Nordic Harbour · Text-to-Scene

An open-ended scene description, an asset catalog and the Unreal Engine editor.

Agent inputText + assets
Task instruction

Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks.

Townhouses, canals, bridges and a waterfront promenade — with detailed requirements for materials, layout and street furniture.

Given: an empty level and the content pack’s asset catalog.

EvaluationEngine-native scene
CASE SCORE0.877
Detailed Alignment
0.757
Overview Alignment
0.909
Physical Safety
0.900

Task fulfillment, artifact integrity and static physical validity.

0.2 × Detailed + 0.6 × Overview + 0.2 × Physical

+ Full task instructionTEXT-TO-SCENE

Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.

20 scenes in the Research Track: affordable, reproducible evaluation for academic and model-development use160 scenes in the Full Benchmark: comprehensive evaluation for leaderboard and final reporting

Leaderboard

Research Track · 20 representative scenes · 14 coding-agent configurations.

Text-to-scene generation with frontier models is a long-horizon agentic task with substantial inference cost and runtime, which makes repeated full-scale evaluation difficult for many research groups. The Research Track supports affordable, reproducible experimentation; the Full Benchmark provides the more comprehensive evaluation for final model comparison and leaderboard reporting.

Code4Scene Research Track

20 scenes · score out of 100

Swipe to view all 14 configurations →

Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.

Code4Scene Pareto Frontier

Quality against the cost of one case · Research Track, 20 scenes

Swipe to explore all configurations →

Solid: the frontier — nothing cheaper scores higher. One point per configuration, at its evaluated reasoning effort. Cost is the mean USD per case.

14 configurations
Ranking

Research Track, 20 scenes. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.

#Agent configurationScoreDetailedOverviewPhysicalCost / case
1Claude Fable 5.1 (max)ParetoAnthropic
0.788
0.7250.7720.898$18.16
2GPT-6 Astra (max)OpenAI
0.724
0.7070.7330.712$22.47
3Claude Opus 5 (max)Anthropic
0.718
0.6310.7160.811$34.50
4GPT-5.6 Sol (high)ParetoOpenAI
0.707
0.6320.7100.772$4.07
5Gemini 3.8 Flash (high)ParetoGoogle
0.657
0.6370.6640.655$2.49
6Muse Spark 1.3 (medium)ParetoMeta
0.646
0.5980.6610.651$0.10
7Grok 4.6 (high)xAI
0.567
0.5220.5760.585$4.33
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.557
0.5070.5300.688$0.49
9Qwen 3.8 27B (thinking on)open weightsAlibaba
0.516
0.5050.4600.694$0.33
10GLM-5.3 Flash (max)open weightsZ.ai
0.509
0.4030.5370.533$0.71
11Gemma 4 31B (thinking on)open weightsParetoGoogle
0.471
0.4120.4160.694$0.03
12Gemma 4 31B (thinking off)open weightsParetoGoogle
0.455
0.4380.3700.725$0.02
13Inkling (high)Thinking Machines
0.425
0.3930.3370.722$0.20
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.243
0.2170.1610.513$0.19

Click a column to sort. Pareto: on the score-against-cost frontier.

Visualizing agent outputs

Drag any scene to inspect the saved 3D output, or open a case to explore every agent.

Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Nordic Harbour

GPT-6 Astra (max) · score 0.877

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Medieval Big Farm Town

Claude Fable 5.1 (max) · score 0.830

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Egyptian Temple

Claude Opus 5 (max) · score 0.777

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Bazaar

Claude Fable 5.1 (max) · score 0.930

Compare all 14 agents ↗
Open the case explorer

Takeaway

What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.

Takeaway

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & Environmentmean 67Content & Quantitymean 54Attributes & Materialsmean 48Spatial Composition, in the agent’s colourmean 36
One group per agent, in Text-to-Scene rank order: Identity & Environment, Content & Quantity, Attributes & Materials, then Spatial Composition, the lowest bar in every group. Scores ×100 (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch when the objects are present
0%10%20%30%40%Spatial relation32.0%Distribution19.1%Composition4.3%
Share of checks that fail although every object they name is in the scene, across the benchmark (paper Figure 6a).

Cite This Work

@misc{ye2026code4scene,
  title         = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
  author        = {Xiaokang Ye and Siddhant Hitesh Mantri and Zimeng Chen and Edward Zhang and Zhaoxu Zheng and Yuanheng Li and Yizhao Chen and Tianyang Huang and Lianhui Qin},
  year          = {2026},
  eprint        = {2609.36777},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2609.36777},
}
arXiv GitHub Repository SimWorld Main Site