Benchmarking coding agents that turn open-ended scene descriptions into engine-native 3D scenes in Unreal Engine. Measuring spatial reasoning, task fulfillment and physical validity.
Task fulfillment, artifact integrity and static physical validity.
0.2 × Detailed + 0.6 × Overview + 0.2 × Physical
+ Full task instructionTEXT-TO-SCENE
Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.
20 scenes in the Research Track: affordable, reproducible evaluation for academic and model-development use160 scenes in the Full Benchmark: comprehensive evaluation for leaderboard and final reporting
Leaderboard
Leaderboard
Research Track · 20 representative scenes · 14 coding-agent configurations.
Text-to-scene generation with frontier models is a long-horizon agentic task with substantial inference cost and runtime, which makes repeated full-scale evaluation difficult for many research groups. The Research Track supports affordable, reproducible experimentation; the Full Benchmark provides the more comprehensive evaluation for final model comparison and leaderboard reporting.
Code4Scene Research Track
20 scenes · score out of 100
Swipe to view all 14 configurations →
Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.
Code4Scene Pareto Frontier
Quality against the cost of one case · Research Track, 20 scenes
Swipe to explore all configurations →
Solid: the frontier — nothing cheaper scores higher. One point per configuration, at its evaluated reasoning effort. Cost is the mean USD per case.
14 configurations
Ranking
Research Track, 20 scenes. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.
What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.
Takeaway
Spatial Composition remains the weakest requirement family for all 14 agents.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.
Requirement-family score per agent
Identity & Environmentmean 67Content & Quantitymean 54Attributes & Materialsmean 48Spatial Composition, in the agent’s colourmean 36
One group per agent, in Text-to-Scene rank order: Identity & Environment, Content & Quantity, Attributes & Materials, then Spatial Composition, the lowest bar in every group. Scores ×100 (paper Figure 14b).
The objects are there; the relation is not
GPT-5.6 Sol (high)Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch when the objects are present
Share of checks that fail although every object they name is in the scene, across the benchmark (paper Figure 6a).
Citation
Cite This Work
@misc{ye2026code4scene,
title = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
author = {Xiaokang Ye and Siddhant Hitesh Mantri and Zimeng Chen and Edward Zhang and Zhaoxu Zheng and Yuanheng Li and Yizhao Chen and Tianyang Huang and Lianhui Qin},
year = {2026},
eprint = {2609.36777},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2609.36777},
}