A team introduced LEGO‑Bench, a test set of 208 photos rendered from simulated 3D scenes, to measure how well coding agents can write Blender scripts from a single image. Six GPT configurations were run; the top model, GPT‑6 Astra, reconstructed indoor scenes correctly 53.4% of the time and outdoor scenes 39.6%. The agents could usually produce a usable scene file, but their ability to judge geometric accuracy was close to random. Adding a plug‑in that anchors the scene to the reference image boosted weaker models by up to 62.7%.
Why it matters
For developers of 3D tools, the results show that AI can automate initial scene creation, but reliable geometry checks are still needed.