ArticleOpen Access http://dx.doi.org/10.26855/ea.2026.09.013
Research on Large Language Model-driven Text-to-3D Virtual Scene Generation Methods
Qingyao Jia
Tandon School of Engineering, New York University, Brooklyn, NY 11201, USA.
*Corresponding author: Qingyao Jia
Published: September 02, 2026
Abstract
Text-driven 3D virtual scene generation requires coordinated handling of object semantics, spatial relations, scale, geometry, asset selection, and cross-view consistency. A three-layer framework is proposed in which a large language model supports prompt normalization, entity and relation extraction, and scene-graph construction, while downstream modules address constraint translation, asset retrieval or generation, scene instantiation, geometric representation, and multi-view inspection. 3D-FRONT and Objaverse are used as reference data sources, and published Holodeck results provide component-level evidence for fixed-object layout evaluation. The framework is designed to expose intermediate states, preserve relation provenance, and route detected errors to the processing stage responsible for them. Semantic omissions, infeasible placements, asset mismatches, visibility errors, and geometric inconsistencies are treated as separate failure types rather than merged into one score. The proposed architecture remains conceptual and does not claim end-to-end experimental validation. Its correction accuracy, runtime behavior, scalability, and reproducibility require implementation, compatible baselines, ablation studies, generated scene records, and complete execution logs.
Keyword
Large language model; Text-to-3D; virtual scene generation; scene graph; spatial constraint; 3D asset retrieval
References
[1] Lee HH, Savva M, Chang AX. Text-to-3D shape generation. Comput Graph Forum. 2024;43(2):e15061.
[2] Zhang J, Li X, Wan Z, et al. Text2NeRF: text-driven 3D scene generation with neural radiance fields. IEEE Trans Vis Comput Graph. 2024;30(12):7749-7762.
[3] Liu Z, Hu J, Hui KH, et al. EXIM: a hybrid explicit-implicit representation for text-guided 3D shape generation. ACM Trans Graph. 2023;42(6):228:1-228:12.
[4] Feng W, Zhu W, Fu TJ, et al. LayoutGPT: compositional visual planning and generation with large language models. In: Advances in Neural Information Processing Systems. 2023;36.
[5] Huang D, Wang N, Huang X, et al. Mesh-controllable multi-level-of-detail text-to-3D generation. Comput Graph. 2024;123:104039.
[6] Gu Z, Gao T, Liu H. Text-to-3D scene generation framework: bridging textual descriptions to high-fidelity 3D scenes. Vis Comput Ind Biomed Art. 2025;8:29.
[7] Fu H, Cai B, Gao L, et al. 3D-FRONT: 3D furnished rooms with layouts and semantics. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021:10933-10942.
[8] Deitke M, Schwenk D, Salvador J, et al. Objaverse: a universe of annotated 3D objects. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023:13142-13153.
[9] Yang Y, Sun FY, Weihs L, et al. Holodeck: language guided generation of 3D embodied AI environments. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024:16227-16237.
Copyright
© 2026 by the author(s).
This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial-NoDerivatives (CC BY-NC-ND) license, which permits non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited and is not modified or adapted.
https://creativecommons.org/licenses/by-nc-nd/4.0/
How to cite this paper
Research on Large Language Model-driven Text-to-3D Virtual Scene Generation Methods
How to cite this paper: Qingyao Jia. (2026). Research on Large Language Model-driven Text-to-3D Virtual Scene Generation Methods. Engineering Advances, 6(3), 214-218.
DOI: http://dx.doi.org/10.26855/ea.2026.09.013