Computational models have emerged as powerful tools for multi-scale energy modeling research at the building and urban scale, supporting data-driven analysis across building and urban energy systems. However, these models require large amounts of building parameter data that is often inaccessible, expensive to collect, or subject to privacy constraints. We introduce a modular framework that applies generative Artificial Intelligence (AI) to construct simulation-ready building datasets from publicly available records and imagery. To improve the reliability of this framework, we evaluate both its AI components and its overall result. Our occlusion analysis demonstrates that for our selected images, LLaVA achieves greater visual focus than a GPT-based alternative for building image processing. We also assess plausibility of our results against a national reference dataset, finding that our synthetic data overlaps more than 95% for three of the four selected variables. This work aims to reduce dependence on costly or restricted data sources, lowering barriers to building-scale energy research and Machine Learning (ML)-driven urban energy modeling, thereby providing simulation-ready datasets intended to support downstream applications such as energy modeling, retrofit analysis, and urban-scale simulation under data scarcity.
Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity
Computational models have emerged as powerful tools for multi-scale energy modeling research at the building and urban scale, supporting data-driven analysis across building and urban energy systems.
- Preview

- Year
- 2025
- Hosting
- Excerpt onlyCC-BY-NC-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2509.09794CC-BY-NC-4.0
- TL;DR
- Semantic Scholar