TY - JOUR
T1 - Challenges in the evaluation of machine learning techniques in generative urban design
AU - Brama, Haya
AU - Grinshpoun, Tal
AU - Dalach, Agata
AU - Dortheimer, Jonathan
N1 - Publisher Copyright:
© The Author(s) 2026. This article is distributed under the terms of the Creative Commons Attribution-NonCommercial 4.0 License (https://creativecommons.org/licenses/by-nc/4.0/) which permits non-commercial use, reproduction and distribution of the work without further permission provided the original work is attributed as specified on the SAGE and Open Access page (https://us.sagepub.com/en-us/nam/open-access-at-sage).
PY - 2026
Y1 - 2026
N2 - Evaluating the quality of generative design machine learning tools is a critical challenge. Existing methods range from human-based assessments to performance-based metrics and statistical comparisons. We focus on generative urban design and critically review the evaluation methods employed in recent literature. We experimentally test and comprehensively analyze these methods. We find that existing approaches favor disrupted designs over well-designed ones, have inherent limitations, and fail to capture the tool performance quality. To address this critical gap, we develop two strategies: (1) modifying the Fréchet Inception Distance (FID) score to align with specific design principles, and (2) leveraging visual language models to assess design outputs. Our experiments show that these approaches provide more robust and comprehensive evaluations. The findings underscore the need for practical and reliable evaluation frameworks for AI in design fields to advance the research in this field.
AB - Evaluating the quality of generative design machine learning tools is a critical challenge. Existing methods range from human-based assessments to performance-based metrics and statistical comparisons. We focus on generative urban design and critically review the evaluation methods employed in recent literature. We experimentally test and comprehensively analyze these methods. We find that existing approaches favor disrupted designs over well-designed ones, have inherent limitations, and fail to capture the tool performance quality. To address this critical gap, we develop two strategies: (1) modifying the Fréchet Inception Distance (FID) score to align with specific design principles, and (2) leveraging visual language models to assess design outputs. Our experiments show that these approaches provide more robust and comprehensive evaluations. The findings underscore the need for practical and reliable evaluation frameworks for AI in design fields to advance the research in this field.
KW - deep learning
KW - FID score
KW - GAN
KW - generative urban design
KW - machine-learning
KW - visual language models
UR - https://www.scopus.com/pages/publications/105042082418
U2 - 10.1177/14780771261459705
DO - 10.1177/14780771261459705
M3 - ???researchoutput.researchoutputtypes.contributiontojournal.article???
AN - SCOPUS:105042082418
SN - 1478-0771
JO - International Journal of Architectural Computing
JF - International Journal of Architectural Computing
ER -