BMAT: A footprint-level building facade material dataset for 73 major cities worldwide
Abstract. Building facade materials are closely associated with urban microclimates, energy consumption, and carbon emissions, yet existing urban building datasets typically capture only building footprints and heights, leaving semantic material information unavailable at the individual-building level. Here we present BMAT, a footprint-level building facade material dataset covering 22.09 million buildings across 73 major cities worldwide. BMAT links individual building footprints with facade material labels inferred from over 147 million temporally stamped street-view images collected between 2007 and 2025, where historical imagery is available. We developed an automated inference pipeline based on a fine-tuned Vision-Language Model trained on 39,405 manually annotated samples, achieving an F1-score of 0.91 against held-out test data. Dataset reliability was further assessed through independent cross-validation against 47,434 Overture Maps building records with known material attributes, yielding an overall accuracy of 0.80. BMAT reveals spatially distinct material signatures associated with climate, geography, and regional building traditions. In cities with sufficient repeated imagery, the temporally stamped records further provide exploratory evidence of facade material transitions, including increasing glass facade adoption in selected rapidly urbanizing regions. By bridging building geometry and semantic surface properties, BMAT supports studies in building energy modelling, embodied carbon accounting, urban microclimate analysis, and urban hazard risk assessment. The dataset is openly available at https://doi.org/10.6084/m9.figshare.31569370 (Yin, 2026).
This manuscript presents BMAT, a large scale facade material dataset covering 22.09 million building footprints across 73 cities, derived from approximately 147 million Google Street View and Baidu Street View images. The dataset addresses a gap in existing building databases. However, the current validation does not sufficiently establish the reliability of the final building level and multi temporal product. The reported F1 score of 0.91 evaluates image classification performance, but the dataset also depends on correct image to building correspondence, adequate facade visibility, consistent aggregation of multiple observations, stable building identities across years, and representative spatial coverage. These components have not been comprehensively validated. In addition, each record appears to describe the dominant material visible from a street facing viewpoint rather than the complete building envelope. The claims concerning whole building characteristics, temporal transitions, embodied carbon, and energy modelling should therefore be moderated. I recommend major revision.
Â
Major comments
1. The meaning of a BMAT material record needs to be defined more precisely
The manuscript describes BMAT as a footprint level building facade material dataset, although each label appears to represent the dominant material visible from one or more street facing images rather than the entire building envelope. This distinction is important for mixed material buildings, where different facades or facade sections may contain glass, stone, concrete, metal, or other materials. The authors should clearly define what each label represents, explain how the dominant material is determined, and describe how mixed facades are handled. The dataset should also include information on facade orientation, number of images and viewpoints, visible facade proportion, and visibility status. If such information is unavailable, the manuscript should consistently describe BMAT as a dataset of dominant visible street facing facade materials.
2. The correspondence between street view images and target building footprints has not been adequately validated
The material classification model assumes that the central building in each image is the intended target, but the geometric retrieval process does not fully guarantee this correspondence. Multiple buildings may appear in one image, while footprint errors, panorama positioning, road geometry, camera parameters, building height uncertainty, and unmodelled obstructions such as trees, walls, vehicles, and elevated roads may affect target visibility. The assumption that partial facade exposure is sufficient also requires empirical validation. The authors should therefore conduct a stratified manual audit of the final dataset to assess whether the intended building is correctly identified, sufficiently visible, and assigned the correct material label. The audit should cover different cities, platforms, building densities, sampling distances, building heights, and visibility conditions, with separate results for fully and partially visible buildings. Without such validation, image classification performance cannot be directly interpreted as the accuracy of the footprint level dataset.
3. The necessity of using a vision language model is insufficiently justified
Facade material identification is formulated as a closed set classification task with nine predefined categories, and the model is required to return only one label from a fixed list. Under this setting, the language generation and instruction reasoning capabilities of Qwen2.5-VL appear to be largely unused, making the model function mainly as a computationally expensive image classifier. Although the fine-tuned model outperforms several conventional vision models, the comparison does not demonstrate that a vision language model is necessary because the performance gain may result from differences in model scale, pretraining, optimization, or data partitioning. This concern is especially important because the model is applied to approximately 92 million images and is substantially slower than traditional classifiers. The authors should justify the use of the Vision Language Model through fair comparisons with modern pretrained visual encoders under consistent experimental settings and report model size, inference speed, memory use, total computing cost, and cross city or cross platform generalization. If no clear advantage beyond classification accuracy is demonstrated, a smaller visual model may be more appropriate for this large scale task.
4. The current test split may contain building, spatial, or temporal information leakage
The manuscript states that 39,405 annotated images were divided into training and testing sets in a ratio of 7 to 3, but the unit used for data partitioning is unclear. A random image level split may place repeated views of the same building or visually similar images from nearby locations in both subsets, leading to overestimated performance. The authors should therefore use building independent and city independent splits, report separate results for Google Street View and Baidu Street View, and consider a temporal split to assess generalization across years.
5. The Overture Maps comparison is not a sufficient independent validation of the dataset
The comparison with 47,434 Overture Maps records should not be described as independent cross validation because the provenance, observation date, and accuracy of these material attributes are unclear. The semantic meaning of the Overture labels may also differ from the dominant visible facade material inferred from street view images. Moreover, BMAT uses Overture Maps building footprints, so the validation is not fully external at the spatial object level. This analysis should therefore be presented as an agreement assessment with an auxiliary reference source. The authors should clarify the origin and definition of the Overture attributes and report agreement by city and material class. They should also develop an independently annotated validation set drawn directly from BMAT and stratified by city, platform, material class, building density, visibility condition, and acquisition year.
6. The historical matching strategy systematically inflates the reported agreement
The manuscript reports an accuracy of 0.72 when the latest BMAT prediction is compared with the Overture Maps record. The value increases to 0.80 when a prediction is considered correct if any available historical BMAT label matches the undated Overture record. This approach introduces an upward bias because the probability of at least one match increases with the number of available observations. It also creates unequal validation conditions across cities because Google Street View and Baidu Street View provide different temporal coverage. The 0.80 value should therefore not be interpreted as a conventional accuracy estimate. The authors should retain the latest observation comparison as the more transparent result and, where possible, compare records from similar dates. If reference dates are unavailable, agreement should be reported according to the number of historical observations, with the inflation caused by chance matching explicitly quantified and discussed.
7. The manuscript validates the classifier but does not adequately validate the complete dataset production pipeline
The final BMAT records are generated through multiple stages, including street view sampling, camera orientation, obstruction detection, image quality filtering, material classification, and footprint association, so errors may accumulate throughout the pipeline. Although the MobileNetV2 filter achieves 88.75 percent validation accuracy, class specific metrics and external validation are not reported, and even a modest false positive rate could introduce many unsuitable images at this scale. The manuscript also does not explain how approximately 92 million valid images are aggregated into records for 22.09 million buildings, including the number of images used per building and year, the treatment of repeated panoramas and conflicting predictions, and the selection of the latest material label. The authors should therefore provide end to end validation of the final dataset and fully document the aggregation procedure, including the handling of low confidence predictions, inconsistent views, and ties.
8. The temporal dimension does not yet provide reliable evidence of facade material transitions
The temporal dimension is presented as a major contribution, but differences between annual labels cannot be directly interpreted as physical facade changes. Historical street view images are linked to contemporary building footprints, although buildings may have been constructed, demolished, expanded, merged, subdivided, or replaced between 2007 and 2025. Label changes may also result from different viewpoints, mixed facades, illumination, image quality, vegetation, occlusion, or incorrect building alignment. The manuscript is therefore conceptually inconsistent in treating label changes as evidence of renovation while also acknowledging that such changes may reflect mixed materials or varying visible surfaces. The authors should establish a stable building cohort, exclude or flag buildings with uncertain historical identities, require repeated or persistent evidence for material transitions, and manually validate a representative sample to distinguish genuine renovation from building replacement, viewpoint effects, mixed materials, association errors, and classification errors.
9. Spatial coverage is incomplete and nonrandom, which may bias cross city comparisons
Material information is available for approximately 54 percent of the buildings, with substantial variation across cities. Because missing records are associated with street accessibility, building density, gated development, platform coverage, and distance from urban centers, the observed buildings are unlikely to represent a random sample. Consequently, city level material proportions and geographic patterns may partly reflect differences in imagery availability, study boundaries, and platform characteristics rather than actual facade distributions. The authors should assess coverage bias using building size, height, density, road distance, urban location, and neighborhood morphology, report coverage rates and uncertainty, and restrict or adjust cross city comparisons accordingly. Furthermore, the statement on page 20, lines 281 to 284 that missing material records can indirectly identify underdeveloped areas, informal housing, and marginalized neighborhoods is insufficiently supported. Missing street view imagery may also result from platform policies, legal or access restrictions, private roads, incomplete updates, and technical limitations. This interpretation should therefore be removed or validated using independent accessibility, settlement, and socioeconomic data.
10. Reproducibility and data documentation need substantial improvement
A data paper should allow users to understand, reproduce, and responsibly reuse the product. The manuscript should report the exact versions, release dates, access dates, or image acquisition periods for Overture Maps, GlobalBuildingAtlas, GADM, OpenStreetMap, Google Street View, and Baidu Street View, as applicable.
Â
Minor comments