the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
SETP_GLI: An annual 10–30 m glacial lake inventory for the southeastern Tibetan Plateau from 1990 to 2025
Abstract. Glacial lakes in the southeastern Tibetan Plateau (SETP) have expanded, increasing the potential for cascading hazards associated with glacial lake outburst floods (GLOFs). However, long-term, annual monitoring data that include micro glacial lakes remain relatively limited for this region. To address this gap, this study integrated Landsat series and Sentinel-2 imagery and used the GLA-RCNN deep learning framework with an embedded Convolutional Block Attention Module to construct and release an annual glacial lake inventory (SETP_GLI). The dataset comprises 36 annual vector layers from 1990 to 2025, recording the annual evolution of regional glacial lake numbers and areas. The use of 10 m resolution imagery and model optimization improved the detection of micro glacial lakes (<0.01 km²). The inventory provides annual vector boundaries and standardized physical attributes—including longitude, latitude, area, perimeter, and mean elevation, together with area uncertainty metrics derived from mixed-pixel theory. Quality assessments indicated that the extraction framework is robust against interference from mountain shadows and turbid water. For model performance, the overall F1 scores for typical years remained above 0.82 (with a maximum of 0.895); cross-validation with existing public databases (Hi-MAG and Glacial lake inventory of high-mountain Asia) showed that the matched polygon-level Intersection over Union (IoU) ranged from 0.54 to 0.80, with spatial agreement increasing with improvements in historical image quality. Spatiotemporal analysis revealed a persistent expansion trend, with the annual area growth rate rising from 3.65 ± 1.12 km² a⁻¹ (1990–2012) to 5.95 ± 2.44 km² a⁻¹ (2016–2025). The dataset is archived at the National Tibetan Plateau Data Center (TPDC) (https://doi.org/10.11888/Cryos.tpdc.303491), with processing code released openly. SETP_GLI serves as a baseline dataset for cryospheric response analysis, hydrological modeling, and GLOF risk assessment.
- Preprint
(3394 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-452', Anonymous Referee #1, 03 Aug 2026
-
RC2: 'Comment on essd-2026-452', Anonymous Referee #2, 08 Sep 2026
This manuscript presents a valuable contribution to glacial lake monitoring in a critical region, with a well-designed methodological framework and a comprehensive dataset. However, several issues require attention before publication. Major comments are as follows.
- GLA-RCNN Architecture Description is unclear. Should provide a detailed architecture diagram showing where CBAM modules are inserted (after which residual blocks). Specify the reduction ratio used in channel attention and kernel sizes in spatial attention.
- Training Data Generation. While the text notes that ground-truth labels were manually produced using Wang et al. (2020) as a reference, the specific methodology for the 1991 training node remains ambiguous, given that Wang et al. only provides inventories for 1990 and 2018. The authors should explicitly clarify how contemporaneous imagery was utilized for 1991 manual digitization to avoid temporal mismatch. Furthermore, to ensure data reliability, the manuscript must provide essential operational details: who performed the digitization, what quality control protocols were implemented, and what inter-operator consistency metrics were achieved.
- Sensor Transition Handling. Lines 186-191 mention training separate models for different resolutions but don't explain how predictions are harmonized for the full time series. Explain whether separate models were applied to each period (1990-2012, 2013-2015, 2016-2025) or a single model with resolution-aware inference was used. How model performance differences across resolutions were addressed in the final product.
- Validation Limitations. Lines 258-262 state that the Hi-MAG and Wang et al. (2020) databases were used for accuracy assessment, but Lines 96-98 indicate these same datasets assisted in label generation. Using the same data to inform training and conduct validation introduces a fundamental circularity bias. This not only artificially inflates spatial agreement metrics but also penalizes the model's high-resolution innovations, as new micro-lakes missed by the older 30m reference datasets will register as false positives. Please ensure a strict separation between training and validation sets. If the dual-use of these datasets is unavoidable, the authors must explicitly acknowledge this limitation in the text and thoroughly discuss how this circularity biases the reported accuracy metrics.
- Ground Truth Quality for Small Lakes. Lines 371-373: Google Earth imagery used as validation for <0.01 km² lakes. But google Earth imagery has varying dates and resolutions in the SETP. The author may provide specific dates of Google Earth imagery used and quantify the resolution and quality, and discuss potential misalignment between Google Earth imagery and Landsat/Sentinel-2 acquisition dates.
- Uncertainty Quantification. Lines 274-284: Area uncertainty calculated using Hanshaw and Bookhagen (2014) boundary pixel method. This method doesn't account for georeferencing errors (especially in early Landsat), seasonal water level fluctuations and shadow-induced boundary shifts. May add a discussion of additional uncertainty sources.
- Temporal Consistency Issues. Lines 102-119 mentioned the primary window for image selection are Sept-Nov, backup August, then Dec-Feb. This introduces seasonal variability that could obscure long-term trends. Lakes in December may be partially frozen; August lakes may be at peak melt. Can quantify the proportion of lakes from each month and analyze whether there's a systematic bias in lake area from different months. Lines 226-233 say sensor-adaptive thresholds of 0.001 km² (10m), 0.002025 km² (15m), and 0.0054 km² (30m). This means the minimum detectable lake size changes over time. Lines 473-485: There's a gap between Landsat and Sentinel-2 coverage. How was 2014-2015 handled? Was Landsat 8 used exclusively? What about 2016 transition year? The author can provide detailed per-year information on primary sensor used, the number of scenes and cloud cover statistics and any data gaps and how they were filled.
Minor comments.
- Lines 287-298: GLA-RCNN compared with Mask R-CNN, U-Net, and DeepLab V3. Were the comparison models optimized for this task? Should specify training details for comparison models (learning rate, epochs, data augmentation). Were hyperparameters tuned for each model? Should ensure a fair comparison with equal training resources.
- Lines 165-170: "CBAM modules were embedded into each fused feature layer (P2–P5)." This is non-standard. How were the CBAM outputs integrated with the FPN features? The authors should explain precisely how CBAM outputs were integrated with FPN features and provide a schematic diagram of the modified architecture.
- Lines 133-134: "input channel number was inferred automatically in the model configuration." Given that Sentinel-2 (10m bands) and Landsat (30m bands) feature different band configurations, the authors must clarify how the model handles these variable input dimensions and discuss any potential spectral inconsistencies introduced.
- Lines 174-179: The loss function description is very technical. Suggest to provide a brief intuitive explanation of what each loss term does to improve readability.
- Lines 521-544 repeats the Introduction. Suggest to focus Discussion on the implications of the new dataset, not reiterating its features.
- Figure 2, correct ‘STREM’ to ‘SRTM’
- Figure 3, ensure all text can be seen clearly in all panels.
- Figure 10, what is raw and MMU trend. Explain in figure caption or main text.
Citation: https://doi.org/10.5194/essd-2026-452-RC2
Data sets
Annual glacial lake inventory dataset for the southeastern Tibetan Plateau from 1990 to 2025 H. Li et al. https://doi.org/10.11888/Cryos.tpdc.303491
Model code and software
GLA-RCNN-SETP-GLI: GLA-RCNN code for SETP_GLI v1.0.1 AI-Geohazard https://doi.org/10.5281/zenodo.20555222
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 407 | 87 | 53 | 547 | 38 | 61 |
- HTML: 407
- PDF: 87
- XML: 53
- Total: 547
- BibTeX: 38
- EndNote: 61
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General comments
Li and Dou et al. present a dataset of glacial lakes on the southeastern Tibetan Plateau derived using a GLA-RCNN deep learning framework they developed and applied to Landsat series and Sentinel-2 imagery. The dataset has the potential to provide a useful contribution and is presented in a suitable format with associated metadata. Issues requiring addressing are detailed below and mainly relate to: (1) the erroneous use/ misunderstanding about the origin and date of the ‘ALOS PALSAR RTC DEM’; (2) lack of details about the model training and notable manual adjustments performed to the outputs at several stages; (3) overstated improvements of the developed model with respect to other models, as there are negligible (<0.01) differences in e.g. F1 scores; (4) the stepped increase in glacial lake area and count resulting from using imagery of 30 m and 10 m (post 2016) resolution, which was identified, but is not addressed; (5) the paper does not define what qualifies as a ‘glacial lake’, as other water bodies including in non-glacierised catchments (e.g. TRK_03213), appear to be included in the dataset; (6) there are inconsistencies in the timeseries for individual lakes, potentially resulting from the availability/quality of imagery in a given year, causing lakes to be missed. I would support publication if the authors address these issues.
Specific comments;
Technical corrections:
L13. Define ‘GLA-RCNN’
L145. Change the DEM and correct references to ‘ALOS DEM’ or similar.
L151. ‘sensor-specific’ not ‘adaptive’.
L222. Clarify ‘full manual verification’. Were all lake polygons in the full dataset checked?
L330 . 41,916.3 m². Check whether the reported precision and number of significant figures are appropriate.
L474. Change ‘apparent numerical growth’ to ‘increase in lake count’ or similar.
L477. Replace/define ‘mathematical level’.
L493. ‘minute glacial lakes’. Make terminology consistent and define area thresholds.
L595. Remove ‘target’.