the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Glacial-Lake-Bench: A Global Multi-Sensor Benchmark Dataset for Evaluating Deep Learning Models for Glacial Lake Mapping
Abstract. Glacial lakes are among the most sensitive indicators of climate change, closely linked to natural hazards and are natural reservoirs of freshwater resources. Thus, automated mapping and monitoring of glacial lakes is imperative. However, most automated approaches either remain regional in scope or show limited performance under challenging conditions such as cloud cover, shadows, and spatially small or frozen lakes. The global scale analysis and comparative evaluation of deep learning models is primarily hindered by the lack of readily available datasets for training data. To address this gap, we present Glacial Lake-Bench (GLB), a multisource remote sensing dataset comprising Sentinel-2, Sentinel-1, and Copernicus DEM-derived terrain (11 channels in total). GLB consists of 19,115 image-label pairs (256x256x11) spanning all Randolph Glacier Inventory (RGI) regions except Antarctica, providing the first-ever global, multi-sensor dataset for glacial lake segmentation. In addition, we compiled Glacial Lake-Bench-Challenge (GLBC), a curated subset of 1,105 image-label pairs representing scenes with cloud cover, shadow, frozen lake surfaces, and small lakes to establish a community standard for evaluating model robustness under difficult conditions. Labels are derived from Zhang et al. (2024); we independently quantify label quality through stratified sampling of 50 image-label pairs from each RGI region, amounting to 900 chips in total. Our quality assessment reveals a mean Intersection over Union (mIoU) of 0.95, precision of 0.99, recall of 0.96, and per-region agreement that is consistent with the known difficulty of small, turbid, shadowed lakes in high-mountain terrain. To demonstrate that the dataset is usable, well-posed, and appropriately challenging, we provide reference baselines from two convolutional networks (U-Net, DeepLabv3+) and two Geo-Foundation Models (GFMs) (DOFA, Prithvi-EO-2.0), evaluated with a recommended leave-one-region-out (LORO) protocol that minimizes spatial autocorrelation, alongside a random split and the GLBC subset. Baseline mIoU reaches 0.80–0.85 on GLB and drops to 0.74–0.79 on GLBC, confirming that the challenge subset isolates genuinely difficult conditions. The GLB dataset is available at https://zenodo.org/records/17917359 (Kaushik, 2026)
- Preprint
(12169 KB) - Metadata XML
-
Supplement
(512 KB) - BibTeX
- EndNote
Status: open (until 10 Oct 2026)
- RC1: 'Comment on essd-2026-474', Xingyu Xu, 12 Aug 2026 reply
-
RC2: 'Comment on essd-2026-474', Renzhe Wu, 14 Sep 2026
reply
Recommendation: Major Revision
Rationale:The dataset is a genuine contribution — the first global, multi-sensor benchmark for glacial lake segmentation, covering all RGI regions with a leave-one-region-out (LORO) protocol. The current problems are mainly about transparency, standardization, and fairness of comparison rather than fundamental flaws, and in principle they can be fixed with clarifications and additional experiments. However, these issues touch several core aspects of evaluation validity (co-registration and label consistency, normalization, the input-information confound, GLBC construction, and statistical support), so a full second-round review is needed. In the second round, I will pay particular attention to how Major Comments 2, 3, 4, 6, and 7 are addressed. Acceptance is recommended if the authors (i) provide channel-matched experiments or clearly justify the limitation, (ii) fully document co-registration and normalization, and (iii) make the GLBC selection criteria quantitative. Major Comment 9 is a constructive suggestion: adding a geometric-distortion mask band, or at least quantifying the distorted area and discussing it as a limitation, would both be acceptable.
Major Comments
1. Lake counts are missing from Table 1 (dataset composition)
Table 1 reports only the number of images and areas. Please add: for each region, the number of lakes in the original Zhang et al. (2024) inventory versus the number actually kept in GLB (in both count and area), and the screening flow — how many lakes were dropped due to cloud cover, chip cropping, or de-duplication. On this basis, please analyze whether the losses are systematically biased by region or by lake type. The manuscript already admits that the 15% cloud threshold leaves only 9.21% coverage in South Asia East and 11.67% in Iceland; whether cloudiness causes systematic under-sampling of particular regions or lake types (small supraglacial lakes, turbid lakes) needs quantitative evidence, not just area percentages.
2. Temporal consistency between labels and images, digitization accuracy, and the handling of known label errors
The labels come from the manual interpretation of Zhang et al., which is highly reliable, but using them directly as a benchmark dataset needs further clarification:
(a) Please report the date-gap distribution between each chip's Sentinel-2 image and the image on which Zhang et al. digitized the label. Are they the same scene? If both are from the ablation season, how many days apart? Lakes change seasonally and between years, so labels and image content may not match. Please report the gap distribution and its effect on label validity.
(b) Manual labels for small lakes typically use simplified boundaries with only a few vertices, which deviate systematically from the true shoreline at 10 m resolution. In my experience, even when several experts independently digitize the same set of well-defined small glacial lakes, the IoU rarely reaches 85%. Therefore, the low IoU of small lakes in Figs. 4 and 5 should not be attributed entirely to label errors. Please discuss the subjectivity of manual digitization, and consider tolerance-based or size-classed metrics for small lakes.
(c) The 900-chip quality check found about 4% omissions and some boundary adjustments. Were these corrections propagated to the full dataset? If the released version intentionally keeps known errors, please explain why. Ideally, provide a corrected label layer or per-chip quality flags (e.g., as a v2 plan). Otherwise, users cannot distinguish label errors from model errors.3. The co-registration workflow for multi-source data must be documented in detail
(a) Sentinel-1 RTC and Sentinel-2 L2A have pixel-level geolocation offsets by nature, yet the model takes both as inputs under the same labels. Please state whether relative co-registration was performed, and how its accuracy was assessed.
(b) Please clarify whether the label source image of Zhang et al. and the GLB image of each chip are the same scene, and how geometric consistency was checked (this links to the date gap in Comment 2a).
(c) The offset between simplified polygons and the true shoreline is likely widespread for small lakes, not exceptional. If precise co-registration was not performed, please list this explicitly as a limitation in the Discussion and explain its impact on the SAR supervision signal and multi-modal fusion. Such offsets may be negligible for global change analysis, but for a pixel-wise segmentation benchmark they are a direct source of label noise that affects both training and evaluation. (The inherent geometric distortion of SAR in mountains is discussed in Comment 9.)4. The GLBC construction criteria are neither reproducible nor exhaustive
Moving cloud and frozen conditions into a separate challenge set (GLBC) is reasonable, but the current manual selection of a subset has two problems:
(a) The selection criteria are not quantitative. Please provide reproducible rules (e.g., per-chip cloud-percentage threshold, frozen fraction, shadow criteria), the screening workflow (automatic screening plus manual check?), and consistency checks.
(b) Exhaustiveness is not demonstrated. How many similarly contaminated chips remain in GLB? Those chips take part in both training and the GLB benchmark, affecting baseline scores and the conclusion that "GLBC isolates genuinely difficult conditions". I suggest an automated contamination screening of the full dataset, moving all chips that meet the criteria into GLBC — or at least reporting the residual fraction.
Please also state explicitly in the text that GLBC is used for evaluation only and that the corresponding chips were excluded from training (currently this is only implied in the Methods).5. Label validity for frozen-lake chips
The labels in GLB/GLBC come from ice-free ablation-season water extents. When a lake is partially or fully frozen, the real water surface may differ, so labels and image content can be inherently inconsistent. Please describe how labels for frozen scenes were checked, or quantify the label uncertainty of this subset as a limitation.
6. The model comparison confounds "model quality" with "input information"
In the current design, DOFA uses all 11 channels, Prithvi is restricted to a fixed 6-channel input (and is the only model for which a band-subset search was done), and U-Net/DeepLabv3+ use all 11 channels directly. Input information and model/pre-training quality are therefore not separated. The final combination Blue-Red-NIR-SWIR1-Slope-VV feeds slope and VV into an optically pre-trained encoder (channel substitution); it works empirically, but it shifts the input distribution away from pre-training. Please:
(a) Add a channel-matched experiment: either run all models on the same 6 channels, or inflate Prithvi's input stem from 6 to 11 channels (copy + noise initialization for the new channels; cheap under full fine-tuning) and retrain. If the authors claim GFM superiority, the untested possibility that "Prithvi might perform even better with all 11 channels" should be excluded;
(b) Explain how the optimal band combination was chosen, and on which data split. If bands were selected by test performance, there is information leakage; selection should be based on the validation set;
(c) Add a Transformer trained from scratch (e.g., SegFormer) to separate the value of GFM pre-training from the value of the Transformer architecture itself;
(d) Add simple baselines (e.g., NDWI/Otsu thresholding) as a lower-bound reference, so the community can gauge the difficulty of the benchmark.7. Normalization granularity is unspecified, and the elevation channel may suffer from dynamic-range imbalance
Section 2.1 only says the stacked bands were min-max stretched to 0–1, without stating whether this was done per 256×256 chip, per region, or globally:
Per-chip normalization removes absolute elevation semantics: the elevation channels of Iceland (hundreds of meters) and the Himalaya (5000 m+) become incomparable, leaving only within-chip relief;
Global normalization squeezes the densely populated elevation band (e.g., 2800–3500 m maps to only about 0.32–0.40) into a very narrow value range, unbalanced against full-range channels such as NDWI. Combined with weight decay's penalty on large weights, the elevation channel may be systematically under-used.Please state the normalization granularity and how the per-channel statistics were computed, discuss the effective dynamic range of the elevation channel (1–99 percentile clipping could prevent outlier elevations from stretching the range), and provide the normalization parameters in the README for reproducibility.
8. Statistical rigor is insufficient to support the model ranking
(a) All experiments use a single random seed (42), while the key gaps between models are only 0.02–0.03 mIoU (Prithvi 0.85 vs. DOFA/U-Net 0.82). Please repeat the experiments with at least 3 seeds, report mean ± std, and comment on the significance of the model rankings.
(b) Hyperparameters are identical for all models (lr 1e-5, batch 16, 100 epochs, StepLR, focal loss). Please show that this learning rate is also optimal for the CNN baselines; using identical hyperparameters is not the same as fair tuning.
(c) For the region- and size-grouped conclusions (Table 5, Fig. 13), please report per-group sample sizes and dispersion. The quality check uses 50 pairs per region (900 pairs in total), while regional image counts differ by more than 30× (Alaska 3,323 vs. Middle East 97); please clarify how the global mIoU of 0.95 was estimated (weighted or not) and provide confidence intervals.9. SAR geometric distortion in mountains affects detection under cloud — suggest adding a geometric-distortion mask band
Mountain SAR images suffer from severe geometric distortion (layover/shadow), determined jointly by the radar incidence angle and the DEM. The RTC product used in the manuscript only corrects radiometry; the geometric distortion itself is not removed. This matters especially in this dataset's use case: when optical imagery is reliable, the model can cross-check optical against SAR and the impact of distortion is limited; but when optical imagery is degraded (dense cloud — exactly the scenario where the manuscript stresses the value of SAR and built GLBC), lake detection relies entirely on SAR features and distortion effects are strongly amplified. This connects directly to the statement in Sect. 5.3 that detection "relies entirely on SAR backscatter" when optical data is unusable, and it may be an important cause of the missed small lakes under cloud in Fig. 12. The authors already have both Sentinel-1 imagery and the DEM, so a geometric-distortion mask can be built fairly easily from the incidence angle and terrain parameters. I suggest: (a) adding a distortion mask band (e.g., a 12th channel marking layover/shadow/normal areas), which can be used in training (loss masking or attention) and in evaluation (reporting metrics separately inside and outside distorted areas); or at least (b) quantifying the distorted fraction of the study regions and discussing, in the Discussion, how distortion affects the conclusions about detection capability under cloud.
Minor Comments
1. Inconsistent naming: "Glacial Lake-Bench" / "Glacial-Lake-Bench" / "Glacial Lake Benchmark" are used interchangeably, and likewise for GLBC (the Fig. 10 caption even says "Glacial Lake-Challenge dataset"). Please standardize throughout.
2. Reference issues: the text cites "Tang et al. (2024b)" but only 2024a is listed in the references; Dosovitskiy et al. (2020) and He et al. (2021) appear in the reference list but are never cited in the text. Please check all entries.
3. Equation (3) is typeset confusingly (multiplication shown as ".", misplaced parentheses, and stray characters), and equations (2) and (3) are never cited by number in the text. Please re-typeset following the standard Boundary F1 formula and cite all equations by number.
4. The text says "R² of 0.99" while Fig. 4a shows R² = 0.998. Also fix missing spaces such as "R2of" throughout.
5. Sentinel-1 images are acquired within a ±5-day window: please report the actual S1–S2 date-gap distribution and discuss how water-level changes or freeze-up affect the SAR channel and label consistency (related to Comment 3).
6. The BF1 tolerance of 2 pixels is relatively large for lakes of 0.002–0.01 km² (only a few to a few dozen pixels). Please provide a tolerance sensitivity analysis, or report boundary accuracy grouped by lake size.
7. A comparison table against existing datasets (e.g., the Glacial Lake Image dataset of Ma et al., 2025) — coverage, sensors, sample size, label source, task type — would support the claim of "the first global multi-sensor benchmark".
8. There are minor typos (e.g., "Zhang et al., (2024)" with an extra comma; "consists of 19,115 image-label pairs, consists of 11 channels" repeats the verb). Language polishing is recommended.Declaration: These review comments were written and finalized by the reviewer; AI tools were used only for polishing the language of review comments.
Citation: https://doi.org/10.5194/essd-2026-474-RC2 -
RC3: 'Comment on essd-2026-474', Anonymous Referee #3, 29 Sep 2026
reply
This manuscript presents a timely and valuable global multi-sensor benchmark (GLB/GLBC) alongside comprehensive baseline evaluations for automated glacial lake mapping. Constructing a standardized global dataset across diverse RGI regions addresses a critical gap in cryospheric Earth observation and aligns nicely with the scope of ESSD. The evaluation framework is ambitious, and the insights on spatial transferability and multi-modal integration provide useful guidance for future model development. To ensure the benchmark serves as an authoritative and dependable community standard, several methodological aspects would benefit from further clarification and refinement. In particular, inherited label uncertainties in complex terrain (especially within the challenge subset) warrant more transparent handling, and the physical rationale behind certain band configurations, such as mixing SAR polarizations across latitudes and assigning optical proxy wavelengths to DEM layers in DOFA, deserves clearer justification. Additionally, reporting regression performance without excluding prediction outliers (z > 3) would offer a more objective view of model robustness, while the impact of spatial sampling underrepresentation in cloud-prone, high-hazard regions like the eastern Himalayas should be discussed more thoroughly. Addressing these points will significantly strengthen the manuscript and data utility.
- In this work, the lake labels are directly inherited from the semi-automated global dataset of Zhang et al. (2024), rather than independently delineated from scratch. The authors’ own quality assessment of 900 stratified chips reveals that ~20% of the examined chips required manual correction. This indicates that the remaining >18,200 chips in the unverified training and evaluation pools inevitably harbor considerable label noise and systematic commission/omission errors. Training and evaluating large-scale deep learning models against noisy ground truth fundamentally undermines the objectivity of model comparison. The authors should clearly isolate and release the manually validated 900 chips (or ideally an expanded pool of 1,200–1,500+ chips) as a distinct, noise-free test set, and re-evaluate all baseline models against this clean reference to verify whether benchmark rankings and error distributions are influenced by label inaccuracies.
- The authors state that for Arctic regions where VV/VH data are unavailable, Horizontal-Horizontal (HH) and Horizontal-Vertical (HV) polarization images are substituted into Bands 10 and 11 (Lines 161–163). In radar remote sensing, copolarized channels (VV vs. HH) exhibit fundamentally distinct scattering physics: VV backscatter is governed by Brewster angle effects and vertical dielectric roughness, whereas HH scattering differs substantially over water surfaces, ice covers, and wind-roughened waves. Furthermore, Sentinel-1 EW mode data (frequently used in the Arctic) have different spatial resolutions and radiometric noise floors (NESZ) compared to IW mode data. Feeding HH/HV into the exact same input channels previously reserved for VV/VH without explicit channel flags or separate architectural conditioning induces severe physical and semantic ambiguity for the neural networks. This polarization mismatch likely explains why model performance drops precipitously in Svalbard and Arctic Canada North in Table 5. The authors could provide an ablation analysis isolating the impact of polarization mode switching and thoroughly discuss the scattering mechanism discrepancies in polar glacial lake delineation.
- In Section 4.1 and Section 4.2 (Lines 367–369, Fig. 7, and Fig. 9), the authors report R2, MAE, MSE, and MPE after removing "outliers (z-score > 3)", which eliminated 24 to 41 samples per model. In machine learning benchmark evaluations, severe prediction discrepancies (e.g., massive false positives over topographic shadows or total omission of lakes where predicted area equals zero, as visible along the axes of Fig. 7 and Fig. 9) represent critical failure modes of the models under evaluation. Discarding these failed predictions artificially inflates R2 (up to 0.98) and deflates error metrics. Benchmarking must report unadulterated metrics across the entire, uncurated test set. If the authors wish to highlight catastrophic outliers, they should do so in a dedicated error-attribution / failure-case analysis rather than excluding them from primary summary statistics.
- The authors enforce a strict cloud-cover threshold (<=15%) during the 2020 ablation season to optimize optical quality. While understandable from an optical perspective, this criterion induces acute spatial representation bias. Table 1 reveals that South Asia East (encompassing the eastern and central Himalayas and southeastern Tibetan Plateau—the global epicenter of catastrophic GLOF events and rapid glacial lake expansion) achieves only 9.21% area coverage. Similarly, Iceland (11.67%), Scandinavia (19.08%), and Western Canada/USA (14.07%) exhibit very poor representation. Consequently, regions characterized by active maritime/monsoonal cryospheres and persistent cloud cover are almost entirely excluded from this "global" benchmark. The authors must explicitly discuss this geographic undersampling bias and its ramifications for global model deployment.
- GLB is strictly constrained to a single timestamp during the 2020 ablation season. However, glacial lakes are dynamic systems characterized by seasonal freeze-thaw cycles, glacier-calving lake ice, and turbidity fluctuations driven by seasonal meltwater runoff. Although the authors acknowledge this limitation in Section 2.3, the inclusion of "frozen lakes" and "partially frozen lakes" in GLBC poses definition dilemmas. In the ablation season, are frozen lakes perennial proglacial features, or do they represent late spring / early autumn snow-covered basins? When lake surfaces are completely frozen and snow-covered, optical contrast vanishes. How did Zhang et al. (2024) ensure lake boundary fidelity under such conditions? Furthermore, does the deep learning model actually identify glacial lakes, or does it merely exploit DEM depressions to memorize topographical basins? The authors should clarify how ground-truth validity was established for frozen lakes and discuss the risk of models memorizing topographical depressions in the absence of distinct optical/radar water signatures.
- In Line 146, the text specifies that Sentinel-1 RTC images were acquired "within a ±5-day temporal window". However, the workflow diagram in Fig. 1 explicitly states "Acquiring Sentinel-1 (RTC) images using Planetary Computer API image (±10 days)". Please eliminate this contradiction and confirm the exact temporal matching window.
- In Fig. 1, under the "Topographic Layers" block, the labels indicate "Band 8(Elevation)" and "Band 9(Elevation)". Line 145 of the text correctly states that slope and elevation are stacked as Bands 8 and 9, respectively. Fig. 1 should be corrected to show "Band 8 (Slope)" and "Band 9 (Elevation)" (or vice versa). Additionally, the top box contains a typographical error: "Acquiring Sentinle-2 data" should be corrected to "Sentinel-2".
- Lines 163–164 mention that stacked images are normalized using min-max stretching to [0, 1]. Please specify whether this stretching was performed patch-by-patch (locally) or globally across the entire collection. Local min-max normalization alters absolute spectral reflectance and SAR backscatter magnitudes, distorting physical comparability across different chips.
- Equation (1) presents the standard multi-class mIoU formula. For binary segmentation tasks dominated by background pixels (>95%), computing the arithmetic mean of the foreground (lake) and background IoU heavily inflates the reported metric due to near-perfect background IoU. Please clarify whether mIoU in Tables 3, 4, and 5 represents the lake-class IoU or the two-class average. If it is the two-class average, foreground lake IoU should be reported separately.
- The manuscript employs conflicting thresholds for defining "small lakes": >0.002 km2 (Line 54), <0.03 km2 (Line 82), <=0.05 km2 (Line 95), 0.002–0.05 km2 (Line 403), and 0.002–0.01 km2 (Line 450, Line 557). Please adopt a unified, hierarchical definition throughout the text and figures (e.g., ultra-small: <0.01 km2; small: 0.01–0.05 km2; medium: 0.05–0.5 km2; large: >0.5 km2).
- In Table 5 (LORO cross-validation), Svalbard exhibits the lowest mIoU across all architectures (0.64–0.66). The discussion on Lines 455–457 briefly mentions this drop without scientific attribution. Please expand the cryospheric discussion regarding high-latitude polar constraints, including widespread permafrost thermokarst water bodies, persistent sea/lake-ice confusion, extreme solar zenith angles, and SAR layover/shadow effects in high-relief coastal fjords.
Citation: https://doi.org/10.5194/essd-2026-474-RC3
Data sets
Glacial-Lake-Bench: A Global Multi-Sensor Benchmark Dataset for Evaluating Deep Learning Models for Glacial Lake Mapping Saurabh Kaushik https://zenodo.org/records/17917359
Model code and software
dl4eo Saurabh Kaushik https://github.com/Sk-2103/dl4eo
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 406 | 150 | 93 | 649 | 59 | 119 | 82 |
- HTML: 406
- PDF: 150
- XML: 93
- Total: 649
- Supplement: 59
- BibTeX: 119
- EndNote: 82
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General comments:
Overall, this is a well-designed and insightful study that makes a valuable contribution to the field of glacial lake mapping. The proposed global glacial lake dataset is particularly meaningful because it helps establish a community benchmark for evaluating model robustness under challenging mapping conditions. I also appreciate the development of the Glacial Lake-Bench-Challenge (GLBC) dataset, which is designed to address several key difficulties in glacial lake mapping, including cloud cover, shadows, frozen surfaces, and small lake size.
One point that would benefit from further clarification concerns the role of the 11-band input in improving small-lake detection. I understand that the SAR band can facilitate glacial lake mapping under cloud-cover conditions. However, it remains unclear how the full set of input bands specifically enhances the detection of small glacial lakes. It would be helpful if the authors could clarify the underlying mechanisms or provide additional evidence, for example by explaining whether these bands improve spectral separability, boundary delineation, or model sensitivity to small water bodies.
A second aspect concerns the evaluation of model generalization. The model demonstrates good spatial transferability, likely benefiting from the globally distributed labels. To further strengthen the evaluation, I suggest adding evidence of the model’s temporal transferability, for example by testing images acquired in years other than those used for training or validation. In addition, a more detailed discussion of the imbalanced model performance across different regions would be valuable. Potential factors may include regional differences in lake morphology, image quality, cloud and shadow conditions, training-sample distribution, and background complexity. Discussing these factors would help readers better understand the limitations of the current model and possible directions for future improvement.
Finally, beyond serving as a baseline dataset for assessing newly proposed models, the GLB dataset has strong potential to function as a high-quality and representative training resource for developing more advanced glacial lake mapping models. I encourage the authors to discuss this possible application more explicitly, particularly how the dataset’s global coverage, challenging scenarios, and annotation quality may support future model training, comparison, and benchmarking.
Specific comments:
Line 24: The statement “Labels are derived from Zhang et al. (2024)” is somewhat ambiguous. Please clarify whether the labels were directly adopted from Zhang et al. (2024), manually delineated in the present study, or further revised based on that reference dataset. A brief description of the label-generation procedure would improve transparency and reproducibility.
Lines 65–83: The literature review would be strengthened by including additional studies on deep-learning-based glacial lake mapping. The following works may be relevant and could help better position the contribution of the present study within the existing body of research:
Qayyum, N., Ghuffar, S., Ahmad, H.M., Yousaf, A., Shahid, I., 2020. Glacial lakes mapping using multi satellite PlanetScopeScope imagery and deep learning. ISPRS Int. J. Geo-Inf. 9 (10), 560.
Wu, R., Liu, G., Zhang, R., Wang, X., Li, Y., Zhang, B., et al., 2020. A deep learning method for mapping glacial lakes from the combined use of synthetic-aperture radar and optical satellite images. Rem. Sens. 12 (24), 4020.
Chen, F., 2021. Comparing methods for segmenting supra-glacial lakes and surface features in the mount everest region of the himalayas using chinese gaofen-3 sar images. Rem. Sens. 13 (13), 2429.
Xu, X., Liu, L., Huang, L., Hu, Y., 2024. Combined use of multi-source satellite imagery and deep learning for automated mapping of glacial lakes in the Bhutan Himalaya. Sci. Remote Sens. 10, 100157.
Line 95: Please specify the minimum mapped lake size for both the Glacial Lake-Bench (GLB) dataset and the Glacial Lake-Bench-Challenge (GLBC) dataset. Clarifying whether the two datasets use the same or different minimum-size thresholds would help readers better understand their comparability and intended applications.
Lines 126–128: Since the glacial lake boundaries were obtained from Zhang et al. (2024), while the 2020 Sentinel-2 images were collected independently in this study, please explain how potential inconsistencies between the labels and image acquisition dates were addressed. For example, were the lake boundaries visually checked or adjusted to match the 2020 imagery, and how were cases of lake expansion, shrinkage, or disappearance handled?
Lines 126–128: Please also clarify the specific months defined as the ablation season. Because the timing of the ablation season varies substantially across glaciated regions worldwide, it would be useful to explain whether a uniform time window was applied globally or whether region-specific seasonal windows were used.
Line 132: Please clarify what is meant by the “specified date range” in this context. If the date range does not correspond exactly to 2020, please indicate which year or time period the labels represent and how this relates to the 2020 imagery used in the analysis.
Lines 178–190: The spatial coverage of the GLB dataset appears to vary substantially among regions, ranging from 9.21% to 89.42%. Please explain the main reasons for this large regional variation. For example, is it primarily caused by cloud-cover conditions during image selection, by imbalanced regional data coverage inherited from Zhang et al. (2024), or by other factors such as image availability, terrain complexity, or quality-control criteria?
Section 2.4: Please provide more information on the time and labor required for the manual correction process. For instance, it would be helpful to report the approximate number of annotators involved, the total or average correction time, and whether any quality-control or cross-checking procedures were applied.
Figure 4: Please clarify whether the term “boundary” in the upper legend refers to the “adjusted boundary.” If so, using consistent terminology between the figure legend and the main text would help avoid confusion.
Figure 5: In the upper-right example, the manually corrected boundary appears not to align well with the background image. Please check whether this is due to a visualization offset, image–label misregistration, or an error in the displayed boundary, and revise the figure or explanation if necessary.
Line 257: Please clarify the input configuration used for the DeepLabv3+ model. Does DeepLabv3+ require a three-band input in this study, and if so, were three bands selected from the previously stacked 11-band input? If only a subset of bands was used, please specify which bands were selected and explain the rationale.
Figure 8: If space allows, please add the corresponding scenario labels to the images in the first column. This would make the figure easier to interpret and help readers more directly compare model performance across different challenging conditions.
Line 313: Because the performance of the DeepLabv3+ model appears broadly comparable to that of the other models, rather than substantially worse, please consider adding leave-one-region-out cross-validation (LORO) results for DeepLabv3+ as well. Including this comparison would make the experimental evaluation more consistent across model architectures and would provide a fairer assessment of spatial transferability.
Figure 11: Please clarify which scenario Figure 11(i) is intended to represent. Adding a label or brief explanation in the caption would help readers understand the purpose of this example and how it relates to the other scenarios shown in the figure.
Lines 581–582: To better demonstrate the model’s temporal transferability, please consider adding tests using images acquired in years other than those used for training or validation. If such experiments are not feasible, it would still be helpful to discuss this limitation and explain how temporal variability may affect model performance.
Technical corrections:
Line 74: Please correct the typographical and punctuation errors in the following sentence: “Similarly, the method proposed by Jiang et al. (2025) also showed severe limitations in spatial transferability.”
Figure 1: Please check whether Band 9 should be labeled as “slope.” If so, please revise the figure or caption accordingly to ensure consistency with the input-band description.