the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Mapping forest canopy height over Europe by integrating Sentinel-1, Sentinel-2, GEDI, and ICESat-2 data
Abstract. Timely, accurate, and spatial explicit information on forest structure, such as canopy height, is important to understand and respond to ongoing changes in forests and to support the mapping of habitat structure. The availability of spaceborne LiDAR data, such as those from GEDI, has stimulated the development of continental to global canopy height maps. Yet, while GEDI data are often used to train canopy height models, these data are lacking in northern areas. In this study, we mapped canopy height over Europe at 10 m resolution by combining Sentinel-1 and Sentinel-2 data and integrating training data from GEDI and ICESat-2. The integration of ICESat-2 and GEDI data mostly enhanced the model performance in the north of Europe, where GEDI data are lacking. The model reached a RMSE of 5.77 m and a MAE of 4.09 m based on an independent validation with ALS data over about 3,700 patches across Europe. The resulting canopy height map and validation dataset have been made publicly available at https://doi.org/10.5281/zenodo.13324731 and https://doi.org/10.5281/zenodo.18471620, respectively.
- Preprint
(2268 KB) - Metadata XML
-
Supplement
(424 KB) - BibTeX
- EndNote
Status: open (until 02 Oct 2026)
- RC1: 'Comment on essd-2026-329', Anonymous Referee #1, 07 Jul 2026 reply
-
RC2: 'Comment on essd-2026-329', Anonymous Referee #2, 21 Aug 2026
reply
The manuscript presents a new 10 m canopy height dataset for Europe by combining GEDI and ICESat-2 observations with Sentinel-1 and Sentinel-2 data. Considerable effort has been devoted to collecting and processing the training data as well as compiling an extensive independent ALS validation dataset across Europe. The resulting dataset is likely to be valuable for forest monitoring and ecosystem studies. However, I have several concerns regarding the methodological description, validation strategy, and interpretation of the results that should be addressed before publication.
Major comments
1. The manuscript states that the objective is to map canopy height of woody vegetation (trees and shrubs), while the title and abstract mainly refer to forests. It would be helpful to clarify this distinction throughout the manuscript, as the mapped product appears to include woody vegetation beyond forests. In addition, the training samples are selected based on the intersection of multiple land-cover products (e.g., ESA WorldCover 2020/2021 and the Copernicus High Resolution Forest Type layer from 2018). These products differ in acquisition year and classification methodology, which may introduce temporal inconsistencies and classification uncertainties.Â
2. The description of the GEDI and ICESat-2 training samples is insufficient. Please report the final number of training, validation, and testing samples for each dataset, together with their spatial distribution. A map similar to Figure 1 showing the spatial density or number of GEDI and ICESat-2 samples would help readers better understand the representativeness of the training data.
3. The manuscript mainly reports the independent validation results using the ALS dataset. Since the canopy height model was trained using GEDI and ICESat-2 observations, the evaluation on the held-out GEDI/ICESat-2 test dataset should also be presented more explicitly.Â
4. The independent validation is based on ALS RH98, whereas the compared canopy height products are not necessarily based on the same canopy height metric (e.g., RH95 in Potapov et al.). It would be helpful to summarize these differences in the Methods or Results to facilitate a fair interpretation of the comparisons.
5. The manuscript frequently refers to the canopy height maps from Pauls et al. in the Introduction. However, their height maps are not included in the quantitative comparison.Â
6. Figure 4 only presents the comparison using all validation samples. Since the manuscript emphasizes the improved performance in northern Europe, it would be helpful to include equivalent comparisons for northern and southern Europe for all reference products. In addition, please clarify whether exactly the same validation samples and filtering criteria (land cover) were applied to all products to ensure a fair comparison.
7. More methodological details are needed regarding the combined training using GEDI and ICESat-2 observations. Since the two datasets differ in footprint size, sampling characteristics, and measurement uncertainty, please clarify how the training samples from the two sensors were combined and balanced during model training. In addition, the key CatBoost hyperparameters should be reported to improve reproducibility.
8. According to Figure 3, the model trained solely with ICESat-2 achieves a very similar accuracy to the model trained with both GEDI and ICESat-2. Therefore, the additional contribution of GEDI is not entirely clear.Â
9. Figure 5 shows a clear saturation effect, with increasing underestimation for canopy heights above approximately 25 m. In this height range, the proposed model performs worse than Lang et al. (2023).Â
10. Figure 6 appears to exhibit noticeable block-like spatial patterns across large parts of Europe. It is unclear whether these patterns originate from the visualization (e.g., map resampling) or from the underlying prediction product. In addition, it would be useful to compare the canopy height distributions of different maps, which may help readers better understand where the differences originate.
11. It is not clear which dataset was used to generate the results shown in Figure 7. The reported accuracy differs from the results presented in Figures 3 and 4.Â
12. Overall, I believe the main strength of this dataset lies in the improved canopy height mapping over northern Europe through the integration of ICESat-2 observations. However, additional clarification of the methodology, more comprehensive validation, and improved comparisons with existing products are needed to fully demonstrate the advantages of the proposed dataset. Providing sufficient methodological details would also substantially improve the reproducibility and long-term value of this dataset.Â
Citation: https://doi.org/10.5194/essd-2026-329-RC2
Data sets
European canopy height map W. De Keersmaecker et al. https://doi.org/10.5281/zenodo.13324731
ALS-based canopy height across Europe L. Bertels et al. https://doi.org/10.5281/zenodo.18471620
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 167 | 102 | 9 | 278 | 115 | 20 | 15 |
- HTML: 167
- PDF: 102
- XML: 9
- Total: 278
- Supplement: 115
- BibTeX: 20
- EndNote: 15
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The authors present a study aimed at producing a 10 m forest canopy height map for Europe by integrating Sentinel-1, Sentinel-2, GEDI and ICESat-2 observations. Overall, this is a good study that complements the growing landscape of continental and global canopy height products. Whilst the overall modelling framework is largely based on established approaches, the work provides an important contribution by extending training to northern Europe through the integration of ICESat-2 data and by delivering an openly available European canopy height dataset.
In my opinion, one of the strongest aspects of the manuscript is the validation strategy. The authors assembled an extensive independent validation dataset from heterogeneous airborne LiDAR datasets collected across multiple European countries. Such a large-scale independent validation effort is relatively uncommon and substantially increases confidence in both the reported accuracy and the resulting data product.
The manuscript is generally well written, the methodology is clearly described, and the data product will likely be valuable for a wide range of ecological and forestry applications. My comments below are intended to further strengthen the manuscript.
Comments:
1. The independent ALS validation dataset is one of the principal strengths of this study. However, the ALS acquisitions span multiple years, while the canopy height map represents conditions in 2020. Although the manuscript acknowledges this limitation, I encourage the authors to discuss more explicitly how temporal mismatches may influence the reported validation statistics.Â
2. Clarify the added value of this dataset. The manuscript compares the proposed product with several existing canopy height datasets and demonstrates modest improvements in validation statistics. While these improvements are encouraging, I suggest that the Discussion better emphasise the broader advantages of the dataset. In particular, the integration of ICESat-2 substantially improves training coverage in northern Europe, the extensive independent validation dataset increases confidence in the results, and the provision of uncertainty estimates enhances the usefulness of the product. These aspects arguably represent the primary contribution of the study and deserve greater emphasis than the relatively small differences in RMSE.
Minor comments:
In the Abstract and the Introduction, "spatial explicit" should be corrected to "spatially explicit"
Please clarify whether the reported RMSE and MAE are computed over all validation pixels or only over pixels classified as woody vegetation.
Table 1 refers to the combined Sentinel-1/Sentinel-2 model as "S1S2-9B", whereas Figure 7 uses "S2S2-B9". This appears to be a typographical inconsistency and should be corrected.
The manuscript occasionally uses the terms height estimate, height prediction, and canopy height map interchangeably. Using consistent terminology throughout would improve readability.
The choice of adding random noise of ±1° to the latitude and longitude variables, together with the decision to saturate the sample weights above 35 m canopy height, would benefit from a brief justification.
A concise table summarising the main characteristics of this product and the existing canopy height datasets used for comparison (e.g. training data, spatial resolution, spatial extent, independent validation, and uncertainty estimates) would provide useful context for readers.
Â