the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A 1 km resolution dataset of Northern Hemisphere permafrost active layer thickness (2000–2024)
Abstract. Active layer thickness (ALT) is a key indicator of permafrost change, with important implications for soil hydrothermal conditions, carbon feedbacks, ecosystem processes, and cold-region infrastructure. However, hemispheric-scale ALT mapping remains challenging because field observations are sparse and unevenly distributed, while thaw depth is controlled by complex and nonlinear environmental interactions. Here, we compiled 2,196 annual ALT observations and developed an ensemble spatiotemporal machine-learning framework to reconstruct a continuous annual ALT dataset at 1 km resolution for the Northern Hemisphere permafrost region from 2000 to 2024. Spatial leave-one-site-out cross-validation yielded an ensemble R2 of 0.76 and an RMSE of 60.48 cm. Independent temporal evaluation showed a significant correlation between observed and predicted Sen’s slopes (r = 0.73, p < 0.001), with 86.8 % agreement in trend direction. The dataset reproduced broad latitudinal and elevational patterns, biome-related differences, and interannual variability. Comparisons with two existing 1 km hemispheric products showed that our reconstruction occupied an intermediate position in hemispheric mean ALT while preserving relatively fine spatial variability. Substantial inter-product differences in both magnitude and spatial structure further highlighted persistent structural uncertainty in large-scale ALT mapping. Pixel-wise uncertainty maps further provide spatially explicit information on prediction confidence. This dataset provides a spatially detailed, temporally continuous, and uncertainty-explicit resource for regional- to hemispheric-scale studies of permafrost dynamics, carbon-cycle feedbacks, ecosystem and hydrological responses, and infrastructure exposure.
- Preprint
(1900 KB) - Metadata XML
-
Supplement
(990 KB) - BibTeX
- EndNote
Status: open (until 27 Sep 2026)
- RC1: 'Comment on essd-2026-447', Anonymous Referee #1, 24 Aug 2026 reply
-
RC2: 'Comment on essd-2026-447', Anonymous Referee #2, 18 Sep 2026
reply
This manuscript presents a new annual active layer thickness (ALT) dataset for the Northern Hemisphere permafrost region from 2000 to 2024. The dataset was developed using observations from 201 monitoring sites and a set of climatic, vegetation, topographic, soil, and ecological predictors, with LightGBM and CatBoost combined in an ensemble framework. The final product provides annual ALT estimates on a 1 km grid, together with model evaluation and uncertainty information. The dataset could be useful for studies of permafrost dynamics and associated environmental changes at regional and hemispheric scales. The public availability of both the data and code is also appreciated. However, I have several concerns regarding the validation strategy, uncertainty characterization, spatial representativeness, and interpretation of the dataset that should be addressed before its reliability and applicability can be fully evaluated. My main comments are as follows:
- I am not convinced that the current temporal validation adequately supports the extension of the dataset to 2021–2024. In the LOSO framework, all observations from one site are excluded, but observations from the same years at other sites remain in the training dataset. Therefore, this validation primarily evaluates whether temporal variations learned from other locations can be transferred to an unseen site, rather than whether a model trained on historical observations can reliably predict future years. This distinction is important because most of the observations used for model development are from 2000–2020, whereas the final product extends to 2024, and the authors themselves acknowledge that the 2021–2024 maps represent temporal extrapolations. I recommend that the authors perform an additional temporally blocked validation, for example, by training the model using observations up to a given year and evaluating it against observations from subsequent years. Such an experiment would provide a more direct assessment of the model's temporal extrapolation capability. Without such a test, the reliability of the 2021–2024 product remains difficult to assess.
- The model appears to substantially underestimate interannual variability. The reported median fluctuation amplitude ratio (FAR) is only 0.628, indicating that the detrended variability of the reconstructed ALT time series is considerably smaller than that of the observations at many validation sites. This is an important limitation for a dataset provided at an annual temporal resolution. In its present form, the manuscript sometimes gives the impression that the dataset can reliably reproduce interannual ALT variability, whereas the validation results indicate substantial smoothing of year-to-year fluctuations. I recommend that the authors revise the relevant statements throughout the manuscript and make it clear that the product appears to be more suitable for characterizing broad spatial patterns and long-term temporal changes than for investigating anomalous individual years or short-term ALT variability.
- The uncertainty analysis does not appear to represent the full predictive uncertainty of the ALT product. The authors randomly selected 80% of the monitoring sites 100 times, refitted the models, and then used the 2.5th and 97.5th percentiles of the resulting predictions as an empirical 95% prediction interval. As implemented, this procedure primarily quantifies the sensitivity of the predictions to changes in the composition of the training-site network. It does not explicitly account for several other potentially important sources of uncertainty, including measurement uncertainty in the ALT observations, residual model error, uncertainty in the predictor datasets, unresolved sub-grid heterogeneity, and uncertainty associated with temporal extrapolation. Therefore, referring to this range as a “95% prediction interval” may give readers the impression that it represents the overall predictive uncertainty of the final product, which is not demonstrated by the current analysis. I recommend that the authors more clearly define the component of uncertainty quantified by this procedure and, if possible, evaluate the empirical coverage of the reported interval using withheld observations. If such calibration is not feasible, the terminology should be revised to more accurately reflect the uncertainty component being represented.
- The overall LOSO-CV statistics may conceal substantial spatial differences in model performance. The authors report an ensemble of 0.761 and an RMSE of 60.48 cm by pooling all annual observations from the withheld sites. However, sites with longer observation records contribute more samples to these pooled statistics, and substantial environmental differences among Arctic, sub-Arctic, and high-elevation permafrost regions may be obscured by a single hemispheric-scale performance metric. The manuscript also shows that large ALT values tend to be underestimated, suggesting that model performance is not uniform across the observed ALT range. I recommend that the authors provide additional validation statistics stratified by major geographic region, permafrost zone, or ALT range. Site-level or site-weighted performance metrics could also be reported in addition to the pooled record-level statistics. For a hemispheric-scale data product, users need to understand not only its overall accuracy but also the geographic and environmental conditions under which prediction errors are relatively large.
- The relationship between the reported spatial uncertainty pattern and the distribution of the observations requires further examination. The authors note that monitoring sites are relatively sparse in regions such as interior Siberia and the Tibetan Plateau. However, the uncertainty maps indicate relatively low uncertainty across extensive areas of northern Siberia, whereas high uncertainty is concentrated in more scattered regions. This pattern is not necessarily incorrect, but low disagreement among repeatedly resampled models does not by itself demonstrate high predictive reliability in data-sparse regions. Different model realizations may produce similar predictions even when all of them are extrapolating beyond well-sampled portions of the predictor space. I therefore suggest that the authors examine the relationship between the uncertainty estimates and observation density or distance to the nearest training sites. An analysis of environmental novelty, extrapolation distance, or area of applicability would also be useful for distinguishing regions dominated by interpolation from those requiring substantial extrapolation and would substantially improve the interpretation of the uncertainty product.
- The comparison with the Westermann et al. and Peng et al. products is useful, but the harmonization procedure should be described more explicitly. The three products show substantial differences in both magnitude and spatial pattern, with pixel-wise RMSE values exceeding 90 cm in both comparisons. The manuscript states that the pixel-wise comparison was conducted using common valid pixels for the 2000–2020 period, but it is less clear whether an equivalent spatial and temporal harmonization was applied when calculating the hemispheric and regional mean ALT values shown in the broader product comparison. Differences in permafrost masks, valid spatial domains, temporal averaging periods, or the treatment of deep thaw and talik conditions could themselves contribute to the reported discrepancies. I agree with the authors that inter-product comparisons should not be interpreted as independent validation. Nevertheless, the procedures used to harmonize the three products should be described in sufficient detail to allow readers to distinguish genuine differences among ALT estimates from differences caused by spatial or temporal domain definitions.
- Several methodological details require further clarification to ensure reproducibility. First, the manuscript should specify which SoilGrids depth interval, or which depth-weighting procedure, was used to derive bulk density, coarse fragment content, sand, silt, and clay. Second, the authors should explain how the categorical biome variable was encoded or handled in each machine-learning algorithm. Third, MATmax and MATmin are defined as the maximum and minimum monthly mean temperatures within a year rather than the absolute annual maximum and minimum air temperatures; therefore, the current terminology may be ambiguous and should be revised or more explicitly defined. Finally, because winter snow depth is calculated from October of the previous year through April of the current year, the authors should clarify how the winter snow-depth predictor for the year 2000 was derived.
Citation: https://doi.org/10.5194/essd-2026-447-RC2
Data sets
A 1 km resolution dataset of Northern Hemisphere permafrost active layer thickness (2000–2024) Y. Wei and Y. Chen https://doi.org/10.5281/zenodo.21667583
Model code and software
Code for reconstructing Northern Hemisphere active layer thickness at 1 km resolution (2000–2024) Y. Wei https://zenodo.org/records/21835259
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 193 | 100 | 49 | 342 | 31 | 31 | 74 |
- HTML: 193
- PDF: 100
- XML: 49
- Total: 342
- Supplement: 31
- BibTeX: 31
- EndNote: 74
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript compiles 2,196 annual active layer thickness (ALT) records from 201 GTN-P/CALM monitoring sites, constructs an environmental dataset with 12 predictor variables, and applies a LightGBM-CatBoost ensemble to produce annual 1 km resolution ALT maps for the Northern Hemisphere permafrost region. The authors have made all code and training data publicly available, which is commendable. My main comments are as follows: