A compiled dataset of Greenland weather station observations and reconstructed surface air temperature
Abstract. Climate variability is a main driver of mass imbalance of the Greenland Ice Sheet. However, quantitative assessments of climate change in this region are constrained by the sparseness, heterogeneity, and discontinuity of in situ measurements, imposing significant challenges to reliable modelling and future projections. To address this, we compile all available weather station observations over Greenland from 1941 to present into a single and standardized dataset at both daily and monthly temporal resolutions. This dataset includes seven climate variables: temperature, wind speed, wind direction, relative humidity, radiation, precipitation, and pressure, which is accessible at https://doi.org/10.11888/Atmos.tpdc.303490 (Wang and Wang, 2026). All data were harmonized and processed by a uniform quality control procedure, with the outliers flagged. For each station, detailed metadata covering source, location, and data availability are supplied. The dataset facilitates a thorough analysis of multi-variable climate characteristics in Greenland and can be used for the validation of global reanalysis products, regional climate models and remote sensing retrievals as well as data assimilation and climate reconstruction. Furthermore, we use the temperature observations from the compiled dataset together with a digital elevation model and ERA5 reanalysis to reconstruct a high-resolution (0.1°) monthly surface air temperature (SAT) dataset for Greenland (1950–2024), based on the optimized method identified after comparing five distinct machine learning models. The resulting SAT reconstruction exhibit a great improvement over ERA5, reducing the root mean square error (RMSE) by more than 60 % and increasing squared correlation coefficient (R2) by 5 %. The high accuracy of this reconstruction is further confirmed by spatial cross-validation across 11 distinct regional divisions. The reconstruction is also publicly available at the Third Pole Environment Data Center (https://doi.org/10.11888/Cryos.tpdc.303026, Wang and Wang, 2025).
Summary
The paper presents a new compilation of Greenland near-surface observations from various stations and uses these to construct a monthly surface-air-temperature reconstruction spanning from 1950 to 2024 on a 0.1° grid. The authors harmonize and compile the rather fragmented observations into a single data set.
The authors train several machine learning models using the station data and ERA5-Land to reconstruct the Greenland-wide temperature. They report substantially lower errors than ERA5 and use their reconstructed product to examine spatial climatologies and long-term temperature trends.
The paper is discussing an important and timely topic. The paper is generally easy to follow and the language is mostly clear but needs to be more scientifically accurate in some parts.
However, I have several major issues with the paper that need to be addressed before being considered for publication.
Major issues
The distinction between the validation and test sets is unclear and should be clarified. For the LSTM-CNN and Transformer-CNN, the data are split into training (86%), validation (4%), and test (10%) sets, whereas the other three models use only a training (86%) and test (14%) split. However, the manuscript subsequently states that all five models underwent iterative training and tuning until predefined performance criteria were reached. It is therefore unclear how the three models without a validation set were tuned.
If the reported test set was inspected or used to guide hyperparameter tuning, model development, or model selection, it no longer constitutes an independent test set and should instead be considered a validation set. A genuinely unseen test set should only be evaluated once model development and model selection are complete.
Relatedly, the manuscript refers to the combined validation and test data for the LSTM-CNN and Transformer-CNN as "unseen data." This terminology is misleading because the validation set is explicitly used for early stopping and therefore contributes to model development. Validation and test performance should be reported separately.
I therefore recommend that the authors clearly describe the role of each data subset and ensure that model tuning and model comparison are performed using training/validation data (or cross-validation), followed by evaluation on a separate, untouched test set.
Specific comments:
L22 Shouldn't it be "a method" not "the method" since you didn't talk about it before
L23 I suggest you avoid terms like "great improvement"
L.25 missing "the"
L.35 There are more recent references than the Lenton papers, e.g. Boers et al. 2025, Bochow et al. 2023, Global Tipping Point Report 2025
L.52 I guess the "at" should be an "and"
L.53 Specifically instead of "specially"
I am missing some discussion on more recent temperature reconstructions from incomplete spatial data. There is quite some advancement in the last years using CNNs and generative methods. Some recent references are Kadow et al. 2020, Bochow et al. 2023, Qian et al. 2026
L.120 how did you merge the information?
L. 133f Do you actually aggregate precipitation as average or sum?
Section 5.1
It is not completely clear to me if you use ERA5 or ERA5-Land.
L. 332 I find it difficult to merge the term validation and testing data here. The validation data is indirectly involved in the training since you tune the model parameters using the validation set. I would suggest the authors keep the distinction between test and validation data clear throughout the manuscript.
L.340 Refers to a previous major comment. I wonder how and if the other variables besides T2M even improve the SAT reconstruction or if T2M is just enough, did the authors test this?
L. 363 Here you make the distinction between validation and test data again but one sentence later, you refer to unseen data again.
L.366 How did you get your definition of "good"?
L.558 The numbers should be in the exponent
Table 2
Why do you exclude radiation values =0 Wm-2?
Figures