HiMIC-Daily: A high-resolution (daily and 1 km) multi-indicator atmospheric moisture collection over China, 2003–2020
Abstract. Near-surface atmospheric moisture is a fundamental component of the hydrological cycle and plays a key role in regulating land-atmosphere exchanges and surface energy partitioning. Reliable daily high-resolution moisture data are essential for regional climate analysis and fine-scale applications, particularly for capturing short-term variability and extreme moisture dynamics. With complex terrain and a dense population, China is highly vulnerable to extreme hydro-meteorological extremes, yet existing moisture products over China are largely constrained by coarse temporal resolution, insufficient spatial detail, and limited indicators. Here, we present HiMIC-Daily, a seamless daily 1-km-resolution near-surface atmospheric moisture dataset for China, 2003–2020. HiMIC-Daily provides a comprehensive suite of six widely used indicators that characterize atmospheric moisture from different perspectives: actual vapor pressure (AVP), dew point temperature (DPT), mixing ratio (MR), relative humidity (RH), specific humidity (SH), and vapor pressure deficit (VPD). This dataset is generated using the Light Gradient Boosting Machine (LightGBM) framework, which integrates in-situ observations from 2419 meteorological stations with multiple environmental and temporal covariates, including ERA5-Land derived near-surface temperature and DPT, AVP, land surface temperature, topography, and day of year. Validation against observations shows that HiMIC-Daily achieves robust performance across all six indicators, with R2 values ranging from 0.877 to 0.989. The strongest performance is obtained for AVP, DPT, MR, and SH, with R2 values exceeding 0.985, and error metrics remain within acceptable ranges for all indicators (e.g., mean absolute error of 0.677 hPa and a root mean square error of 0.933 hPa for AVP). Compared with two existing coarse resolution products, HiMIC-Daily provides finer spatial detail, higher accuracy, and more realistic temporal variability across different climatic regions. These capabilities support spatially explicit studies of climate variability and environmental processes. The HiMIC-Daily dataset is publicly available at https://doi.org/10.11888/Atmos.tpdc.303449.
This manuscript presents a daily 1km dataset of six near-surface atmospheric moisture indicators over China for 2003-2020 by integrating a variety of data sources, such as meteorological stations, remote sensing, climate reanalysis, and population data. The topic is highly relevant to the journal, and the data itself will benefit broad research applications. The manuscript is clearly organized and well written. I have some comments/suggestions for improving the quality of this work.
Major comments.
Model validation data. The authors split daily samples randomly into 80% training and 20% validation subsets. Because daily observations from the same station, nearby stations, and adjacent dates are strongly autocorrelated, a sample-level random split is likely to place observations from the same stations and closely related dates in both subsets. Consequently, the reported R2, MAE, and RMSE values may characterize interpolation among familiar stations and dates rather than performance at other unseen locations. Have the authors tried adding station-held-out or spatially blocked validation, in which all dates from a few selected stations are excluded from training? Additionally, the authors should explain how errors and systematic biases in LST, ERA5 and other inputs may influence the model.
Baseline comparison. I suggest the authors do some baseline comparison against the input datasets. For example, compare native ERA5-Land values with ground stations; native ERA5 with resampled ERA5 values; resampled ERA5 with ground stations. These will help examine whether the main feature-importance results that show ERA-derived moisture variables are the dominant predictors are partly biased by the close mathematical and statistical relationships between these highly correlated variables.
More details are needed regarding harmonization of heterogeneous input datasets for reproducibility. For example, what are the exact spatial resampling methods for each variable; how hourly data were converted to daily values; how missing MODIS observations were handled; how station and grid-cell elevation differences were treated.
Minor comments.
The authors should report quantitative comparisons in the text. The authors state in several places that “closest agreement”, “larger deviations”, etc. The exact values (especially for MAE and RMSE) and significance tests are needed when comparing different products.
The manuscript should describe HiMIC-Daily more explicitly as a machine learning downscaling/bias-correction product, to avoid claiming that 1-km information is independently observed.
The full dataset is not shared on Zenodo. Now it only includes some sample data.
L21: remove “extreme”. It is repetitive with “extremes”
L39-40: claims of higher accuracy and more realistic temporal variability should be supported by quantitative comparisons.
L64: “has” to “have”
L111: “of assessing” to “for assessing”
L180: “daily estimates” to “estimates”
Table 2 and L232: version of MODIS doesn’t seem consistent.
Table 4: is cited several times in the text but absent.
Figure 5 caption: “Deeper red indicates larger values, reflecting better model performance” is true only for R2. Larger MAE and RMSE indicate worse performance.