the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
MLAWind: A Monthly Sea Surface Wind Dataset Derived from an Interpretable Machine Learning Approach Integrating In-Situ Observations and Satellite Data
Abstract. A gridded sea surface wind dataset with long temporal coverage is crucial for understanding atmospheric circulation changes and air-sea interactions at different time scales. This study employs an interpretable machine learning model based on random forest algorithm to generate a 1°×1° monthly sea surface wind dataset (MLAWind) from 1950 to 2023, covering the near-global ocean within 60° S–60° N. The data reconstruction model integrates the Cross-Calibrated Multi-Platform (CCMP) satellite data and the spatially sparse long-term International Comprehensive Ocean-Atmosphere Data Set (ICOADS), exhibiting robust interpretability and generalization capability. Evaluations demonstrate that the MLAWind dataset exhibits better agreement with remote sensing observations than existing reanalysis datasets during the training period (1993–2022), while maintaining robust performance during the independent testing period in 2023. Moreover, the performance of MLAWind since 1950 is assessed across multiple time scales. Its characteristics in climatology, annual cycle, and inter-annual variability are comparable to those of existing reanalysis datasets, even during the non-satellite period prior to 1993. Uncertainties remain in the long-term trends of different datasets. The trend derived from MLAWind is corroborated by independent coral records during 1950–1982, which demonstrates its strong capability in reconstructing historical sea surface wind variations. The results indicate that MLAWind serves as a reliable data resource for global climate change research. The reconstructed MLAWind dataset is publicly accessible at https://doi.org/10.5281/zenodo.17354864 (Guo et al., 2025b).
- Preprint
(6795 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2025-725', Anonymous Referee #1, 17 Feb 2026
-
RC2: 'Comment on essd-2025-725', Anonymous Referee #2, 03 Aug 2026
The paper applies an interpretable machine learning model based on the random forest algorithm to generate a 1° × 1° monthly sea surface wind dataset (MLAWind) spanning 1950-2023 and covering the near-global ocean between 60°S and 60°N. I think the authors present a novel and valuable approach for estimating sea surface winds directly from observed ICOADS data, rather than relying on winds derived from atmospheric data assimilation products such as ERA5.
Although the paper presents climatological 10-m winds for the period 1981-2010 (Fig. 7) and devotes substantial discussion to interannual variability and long-term trends (Sections 4.2 and 4.3), these analyses do not fully address the reliability of the dataset in earlier decades and data-sparse regions. The climatological period of 1981-2010 largely overlaps with the model training period (1993-2022), during which observational coverage is relatively dense. In addition, sea surface winds in tropical and low-latitude regions are generally easier to reconstruct because of the stronger relationship between SST and surface winds compared with mid- and high-latitude regions.
My primary concern is the quality and reliability of the dataset during data-sparse periods and in poorly observed regions, particularly prior to 1970 and over the Southern Ocean. To increase confidence among readers and potential users of the dataset, I recommend that the authors provide additional evidence demonstrating the robustness and accuracy of MLAWind under these challenging conditions.
Specifically, I suggest a comparative evaluation against the ERA5 reanalysis for three distinct periods: 1950-1969, 1970-1992, and 1993-2023. ERA5 is a widely used and well-validated reanalysis product and would provide a useful benchmark for assessing the temporal consistency of MLAWind. To facilitate a direct comparison, MLAWind could first be interpolated onto the ERA5 grid.
The following diagnostics would be particularly helpful:
- Global maps of the climatological mean of the 10-m zonal (10u) and meridional (10v) wind components, together with maps of their differences relative to ERA5.
- Global maps of the standard deviation of 10u and 10v anomalies (removed the seasonal cycle), together with maps of their differences relative to ERA5.
- Global maps of the correlation coefficient and root-mean-square deviation (RMSD) of 10u and 10v relative to ERA5.
In addition, for three broad regions (20°S-60°S, 20°S-20°N, and 20°N-60°N), I recommend presenting time series from January 1950 to December 2023 showing:
4. The differences for regional mean 10u and 10v between MLAWind and ERA5.
These analyses would provide a more comprehensive assessment of the spatial and temporal consistency of MLAWind and would help demonstrate its reliability, particularly during periods and in regions with limited observational coverage.
Minor comments:
Lines 114–115: Please provide appropriate references for SHAP.
Figure 2: The details in the right section of the "Pre-trained model" box lack clarity.
Line 221: Were the Ta and SLA variables in ICOADS excluded from the training data?
Lines 273–274: The formulas (4) and (5) for skewness and are incorrect; the index should start at i=1 rather than i-1.
Citation: https://doi.org/10.5194/essd-2025-725-RC2
Data sets
Machine learning-assisted sea surface wind dataset (MLAWind) W. Guo et al. https://doi.org/10.5281/zenodo.17354864
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 330 | 164 | 29 | 523 | 45 | 60 |
- HTML: 330
- PDF: 164
- XML: 29
- Total: 523
- BibTeX: 45
- EndNote: 60
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The authors used a random forest and SHAP algorithms reconstructed a monthly sea surface wind in 1°×1° grids between 60°S-60°N from 1950-2023 based on ICOADS observations. The results were validated by ERA5, JRA-55, NOAA-20C, NCEP1, NCEP2, and Mn/Ca records in climatology, annual cycle, interannual variability, and long-term trend. The paper is well written and the dataset should be a good reference to user communities, and can be published in ESSD after a major revision. My major concerns are (a) lack of clarification on the use of CCMP as label variable, (b) lack of inter-comparisons against independent observations, and (c) lack of validation on wind direction.
L28-30, it is more important about the evaluation during the non-training period, as stated later in L32-35 the results are comparable. Therefore it should be clearly stated about the improvement of the current study.
L90, “It” is not clear, is it “WASWind”?
L121, “an interpretable machine learning algorithm” need a reference.
L136, WD, it is not clear why U, V, and WS were used in validation data while the inputs from ICOADS used WS and WD.
L156, “Smith (1980)” is messing in references.
L160-171, Have these satellite-based wind gone through a “bias-adjustment” process as for the in situ observations? How do we know these satellite observations are consistent with in situ observations?
L208-211, I am not clear how the CCMP (1993-2023) is used as a label variable while ICOADS (1950-2023) used as feature input. What is the label variable during 1950-1992?
Figure 3a,c, I assume the color shading represent wind speed, my question is: what are the wind directions. Is there any metrics to measure the successes of the reconstruction? Clear differences can be found in wind speed in 2000 in the northern North Pacific and northern North Atlantic.
L309-316, I suggest adding a table to compare these RMSE, R-squared, bias etc so that readers can easily understand the results.
Figure 4, is it possible to check the wind direction v/u as a verification metrics in Eqs (1)-(5)?
Figures 7-10 are good for the intercomparisons. However, I am wondering the direct comparisons against observations particularly independent observations to see whether MLWind’s performance is comparable with those reference datasets?
Table 2, I have difficulty to understand how the Mn/Ca trend can be quantitatively compared with wind because they are different variables in different units. Further, the difference among MLWind, ERA5, NCEP1, and NOAA-20C are large, why?