the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Machine Learning-Based Fusion of Multi-Source Daily Precipitation Products for the Tibetan Plateau Rainy Season
Abstract. As the "Asian Water Tower," the Tibetan Plateau (TP) critically influences regional water security and global climate. Yet, due to its complex terrain and scarce observations, existing precipitation products poorly represent precipitation characteristics over the western.TP. Here, we present 3DMergePrec (3DM), a 0.25° daily precipitation dataset for the TP rainy seasons (1961–2021), generated by fusing 12 mainstream products using a deep learning framework combining Graph Attention Networks and 3D Convolutional Neural Networks. Validated against long-term observations (CMA stations) and independently verified with automatic stations in the western and central-western TP, 3DM demonstrates robust performance: Overall, it reduces mean squared error by 30–40% compared to satellite-only products (e.g., TRMM, GPM) and effectively mitigates the high-error belt in the southeastern TP. Crucially, in the data-sparse western TP, 3DM achieves RMSE reductions of 25–40% (e.g., mean squared error of 10.78 mm in the Qiangtang region), outperforming existing products. Its long-term precipitation trends closely align with observations, surpassing most counterparts. Limitations include underestimation of extreme precipitation frequency and overestimation of light precipitation days, with limited improvement in precipitation detection—likely due to the lack of dynamical constraints. Overall, 3DM offers stable, spatially continuous, and accurate precipitation estimates, particularly in the western TP, providing a valuable long-term dataset to support climate change studies across the region.
- Preprint
(4124 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 08 Aug 2026)
- RC1: 'Comment on essd-2026-278', Anonymous Referee #1, 22 Jul 2026 reply
-
RC2: 'Comment on essd-2026-278', Anonymous Referee #2, 31 Jul 2026
reply
Obtaining reliable precipitation data for the Tibetan Plateau is of great importance for various fields but is challenging due to its harsh environment and complex terrain. This manuscript developed a long-term precipitation dataset for the Tibetan Plateau by merging multi-source data using a deep learning method. While this research topic holds solid scientific value, the newly generated dataset has limited competitive advantages with respect to the horizontal resolution and accuracy. Additionally, the deep learning framework adopted lacks sufficient methodological documentation and rigorous validation. My detailed concerns are listed below
1) The developed precipitation dataset has a horizontal resolution of 0.25°, which is coarser than many mainstream datasets, such as IMERG, CMFD, MSWEP, HAR,HAR V2 and TPMFD. Furthermore, the results presented herein show that the developed dataset displays no notable strengths compared with existing datasets in multiple metrics: precipitation frequency (Figs 7 and 12), detection skills (Figs 4, 8 and 13), mean square error (Fig 5), and other evaluation indicators.
2) The deep learning approach is inadequately documented and validated. For instance, the structures of the GAT and 3D-CNN models are not specified. Critical training configurations and hyperparameters (e.g., learning rate, training epochs, optimizer type) remain undescribed. Moreover, the role of each module (GAT, 3D-CNN, and post-processing) is not well investigated.
3) The observations from the 134 CMA (China Meteorological Administration) stations have been incorporated in the development of many datasets used in this study (e.g. APHRODITE, CN05, CMFD, and CHM_PRE). These same station observations are also utilized to train the deep learning model in the current study. Therefore, the validations of these datasets and the merged product using these observations in sections 4.1 and 4.3 are not independent and carry no significance.
4) Lines 51-52, ‘These products, …, cloud-top brightness temperature’: Among the many satellite-based products, only the infrared-based products rely on cloud-top brightness temperature, while passive microwave and spaceborne radar are both capable of providing vertical profile information.
5) Lines 69-70, ‘precipitation is often strongly parameterized as a function of elevation’: Precipitation formation involves many cloud microphysical and convective parameterization schemes in atmospheric models, rather than being merely a function of terrain.
6) Lines 152-153: Details of the percentage of precipitation anomaly are needed.
7) Section 2.1.3: The rationale behind post-processing is insufficiently explained. Several questions remain unanswered: Why do deep learning predictions require further combination with the optimal data source? Why cannot raw model outputs be directly used as the final merged precipitation? What is the objective of selecting an optimal data source, and how is this optimal source determined? Moreover, the method for determining the weights should be described.
8) A comparison of precipitation spatial distributions derived from different datasets is recommended.
9) Lines 387-388, ‘The evaluation based, …, those observed over the central TP’: This sentence is confusing.
Citation: https://doi.org/10.5194/essd-2026-278-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 40 | 23 | 5 | 68 | 7 | 8 |
- HTML: 40
- PDF: 23
- XML: 5
- Total: 68
- BibTeX: 7
- EndNote: 8
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This study addresses an important problem and the proposed GAT–3D-CNN framework, together with the western-station evaluation, has potential value. However, several fundamental issues remain unresolved. The changing availability of input products raises serious concerns about the temporal consistency of the 1961–2021 record; observational information may overlap among model training, graph construction, post-processing, and evaluation; essential methodological details are insufficient for reproducibility; and the evidence for long-term trends and extreme precipitation performance remains limited. In addition, several claims of robustness, reliability, and superiority are stronger than the reported results support, while the quality control, uncertainty, metadata, and long-term usability of the released product are not adequately documented. The independence of the validation design, the temporal stability of the dataset, the reproducibility of the framework, and the validity of the main conclusions must be convincingly established before the manuscript can be considered suitable for ESSD.
Major Comments:
1. 3DM covers 1961–2021, whereas the 12 input products have substantially different temporal coverages. How is methodological consistency maintained when the number and type of available inputs change over time, and could these changes introduce artificial discontinuities into the long-term record?
2. The 134 long-term stations are used both to construct the graph/provide training labels and to perform the main evaluation for 2001–2015. To what extent do these results represent independent generalization rather than performance at stations already involved in model development?
3. The post-processing procedure selects a “daily optimal data source” based on in-situ observations and then applies dynamic weighting. Are observations from the target day used to determine the selected source or weights? If so, this would constitute direct information leakage into the final product.
4. The manuscript states that the predictor set contains nine variables, including basin descriptors, rainfall characteristics, centroid timing, skewness, and one- and five-day antecedent precipitation. However, the MLP architecture is described as having six input nodes, and Table 1 lists six variables.
5. Essential details are missing, including graph adjacency rules, virtual-node connections, Transformer/GAT/3D-CNN configurations, training settings, treatment of unavailable products, and the dynamic-weighting equations. The current description is insufficient to independently reproduce 3DM.
6. 3DM is not consistently the best-performing product in the independent western evaluation: it ranks fourth for the automatic stations, seventh for the north–south transect, has a CC of only 0.41, and a median CSI of about 0.45.
7. The trend reliability of a 61-year dataset is mainly assessed over the short 2001–2015 period, using stations involved in model construction. Can a 15-year evaluation establish the stability of trends over the full 1961–2021 record?
8. 3DM overestimates light-precipitation frequency and systematically underestimates precipitation above 20 mm d⁻¹ and some extreme-event frequencies, yet SWMSE is claimed to improve tail sensitivity.
9. The study claims to release a spatially continuous gridded precipitation dataset, yet the results focus mainly on station-based metrics and error distributions, without directly showing the climatology, representative daily fields, or major spatial structures of 3DM. How can the physical plausibility of the generated fields in ungauged regions be assessed, rather than merely their accuracy at station locations?
10. The evaluation mainly compares 3DM with its input datasets and conventional products, while recent merged datasets such as MSWEP, GMCP, and TPHiPr are not considered.
11. POD, FAR, and CSI are all calculated using a 0.1 mm d⁻¹ threshold, although 3DM noticeably overestimates very light precipitation and this threshold is sensitive to near-zero retrieval noise. Do the product rankings and conclusions regarding improved event detection remain valid under thresholds of 0.5 or 1.0 mm d⁻¹
Minor comments:
1) The CC formula in Table 3 uses absolute anomalies and assigns units of mm to the correlation coefficient.
2) The abstract reports 10.78 as MSE while discussing RMSE reduction, and MSE is sometimes expressed in mm rather than mm².
3) The input products may use different daily boundaries and native temporal resolutions. Are the time zone, daily accumulation period, and treatment of missing hours consistent across products?
4) A 0.25° grid cell differs substantially in scale and elevation from a mountain gauge. How were stations matched to grid cells, and was point-to-grid representativeness error considered
5) Table 2 provides only station identifiers and record periods, without coordinates, elevation, instrument type, missing-data rate, or quality-control information. The representativeness and reliability of the independent validation samples therefore remain unclear.
6) Figure 2 indicates that 61 years were randomly divided into 56 training years and 5 validation years, but the selected years, random seed, and variability across different splits are not reported.
7) Extreme precipitation is defined using the observed 95th percentile, but it is unclear whether the threshold is station-specific, regionally uniform, or calculated from pooled samples. This definition directly affects the estimated number and trends of extreme precipitation days.
8) Figures 16 and 17 evaluate extreme-precipitation changes using the ratio of the product trend to the observed trend. How are cases handled when the observed trend is close to zero or statistically insignificant, for which this ratio may become extremely large or unstable.