the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A physically guided deep learning reconstruction of terrestrial water storage anomalies at 0.1° across China
Abstract. Terrestrial water storage (TWS), comprising all surface and subsurface water components, is a key indicator of water availability. The Gravity Recovery and Climate Experiment (GRACE) satellite mission provides large-scale estimates of TWS anomalies (TWSA), but its coarse spatial resolution (3°, approximately 300 km) limits the analysis of hydrologic processes at sub-regional scales. Using a physically-guided deep learning framework, we downscale TWSA from the original 3° GRACE mascons to 0.1° (approximately 10 km) across China, generating a standard version (2002–2019) with comprehensive observations used for model constraints and independent evaluation and an extended version (2020–2023) to support more recent hydrologic analyses. The downscaled TWSA preserves large-scale GRACE signals at the 3° grid scale (median correlation coefficient (CC): 0.95; root-mean-square error (RMSE): 1.38 cm) and basin scale (median CC: 0.94; RMSE: 1.72 cm), with a low median uncertainty (0.88 cm) across China. Its reliability is supported by high consistency with physically informed TWSA spatial patterns at the 0.1° resolution (median CC: 0.91) and internally consistent water balance closure beyond the native GRACE resolution (median CC: 0.80; RMSE: 1.44 cm). Evaluation against independent observations demonstrates that the downscaled TWSA agrees well with groundwater variations in intensively irrigated regions (CC: 0.65 for irrigation intensity > 50 %) and annual glacier elevation change in cryospheric areas (CC: 0.97). The datasets improve fine-scale characterization of TWS variability and associated hydrologic processes in China, and can be used as a reference for evaluating performance of high-resolution hydrologic models. The two versions of the dataset are available at https://doi.org/10.5281/zenodo.19502906.
- Preprint
(2650 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-282', Anonymous Referee #1, 04 Sep 2026
-
RC2: 'Comment on essd-2026-282', Anonymous Referee #2, 14 Sep 2026
The authors developed a physically-guided deep learning framework based on a U-Net architecture to downscale GRACE/-FO Terrestrial Water Storage Anomalies (TWSA) from a native resolution of 3° to a high resolution of 0.1° across China. The incorporation of a dual-loss function—combining large-scale mass conservation from GRACE mascons and pixel-scale spatial structural constraints from hydrological simulations (PCR-GLOBWB 2)—is methodologically sound and represents an encouraging attempt to resolve cryospheric and anthropogenic signals. Overall, the manuscript is well-written, clearly organized, and the produced high-resolution dataset will be of great value to the regional hydrological community. To further strengthen the manuscript and improve the clarity, transparency, and robustness of the study, a few constructive suggestions are provided below for the authors’ consideration:
- Line 330-331. The authors partitioned the spatial domain into 70% training and 30% validation sets within each monthly time step rather than across the temporal axis during 2002–2019. While this spatial-within-month splitting strategy is reasonable for spatial reconstruction and inpainting tasks, it is suggested to enrich the discussion regarding potential temporal information sharing when evaluating long-term basin-scale time series (e.g., Fig. 6) and spatial metrics (e.g., Fig. 4). In addition, it would be beneficial to add an independent test set.
- Lines 375–381. It would be helpful to provide the explicit mathematical formula used to calculate the pixel-level uncertainty from the three independent training runs. Furthermore, while the substantial computational burden of training deep networks for 800 epochs is well appreciated, an ensemble size of N=3 is statistically on the smaller side for estimating standard deviations. If computational resources permit, increasing the ensemble size (e.g., to 5–10 runs) or providing a brief discussion on the variance stability with N=3 would make the uncertainty quantification even more convincing.
- Line 367. Could the authors provide more details on how the benchmark uncertainty value of “1.96 cm” for the JPL-M product was obtained? Specifically, it would be valuable to clarify whether this value represents the formal satellite measurement error provided in the JPL-M product or an empirical metric, so that readers can better understand the baseline comparison between this value and the AE loss of the downscaled TWSA.
- The definition and description of TWSAPI could be made more consistent throughout the manuscript. In some parts, it is defined as the sum of PCR-GLOBWB 2 simulations and satellite-derived glacier mass changes, while in others, it is referred to mainly as simulations from PCR-GLOBWB 2. It is suggested to introduce a clear mathematical equation explicitly defining TWSAPI and ensure that the terminology and notations remain unified across all sections.
- Line 185. It would be helpful to add a brief physical explanation for upscaling the JPL-M mascon product from its nominal 0.5° resolution back to its native 3° resolution (e.g., ensuring data independence by excluding land-grid scaling factors and respecting the effective spatial resolution of GRACE), explaining why using the 0.5° grid directly was avoided.
- Line 324: The bounding coordinates (15∘–57∘N, 69∘–138∘E) do not cover the entire territorial waters of China (such as the southernmost islands in the South China Sea). It is suggested to rephrase this as “…covering mainland China and surrounding regions…” to ensure geographic and cartographic precision.
- Lines 108, 178, and 550. The dates defining the gap period between GRACE and GRACE-FO conflict.
- Figure 1: The locations of the in-situ groundwater wells in the North China Plain inset are difficult to distinguish. Please enlarge the marker size and/or select a higher-contrast color to improve legibility.
- Figure 2: In the figure schematic and caption, define the abbreviations “AE” (Absolute Error) and “MAE” (Mean Absolute Error). Figure captions should be self-contained so readers can interpret the workflow without referring back and forth to the main text.
- Line 686. The sub-grid variability is displayed in the fourth row (d1–d3), not the “fourth column”. Please correct this typographical error.
Citation: https://doi.org/10.5194/essd-2026-282-RC2
Data sets
A 0.1° terrestrial water storage anomaly dataset over China Xueying Li and Yan Sun https://doi.org/10.5281/zenodo.19502906
Model code and software
Code for: A physically guided deep learning reconstruction of terrestrial water storage anomalies at 0.1° across China Xueying Li and Yan Sun https://doi.org/10.5281/zenodo.19502631
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 443 | 197 | 84 | 724 | 95 | 78 |
- HTML: 443
- PDF: 197
- XML: 84
- Total: 724
- BibTeX: 95
- EndNote: 78
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review of A physically guided deep learning reconstruction of terrestrial water storage anomalies at 0.1° across China by Li et al. (2026).
In this study, Li et al. (2026) applied a deep learning model to downscale the JPL mascon to 0.1 degrees, incorporating spatial details from a hydrological model (PCR-GLOBWB). Evaluations across different scales and comparisons with two other publicly available products show that the proposed products perform better. However, after a detailed read, several major points need to be carefully reconsidered and revised before this manuscript can be published in ESSD. One general comment on the dataset is that the “main product” described by the authors is only available through 2019 (about 7 years ago). For a data description paper being considered in 2026, the data record is a bit out of date, and I highly recommend that the authors extend it so that the community can benefit more from it. There are several other issues regarding method definition, description, and product evaluations. Please refer to my detailed comments below.
Major comments
Line 190: The referenced glacier products are only available annually. How did you use them for your monthly downscaling? Please mention it clearly. On line 205, you mentioned that they provide the monthly products, but the link does not refer to them. Please correct.
Line 228 and the title: As you are writing a paper with a strong focus on DL, I highly recommend avoiding the term "physics-informed," as it causes considerable confusion with PINN.
Line 243: Yes, accumulated meteorological fluxes are typically comparable to TWS changes. However, the errors are also accumulated. This is the reason that comparisons between GRACE data and P, ET, R are usually in the differential domain. Can you comment on this? And please clarify what n is in your case?
Major issues with Section 3 Methodology
On line 302, you clearly mention that your CNN takes inputs with a spatial size of 3 degrees, but Fig. 2 clearly does not show a 3-degree size; rather, it shows around 15 degrees (5 JPL mascons). Please clarify and keep it consistent.
Section 3.1 Training strategy
Major issue of this section: The whole methodology, from model architecture, training design (especially the loss function), to the uncertainty quantification approach, is too similar to the one proposed by Gou and Soja 2024, which the authors also referred to in the introduction section. It is fine to further improve the methods and/or apply them to another set of data/target regions, but you need to:
Line 374: You actually spent two paragraphs explaining how to determine the lambdas, and the conclusion is that they are not necessary at all. I do acknowledge your efforts in this hyperparameter tuning. Still, you should comment on why lambda=1 provides good results and, ultimately, may consider reducing it in the main text to provide more concise information.
Line 375: So you basically just get the ensemble of three trainings. So, it's just "half" of the deep ensembles (see Lakshminarayanan et al., 2017) and can therefore only capture epistemic uncertainties (another issue: 3 repeated runs might be too limited for approximating epistemic uncertainties). So how about aleatoric uncertainties? Actually, Gou and Soja 2024 also devote quite some text to discussing this issue and mention it as one of their main limitations. To my understanding, they tried to consider the aleatoric uncertainty to a certain extend. So, regarding the uncertainty quantification part, your method does not advance on the previous study, right?
Fig. 3: The validation AE loss is around double the training AE loss. This is a risky sign, indicating that the model cannot maintain agreement with GRACE data. I guess the issue is associated with the small patch size and the choice of 3-deg JPLM. The sharp mascon boundary may have a remarkable negative impact in this case; see your Fig. 7.
Fig. 7: Both of the maps show clear “mascon-like” shapes. These shapes indicate a crucial point: the downscaled products may not overcome the limitation imposed by a 3-degree mascon boundary. It is a rather fatal problem for a product that is argued to have 0.1-degree resolution. Please comment on it.
Problem with evaluations and discussions.
Line 695: Xiong et al. (2026) used SHC-based products and did not provide anything based on the mascon product. Therefore, all the following comparisons that use the JPL mascon as the ground truth are unfair and cannot yield a convincing conclusion.
Line 720: Did you also examine if the correlations are statistically significant? And although groundwater dominates NCP. You should also carefully consider other components, especially root-zone soil moisture.
Minor comments
Line 74: The study by Vishwakarma et al. (2021) does not belong to ML methods.
Fig. 1: This figure does not show the necessary information clearly. For example, the positions of wells are not readable.
Line 172: Is it RL06.3 or? Please specify.
Line 324: Then how did you handle the ocean pixels? Did you remove them or mask them out?
Line 330: But the problem is that the same location may also be used for training in another month. As you highlighted, you want to learn the spatial details, so it's not really a fair "training-validation" strategy, right?
Line 334: Please first cite Adam, then report more details on your LR scheduling so that the audience can fully reproduce it.
Line 398: Well, you can argue that using ERA5L here is for consistency. But then, you can't judge whether your method is better than their other/original GRACE methods anymore,, since you explicitly incorporate this information. They no longer provide independent evaluation.
Fig 3: It's hard to gain valuable information from this figure as the point density is missing. Please consider including this density information so that the distributions of the points can be better analyzed.
Line 451: I doubt that the issue is due to land-ocean leakages. You used the CRI version, which explicitly accounts for the ocean-land leakage issue, right? So, it should be fine for your case, given the current accuracy.
Fig. 5: Didn't you mention that you chose Hydro Basin L5 (Line 403)? They are certainly not in this figure.
About Section 4.3: I’m a bit confused by the design of this manuscript: You explicitly exclude the highly irrigated glacierized regions in your previous evaluations. But now you're even focusing on them and providing quite a lot of discussion. They actually contradict each other. Could you please explain your intention a bit more? Thanks!
Fig. 9: Again, the readability of these figures is not good. And please choose different color maps for the absolute correlation and the differences, respectively.
Line 709: But Gou and Soja 2024 clearly mentioned that their uncertainty information is underestimated. So being close to theirs is not a good sign for your product. Instead, it would be a meaningful contribution if you could provide more realistic uncertainty information.