A physically guided deep learning reconstruction of terrestrial water storage anomalies at 0.1° across China
Abstract. Terrestrial water storage (TWS), comprising all surface and subsurface water components, is a key indicator of water availability. The Gravity Recovery and Climate Experiment (GRACE) satellite mission provides large-scale estimates of TWS anomalies (TWSA), but its coarse spatial resolution (3°, approximately 300 km) limits the analysis of hydrologic processes at sub-regional scales. Using a physically-guided deep learning framework, we downscale TWSA from the original 3° GRACE mascons to 0.1° (approximately 10 km) across China, generating a standard version (2002–2019) with comprehensive observations used for model constraints and independent evaluation and an extended version (2020–2023) to support more recent hydrologic analyses. The downscaled TWSA preserves large-scale GRACE signals at the 3° grid scale (median correlation coefficient (CC): 0.95; root-mean-square error (RMSE): 1.38 cm) and basin scale (median CC: 0.94; RMSE: 1.72 cm), with a low median uncertainty (0.88 cm) across China. Its reliability is supported by high consistency with physically informed TWSA spatial patterns at the 0.1° resolution (median CC: 0.91) and internally consistent water balance closure beyond the native GRACE resolution (median CC: 0.80; RMSE: 1.44 cm). Evaluation against independent observations demonstrates that the downscaled TWSA agrees well with groundwater variations in intensively irrigated regions (CC: 0.65 for irrigation intensity > 50 %) and annual glacier elevation change in cryospheric areas (CC: 0.97). The datasets improve fine-scale characterization of TWS variability and associated hydrologic processes in China, and can be used as a reference for evaluating performance of high-resolution hydrologic models. The two versions of the dataset are available at https://doi.org/10.5281/zenodo.19502906.
Review of A physically guided deep learning reconstruction of terrestrial water storage anomalies at 0.1° across China by Li et al. (2026).
In this study, Li et al. (2026) applied a deep learning model to downscale the JPL mascon to 0.1 degrees, incorporating spatial details from a hydrological model (PCR-GLOBWB). Evaluations across different scales and comparisons with two other publicly available products show that the proposed products perform better. However, after a detailed read, several major points need to be carefully reconsidered and revised before this manuscript can be published in ESSD. One general comment on the dataset is that the “main product” described by the authors is only available through 2019 (about 7 years ago). For a data description paper being considered in 2026, the data record is a bit out of date, and I highly recommend that the authors extend it so that the community can benefit more from it. There are several other issues regarding method definition, description, and product evaluations. Please refer to my detailed comments below.
Major comments
Line 190: The referenced glacier products are only available annually. How did you use them for your monthly downscaling? Please mention it clearly. On line 205, you mentioned that they provide the monthly products, but the link does not refer to them. Please correct.
Line 228 and the title: As you are writing a paper with a strong focus on DL, I highly recommend avoiding the term "physics-informed," as it causes considerable confusion with PINN.
Line 243: Yes, accumulated meteorological fluxes are typically comparable to TWS changes. However, the errors are also accumulated. This is the reason that comparisons between GRACE data and P, ET, R are usually in the differential domain. Can you comment on this? And please clarify what n is in your case?
Major issues with Section 3 Methodology
On line 302, you clearly mention that your CNN takes inputs with a spatial size of 3 degrees, but Fig. 2 clearly does not show a 3-degree size; rather, it shows around 15 degrees (5 JPL mascons). Please clarify and keep it consistent.
Section 3.1 Training strategy
Major issue of this section: The whole methodology, from model architecture, training design (especially the loss function), to the uncertainty quantification approach, is too similar to the one proposed by Gou and Soja 2024, which the authors also referred to in the introduction section. It is fine to further improve the methods and/or apply them to another set of data/target regions, but you need to:
Line 374: You actually spent two paragraphs explaining how to determine the lambdas, and the conclusion is that they are not necessary at all. I do acknowledge your efforts in this hyperparameter tuning. Still, you should comment on why lambda=1 provides good results and, ultimately, may consider reducing it in the main text to provide more concise information.
Line 375: So you basically just get the ensemble of three trainings. So, it's just "half" of the deep ensembles (see Lakshminarayanan et al., 2017) and can therefore only capture epistemic uncertainties (another issue: 3 repeated runs might be too limited for approximating epistemic uncertainties). So how about aleatoric uncertainties? Actually, Gou and Soja 2024 also devote quite some text to discussing this issue and mention it as one of their main limitations. To my understanding, they tried to consider the aleatoric uncertainty to a certain extend. So, regarding the uncertainty quantification part, your method does not advance on the previous study, right?
Fig. 3: The validation AE loss is around double the training AE loss. This is a risky sign, indicating that the model cannot maintain agreement with GRACE data. I guess the issue is associated with the small patch size and the choice of 3-deg JPLM. The sharp mascon boundary may have a remarkable negative impact in this case; see your Fig. 7.
Fig. 7: Both of the maps show clear “mascon-like” shapes. These shapes indicate a crucial point: the downscaled products may not overcome the limitation imposed by a 3-degree mascon boundary. It is a rather fatal problem for a product that is argued to have 0.1-degree resolution. Please comment on it.
Problem with evaluations and discussions.
Line 695: Xiong et al. (2026) used SHC-based products and did not provide anything based on the mascon product. Therefore, all the following comparisons that use the JPL mascon as the ground truth are unfair and cannot yield a convincing conclusion.
Line 720: Did you also examine if the correlations are statistically significant? And although groundwater dominates NCP. You should also carefully consider other components, especially root-zone soil moisture.
Minor comments
Line 74: The study by Vishwakarma et al. (2021) does not belong to ML methods.
Fig. 1: This figure does not show the necessary information clearly. For example, the positions of wells are not readable.
Line 172: Is it RL06.3 or? Please specify.
Line 324: Then how did you handle the ocean pixels? Did you remove them or mask them out?
Line 330: But the problem is that the same location may also be used for training in another month. As you highlighted, you want to learn the spatial details, so it's not really a fair "training-validation" strategy, right?
Line 334: Please first cite Adam, then report more details on your LR scheduling so that the audience can fully reproduce it.
Line 398: Well, you can argue that using ERA5L here is for consistency. But then, you can't judge whether your method is better than their other/original GRACE methods anymore,, since you explicitly incorporate this information. They no longer provide independent evaluation.
Fig 3: It's hard to gain valuable information from this figure as the point density is missing. Please consider including this density information so that the distributions of the points can be better analyzed.
Line 451: I doubt that the issue is due to land-ocean leakages. You used the CRI version, which explicitly accounts for the ocean-land leakage issue, right? So, it should be fine for your case, given the current accuracy.
Fig. 5: Didn't you mention that you chose Hydro Basin L5 (Line 403)? They are certainly not in this figure.
About Section 4.3: I’m a bit confused by the design of this manuscript: You explicitly exclude the highly irrigated glacierized regions in your previous evaluations. But now you're even focusing on them and providing quite a lot of discussion. They actually contradict each other. Could you please explain your intention a bit more? Thanks!
Fig. 9: Again, the readability of these figures is not good. And please choose different color maps for the absolute correlation and the differences, respectively.
Line 709: But Gou and Soja 2024 clearly mentioned that their uncertainty information is underestimated. So being close to theirs is not a good sign for your product. Instead, it would be a meaningful contribution if you could provide more realistic uncertainty information.