A dataset of soil water measurements at multiple scales and depths on China’s Loess Plateau
Abstract. Soil water content (SWC) is a key state variable in land-surface hydrological processes and the primary water source for vegetation, particularly in arid and semi-arid regions. However, long-term, multi-scale in situ measurements of deep SWC remain insufficient on the Loess Plateau of China (LPC), the world’s largest and deepest loess deposit area. To address this gap, the Loess-Obs network, a long-term multi-scale SWC observation system, was established on the LPC about 20 years ago, covering five spatial scales including plot, hillslope, local-landscape, plateau-wide transect, and regional-survey scale. Here, we present a long-term SWC profile dataset derived from the Loess-Obs network, including continuous measurements (2004–2019) of four plots with distinct land uses, two hillslope transects of 300 m and 270 m (2004–2016), a 1340 m local-landscape transect spanning multiple hillslopes (2012–2013), an 860 km plateau-wide transect (2013–2016), and a regional survey conducted in 2015. Volumetric SWC in the 0–5 m layer was measured at multiple depths using calibrated neutron probes (sampling frequency: weekly to monthly). We detail the network design, field protocols, experimental setup, sampling strategies, calibration equations, maintenance procedures, and quality control measures, alongside associated environmental datasets (land use, soil properties, topography, meteorology) to support extended applications. Key findings include: (1) exotic shrubs (Caragana korshinskii) and grasses (Medicago sativa) accelerated deep SWC depletion compared to cropland and naturally restored fallow land; (2) temporal variability of SWC decreased, while its spatial variability increased with soil depth; (3) time-stable sampling sites could be effectively used to characterize mean SWC of different soil layers, with accuracy improving with depth; (4) regional soil water storage decreased from southeast to northwest; (5) the discrepancy between SWC products and Loess-Obs observations generally increased with depth, reflecting the scarcity of deep in situ observations in model development. This dataset is valuable for validating SWC products, developing upscaling methods, and understanding deep soil hydrological responses to land use and climate change. It is publicly accessible at https://zenodo.org/records/21859310.
Overall assessment
The dataset is potentially valuable because it combines deep soil profile measurements with observations across multiple spatial extents in a data scarce region. However, long term records are available primarily for the experimental plots and hillslope HB, both covering relatively small spatial areas, while the remaining components consist of shorter term campaigns or a single regional survey. The measurements are also periodic rather than continuous. Moreover, the procedures for data processing, validation, and uncertainty assessment are not described with sufficient clarity or quantitative evidence to support the manuscript’s strong claims regarding data accuracy, reliability, and comparability among scales. In particular, the calibration and independent validation of the neutron probe measurements, long term instrument stability, reproducible quality control procedures, treatment of missing and anomalous records, and uncertainty estimates require substantial clarification. The hierarchical structure of the component datasets and their differences in monitoring duration, frequency, depth, and sampling design should also be described more accurately, and several scientific interpretations should be moderated. I therefore consider the dataset potentially suitable for ESSD, but substantial revision of both the data documentation and the manuscript is required. My detailed comments and suggestions are provided below.
Major comments
1. Abstract: For an ESSD data paper, the Abstract should provide clearer quantitative evidence of data quality and validation. Please condense the five application oriented findings and report key information on calibration accuracy or quality control outcomes, enabling readers to assess the reliability of the dataset. The description of Loess-Obs as a long term multiscale observation system providing continuous measurements should be revised. Long term monitoring was primarily conducted at the plot scale and on hillslope HB, whereas the other components represent shorter term repeated campaigns or a single regional survey. In addition, the monitoring periods, sampling frequencies, profile depths, and measurement schemes were not consistent among scales or over time. Please describe the duration, continuity, and design of each component accurately and avoid applying terms such as “long term” or “continuous” to the entire dataset without qualification.
2. Introduction: The data gap and novelty of Loess-Obs are not sufficiently demonstrated. The cited studies cover different research topics but do not provide a systematic comparison of existing datasets in terms of profile depth, duration, spatial coverage, sampling frequency, and land use representation. Please compare Loess-Obs with representative existing soil water datasets and provide evidence for the claim that it is the first operational SWC observatory on the Loess Plateau. Line 75: Zreda et al. (2008) and Desilets et al. (2010) concern cosmic ray neutron sensing, whereas Loess-Obs uses conventional access tube neutron probes for soil profile measurements. Please verify whether these references are appropriate for the measurement technique used in this study.
3. Section 2: The study area description should better reflect the multiscale structure of the dataset. Figure 1 shows that the regional survey provides broad coverage of the typical loess region, while the plateau wide transect, local landscape, hillslope, and plot observations represent progressively finer spatial extents. Please clarify this spatial hierarchy and provide scale specific information on geographical extent, climate, topography, soil characteristics, land use composition, and monitoring design, preferably in a summary table. Appropriate references should also be provided for the specific climatic, environmental, and vegetation information reported in this section. This would help users understand the relationship among the component datasets and evaluate their representativeness and potential applications.
4. Section 3.1.1: It appears that a single calibration equation derived from eight sites was applied to all 243 regional survey sites across different land use types and to the full 0–5 m profile. However, the criteria used to select the eight representative sites are unclear. Please explain how “mean neutron counts” were calculated and used for site selection, and report the geographical distribution, land use, soil properties, depth coverage, and SWC range of the calibration samples. The authors should also demonstrate that a pooled calibration is appropriate across land use types, sites, and depths, considering possible differences in bulk density, soil texture, organic matter, and other factors affecting neutron probe response. In particular, the applicability of a calibration based on samples from the upper 1 m to measurements extending to 5 m requires explicit validation.
5. Section 3.1.2: Each plateau wide transect campaign required three to four days, but only the first day was recorded as the observation date. This treatment may introduce substantial temporal uncertainty, particularly when rainfall occurs during a campaign. Please provide the original measurement date and, if available, time and sampling sequence for each site. If site specific timing information is unavailable, this limitation and its implications for temporal stability analysis and comparison with external products should be explicitly discussed.
6. Section 3.1.3: The rationale for using different calibration strategies across spatial scales is unclear. A single pooled calibration equation appears to have been applied across the much larger and more heterogeneous regional survey and plateau wide transect, whereas a separate equation was developed for the relatively small local landscape transect. Please explain the criteria used to determine when calibration data could be pooled and when a scale specific equation was required. The calibration datasets, soil conditions, SWC ranges, sample sizes, and fitting uncertainties should be compared among scales, and differences in the slopes and intercepts of the calibration relationships should be statistically assessed. Please also explain how the 12 calibration sites for the local landscape transect were selected and whether they adequately represent its main vegetation types and landscape positions.
7. Section 3.1.4: The vertical sampling scheme for hillslope HB changed in April 2013, resulting in inconsistent depth resolution over time. Please explain the reason for this change and describe how measurements collected under the two schemes were harmonized for temporal analysis and data publication. The affected depths and periods should be clearly identified in the dataset and metadata.
8. Section 3.1.5: Each land use treatment appears to be represented by only one plot, with the 11 access tubes serving as spatial subsamples rather than independent treatment replicates. However, the published dataset contains only plot means, making it impossible to evaluate within plot variability or determine how the subsampling structure was considered in the statistical analyses. Please clarify the experimental unit and level of replication, quantify the variability among tubes across depths and sampling dates, and explain why only averaged values are provided. Individual tube measurements should be published where available, with plot means retained as a derived product; otherwise, the means should at least be accompanied by the standard deviation, range, and number of valid tubes. Without independent plot replication, differences among the four plots should not be attributed conclusively to land use, and the individual tubes should not be treated as independent replicates in statistical tests.
9. Section 3.3: The generation and uncertainty of the ancillary environmental data are not described in sufficient detail for reproducible reuse. Please specify the source, reference period, definitions, interpolation method, spatial resolution, and validation accuracy of the climate variables; provide the version, temporal coverage, compositing, and quality control procedures for MOD13Q1 NDVI; and quantify the positional accuracy of coordinates derived from GPS and satellite imagery, particularly for sampling points only 10 m apart. The authors should also clarify that the vegetation measurements represent single survey dates rather than conditions throughout the full SWC monitoring period.
10. Section 4: Although this section is titled “Uncertainty analysis,” it mainly provides a qualitative discussion of potential error sources and quality control procedures. It does not quantify the uncertainty of the published SWC values or demonstrate how uncertainties from neutron counting, calibration, soil properties, temporal instrument stability, and data processing propagate into the final dataset. Please provide a quantitative uncertainty assessment, including calibration errors, repeated measurement precision, and uncertainty variation across depths and sites. The derivation and temporal stability of the standard count of 685, the thresholds used for anomaly screening, the number of records removed or corrected, and the treatment of missing values should also be documented. Appropriate quality flags or uncertainty estimates should be included in the published data.
11. Section 5: This section contains several useful demonstrations of the dataset, but many analytical methods are not described in sufficient detail, and some interpretations extend beyond what can be supported by the observational design. The following issues should be addressed: a. Section 5.1: The identification of the three temporal phases, calculation of the reported depletion and recovery percentages, and statistical tests used to compare the four plots are not described. Please provide the corresponding methods and account for temporal autocorrelation and the experimental limitations identified in Comment 8. Statements that SWC reached a level inaccessible to plants or recovered because of climatic humidification should be presented as possible explanations unless supported by direct analyses of plant available water, precipitation, evapotranspiration, or other relevant variables.
b. Section 5.2.2: The time stable sites appear to be selected and evaluated using the same observations, so the reported representation errors are not based on independent validation. Please assess their predictive performance using a temporal holdout period, cross validation, or another independent procedure. The criteria were also relaxed after no sites met the original thresholds in the shallow layers, and sites with ITS values of 10–11% were subsequently recommended despite the stated criterion of ITS < 10%. Please justify these changes and distinguish exploratory site selection from validated upscaling performance.
c. Section 5.3: Please describe and validate the spatial interpolation method used to produce Figure 8. The comparison of PAWS among land use types should also account for their uneven spatial distribution and potential confounding by climate, soil, and topography.
d. Section 5.4: The comparison with external SWC products is not reproducible because the product versions, spatial and temporal resolutions, grid extraction, temporal matching, and soil layer harmonization procedures are not described. These issues are particularly important because each transect campaign lasted three to four days and the products have different layer definitions and spatial supports. Please provide complete matching procedures and quantitative performance metrics, including bias, RMSE, correlation, and uncertainty. Conclusions regarding increasing errors with depth and the relative performance of individual products should be supported by these analyses and should account for point to grid representativeness differences.