the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
GSIM-PLUS: A Gap-Filled Global Monthly Streamflow Dataset for 1995–2015
Abstract. Global monthly streamflow observations are fundamental for understanding changes in the water cycle, supporting large-sample hydrology, and informing water-resources assessments. However, currently available open station archives still suffer from substantial limitations in temporal continuity and spatial coverage. Here we present GSIM-PLUS, a gap-filled global monthly streamflow dataset for 1995–2015 designed to improve the completeness and reusability of global runoff records. Using the GSIM monthly archive as the basis, we identified 7,323 high-completeness anchor stations and 8,731 target stations from 30,959 gauges. Basin descriptors from five groups – climate, topography, soil, spatial location, and hydrology – were used to identify the most similar donor stations for each target site. Donor-Trend Recursive Regression (DTRR) was adopted as the default imputation method, with a guarded fallback to baseline MAML for a limited subset of very-low-flow stations in order to improve production stability under long recursive gaps. Multi-scenario validation shows that DTRR achieved the best overall performance under random 30 % masking (NSE = 0.865; KGE = 0.920) and remained robust for both 12-month continuous gaps (NSE = 0.795) and very long gaps exceeding 25 months (NSE = 0.511). Independent validation using 16 GRDC stations across six regions further confirmed good transferability, while indicating that temporal agreement was generally more robust than exact magnitude reconstruction under donor-limited or long-gap conditions. Under the guarded DTRR production scheme, GSIM-PLUS fills 303,271 missing monthly records for 16,054 stations, increasing the median completeness of target stations from 66.3 % to 81.0 % and that of the full dataset from 86.9 % to 95.2 %. Each released record is accompanied by quality and context metadata, including reconstruction class, gap length, fill method, and basin-context flags. GSIM-PLUS provides a more continuous and traceable global monthly streamflow resource for regional hydrological analysis, large-sample studies, model evaluation, and related monthly-scale applications. The GSIM-PLUS dataset is publicly available through Zenodo at https://doi.org/10.5281/zenodo.21425702.
- Preprint
(2115 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 08 Oct 2026)
-
RC1: 'Comment on essd-2026-316', Anonymous Referee #1, 18 Aug 2026
reply
-
AC1: 'Reply on RC1', Mingrui Chen, 01 Sep 2026
reply
We agree that this question concerns the central purpose of the dataset. The "PLUS" in GSIM-PLUS does not mean that we have created a longer observational history or enlarged the original gauge network. It refers to added continuity, traceability, and common-period usability within the observation-supported 1995-2015 window. This distinction is important: a 21-year series is not, on its own, sufficient for robust multi-decadal climate-trend attribution, but fragmentation within those 21 years still substantially limits many monthly-scale uses.
This limitation is especially relevant to data-driven modelling. Deep-learning and other sequence-based methods commonly form continuous input and target windows. A missing month can invalidate an entire surrounding window, force users to introduce their own undocumented preprocessing, or reduce the analysis to a smaller and potentially non-representative subset of stations. The same problem affects common-period model evaluation, seasonal climatology, and cross-basin comparison, because records with different internal gaps are not evaluated over the same months. Although some modern models can represent missing inputs explicitly, substantial fragmentation still reduces the amount of directly usable training and evaluation information. Improving continuity is therefore a practical data gain even when the beginning and end dates remain unchanged.
The scale of this gain is measurable. GSIM-PLUS reconstructs 303,271 internal missing months across 16,054 stations. For the 8,731 incomplete target stations, median completeness within the common 252-month window rises from 66.3% to 81.0%. Across the released station set, the number of records complete for all 252 months increases from 2,605 to 7,674; 2,325 stations newly attain at least 10 complete calendar years and 3,487 newly attain at least 15 complete years. These gains create more continuous monthly windows and a larger consistently aligned station sample. They do not create additional gauge observations, and they do not make every reconstructed station appropriate for every application.
The referee is also correct that reconstructed values contain error. Our position is not that an estimated value is equivalent to an observation. Rather, GSIM-PLUS offers a single reproducible reconstruction workflow whose behavior has been tested under several missingness structures. DTRR was the strongest method among the eight candidates evaluated in our target-population stress tests, while the new leakage-controlled anchor experiment provides a more conservative assessment of the actual production rule. The product now retains the observed/reconstructed status, the applied method, gap-length class, hydroclimatic and low-flow context, and an empirical 80% reconstruction range for each reconstructed value. These fields do not eliminate error, but they allow users to identify it, screen higher-risk cases, and carry reconstruction uncertainty into sensitivity analysis. In this sense, the added value is not only that more values are available; it is that the estimates are generated and documented consistently rather than produced independently by each downstream user without shared validation or provenance.
We also agree that internal gap reconstruction and extension beyond 2015 are fundamentally different tasks. Internal reconstruction operates within a station's archived temporal support. The model can learn from that station's available observations and from simultaneous donor records, and its assumptions can be tested by hiding known values. Extension beyond the observed period is temporal extrapolation: the target values are unavailable by definition, the future relationship may not remain stationary, and additional external information is required. The uncertainty of that external information then becomes part of the extended product.
Gridded runoff, reanalysis, or precipitation-runoff estimates are valuable possible sources for such an extension, but they are not interchangeable with outlet-gauge discharge. A gridded runoff value represents an areal model estimate at a particular spatial support, whereas a gauge measures the routed response of its upstream catchment and may include reservoir operations, abstractions, and local controls. Grid cells and gauge catchments generally do not have a one-to-one spatial correspondence. Directly attaching grid values to station records can smooth peaks or zero-flow behavior and can produce particularly large relative errors in intermittent or very-low-flow rivers. A rainfall-runoff extension would therefore require explicit scale matching, routing and regulation treatment, non-stationarity assessment, and independent post-2015 validation. We regard this as a valuable future research direction, but it should be developed and validated as a distinct model-assisted extension rather than presented as if it were the same operation as filling internal gaps.
Accordingly, we will sharpen the dataset's role in the revised submission: GSIM-PLUS is a traceable, uncertainty-aware continuity layer for monthly station observations, complementary to the original GSIM archive. It increases the amount of analysis-ready information within a fixed period while preserving clear boundaries between observations, internal reconstructions, and any future temporal extrapolation.
## Reliability and the role of GRDC (Major comment 2 and Minor comment 3)
We agree that controlled masking of high-completeness stations provides a stronger test than relying principally on the small GRDC comparison. Following your suggestion, we have completed a fivefold station-wise experiment using all 7,323 anchor stations. Only originally observed months are masked. Each test fold is excluded from the donor pool and MAML meta-training, while the tested station's flow scaling, mean-flow matching attribute, and low-flow criterion are recomputed from its remaining observations. The experiment therefore evaluates reconstruction without using the hidden target values to select or fit the reconstruction.
The tests cover random, continuous, very-long, and mixed missingness. Median station-level NSE is 0.670 under random 30% masking and 0.622 for the 25-48-month block scenario; across the seven scenarios, 77.3%-95.7% of stations have positive NSE. These are station-level summaries, not scores calculated after pooling all stations. The results also reveal important failures: for the 12-month scenario, stations with median support-period flow below 0.02 m^3 s^-1 have median NSE = -0.251 despite the MAML alternative. Thus, the additional evidence supports useful reconstruction at many stations, but not a universal accuracy claim. Synthetic masking also cannot reproduce every mechanism responsible for real archive gaps.
We will use this experiment as the principal validation and retain GRDC as supplementary cross-archive evidence. We accept the concern about small samples: US_0005774 has only three paired observations, and its high correlation is not statistically significant (p = 0.0897). It should not be presented as convincing evidence of reconstruction accuracy. The revised comparison reports sample sizes and correlation p values and qualifies the interpretation accordingly.
## Consequences for downstream analyses (Major comment 3)
We share the concern that filling gaps can introduce artifacts and should not automatically be assumed to improve scientific conclusions. To test this directly, we compared hydrological statistics from the original high-completeness anchor records, their synthetically incomplete versions, and their reconstructions. The diagnostics include mean flow, seasonal behavior, variability, monthly-flow quantiles, and annual-mean slopes.
The findings are mixed, which we consider important to report. Under random masking, median normalized mean-flow error decreases from 0.0276 to 0.0143, whereas median absolute error in the coefficient of variation increases from 0.0263 to 0.0381. Lower monthly-flow quantiles also become less accurate in that scenario. Annual-mean slope estimates improve on average, but this only measures reconstruction-induced distortion within the 21-year window; it does not establish the adequacy of that window for long-term attribution.
These results change how we frame appropriate use: improved completeness can benefit some statistics while degrading others. We have also calculated empirical 80% reconstruction ranges, but these pointwise ranges are not uncertainty intervals for an entire time series or its trend. The revised submission will present the favorable and unfavorable results together and recommend checking the sensitivity of conclusions to reconstructed values, gap length, and basin context rather than treating filled values as equivalent to observations.
## Method clarity and remaining technical points (Major comment 4 and Minor comments 1-2)
We agree that the method description was too compressed. We have prepared an expanded explanation and workflow diagram distinguishing attribute-based donor matching from station-level reconstruction. DTRR combines seasonal terms, target persistence, and donor levels and changes in a station-specific ridge regression. MAML uses an initialization learned from anchor tasks, then adapted with matched-donor and target observations. Both predict recursively. The full operational steps, fixed similarity coefficients, and numerical example will accompany the revised submission; the low-flow alternative will not be described as eliminating low-flow uncertainty.
Two specific clarifications can be made here. First, ARMA does not require catchment similarity; our original wording conflated temporal models with spatial-transfer methods and has been corrected in the draft. Second, attribute standardization uses anchor-based z scores, not bounded min-max scaling. Target attributes beyond the anchor range are not clipped, although their nearest available donors may still be poor analogues. Reconstruction separately uses the target's own observed-flow scale. Standardization therefore permits the calculation but does not solve inadequate donor representation.
We appreciate these comments because they help distinguish what the dataset provides from what it cannot establish. Our aim in the revision is to substantiate its practical, bounded value while making the remaining risks visible to users.
Citation: https://doi.org/10.5194/essd-2026-316-AC1
-
AC1: 'Reply on RC1', Mingrui Chen, 01 Sep 2026
reply
-
RC2: 'Comment on essd-2026-316', Anonymous Referee #2, 20 Aug 2026
reply
General statement
This is a well-presented manuscript, that takes an existing streamflow dataset and applies a gap-filling strategy in order to help potential users that would not have the hydrological expertise to do it. Gap-filling can be seen as a dangerous practice: those who master the technique know how much it can be uncertain, and those who do not master it can be too confident. Anyway, it is useful and this gap-filled dataset is extremely valuable.
General impression from the paper
I found the paper sometimes not very well organized, and sometimes not detailed enough.
Recommendations
I have a few concerns:
- First of all, the concept of catchment similarity assumes implicitly that human activity is negligeable, i.e. that there is no notable regulation/abstraction/derivation impacting the behavior of the catchment. I was surprised to see that there are no discussions of this issue. Did I miss it? I believe that you need either to work on a subset of minimally-impacted stations, or to at least verify that your donor stations are minimally impacted.
- Jargon: I believe any author of a scientific paper should try to avoid jargon (I am not sure to respect this recommendation myself, but nonetheless, I believe we should aim at it…) I found a lot of jargon that I was not used to. Could you try to reduce it or at least to provide a list of it with definitions at the beginning of the paper? For example, ‘donor’ (I use it a lot too) would deserve to be defined even if it seems rather easy to understand, the same for ‘anchor’, ‘target’, ‘guarded fallback’… Last I believe you should try to avoid completely jargon in the abstract (you can still add all the specialized vocabulary in your list of keywords).
- The notion of distance is central in your approach, and the distance you choose combines the geographical Euclidean distance with a distance in the physical feature space: ‘Similarity between a target station and a candidate donor was then defined as a monotonic transformation of weighted Euclidean distance in the standardized feature space.’ Honnestly, I am afraid that you remained too abstract in the explanation you give in section 3.2,
- I believe that your explanation of the DTRR method is so minimalist that nobody will be able to reproduce it! You need to detail more the procedure, perhaps introduce a flow-graph or something equivalent to help the reader visualize what this equation represents.
- Also, the “Validation” part (section 4) should be part of the Method section. I liked the fact that you used several benchmark methods, in particular the two simplest ones (IDW interpolation and infilling by mean seasonal flow). It would have been challenging to use these two methods as reference in your numerical assessment, for example in using the ration of the RMSE of DTRR with the RMSE of one or two of the benchmarks (even perhaps with the lowest of both).
- Last, although I know that this represents a lot of work, I regret that you limited your uncertainty assessment to a quality flag. As I mentioned in introduction, users are largely unaware of the uncertainties of streamflow reconstruction. If you had produced a numerical 80% uncertainty range, they would have seen the actual uncertainties!
Citation: https://doi.org/10.5194/essd-2026-316-RC2 -
AC2: 'Reply on RC2', Mingrui Chen, 01 Sep 2026
reply
We thank the referee for recognizing the potential value of the dataset and for highlighting the danger of overconfidence in reconstructed streamflow. We share this concern. Making a record more continuous is useful only if users can understand how the estimates were obtained and where they remain uncertain. Below, we address the main scientific concerns and summarize the additional evidence already obtained, followed by clarifications of the method and presentation.
Human influences on catchment similarity (Comment 1)
Physical similarity alone cannot establish comparable discharge behavior where reservoir operations, abstractions, or diversions differ. This limitation was insufficiently discussed in the manuscript. Our intention is to reconstruct observed flow regimes, which may include human influences, rather than to produce naturalized streamflow. However, this does not remove the risk that physically similar basins have incompatible flow behavior.
We have therefore examined the regulation context of the selected stations using RiverATLAS, with matches available for 99.6% of the released station set. We use degree of regulation (DOR), which relates cumulative upstream reservoir storage to mean annual flow volume, as a screening indicator. At a 10% DOR boundary, 84.0% of the 43,545 target-donor pairs with known context use a donor below that boundary, and 80.2% place the target and donor in the same broad category. These percentages describe pairs, not unique donor stations. Sensitivity checks at 5% and 20% also show that most pairs use lower-regulation donors.
This audit provides evidence about reservoir-storage context, but it does not prove that all donors are minimally impacted. DOR cannot recover operating rules or all abstractions and diversions. We have therefore retained the observational station set rather than relabel it as natural, and added regulation-context information to the prepared revised products. We have also compared validation performance across these groups. We will present this as an additional diagnostic and screening aid, not as evidence that physical similarity fully captures human influences. Applications requiring natural-flow records will still need more specific screening.
Making uncertainty visible (Comment 6)
We agree that a quality flag alone does not communicate the magnitude of uncertainty. In response, we have calculated empirical central 80% reconstruction ranges and added lower and upper bounds to the reconstructed values in the prepared revised products. The original observations remain distinguishable from these estimates.
The bounds use prediction residuals from a new fivefold anchor-station masking experiment. When evaluating a fold, its residuals are excluded from the quantiles used to construct its bounds. Calibration distinguishes method, gap class, aridity, and low-flow context where sample support permits. Empirical coverage is 80.0% overall and 79.6%-80.7% across the seven masking scenarios. Relative ranges are wider in arid and very-low-flow settings, making those risks visible numerically rather than through a categorical warning alone.
We will also state the limits of this information. Aggregate coverage in the masking experiment does not guarantee the same coverage at every target station. These are reconstruction-error ranges, not a complete account of measurement error, human-regulation uncertainty, or temporally dependent errors. In particular, pointwise 80% ranges do not establish an 80% uncertainty interval for a derived trend. Our additional downstream analysis shows that reconstruction can improve mean-flow and seasonal estimates while worsening variability or lower-flow quantiles. The revised submission will present these limitations alongside the numerical bounds, rather than suggest that adding an interval resolves all uncertainty.
A transparent and understandable workflow (Comments 2-4)
We accept that readers should not have to infer the meaning of our terminology. An anchor is a high-completeness reference station, a target is the incomplete station being reconstructed, and a donor is an anchor selected for that target. We have introduced these definitions together in the draft and simplified the Abstract. We also replace "guarded fallback" with "predefined low-flow alternative," explaining the rule rather than relying on the label.
The distance calculation requires an important clarification: the implementation does not add a separate geographical distance to a physical-feature distance. Latitude and longitude are two coordinates in the same standardized 17-attribute vector. Fixed coefficients scale the coordinate differences before Euclidean distance is calculated, and similarity is 1/(1+distance). Mean flow and basin areas are log-transformed first. We have prepared the complete coefficients and a worked example so that these operations, including the role of geography, are explicit rather than abstract.
For reconstruction, DTRR combines seasonal variation, the preceding target flow, and the current levels and changes of selected donors in a station-specific ridge regression. During a gap, each prediction becomes the lagged target input for the next month. MAML provides the predefined alternative for very-low-flow stations; it learns an initialization from anchor tasks and adapts it using donor and target observations. It also predicts recursively and does not eliminate low-flow uncertainty. The expanded explanation and new workflow diagram will show these steps and their inputs directly. Detailed settings will be provided without requiring readers to reconstruct the procedure from one equation.
Simple benchmarks and validation organization (Comment 5)
We agree that the experimental design belongs in Methods, with the resulting performance discussed separately. The draft now follows that distinction. We have also implemented the suggested comparison with the better of the two simple methods, using exactly the same hidden months and station-fold exclusions as the production evaluation. The implemented IDW uses similarity-based donor weights, while Seasonal Mean pools donor calendar-month observations; these choices will be defined explicitly.
For each station, we divide production RMSE by the lower RMSE of IDW and Seasonal Mean. Across the seven scenarios, the median station-wise ratio ranges from 0.829 to 0.867, corresponding to a median relative reduction of 13.3%-17.1%. The numerator follows the actual production rule, including the low-flow alternative, rather than DTRR alone. The production workflow does not outperform both simple methods everywhere, and we will show that variation rather than report only the median improvement.
We appreciate the emphasis on making the product usable without encouraging uncritical reliance on filled values. The additional evidence and clearer explanations are intended to support that distinction, and the full analyses will accompany the revised submission.
Citation: https://doi.org/10.5194/essd-2026-316-AC2
Data sets
GSIM-PLUS-v3 Mingrui Chen https://doi.org/10.5281/zenodo.21425702
Model code and software
GSIM-PLUS Mingrui Chen https://github.com/skyhhu2024-create/GSIM-PLUS
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 157 | 46 | 26 | 229 | 27 | 21 |
- HTML: 157
- PDF: 46
- XML: 26
- Total: 229
- BibTeX: 27
- EndNote: 21
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This study presents GSIM‑PLUS, a globally gap‑filled streamflow dataset built upon the GSIM archive using a donor‑station‑based recursive imputation workflow (DTRR) with a MAML fallback for challenging low‑flow conditions. The work increases temporal completeness for thousands of streamflow records and delivers a potentially valuable resource for large‑sample hydrological analyses. Nevertheless, several important methodological, validation and uncertainty limitations remain that constrain the dataset’s real‑world applicability. Below are the detailed major and minor comments.
Major comments
1. Temporal and spatial coverage remains limited. The authors argue that the pervasive gaps in streamflow time series reduce their reuse value for trend attribution, extreme-event analysis, and longterm modelling. Yet the new dataset only spans 21 years, which raises a key question: how can this new dataset support robust long-term trend assessment? In addition, GSIM-PLUS only includes about half of GSIM. This prompts reflection on what incremental value the “‑PLUS” extension actually delivers. These limitations might constrain the practical utility of GSIM-PLUS given that GSIM has relatively balanced spatio-temporal coverage. The authors have to carefully think about its added value. Maybe the records of these target stations are already informative for most studies.
2. The valiation experiment design is insufficiently rigorous, and the independent validation against GRDC does not yield convincing results. A well‑designed synthetic validation experiment is essential to fully demonstrate the reliability and effectiveness of the proposed gap-filling approach. Specifically, the authors should conduct controlled masking tests using only high-quality anchor stations, where known observational records are intentionally masked and subsequently reconstructed. This additional synthetic validation can provide solid, quantitative evidence to further verify the method’s performance and add substantial credibility to the dataset.
4. The method section is overly concise and lacks sufficient technnical detail. The operational principles of the proposed DTRR method and the supplementary MAML approach should be substantially expanded and clarified. Specifically, it remains unclear how the two algorithms process streamflow time series and achieve gap-filling. In addition, the criteria used to quantify hydrological similarity and select donor catchments for each target station need explicit elaboration.
Minor comments
1. L50: why does ARMA need catchment similarity?
2. L140: Here the standardization is based on anchor stations. What if the target stations have much large ranges that beyond the that of the anchor stations?
3. L270-275: Here is the evaluation is not very meaningful given the very limited sample sizes. As I commented above, other validations are needed.