the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
GRIT-ADB: A Global Attribute Database for the GRIT Hydrography
Abstract. Global hydro-environmental databases provide essential information for large-scale hydrological, ecological, geomorphological, and Earth system analyses. Most existing global databases are built upon convergent river representations that do not explicitly capture bifurcating, multi-channel river systems. In addition, these databases primarily characterise long-term climatological means or static representations of environmental conditions derived from earlier-generation global datasets, limiting their applicability for time-varying analyses of hydroclimatic and geomorphological processes. Here we present GRIT-ADB, a new attribute database for the vectorised Global River Topology (GRIT), a new river hydrography dataset created from a 30 m resolution river mask and terrain data that provides a topology-explicit and physically realistic representation of river networks including divergent flow pathways. GRIT-ADB provides standardised hydro-environmental information for 19.6 million km of rivers and streams. It currently comprises 64 time-varying (multi-dimensional embeddings) and 35 static variables (>300 attributes), spanning five categories: hydrology, physiography, climate, land cover and use, and soils and geology. Attributes are derived by aggregating and harmonising data from state-of-the-art global datasets and are accumulated along the river network from headwaters to basin outlets, while preserving the topology of divergent and complex flow pathways. The attributes are linked to multiple GRIT scales, including hierarchically-nested subbasins, individual river reaches of up to 1 km long, and coarser-scale river segments of several kilometres long, providing a flexible framework that can accommodate future extensions of GRIT and additional attributes. By combining a standardised attribute framework with explicit representation of bifurcating river hydrography, GRIT-ADB enables improved large-scale yet high-resolution analyses of river connectivity, hydrological extremes, hydro-ecological processes, and climate impacts in complex river systems, supporting a wide range of global hydrological and environmental applications.
- Preprint
(15937 KB) - Metadata XML
-
Supplement
(3183 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-279', Anonymous Referee #1, 13 Jul 2026
-
RC2: 'Comment on essd-2026-279', Anonymous Referee #2, 06 Aug 2026
General comments
GRIT-ADB is a useful and timely contribution. The community has needed an attribute layer on top of GRIT since the hydrography was published, and building it as a general, extensible framework is the right design decision. The scope of harmonised source data is respectable, the multi-scale structure (reach / segment / subbasin, local / upstream) is sensible, and the inclusion of AlphaEarth embeddings is forward-looking. I expect this dataset to be useful to the community.
The most immediate problem is that the global dataset described by the manuscript is not available for review. Section 6 provides only a Congo River basin example and states that the full product will be deposited before final publication. This prevents assessment of global completeness, internal consistency, regional quality variation, and practical usability. The ESSD data policy requires the data described by a submission to be accessible during review, and the journal asks reviewers to assess the dataset itself, not only the manuscript.
Beyond accessibility, there is no dedicated validation or quantified quality assessment of the derived product. The Discussion acknowledges that biases in the source datasets will propagate into GRIT-ADB, but provides no uncertainty characterisation, independent cross-check, internal consistency assessment, or sensitivity analysis on consequential processing choices. Figure 7 shows that GRIT-ADB differs from HydroATLAS; it does not show that GRIT-ADB is closer to reality.
Second, the manuscript repeatedly claims that bifurcation-aware accumulation materially improves attribute estimates, but never quantifies the magnitude or spatial distribution of the effect. This is the claimed advantage of the product and can be tested using the GRIT network and source data already available to the authors.
I therefore recommend major revision and re-review once the complete dataset is accessible and a substantive quality assessment has been added.
Major comments
1. Make the complete global dataset available for review
The Data Availability section provides only a Congo River basin example and promises the global product before final publication. That is not sufficient for a data-description paper whose subject is a global database. The complete release should be made accessible through a functional DOI or an anonymous repository review link before the review can be completed.
2. Add a data-quality, validation, and uncertainty assessment
The qualitative statement that source-data biases propagate into GRIT-ADB is not an assessment of the derived product. The ESSD criteria for data-description articles call for validation against independent reference data or, if that is unavailable, a structured plausibility assessment. Please add a dedicated section that addresses the following:
- Completeness. Report valid, missing, and imputed observations by attribute and geographic domain, including the spatial pattern of failures and coastal filling.
- Internal invariants. Test conservation of extensive quantities across bifurcations and reconvergences.
- Source-to-product checks. Re-extract a documented sample of raster and vector inputs and compare the independently calculated values with the released fields.
- Independent or semi-independent plausibility checks. Examples could include GRIT drainage area against gauge metadata, known reservoir or lake totals in selected basins.
- Sensitivity. Quantify sensitivity to bifurcation weights, midpoint versus catchment extraction and resampling choices.
Specific comments
- l. 8-10 / Table 2 / repository metadata: 35 static variables plus 64 embedding dimensions equals 99, whereas Table 2 reports 98 and the example repository describes approximately 60 attributes. Reconcile all totals and define consistently what counts as a variable, embedding dimension, stored field, and expanded attribute.
- Abstract l. 13-14 versus l. 63-64: The statements that reaches are up to 1 km long and that each segment is divided into equal-length reaches shorter than 1 km are compatible, but they do not uniquely specify the subdivision algorithm. State the exact rule, including treatment of short remainders and segment nodes, and provide the distribution of reach lengths.
- l. 71-74: State the raster reprojection and resampling method for every data type. Nearest-neighbour, bilinear, and area-conservative resampling have materially different implications for categorical classes, continuous variables, and totals.
- Table 1, "Year" column: "Most recent" is not a reproducible date. Give the exact source version, release date, and retrieval date used for GRanD v1.3, SoilGrids v2, FABDEM, GRIT, and any other rolling or versioned source.
- Table 1, air-temperature seasonality: "Climatology" is capitalised and the citation is parenthesised inconsistently with adjacent rows.
- Units across figures and Table A1: Sand fraction appears as percent in Figures 2 and 7 but as g/kg in Figure 5 and Table A1. Stream gradient appears as m/m in Figures 3 and 5 but m/km in Table A1. Harmonise units or explicitly annotate every conversion.
- Figure 5 and l. 144-145: The statement that no visible differences occur in sand fraction across climate regions appears stronger than the boxplots support. Quantify the comparison or soften the statement.
- Data availability and licensing: State the licence of the derived product and substantiate the assertion at l. 77-79 by listing the relevant licence and redistribution terms for each source dataset. Confirm that redistribution of the derived attributes is compatible with the entire source chain.
Citation: https://doi.org/10.5194/essd-2026-279-RC2 -
RC3: 'Comment on essd-2026-279', Anonymous Referee #3, 28 Aug 2026
============================================================GENERAL ASSESSMENT AND SUMMARY============================================================The manuscript presents GRIT-ADB, a global hydro-environmental attribute database built on the GRIT bifurcating-river hydrography. For roughly 20 million river reaches (19.6 million km of rivers), it harmonises 35 static and a set of time-varying attributes from state-of-the-art global sources across five categories (hydrology, physiography, climate, land cover and use, soils and geology), and reports them as local and upstream-accumulated statistics at reach, segment and catchment scales. Its distinguishing feature is that the accumulation respects divergent, multi-channel flow, so attributes are propagated correctly through bifurcations and deltas that single-thread networks such as HydroSHEDS/HydroATLAS cannot represent.The core idea is genuinely novel and, by ESSD's standards, unique: a topology-aware attribute database that follows bifurcating rivers does not exist elsewhere, and it is not something a user could reproduce on a routine basis. The choice to build on GRIT, to draw only from openly licensed, actively maintained sources, and to provide both reach- and segment-scale statistics is sensible and clearly useful for large-sample and network-based hydrology. Adding annual satellite embeddings as a time-varying layer is forward-looking. The topic sits squarely within ESSD's scope, and the underlying hydrography (Wortmann et al., 2025) is a strong foundation.My reservations are about the paper as a data paper rather than the idea. The main one is close to blocking until addressed: there is no validation or quality assessment of the attributes, so a reader cannot yet judge how far to trust them. A second concern is how the manuscript currently reads: the text is weighted toward motivation and illustrative science, and comparatively little of it does what a data descriptor should do, namely describe the dataset's contents, structure, reference periods and quality so that a reader can pick it up and use it. None of this is fatal; the dataset looks valuable and the fixes are mostly a matter of adding a validation section and rebalancing the writing. The comments below start with that quality gap, then the readability point, then more specific matters.============================================================MAJOR COMMENTS============================================================1. A data paper needs a quality assessment (Sect. 4).A data descriptor of this kind needs validation against independent references or, where that is impossible, a plausibility assessment that characterises uncertainty and flags extrapolation zones and data-sparse regions. The manuscript currently offers neither. The nearest thing is a visual comparison with HydroATLAS in the Pearl River Delta and the Congo (Fig. 7), which shows that the two products differ. Since every attribute is inherited from an external source and then aggregated and accumulated along the network, two questions matter and go unanswered: how much uncertainty each source contributes (the SoilGrids spatial bias is noted at l. 200–202 but not quantified), and how the accumulation step propagates, and possibly amplifies, that uncertainty downstream. CAMELS (Addor et al., 2017) is a good model here, since much of its value came from documenting the limitations of each attribute. A dedicated validation and uncertainty section, even a plausibility assessment with a few independent cross-checks (gauged basins, RiverATLAS, in-situ soil or lake records), would move the paper from "here are some maps" to "here is a dataset you can trust."2. Reads more like a research paper, less like a data description.It is worth looking at how the text is distributed, since a data descriptor is consulted more often than it is read straight through. Three things stand out. The Introduction is the single longest section and spends much of its length re-motivating large-sample hydrology; it can lose a third. A reader would come to "Dataset overview" (Sect. 3) to learn what is in GRIT-ADB and how to use it. For that currently the reader has to go to the Appendix to find some of the reference material they actually need, namely variable definitions, file formats and naming, the keys that join attributes to GRIT geometries, per-attribute reference periods, and quality. The vignettes provided currently in the section are interesting, but not primary. Meanwhile the Conclusions largely restate the Abstract. A data descriptor should be concise and should say plainly what is in the dataset, where, when, and how. I would rebalance: tighten the Introduction, turn Sect. 3 into a genuine description of the dataset's contents and structure (moving the exploratory analyses into a short "example applications" subsection or the supplement), add the validation section from comment 1. The paper would end up both shorter and easier to use.3. Clarify the attribute count.The headline count, "64 time-varying (multi-dimensional embeddings) and 35 static variables (>300 attributes)," reads as if the database holds 64 independent environmental variables that change through time. In fact the 64 are the latent dimensions of a single product, the AlphaEarth satellite embedding (Brown et al., 2025); essentially the only other time-resolved quantity is lake area. Table 2 then lists GRIT-ADB as "98 (>300)" variables, which matches neither "35 + 64 = 99" nor the abstract. I would state plainly that the time-varying content is one embedding product of 64 dimensions and reconcile the counts. Also, the embeddings are abstract features whose value for hydrology is currently asserted in this manuscript (Fig. 6) rather than demonstrated or previous application cited. I would suggest to add few sentences on what a user can and cannot do with them would set expectations straight.4. Engage the large-sample datasets the paper is built for.GRIT-ADB is repeatedly motivated by large-sample and machine-learning hydrology (Abstract; l. 162–164), yet it does not cite the reference dataset of exactly that field: Caravan (Kratzert et al., 2023), which is a co-author's own work and the standard integration of CAMELS-style static and dynamic attributes for global ML hydrology. CAMELS itself (Addor et al., 2017) is also absent. A short comparison with Caravan, particularly on the time-varying-attribute angle where it is the obvious benchmark, would sharpen the novelty claim. Two further omissions bear on specific analyses: the lake-connectivity results (Sect. 3.2) use GLAKES but mention neither HydroLAKES (Messager et al., 2016), the standard global lake inventory, nor Lake-TopoCat (Sikder et al., 2023), an ESSD dataset that already provides global lake drainage topology and catchments and is the natural comparison for GRIT-ADB's lake-connectivity attributes.5. Climatologies remain outdated (Table 1; l. 38–41).A central argument of the Introduction is that existing databases rely on outdated climatologies (for example HydroATLAS discharge from 1971–2000). But GRIT-ADB's climate layer is itself a mix of periods: aridity and PET come from a 1970–2000 climatology (Global-AI_PET_v3), while temperature and precipitation use 1991–2020 (MSWX/MSWEP) and land cover is a single 2020 snapshot. So the gain in temporal currency is real for some attributes and absent for others. I would harmonise the reference periods where possible or, at least, state the mismatch openly, so that users are not misled by the "present-day" framing.============================================================MINOR AND TECHNICAL COMMENTS============================================================1. Broken DOI (l. 227).The GRIT vector DOI is printed as "10.528 1/zenodo.17435232" with a space; correct to 10.5281/zenodo.17435232.2. Dataset vs paper authorship.The archived-dataset citation (Zhang et al., 2026; l. 225 and References) lists Zhang, Slater, Wortmann, Liu and Moulds, omitting Kratzert, who is a manuscript co-author. Please reconcile the authorship if this is unintended.3. Table 1 citations.Parenthesis typo on "Air temperature seasonality / (Beck et al., 2022)."4. No Code availability statement.The Methods describe an "efficient, automated workflow" (l. 203) and the Conclusions call the framework "reproducible" (l. 221). Do the authors intend to make the code open source? Please add a Code availability section and comment on this.5. Replace "most recent" with a year (Table 1).This is vague for GRanD v1.3, GRWL, GRIT and SoilGrids; give explicit year so users know how current each layer is.6. Repetition.Several sentences recur almost verbatim across the Abstract, Introduction and Conclusions; trimming them would cut length and aid readability.============================================================REFERENCES============================================================- Addor, N., Newman, A. J., Mizukami, N., & Clark, M. P. (2017). The CAMELS data set: catchment attributes and meteorology for large-sample studies. Hydrology and Earth System Sciences, 21, 5293–5313. doi:10.5194/hess-21-5293-2017.- Kratzert, F., Nearing, G., Addor, N., et al. (2023). Caravan — A global community dataset for large-sample hydrology. Scientific Data, 10, 61. doi:10.1038/s41597-023-01975-w.- Messager, M. L., Lehner, B., Grill, G., Nedeva, I., & Schmitt, O. (2016). Estimating the volume and age of water stored in global lakes using a geo-statistical approach (HydroLAKES). Nature Communications, 7, 13603. doi:10.1038/ncomms13603.- Sikder, M. S., Wang, J., Allen, G. H., Sheng, Y., Yamazaki, D., Song, C., Ding, M., Crétaux, J.-F., & Pavelsky, T. M. (2023). Lake-TopoCat: a global lake drainage topology and catchment database. Earth System Science Data, 15, 3483–3511. doi:10.5194/essd-15-3483-2023.Citation: https://doi.org/
10.5194/essd-2026-279-RC3
Data sets
GRIT-ADB: A Global Hydro-Environmental Attribute Database for the GRIT Hydrography B. Zhang et al. https://doi.org/10.5281/zenodo.19363178
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 203 | 94 | 29 | 326 | 53 | 31 | 41 |
- HTML: 203
- PDF: 94
- XML: 29
- Total: 326
- Supplement: 53
- BibTeX: 31
- EndNote: 41
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Zhang and colleagues provide an extremely exciting and useful new dataset for the large-sample hydrology community. Overall the paper is well-written, structured and following a nice flow, which is greatly appreciated from a reviewer’s perspective. I do have some minor comments that should be addressed/answered before the manuscript could be accepted for publication.