the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
CAMELS-KR: Catchment attributes, meteorology, and reconstructed streamflow for large-sample hydrology in South Korea
Abstract. Large-sample hydrology has increasingly relied on harmonized datasets that integrate hydrometeorological observations with catchment attributes across diverse environmental settings. Although the Catchment Attributes and MEteorology for Large-sample Studies (CAMELS) initiative has expanded rapidly worldwide, South Korea remains underrepresented despite its distinctive hydroclimatic characteristics, including a monsoon-dominated climate, steep mountainous terrain, rapid runoff generation, and extensive anthropogenic water regulation. Here, we present CAMELS-KR, the first CAMELS-style dataset developed for South Korea. CAMELS-KR provides harmonized hydrometeorological time series, catchment boundaries, and catchment attributes for 282 quality-controlled catchments distributed across the major river basins of the country. The dataset includes daily streamflow and water-level observations, catchment-scale meteorological forcing data from 1981 to 2025, and a comprehensive set of static attributes describing topography, climate, hydrology, land cover, soils, and water infrastructure. CAMELS-KR also includes reconstructed streamflow generated using a regionally trained long short-term memory (LSTM) model and a locally calibrated conceptual HBV model, together with model performance metrics and calibrated parameter sets. By providing open-access, analysis-ready hydrological data from a monsoon-dominated and highly regulated environment, CAMELS-KR fills an important geographic gap within the global CAMELS network. The dataset is expected to support comparative hydrology, prediction in ungauged basins, climate-impact assessments, and the development and benchmarking of next-generation data-driven and process-based hydrological models.
- Preprint
(2557 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 11 Oct 2026)
- CC1: 'Comment on essd-2026-544', Iiro Seppä, 21 Aug 2026 reply
-
RC1: 'Comment on essd-2026-544', Nikunj K. Mangukiya, 27 Aug 2026
reply
The compilation of CAMELS-KR represents a commendable effort to harmonize fragmented national datasets in East Asia. However, the manuscript currently lacks transparency regarding GIS watershed delineation workflows, exhibits unaddressed water-balance violations in Budyko space, contains some inconsistencies, and provides incomplete information on the streamflow modeling setup and hydrogeological characterization.
Specific Comments:
1) The manuscript describes the 20% area discrepancy filtering criterion in Section 2.2, but nowhere does it specify the underlying digital elevation model processing, flow direction algorithm, or gauge snapping procedures used to delineate the 282 catchments from the 30-m ALOS DEM. The authors should explicitly document the delineation pipeline, software/tools used, snapping search radius tolerances, and how boundaries in flat coastal or highly urbanized alluvial plains were quality-checked against official national drainage boundaries.
2) In Figure 7b, multiple catchments exhibit long-term runoff coefficients substantially exceeding unity (Q/P > 1.0) or plot well below the lower Budyko demand limit (Eactual >> PET). The statement on lines 441-443 ("Most catchments fall within the theoretical water and energy limits...") overlooks these clear physical discrepancies. The authors must quantify the number of anomalous basins and investigate the underlying causes, such as orographic precipitation underestimation from ordinary kriging, trans-basin water transfers, deep inter-catchment groundwater flow, artificial drainage, or unmodeled dam/reservoir release regimes, and provide a quality/reliability flag in the dataset for basins with severe water balance closure errors.
3) The meteorological forcing dataset was developed using 2D Ordinary Kriging of ASOS/AWS weather stations onto a 0.1° grid. In topographically rugged terrain like the Korean Peninsula, standard spatial interpolation without elevation covariates (e.g., Kriging with External Drift, regression kriging, or PRISM-type elevation lapse rates) typically underestimates ridge-line and high-elevation precipitation. The authors should discuss whether elevation lapse rates were considered and compare the gridded precipitation against satellite/gauge-corrected benchmarks or local high-elevation gauges to assess potential precipitation underestimation.
4) A network of 282 catchments across South Korea contains significant spatial overlap and nested river networks. It would be beneficial for the users if authors can provide a topological routing matrix or metadata attributes (e.g., is_headwater, upstream_basin_ids, nested_area_fraction).
5) Observed streamflow records vary substantially in length (10 to 35 years; Fig 3c) and include non-continuous gaps. The authors should clarify whether hydrologic signatures were computed over the full variable record or harmonized common benchmark window, and address how structural regime shifts (e.g., post-2010 construction from the Four Major River Project) affect static signature stability.
6) While soil characteristics from SoilGrid 2.0 are well documented, the dataset lacks attributes describing underlying bedrock geology, lithology, and hydrogeological properties (e.g., hydraulic conductivity and porosity from GLHYMPS or global lithological datasets like GLiM). Incorporating standard lithological classes and bedrock permeability attributes would bring CAMELS-KR into complete alignment with other CAMELS datasets.
7) The streamflow reconstruction section requires additional technical specifics. For the regional LSTM model, which specific static attributes and meteorological forcings were fed into the feature vector? was hyperparameter tuning performed via k-fold cross-validation or an out-of-basin spatial split? For the HBV model, which optimization algorithm and objective function formulation were applied during local calibration? what parameter search ranges were defined? The authors can also include diagnostic attribution to clearly show where and why model performance degrades.
8) Section 8 (Data Availability) states that daily forcing includes vapor pressure and sunshine duration, but these variables are not present in the dataset.Â
9) The caption for Figure 1 (lines 106–111) erroneously includes a duplicated description belonging to Figure 4.
10) The time series data include negative water level values for certain gauges, but the manuscript lacks an explanation of the underlying datum definition beyond describing it in Table 1 as "relative to the station-specific zero datum". In hydrometric networks, negative readings typically arise from channel bed degradation/scour below an established gauge zero or arbitrary local datum definitions. The authors should explicitly clarify in Section 3.1 why negative stages occur, state whether gauge zeros have shifted over time, and include the official gauge zero datum elevation (in above mean sea level or the national geodetic datum) in the static attribute table so users can convert relative stage into absolute hydraulic head.
11) In Figure 8a (number of dams and reservoirs), the color ramp spans up to ∼800, but the inset histogram shows that the overwhelming majority of catchments have counts concentrated below 0–100. As a result, almost all catchment points appear uniformly pale yellow with little to no visible distinction across moderate differences. Similar issue is there in Figure 8b. Also, in Figure 8c, the colorbar map represent population density. The unit on the colorbar is labeled as [km-2], which is missing the actual quantity (e.g., [persons km-2]).
Citation: https://doi.org/10.5194/essd-2026-544-RC1
Data sets
CAMELS-KR: Catchment attributes, meteorology, and reconstructed streamflow for large-sample hydrology in South Korea (Version version1.0) Lee et al. https://doi.org/10.5281/zenodo.21930882
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 83 | 40 | 11 | 134 | 14 | 16 |
- HTML: 83
- PDF: 40
- XML: 11
- Total: 134
- BibTeX: 14
- EndNote: 16
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,Â
Always great to see a new CAMELS dataset, congratulations on the hard work!
I noticed from figure 7 that there are quite many catchments where the runoff exceeds rainfall (n=~15) and catchments where more of the precipitation is lost before runoff measurement than PET alone explains (n=~25). It would be very beneficial to investigate where these issues likely arise, because there are many possible explanations, such as changes in storages, short timeseries, extractions for human uses, leaky catchments (subsurface flow or bifurcations removing or adding water to the catchment) or quality issues in the source products or catchment delineation. Knowing a possible explanation for the discrepancy would help increase trust in the data, or could direct future improvements if the source data is revealed faulty.
Additionally, I noticed that you did not explain the catchment delineation and quality control procedure or source data anywhere. This is crucial information, and directly affects the quality of the rest of the data.
Hopefully these will help improve the manuscript.
Best Regards,Â
Iiro Seppä