the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
CAMELS-KR: Catchment attributes, meteorology, and reconstructed streamflow for large-sample hydrology in South Korea
Abstract. Large-sample hydrology has increasingly relied on harmonized datasets that integrate hydrometeorological observations with catchment attributes across diverse environmental settings. Although the Catchment Attributes and MEteorology for Large-sample Studies (CAMELS) initiative has expanded rapidly worldwide, South Korea remains underrepresented despite its distinctive hydroclimatic characteristics, including a monsoon-dominated climate, steep mountainous terrain, rapid runoff generation, and extensive anthropogenic water regulation. Here, we present CAMELS-KR, the first CAMELS-style dataset developed for South Korea. CAMELS-KR provides harmonized hydrometeorological time series, catchment boundaries, and catchment attributes for 282 quality-controlled catchments distributed across the major river basins of the country. The dataset includes daily streamflow and water-level observations, catchment-scale meteorological forcing data from 1981 to 2025, and a comprehensive set of static attributes describing topography, climate, hydrology, land cover, soils, and water infrastructure. CAMELS-KR also includes reconstructed streamflow generated using a regionally trained long short-term memory (LSTM) model and a locally calibrated conceptual HBV model, together with model performance metrics and calibrated parameter sets. By providing open-access, analysis-ready hydrological data from a monsoon-dominated and highly regulated environment, CAMELS-KR fills an important geographic gap within the global CAMELS network. The dataset is expected to support comparative hydrology, prediction in ungauged basins, climate-impact assessments, and the development and benchmarking of next-generation data-driven and process-based hydrological models.
- Preprint
(2557 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 16 Oct 2026)
-
CC1: 'Comment on essd-2026-544', Iiro Seppä, 21 Aug 2026
reply
-
CC2: 'Reply on CC1', Kuk Hyun Ahn, 30 Sep 2026
reply
Dear Dr. Iiro Seppä,
We thank you for your insightful comments, which have helped us improve both the manuscript and the dataset.Â
Please find attached our detailed responses to your comments.
Kuk-Hyun Ahn and co-authors
-
CC2: 'Reply on CC1', Kuk Hyun Ahn, 30 Sep 2026
reply
-
RC1: 'Comment on essd-2026-544', Nikunj K. Mangukiya, 27 Aug 2026
reply
The compilation of CAMELS-KR represents a commendable effort to harmonize fragmented national datasets in East Asia. However, the manuscript currently lacks transparency regarding GIS watershed delineation workflows, exhibits unaddressed water-balance violations in Budyko space, contains some inconsistencies, and provides incomplete information on the streamflow modeling setup and hydrogeological characterization.
Specific Comments:
1) The manuscript describes the 20% area discrepancy filtering criterion in Section 2.2, but nowhere does it specify the underlying digital elevation model processing, flow direction algorithm, or gauge snapping procedures used to delineate the 282 catchments from the 30-m ALOS DEM. The authors should explicitly document the delineation pipeline, software/tools used, snapping search radius tolerances, and how boundaries in flat coastal or highly urbanized alluvial plains were quality-checked against official national drainage boundaries.
2) In Figure 7b, multiple catchments exhibit long-term runoff coefficients substantially exceeding unity (Q/P > 1.0) or plot well below the lower Budyko demand limit (Eactual >> PET). The statement on lines 441-443 ("Most catchments fall within the theoretical water and energy limits...") overlooks these clear physical discrepancies. The authors must quantify the number of anomalous basins and investigate the underlying causes, such as orographic precipitation underestimation from ordinary kriging, trans-basin water transfers, deep inter-catchment groundwater flow, artificial drainage, or unmodeled dam/reservoir release regimes, and provide a quality/reliability flag in the dataset for basins with severe water balance closure errors.
3) The meteorological forcing dataset was developed using 2D Ordinary Kriging of ASOS/AWS weather stations onto a 0.1° grid. In topographically rugged terrain like the Korean Peninsula, standard spatial interpolation without elevation covariates (e.g., Kriging with External Drift, regression kriging, or PRISM-type elevation lapse rates) typically underestimates ridge-line and high-elevation precipitation. The authors should discuss whether elevation lapse rates were considered and compare the gridded precipitation against satellite/gauge-corrected benchmarks or local high-elevation gauges to assess potential precipitation underestimation.
4) A network of 282 catchments across South Korea contains significant spatial overlap and nested river networks. It would be beneficial for the users if authors can provide a topological routing matrix or metadata attributes (e.g., is_headwater, upstream_basin_ids, nested_area_fraction).
5) Observed streamflow records vary substantially in length (10 to 35 years; Fig 3c) and include non-continuous gaps. The authors should clarify whether hydrologic signatures were computed over the full variable record or harmonized common benchmark window, and address how structural regime shifts (e.g., post-2010 construction from the Four Major River Project) affect static signature stability.
6) While soil characteristics from SoilGrid 2.0 are well documented, the dataset lacks attributes describing underlying bedrock geology, lithology, and hydrogeological properties (e.g., hydraulic conductivity and porosity from GLHYMPS or global lithological datasets like GLiM). Incorporating standard lithological classes and bedrock permeability attributes would bring CAMELS-KR into complete alignment with other CAMELS datasets.
7) The streamflow reconstruction section requires additional technical specifics. For the regional LSTM model, which specific static attributes and meteorological forcings were fed into the feature vector? was hyperparameter tuning performed via k-fold cross-validation or an out-of-basin spatial split? For the HBV model, which optimization algorithm and objective function formulation were applied during local calibration? what parameter search ranges were defined? The authors can also include diagnostic attribution to clearly show where and why model performance degrades.
8) Section 8 (Data Availability) states that daily forcing includes vapor pressure and sunshine duration, but these variables are not present in the dataset.Â
9) The caption for Figure 1 (lines 106–111) erroneously includes a duplicated description belonging to Figure 4.
10) The time series data include negative water level values for certain gauges, but the manuscript lacks an explanation of the underlying datum definition beyond describing it in Table 1 as "relative to the station-specific zero datum". In hydrometric networks, negative readings typically arise from channel bed degradation/scour below an established gauge zero or arbitrary local datum definitions. The authors should explicitly clarify in Section 3.1 why negative stages occur, state whether gauge zeros have shifted over time, and include the official gauge zero datum elevation (in above mean sea level or the national geodetic datum) in the static attribute table so users can convert relative stage into absolute hydraulic head.
11) In Figure 8a (number of dams and reservoirs), the color ramp spans up to ∼800, but the inset histogram shows that the overwhelming majority of catchments have counts concentrated below 0–100. As a result, almost all catchment points appear uniformly pale yellow with little to no visible distinction across moderate differences. Similar issue is there in Figure 8b. Also, in Figure 8c, the colorbar map represent population density. The unit on the colorbar is labeled as [km-2], which is missing the actual quantity (e.g., [persons km-2]).
Citation: https://doi.org/10.5194/essd-2026-544-RC1 -
AC1: 'Reply on RC1', Kuk Hyun Ahn, 30 Sep 2026
reply
Dear Dr. Nikunj K. Mangukiya,
We appreciate your thoughtful comments and constructive suggestions. They have helped us refine the manuscript, improve the dataset, and provide a clearer account of its quality.
Please find attached our detailed replies to your comments.
Kuk-Hyun Ahn and co-authors
-
AC1: 'Reply on RC1', Kuk Hyun Ahn, 30 Sep 2026
reply
-
RC2: 'Comment on essd-2026-544', Anonymous Referee #2, 28 Sep 2026
reply
This is a very welcome addition to the CAMELS family of datasets. Â As stated in the manuscript, South Korea's region and hydroclimate are both under-represented in the CAMELS offerings to date. Â The dataset is extensive and has been thoughtfully compiled, with efforts to make it consistent with other CAMELS. The manuscript itself is well written and nicely presented. I have the following comments:Â
MAJOR COMMENTS
1. Title, specifically the phrase "reconstructed streamflow". As with many similar datasets, CAMELS-KR contains modelled streamflow, and I guess this is what is meant by "reconstructed". However, a key point of any CAMELS dataset is making the observed streamflow available (since this is not usually available via other means). Using the phrase "reconstructed streamflow" gives the impression that the dataset doesn't contain observed streamflow, which is misleading (because it isn't true!).
2. Level of nestedness. Â A key reason the authors have been able to include such a large set of catchments from such a small space is that much of the dataset is composed of sets of gauges moving downstream along the same river system. All CAMELS datasets have this issue to some extent, so it is not a serious issue. However, it is a problem that it is not emphasised in the manuscript, nor the issues discussed (e.g. it means the timeseries of these catchments have a lot in common and are not independent sources of information). I suggest the authors: (i) make this clearer, including in the abstract; (ii) quantify it via attributes in the dataset - see e.g. the CAMELS-Australia or CAMELS-Peru datasets for examples; (iii) include a figure that summarises the nestedness statistics; and (iv) devote a paragraph to discuss the implications.Â
3. There is clearly a lot more water-level data than discharge data. I have two suggestions, listed below:Â
(i) Figure 3 is fantastic and very useful, but it doesn't help the reader to understand the catchment-by-catchment availability of discharge vs water levels. I suggest to add a panel (d) here based on the data in panel (c) but plotting "years of available water level data" on one axis and "years of available discharge data" on the other axis, and each scatter point is a catchment.Â
(ii) Assuming many catchments where water level data extends backwards more than discharge, surely an opportunity might be to develop a method to estimate discharge for these earlier periods, based on water level. Â I am not suggesting the authors actually do this here, but if the authors agree it is an opportunity, then perhaps discuss it in Section 7 and note any potential issues that would need to be overcome or that would increase uncertainty in the outcomes.Â4. Catchment delineation method is not described, or at least, I can't find it. Â Line 248 claims it is described in Section 2, but it doesn't seem to be? On a related topic, catchment area calculation is very important for streamflow conversions to mm/d but is not given sufficient detail. Â Specifically, a spatial projection would have been needed to express area in square kilometres, but this does not seem to be stated. Â I suggest this be rectified in the text and also specified in Table 2 "basin_area" row.
5. Data quality. Â There seems to be several indications of data quality issues, possibly associated with errors and/or human activities. Â For example, the LSTM and rainfall-runoff results exhibit quite low performance (see separate comment below) while Figure 7 seems to indicate that an unexpectedly large proportion of catchments do not meet basic water balance constraints, with streamflow seemingly exceeding precipitation or, alternatively, implied AET (P-Q) exceeding PET. Again, this is not a problem per se, since data quality is difficult to assure across large datasets. However, it needs to be discussed in more detail. Â Why do the catchments not meet water balance constraints? In particular, how much of it is due to human activities shifting water from place to place, how much of it is due to data errors (etc.)?Â
MINOR COMMENTS
A few parts of the text unnecessarily repeat lists, at the cost of brevity. Â Examples include lines 310-312, which should be combined into lines 308-309. Â Similarly, the information on catchment selection criteria is presented once as a list of definitions (starting L134), then again as a list of reasons (starting L142)--it is recommended that this would be better suited to a table (Criterion || Description || Reason || Number of catchments that meet this rule). Â On a related note, at Line 135: the provision of "n = xxx stations" here is very useful. I suggest to clarify that decrease in numbers from one rule to the other is the result of application of the new rule to the set of catchments proceeding from the previous rules. That is, the text "[stations must] provide at least 10 years of discharge data (314 stations)" should not be read as "there are 314 stations in South Korea with >10 years of data" but instead that the 314 stations out of the 850 identified in the previous step qualify.
Line 60: Please rephrase--given the context is that the hydrology has undergone much human alteration, it's incongruous to call it a "natural laboratory" (but I agree with the sentiment!)
Figure 1 caption appears to contain the Figure 4 caption too due to an error.
Figure 1 domain appears to include parts of North Korea, so please alter wording accordingly (or change the figure). Â Namely, the wording is currently "Major river basins and climate zones across South Korea"--perhaps repeat the phrase "study domain"?
Figure 3a. Â I think this panel is meant to show all the catchment boundaries and how they fit into each other, but the low quality interferes. Â Please alter the figure to ensure all boundaries are visible.Â
Figure 3a. Â Please review the document to ensure the country naming is consistent. Â Since "South Korea" is used in the title, it should be used throughout, unlike Figure 3a in which the boundary is labelled "Korea".
Table 1. For the GLEAM PET, please specify the PET formulation relevant to the GLEAM data.
Figure 4. Â Regarding titles naming the variable being mapped: given there are 11 sub-plots, it is inconvenient for the reader to be forced to refer to the caption to learn what is being mapped in each case. Â Please add titles to each sub plot indicating the variable.
Table 2. Specify what is the difference between designed storage capacity and effective storage capacity.
Line 343. Please change the order of material so that the motivation is stated prior to the specifics of the method. Â This is because the motivation often plays a role in what method is chosen. Â Please review the remainder of the manuscript for similar occurrences.
Line 356. Given how short some of the discharge observations records are, surely an additional motivation is to be able to infill the short records backwards in time?
Figure 5. The performance of the models is relatively low. There's nothing wrong with that, but can the authors provide some initial comments as to why? Future users of this dataset will get similarly low results, but the difference is that they don't have the detailed local knowledge of the dataset authors to inform insights as to why. I note that the authors have provided this guidance already but only for the coastal catchments, which are a small minority. What about the non-coastal catchments? Line 357 states potential reasons for low performance, but this is a general list provided prior to the modelling results--which of these turn out to be responsible? This discussion may be related to Fig. 7, which shows systematic issues with data (e.g. water balance issues) for some catchments. It is ok for this to be brief, as modelling results are not the subject of this paper.
Citation: https://doi.org/10.5194/essd-2026-544-RC2 -
AC2: 'Reply on RC2', Kuk Hyun Ahn, 30 Sep 2026
reply
Dear Second Referee,
We thank you for your careful review and valuable comments, which have contributed to further improvements in the manuscript and the dataset.Â
Please find attached our detailed responses to your comments.Kuk-Hyun Ahn and co-authors
-
AC2: 'Reply on RC2', Kuk Hyun Ahn, 30 Sep 2026
reply
Data sets
CAMELS-KR: Catchment attributes, meteorology, and reconstructed streamflow for large-sample hydrology in South Korea (Version version1.0) Lee et al. https://doi.org/10.5281/zenodo.21930882
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 287 | 102 | 95 | 484 | 124 | 90 |
- HTML: 287
- PDF: 102
- XML: 95
- Total: 484
- BibTeX: 124
- EndNote: 90
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,Â
Always great to see a new CAMELS dataset, congratulations on the hard work!
I noticed from figure 7 that there are quite many catchments where the runoff exceeds rainfall (n=~15) and catchments where more of the precipitation is lost before runoff measurement than PET alone explains (n=~25). It would be very beneficial to investigate where these issues likely arise, because there are many possible explanations, such as changes in storages, short timeseries, extractions for human uses, leaky catchments (subsurface flow or bifurcations removing or adding water to the catchment) or quality issues in the source products or catchment delineation. Knowing a possible explanation for the discrepancy would help increase trust in the data, or could direct future improvements if the source data is revealed faulty.
Additionally, I noticed that you did not explain the catchment delineation and quality control procedure or source data anywhere. This is crucial information, and directly affects the quality of the rest of the data.
Hopefully these will help improve the manuscript.
Best Regards,Â
Iiro Seppä