the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
CFDID: a China Flood Disaster Impact Database from 1994 to 2023
Abstract. China is one of the countries most severely affected by flood disasters worldwide. Long-term and systematic flood impact data are essential for flood risk assessment, flood impact model calibration, disaster prevention and reduction. However, an open, long-term, China-focused event-scale database with detailed flood disaster impact indicators remains lacking. Here we present the China Flood Disaster Impact Database (CFDID), which contains 1716 flood disaster event records from 1994 to 2023, including 1391 top-level (L1) event records. CFDID is compiled from 66 official documents and covers five flood-related disaster categories: rainstorm flood, typhoon, snowmelt flood, ice-jam flood, and dam-break flood. Each record includes date, location, disaster category, eight impact indicators, and data source. The database was constructed using a workflow that combines optical character recognition (OCR), large language model (LLM)-based information extraction, and manual verification, thereby improving the efficiency of converting unstructured text information into structured event records while ensuring data quality. Based on these L1 records, CFDID achieves total coverage ratios of 67.22 % for deaths and missing persons, 72.25 % for affected population, 60.70 % for affected cropland area, 65.68 % for collapsed dwellings, and 74.86 % for direct economic losses relative to official annual flood disaster impact totals during 1994–2023. Compared with commonly used international flood impact databases, CFDID contains more flood disaster event records and has higher coverage for most impact indicators. CFDID also includes agricultural impact indicators, which are rarely available in international hazard impact databases. CFDID is available at https://doi.org/10.5281/zenodo.21699973 and will be continuously updated by year.
- Preprint
(1297 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 18 Sep 2026)
-
RC1: 'Comment on essd-2026-607', Anonymous Referee #1, 05 Sep 2026
reply
The comment was uploaded in the form of a supplement: https://essd.copernicus.org/preprints/essd-2026-607/essd-2026-607-RC1-supplement.pdfReplyCitation: https://doi.org/
10.5194/essd-2026-607-RC1 -
RC2: 'Comment on essd-2026-607', Anonymous Referee #2, 09 Sep 2026
reply
Paper Review
Title: CFDID: a China Flood Disaster Impact Database from 1994 to 2023
Author(s): Shibo Cui et al.
MS No.: essd-2026-607
MS type: Data description article
# General comments
CFDID makes available, in open and structured form, event-scale flood impact information that until now existed only in printed Chinese official yearbooks — the complete annual runs of three national series for 1994–2023, for a country that is both heavily flood-affected and likely under-represented in existing international impact databases, notably because of entry criteria and language biases. The resulting resource can be very useful for disaster risk research, management, and planning. To my knowledge—which is quite limited regarding the Chinese disaster loss data landscape—no alternative is openly available for China, except perhaps for urban floods. It also provides indicators that are missing from global databases, such as the agricultural impact indicators.
The construction workflow (OCR, LLM-based extraction against a fixed output schema, cross-source fusion, and manual verification) is a careful and well-documented application of components that can be considered state-of-the-art practices, rather than a methodological advance. The dataset shares some design choices with EM-DAT and Wikimpacts; Hence, the paper is primarily valued as a data contribution, as expected for a data paper, and the workflow description will still be instructive to others attempting similar transcription projects, though not sufficiently evaluated.
The deposited file is structurally sound but would benefit from a clear README to make the Zenodo repository more self-contained and less dependent on the paper. Event IDs are unique and well-formed; every L2 and L3 record resolves to a parent; no sub-event impact exceeds its parent; there are no negative values or out-of-range dates; the 31 province names are used consistently, and the constant-price column is internally consistent with the CPI adjustment described in Eq. (1). The reference list is likewise sound: every reference is cited, and every citation resolves. I congratulate the authors for the care and rigor they put into the checks. The paper reads really well, without any grammar or spelling issues identified.
While I endorse the paper for publication in ESSD and congratulate the authors on their work, I have 2 general comments that stand in the way of a direct recommendation. None of them require major changes to the dataset or its workflow; rather, they relate to information disclosure and improved comparisons. I therefore recommend a minor revision.
GC1. Lack of extraction evaluation
The manuscript provides no quantitative estimate of extraction accuracy for the DeepSeek model or inter-author agreement, while ESSD explicitly asks for error estimates and documented sources of uncertainty. I believe this could be reported easily based on the manual revisions (section 2.4.3) that were executed, to provide a benchmark for comparison with similar works. Since the workflow is offered as a methodological reference for others, its reliability is what readers need to reuse it.
GC2. Lack of detail, fairness, and reproducibility for the comparison with other data sources
Versions, endpoints, and access dates for external datasets are not disclosed. Some figures could not be reproduced from the archived file, while the comparison with international databases rests on unstated assumptions — including which disaster types were selected, which database version was used, and whether the entry criteria were met. Moreover, the comparison relies on event counts, an indicator that is neither robust nor unbiased due to differences in reporting and aggregation practices. This GC maps to several SCs reported in fragments in my specific comments, providing more specific pathways for corrections.
# Specific comments
SC1. Terminology: coverage ratios
While Eqs. 2 and 3 make it explicit, I, in the first place, was not able to understand, in the abstract and at several places in the manuscript, what was meant by coverage ratios. First, the indicator is a percentage, not a ratio; second, “coverage” can ambiguously refer to whether an event is covered in a database (i.e., its occurrence) or to spatial coverage (e.g., L35 uses coverage to mean geographic coverage). Hence, I would suggest something more explicit and accurate, like “impact reporting percentages relative to China’s official totals…”
SC2. Minor figures corrections
I found many (which is why I prefer to report it as specific comments, rather than technical corrections) small differences in several figures in sections 3.1 and 3.4 (L202-245, e.g., I count 1418 records with affected population, not 1419; 88 L1 in 2005, not 94; …). I suggest checking the counting method, making it scriptable if not yet the case, recomputing the statistics from the archived dataset, and carefully revising the numbers in the manuscript.
SC3. State the filtering details and assumptions behind the comparison with international databases
- Which EM-DAT and Wikimpacts disaster types were counted. The selection is not stated, and it materially changes the comparison: flood-only and flood-plus-tropical-cyclone give very different totals for China.
- Which version of each database was used, another important disclosure for reproducibility. Could also be part of section 5.
- Whether entry criteria were matched. EM-DAT admits an event only if it meets at least one threshold — ≥10 deaths, ≥100 affected, a declared state of emergency, or a call for international assistance. The author might consider a fair comparison, though it will not map the CFDID actual content. I would suggest at least disclosing the number of events and impact shares that would not be reported in EM-DAT due to the entry-criteria gate.
SC4. Event count indicators are misleading
Event counts are reported in several places (e.g., Figure 4, Table 3), however, when it comes to social sensing of floods (i.e., sourced from societal sources with no control over what defines an event), there are no spatiotemporal criteria to frame the definition and compare event counts consistently. This fact is acknowledged by the authors for temporal trends at L247-249, and CFDID's own L1/L2/L3 hierarchy is itself an aggregation choice.
My point is that the same caution should be acknowledged when comparing different databases (Table 3). The number of events is more an indicator of the level of detail in a database than of coverage. In this case, however, it is possible to define aggregation-invariant indicators, e.g., the percentage of events that do not overlap with EM-DAT events, or to compare flood events based on overlapping days.
SC5. The text concludes that CFDID, EM-DAT and the official totals “show broadly consistent interannual variability, indicating that both databases capture major fluctuations in flood impacts” (L307). However, disaster impact distributions are highly concentrated (Pareto-distributed). For instance, the CFDID top 10% deadliest events account for 66% of mortality in the database. Annual totals are therefore mostly driven by a handful of large events that are likely present in the other databases as well. Agreement (or correlation between Fig.7’s series) at that level is close to guaranteed, and says little about the rest of the record. This matters because the paper's own argument for CFDID's added value lies exactly where the databases should differ: at l. 34–35, EM-DAT is said to omit small-scale and localized events. The paper would be improved, in my opinion, if the authors provided statistics that explicitly support the implicitly claimed added value of CFDID (in line with SC4). Comparing the databases in the body of the distribution rather than in the total (for example, after excluding the largest events) or comparing counts and impacts below a magnitude threshold could do this.
SC6. References and potential data sources exhaustiveness
Three dataset references are mentioned in L336. The differences between CFDID and these datasets should be discussed or introduced more clearly. This is a gap in the paper context. The China Flood and Drought Bulletin is not mentioned anywhere in the manuscript and does not appear in Table 1. Fu et al. (2025), one of these three references, benchmarks its own dataset against this Bulletin throughout and describes it as offering “authoritative data on flood disasters, focusing on economic losses, casualties, and agricultural damages at the provincial level”. I don't have sufficient knowledge or language skills to evaluate whether this source should be used, and I leave that responsibility to the authors. Likewise, the paper may lack references and context on LLM-based impact extraction and refers to Li et al. (2026) as the sole example. A few more references and a bit more context would help better anchor the paper’s methodology in existing work.
SC7. Complete the data and code repository
A few things are missing to better align the data repository with FAIR principles, ESSD requirements, and make the reviewer's work easier: (i) The official annual national totals for the five core indicators are the denominator of every “coverage” figure in the paper but are not deposited, which impedes paper reproducibility. (ii) The repository has no README and no data dictionary, so Table 2's field definitions, units, and null semantics exist only in the paper. This is a restriction for machine readability. Besides, (iii) no machine-readable schema or validation script accompanies the data, so conformity with Table 2 and with the retention rule stated at L178–179 cannot be confirmed without writing one, which is especially important for a database intended to be updated annually. (iv) No requirements.txt or stated Python version. A short dependency list would make the code runnable. (v) Lastly, the ESSD paper DOI is not registered as a related identifier on the Zenodo record, with creator ORCIDs and affiliations also absent, with no corresponding author or contact point.
# Technical corrections
# # In the manuscriptCongratulations on the remarkable care in writing.
TC1. Table 2: “semi-required” is not standard terminology. Consider “conditional”
TC2. Figure 2: There is a rounding mismatch between the figures on the top of the bars in panel (a) and the figures mentioned in the text at L215
TC3. L340-343 present the conversion of known spatial information into spatial layers as future work, but Fig. 6 already maps L1 record counts onto all 31 provinces, so the provincial join has evidently been made. Consider narrowing the sentence to the levels that genuinely remain (prefectures, counties, river basins, and precise coordinates) and stating how the provincial join was performed.
## In the dataset
TC4. Record 2007-083: reports only damaged dwellings (i.e., 8), which is not among the five core indicators of L178–179. Either the rule needs rewording, or the record should be removed.
TC5. Record 2006-068: carries an end date (15 September) earlier than its start date (19 September). One of the two is presumably a typo.
TC6. Consider adding one more decimal place to the two Direct economic losses columns; for small losses, the adjustment is not apparent due to rounding (e.g., 2011-001 or 2011-028, can be checked with an equal test)Citation: https://doi.org/10.5194/essd-2026-607-RC2
Data sets
CFDID: China Flood Disaster Impact Database (1994–2023) Shibo Cui et al. https://doi.org/10.5281/zenodo.21699973
Model code and software
CFDID: China Flood Disaster Impact Database (1994–2023) Shibo Cui et al. https://doi.org/10.5281/zenodo.21699973
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 153 | 72 | 20 | 245 | 20 | 17 |
- HTML: 153
- PDF: 72
- XML: 20
- Total: 245
- BibTeX: 20
- EndNote: 17
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1