the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Agreement, opposition, and dataset influence in global evapotranspiration trends
Abstract. Evapotranspiration (ET) is a key component of the terrestrial water and energy balance, and numerous global gridded ET products are routinely used to assess historical variability and trends. However, differences in forcing data, model structure and physics in these products complicate robust ET trend analyses. Here, we present a systematic intercomparison of 14 global terrestrial ET datasets for the period 2000–2019. We introduce a topology framework that categorizes ET datasets according to their trend signatures within multi-product ensembles, providing insight into the structural role of each dataset and revealing how certain products consistently amplify or oppose dominant trends, patterns that are not evident from standard ensemble statistics. We find that products which amplify negative trends consistently oppose the dominant ensemble trend direction, whereas products that amplify positive trends tend to produce statistically significant trends where most datasets indicate weak or non-significant change. We quantify the magnitude, direction, and statistical significance of ET trends across products and evaluate their spatial consistency. The analysis reveals substantial divergence among datasets. While many products indicate predominantly positive ET trends, agreement on the magnitude and direction of change is lacking across many regions. In many regions, trends differ by more than an order of magnitude, and the spatial patterns of significant trends are highly product-dependent. The resulting harmonized trend estimates and classification provide a reference resource for evaluating current and future ET products, assessing uncertainty in trend studies, and guiding the use and improvement of ET datasets. More broadly, the topology framework can be extended beyond ET to geoscientific data product ensembles in general, enabling fitness for purpose evaluation, uncertainty assessment, and more systematic intercomparison across datasets.
- Preprint
(4776 KB) - Metadata XML
-
Supplement
(39426 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-334', Anonymous Referee #1, 15 Jul 2026
-
RC2: 'Comment on essd-2026-334', Anonymous Referee #2, 20 Jul 2026
The manuscript entitled "Agreement, opposition, and dataset influence in global evapotranspiration trends" presents a novel framework to characterise the influence of individual evapotranspiration datasets on ensemble trend estimates. The topic is timely and relevant, and the manuscript is well written. I particularly appreciate the effort to move beyond traditional ensemble statistics and identify recurring behavioural patterns among ET products. I have several comments that I believe would help in further strengthen the manuscript.
Table 1 is highly informative. I suggest adding a column at the beginning indicating the methodological category of each dataset, together with an additional column summarising the main methodological assumptions that may influence long-term trends. This would provide valuable context for readers and would also allow the bullet-point descriptions in Lines 87–94 to be removed.
Regarding dataset selection, I encourage the authors to consider including additional widely used ET products. For example, PMLv2.2 has become a commonly used global dataset. Similarly, GLEAM4.1a has been superseded by GLEAM4.3a, while the GLEAM "b" version represents the most observation-constrained product. It would also be valuable to include thermal infrared/LST-constrained surface energy balance products, as well as machine-learning products such as FLUXCOM or X-BASE, to better represent the diversity of current global ET datasets.
Section 2.5 requires further clarification. It is not entirely clear whether the proposed topology metrics are computed globally as single summary values or whether spatial information is retained during the calculations. A conceptual figure or a summary table describing each metric, their mathematical formulation, and its intended interpretation would substantially improve readability. In particular, the ranking procedure and the computation of each signature should be described more explicitly. I also found the definitions of the "signal dampener" and "opposition contributor" classes somewhat difficult to interpret from Figure 1 alone.
I recommend using the most recent version of the Köppen–Geiger climate classification (Beck et al., 2023), which provides an updated global dataset.
The computation of the DCI considers only the direction of statistically significant trends and ignores trend magnitude. Please discuss the implications of this methodological choice, particularly in cases where products agree on trend direction but differ substantially in trend magnitude.
Lines 184–185 state that "An index value of 1 or −1 indicates complete agreement on positive or negative trends, respectively." I found this wording somewhat misleading because the DCI as formulated in Eq. 2, only considers significant trends. Therefore, all datasets could exhibit positive or negative trends, but if only a subset is statistically significant, the DCI would not equal one. I suggest revising this statement accordingly. Additionally, DCI is defined using statistically significant positive and negative trends. However, Figure 2b appears to show DCI without considering statistical significance. This is confusing and should either be clarified or revised for consistency.
The authors state that Figure 3a shows complete agreement in trend direction for only four of the 44 IPCC reference regions. Although this is an interesting result, it is difficult to identify directly from the chosen colour palette. Consider using a colour scheme that better highlights complete agreement.
I suggest adding a figure summarising the upper and lower quartile uncertainty for each IPCC reference region. This would provide readers with a clearer overview of regional deviations across the ensemble.
How sensitive are the topology classes to ensemble composition? For example, would the rankings remain stable if additional datasets were included or certain products were removed? Similarly, are the topology rankings robust across different analysis periods?
The topology analysis is primarily presented at the global scale. It would be valuable to extend this analysis to all IPCC reference regions to assess whether the identified behaviours are spatially consistent. Currently, only three regions (SAM, TIB, and WCE) are discussed in detail, while the remaining regional results are relegated to the Supplementary Material. I encourage the authors to consider a summary figure in the main manuscript that synthesises the topology classes across all regions.
The Results section could be condensed without losing scientific content, as several findings are repeated across subsections.
The manuscript successfully classifies datasets into distinct behavioural categories but provides relatively little discussion of the physical or methodological reasons underlying these behaviours. For example, why do certain datasets consistently emerge as negative signal boosters or trend opposers? Discussing differences in forcing data, model structure, retrieval methodology, or parameterisations would considerably strengthen the manuscript.
In Line 436 the authors state that the results demonstrate dataset dependence "even after harmonization to a common spatial and temporal framework." However, the temporal framework was fixed to a single analysis period rather than harmonised across multiple periods. I recommend revising this statement to avoid potential ambiguity.
The manuscript frequently interprets agreement within the ensemble as an indicator of robustness. However, agreement among datasets should not be confused with accuracy, particularly given the shared forcing data, model structures, and ancestry among several products. I encourage the authors to emphasise more clearly throughout the manuscript that the proposed topology framework characterises agreement within the ensemble rather than product correctness.
Minor points:
- Extra space before the reference at Line 34.
- The statement "...ET estimation remains challenging" would benefit from additional recent references.
- Line 113: please cite the R software.
- Line 118: briefly describe how the elevation classes were adapted from Hersbach et al. (2020).
- Use en dashes consistently for numerical ranges throughout the manuscript.
- Line 184: please clarify the meaning of "pockets of negative majority trend direction."
- Use acronyms consistently after their first definition (e.g., ET, DCI, R).
- Line 307: use the spelling Köppen–Geiger consistently.
- Line 458: when discussing the contrasting conclusions of Yang et al. (2023) and Kim et al. (2021), please indicate which ET datasets and analysis periods were used, as these differences are directly relevant to the conclusions of this study.
Citation: https://doi.org/10.5194/essd-2026-334-RC2
Data sets
Agreement, opposition, and dataset influence in global evapotranspiration trends Johanna Thomson https://doi.org/10.5281/zenodo.19843461
Model code and software
https://github.com/Jorub/ithaca/tree/main/projects/trend_evap Johanna R. Thomson, Riya Dutta, Yannis Markonis, and Mijael Rodrigo Vargas https://github.com/Jorub/ithaca/tree/main/projects/trend_evap
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 102 | 29 | 8 | 139 | 71 | 14 | 16 |
- HTML: 102
- PDF: 29
- XML: 8
- Total: 139
- Supplement: 71
- BibTeX: 14
- EndNote: 16
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
A well-written and useful study, and a clear contribution that fits ESSD well. The harmonized 14-product dataset, the trend intercomparison, and the trend-signature topology framework are genuinely valuable, and the central message, that ET trend estimates are strongly dataset-dependent and that the disagreement is structured rather than random, is well supported. I have little to add: my comments are all minor and none affects the validity of the analysis.
Comments
1. The role of FLUXNET (l. 212-213, l. 452-454 and l. 472-474). These passages rest on FLUXNET influencing the products more than it does, and I think the statements should be revised.
At l. 212-213 regions lacking flux towers are said to be "likely weakly constrained by direct measurements", and at l. 452-454 the disagreement in "data-rich regions" is called striking. Both assume that flux-tower density constrains the products, but in-situ ET is not an input to any of them, so it is not clear that it does. Please clarify what "data-rich" means here, and state the mechanism by which tower density would constrain the ensemble, or revise the claim.
At l. 472-474, shared FLUXNET use is said to make products converge. Convergence would follow from shared calibration, not from shared evaluation, and the "or evaluation" makes the claim almost always true. The four products named are process-based or semi-empirical and, as far as I know, use FLUXNET mainly for evaluation, so the convergence argument does not hold for them.
2. DCI definition (Eq. 2). Eq. 2 defines the index on significant trends only (subscript s). But Fig. 2b is described as the DCI (l. 183-184) and mapped "irrespective of significance" (Fig. 2 caption), and the text says "regardless of significance" (l. 186). Please make Eq. 2 match the way the index is used; the subscript s may be a slip.
3. "Onset of directional opposition" (Fig. 2d) is not defined. The metric appears only in the figure caption and in the Results (l. 208-215), and Sect. 2 does not say how it is built. The panel is hard to read as it stands: the reader has to work out the procedure from the caption, and the meaning of the dominant "≤ 1" class is not clear. Please define the metric in Sect. 2 and give the thresholds used.
Editorial