the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A continental benchmark dataset for evaluating ecosystem gradient-flux approaches across 47 NEON flux towers
Abstract. Benchmark datasets for evaluating profile-based ecosystem flux methods across environmental gradients are currently lacking. This limits method evaluation, cross-site intercomparison, and the expansion of tower-based monitoring for gases not routinely measured by the eddy covariance (EC) method. Here, we present a benchmark dataset derived from 47 terrestrial towers in the National Ecological Observatory Network (NEON), integrating co-located EC, concentration profiles, tower geometry, and canopy structural metrics. We developed a dataset to evaluate the performance of three widely used gradient flux approaches – the modified Bowen ratio (MBR), aerodynamic (AE), and wind-profile (WP) methods – against co-located EC measurements of CO₂ and H₂O. We evaluate how canopy structure, sensor height configuration, and data filtering influence agreement with EC across ecosystems, with canopy heights ranging from 0.15 to 53 m. Performance varied strongly among approaches and measurement configurations. Across all ecosystems, 11% of height-pair combinations for CO₂ and 19% for H₂O achieved moderate-to-strong concordance with EC, with the MBR approach providing the most consistent performance. Reliable estimates most often occurred when both sampling heights were above the canopy (AA) or when one level was above the canopy and the other within sparse canopies (AW). Concordance declined as canopy complexity increased, highlighting persistent challenges in tall, complex forests. Since filtering substantially reduced data availability, we developed an ensemble framework that combined reliable estimates across the three GF approaches and height pairs. Ensemble GF fluxes reproduced seasonal diel dynamics measured by EC across ecosystems (R2 = 0.58–0.99), capturing both the seasonal magnitude and direction of net ecosystem exchange. Benchmark analyses show that the GF method can provide robust ecosystem-scale flux information when deployed under favorable structural and micrometeorological conditions. The released benchmark dataset provides a resource for testing new flux algorithms, optimizing sensor placement, benchmarking profile methods, and extending multi-gas tower observations to be more inclusive of gases that are difficult to measure with fast-response EC instrumentation, including CH₄, N₂O, volatile organic compounds, and reactive pollutants.
- Preprint
(6322 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 10 Oct 2026)
- RC1: 'Comment on essd-2026-413', Anonymous Referee #1, 09 Sep 2026 reply
Data sets
Benchmark dataset for comparing eddy covariance measurements and gradient flux calculations across 47 National Ecological Observatory Network (NEON) flux towers, 2021–2024 Sparkle L. Malone et al. https://portal.edirepository.org/nis/mapbrowse?scope=edi&identifier=2369&revision=1&key=_LHRQgDvJXjj7acnumjrl-j3mEQ
Model code and software
lterwg-flux-gradient Sparkle L. Malone et al. https://github.com/lter/lterwg-flux-gradient
Interactive computing environment
lterwg-flux-gradient-eval Sparkle L. Malone et al. https://github.com/lter/lterwg-flux-gradient-eval
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 175 | 80 | 35 | 290 | 43 | 43 |
- HTML: 175
- PDF: 80
- XML: 35
- Total: 290
- BibTeX: 43
- EndNote: 43
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript uses 47 NEON sites to compare the eddy covariance and gradient flux methods for estimating CO2 and H2O fluxes. This analysis is valuable for considering the conditions under which the two methods agree and disagree and is unique in the sense that NEON infrastructure and data makes a systematic comparison possible. A better understanding of how well gradient fluxes capture ecosystem gas exchange is particularly valuable for trace gas fluxes which are more difficult or impossible to measure with eddy covariance. Overall the analysis is extremely thorough and contains multitudes of information about how seasonality, and ecosystem structure affect agreement between the two methods, as well as technically useful information such as measurement height for gradient fluxes. One conclusion that I found particularly interesting is the ensemble gradient flux approach in which multiple gradient flux methods, each with different assumptions or limitations, can be combined to produce a much more temporally complete and informative dataset. Rather than selecting a single method, the ensemble approach is intriguing and, in my opinion, a good way to approach gappy data.
The analyses and figures in this manuscript are excellent, and I am impressed with the figures that synthesise huge amounts of information and add conceptual or diagrammatic elements that help with a more intuitive understanding of the information. I did struggle with the many abbreviations and acronyms for the flux methods, statistical methods, sensor placement, particularly when a sentence had multiple acronyms. I am not sure the best way around this, but two ideas are to either drop some of the acronyms and write out the term, or include a glossary table. My recommendation is to publish this manuscript with minor revisions.
Minor and additional comments and suggestions are listed below.
Line 100: ‘With the AE, K is derived from wind and turbulence statistics’ – this was one early sentence I struggled with
Line 129-130: I don’t fully understand this. I assume that for gradient flux the gas concentrations are measured at multiple heights (between 4-8 sampling levels) while eddy covariance is above the canopy. I am therefore confused which sensors are both above the canopy (AA) and which sensor would be within the canopy for (AW). Since gradient has multiple sensors, while eddy covariance is one sensor I assume that the eddy covariance measure is always above (A), but I don’t understand what the other paired measurement is. In figure 1B which category do JORN, KONZ, GUAN, and HARV fall into? Are JORN and KONZ both AA because there is not enough canopy to be ‘within’ while GUAN and HARV are AW because some instruments are above and some within?
Section 2.4.1
In this dataset are the target gas and tracer different combinations of CO and H2O? For example to calculate CO2 flux with MBR, the H2O is the tracer, and to calculate H2O flux, CO2 is the tracer?
Figure 2: This figure is helpful! A very minor aesthetic comment, but I suggest aligning the edge of the Q3 box with the Q2 box to show that they both emerge from Q1. I presume the idea is that Q2 and Q3 are based on the outcomes of Q1?
Line 357: I don’t understand how the ability of GF to capture different ecosystem components was evaluated. Does the footprint of EC and GF differ? Why would GF capture different ecosystem components from EC?
Figure 3: It would be helpful to remind readers in the caption that CCC varies from 1 to -1 which represent strong concordance, while values close to 0 represent weak concordance.
Line 408: very minor but the font size changed in this section
Figure 6F: it’s interesting that TOOL is the only site with low canopy and CCC<0.5. Do the authors have any thoughts on why this site has lower concordance, given the low canopy?
Figure 8: the figure appears to be missing from the main text, I can only see the caption.
Lines 515-525: I think I understand the AA and AW terminology better from this section. How I understand it now is that AA and AW is not refering to EC vs GF, it’s about the paired sample height comparisons for a GF estimate. AA and AW is referring to pairs of sampling heights for gradient flux calculations where the paired gradient measurements are either both above or above and within the canopy.
Line 554-555: It’s unclear to me how GF in itself helps with decomposition into source contributions? Is this because below canopy gradient heights can represent primarily respiration/evaporation contributions, while above-canopy would include both respiration and photosynthesis/transpiration? I am not sure that the analysis in this paper particularly helps highlight that possibility with GF and found the statements about partitioning here, and above (line 357), more confusing than helpful.
Figure A11 and Figure 6 look almost like the same figure to me, though there seems to be some slight difference in the underlying data as panels C, D, E, F are slightly different.
Figure A13: some of the ranges of flux values seem very large eg: CPER with a values at -1000, but even other sites with axis values at +/- 200.
Figure A15: I was curious to see the relationship for TOOL but it’s missing entirely from this graph and A13. It does appear in the H2O fluxes so perhaps all CO2 flux measurements were lost to filtering.