Articles | Volume 18, issue 10
https://doi.org/10.5194/essd-18-7367-2026
© Author(s) 2026. This work is distributed under the Creative Commons Attribution 4.0 License.
Improved estimation of net ecosystem CO2 exchange over North America using LSTM-based flux upscaling (2001–2021)
Download
- Final revised paper (published on 07 Oct 2026)
- Supplement to the final revised paper
- Preprint (discussion started on 24 Jul 2026)
- Supplement to the preprint
Interactive discussion
Status: closed
Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor
| : Report abuse
-
RC1: 'Comment on essd-2026-418', Anonymous Referee #1, 08 Aug 2026
- AC1: 'Reply on RC1', Wei He, 31 Aug 2026
-
RC2: 'Comment on essd-2026-418', Anonymous Referee #2, 08 Aug 2026
- AC2: 'Reply on RC2', Wei He, 31 Aug 2026
Peer review completion
AR – Author's response | RR – Referee report | ED – Editor decision | EF – Editorial file upload
AR by Wei He on behalf of the Authors (05 Sep 2026)
Author's response
Author's tracked changes
Manuscript
ED: Referee Nomination & Report Request started (19 Sep 2026) by Yuqiang Zhang
RR by Anonymous Referee #1 (23 Sep 2026)
ED: Publish as is (24 Sep 2026) by Yuqiang Zhang
AR by Wei He on behalf of the Authors (24 Sep 2026)
Manuscript
This study introduces MemoryFlux, an LSTM-based monthly NEE dataset for North America at 0.1° spatial resolution covering the period from 2001 to 2021. By incorporating multiple vegetation and environmental variables with historical information from a six-month look-back window, the authors aim to improve the representation of ecosystem temporal memory in regional carbon flux upscaling. The resulting product reproduces several well-known spatial and seasonal features, including strong growing season carbon uptake in the U.S. Midwest, and shows greater agreement with selected atmospheric inversion products than several existing empirical upscaling products in terms of seasonal amplitude and continental carbon sink magnitude. The comparison between MemoryFlux and nonMemoryFlux, together with the analyses of major drought and flood events, provides an interesting demonstration of the potential importance of temporal information in data-driven carbon cycle modelling. I consider the dataset suitable for publication in ESSD after minor revision; however, several aspects related tothe attribution of improvements to “memory effects” and the benchmarking strategy should be strengthened and interpreted more cautiously.
The details are as follows:
The central scientific claim of this study is that incorporating memory effects improves the representation of NEE IAVs and climate-extreme-related anomalies. The manuscript states that MemoryFluxincorporates information from previous time steps, whereas nonMemoryFlux relies only on synchronous explanatory variables. However, the definition and implementation of nonMemoryFlux are not described in sufficient detail. Is the temporal look-back structure the only difference between the two model configurations? In addition, variables such as NDVI, LAI, FAPAR, and SIF inherently contain temporal persistence and may themselves reflect legacy effects. Therefore, the implication that one model incorporates ecosystem memory while the other contains no memory may be overly simplistic or potentially misleading. Please clarify this issue and more clearly demonstrate that the differences between MemoryFlux and nonMemoryFluxcan be specifically attributed to the representation of temporal memory effects. It would also be valuable to report quantitative differences in model performance between MemoryFlux and nonMemoryFlux at the site or PFT level, rather than relying primarily on representative time-series examples shown in Fig. S6.
A major narrative throughout the paper is that MemoryFlux “narrows the gap” between bottom-up estimates and atmospheric inversions products. While this result is interesting, numerical agreement with an inversion product should not automatically be interpreted as evidence of higher accuracy. EC-based upscaling does not explicitly account for harvest, some disturbance emissions, and other processes that may be included in top-down net land–atmosphere carbon flux estimates.Therefore, the difference between the approximately −1.27 Pg C yr-1 estimated by MemoryFlux and the approximately −0.7 to −0.8 Pg C yr-1 estimated by atmospheric inversion products should not be interpreted solely as model error; part of this discrepancy may arise from differences in flux definitions and system boundaries. I recommend that the authors precisely define the flux components extracted from CT2022 and OCO-2 MIP and describe in sufficient detail how fire emissions are removed. Furthermore, statements such as “more accurate” or “overestimate” should be replaced with more appropriate terms, such as “closer to,” “larger uptake than,” or “more consistent with,”.
The evaluation of extreme event performance is largely qualitative and should be supported by quantitative metrics. Figure 7 is visually compelling; however, the conclusion that MemoryFlux provides “more accurate constraints” on drought- and flood-related anomalies appears to rely primarily on visual agreement among spatial patterns and predefined event regions. Given the strength of this conclusion, additional quantitative evaluation would be preferable. Such analyses would transform the extreme-event section from a predominantly qualitative demonstration into a more robust assessment of the product’s performance. In addition, anomaly baselines are calculated using the individual temporal coverage of each product. Because the products have different temporal periods, this choice may influence anomaly magnitudes and complicate inter-product comparisons. A common reference period should be adopted for products with overlapping temporal coverage.
Specific comments
L136: “terrestrial vegetation biomes”? These classes are actually IGBP land-cover/PFT categories. Please consider using more precise terminology throughout the manuscript.
L268-270: The manuscript repeatedly emphasizes that MemoryFlux captures the strongest carbon uptake in the Midwest Corn Belt during the growing season. However, on an annual basis, MemoryFlux identifies the strongestcarbon sink in the southeastern U.S., whereas the atmospheric inversion products still indicate stronger uptake in the Midwest. Please explain more clearly how the successful reproduction of the peak growing-season Corn Belt signal can be reconciled with the disagreement in annual spatial patterns relative to the inversions.
L499: Several PFT-specific models are based on a very limited number of sites. Table S1 includes only two sites for SAV, with similarly sparse representation for MF, WSA, and CSH. Although the manuscript already discusses SAV as an example, I think it is necessary to report either the number of sites or the number of monthly NEE observations for each PFT, as this information is critical for assessing the reliability of the regional upscaling results.
L250-252: Please clarify whether the gridded SAV model uses a different set of predictors from the other PFT-specific models and, if so, whether this difference affects the comparability of model performance among PFTs.
Please avoid interpreting standard deviations across years as uncertainty intervals. In Fig. 4 and Tables S2–S3, please clearly specify whether the shading and error bars represent interannual variability, model ensemble spread, or another source of variation.
L549-550: The increasing carbon sink trend is not statistically significant and should therefore not be described as an established long-term increase.
The authors acknowledge in the Discussion that atmospheric inversions are not direct measurements of NEE. In this context, statements such as “overestimate JJA carbon uptake” and “overestimate carbon uptake in recent years” may be overly strong, as they implicitly treat the reference estimates as ground truth. Please revise similar wording throughout the manuscript.Furthermore, the phrase “monitoring climate extremes” may be somewhat too strong for these model-generated NEE products. Terms such as “representing” or “capturing” may be more appropriate.
The manuscript refers to the “2020–2021 southwestern North American drought,” whereas the event analysis presented in the figures appears to focus on a specific summer period. Please clarify which summer or summers were used to calculate the anomaly and provide the rationale for this choice.