the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Improved estimation of net ecosystem CO2 exchange over North America using LSTM-based flux upscaling (2001–2021)
Abstract. Accurate estimation of regional-scale terrestrial carbon budgets is of great importance but remains challenging. With particular advantages, the Long Short-Term Memory (LSTM) networks method shows potential in improving regional carbon budget upscaling estimations. Here, based on LSTM, we upscale regional net ecosystem carbon exchange (NEE) with available flux tower measurements and satellite land surface observations in North America. With well-established ecosystem-specific LSTMs, we produced monthly NEE at a spatial resolution of 0.1° × 0.1° over 2001–2021 (labelled as MemoryFlux). Unlike existing upscaling estimates, our dataset properly identified the Midwest Corn Belt as a region of large seasonal carbon uptake during peak growing seasons, a feature revealed by previous top-down studies and recognized as a model benchmark. Moreover, the estimated seasonal variations of NEE by MemoryFlux coincided well with those by atmospheric inversions, i.e., the ensemble mean of Orbiting Carbon Observatory-2 Model Intercomparison Project (OCO-2 v10 MIP; r = 0.96, p < 0.001) and CarbonTracker2022 (CT2022) (r = 0.97, p < 0.001). The mean annual NEE was estimated at -1.27 ± 0.12 Pg C yr-1, aligning more closely with the inversions (-0.83 to -0.70 Pg C yr-1) than existing upscaling estimates (-3.30 to -1.68 Pg C yr-1) do. In addition, our estimate plausibly captured the NEE spatial anomalies caused by all the recent extreme drought and flood events. We further confirmed that considering memory effects was critical for better indicating interannual variability and spatial anomalies of NEE induced by climate extremes. MemoryFlux provides an improved bottom-up estimation of North American NEE, largely narrowing the gap with top-down inversions. This dataset can be downloaded at https://doi.org/10.5281/zenodo.20482274 (Huang and He, 2026).
Competing interests: At least one of the (co-)authors is a member of the editorial board of Earth System Science Data.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(1855 KB) - Metadata XML
-
Supplement
(1045 KB) - BibTeX
- EndNote
Status: open (until 03 Sep 2026)
- RC1: 'Comment on essd-2026-418', Anonymous Referee #1, 08 Aug 2026 reply
-
RC2: 'Comment on essd-2026-418', Anonymous Referee #2, 08 Aug 2026
reply
The manuscript by Huang et al. presents a new monthly, 0.1° × 0.1° gridded NEE dataset for North America spanning 2001–2021, termed MemoryFlux. This dataset was developed using plant-functional-type-specific LSTM models trained with eddy-covariance observations and multiple environmental and satellite-derived predictors. The authors show that MemoryFlux reproduces the strong summertime carbon uptake over the U.S. Midwest Corn Belt, exhibits seasonal variations broadly consistent with those from OCO-2 v10 MIP and CT2022 atmospheric inversions, and estimates a continental-scale annual carbon sink that is substantially weaker than those reported by many existing machine-learning upscaling products. The manuscript further investigates the potential contribution of temporal memory by comparing MemoryFlux with a non-memory configuration and examines NEE anomalies associated with several major drought and flood events. Overall, the dataset represents a potentially valuable contribution to regional carbon-cycle research, and the manuscript is generally well organized. I support publication after minor revision. However, several methodological and interpretive issues should be clarified to further improve the robustness, reproducibility, and appropriate application of the dataset.
Major comments:
(1) The manuscript states that the site-level LSTM models were developed using environmental variables derived from EC tower observations, whereas regional upscaling was performed using ERA5-Land meteorological variables. If tower-based meteorological measurements were indeed used during model training while ERA5-Land variables were used for regional inference, this may introduce a potential predictor-domain mismatch. Differences in air temperature, precipitation, soil moisture, radiation, and VPD between tower measurements and ERA5-Land estimates could propagate into the regional predictions. Please clarify whether ERA5-Land predictors were extracted at flux-tower locations and used to train the final upscaling models, or whether tower-measured meteorological variables were used during model training. Relatedly, the description of the EC-derived target variable requires further detail. The authors should specify the exact FLUXNET/AmeriFlux predictors used, the quality-control criteria applied, the gap-filling procedures, and the minimum monthly data-coverage requirement for retaining observations. These details are important for an ESSD data paper and should be described explicitly rather than relying solely on a previous publication.
(2) The manuscript emphasizes that MemoryFlux narrows the gap between bottom-up estimates and atmospheric inversion products such as OCO-2 v10 MIP and CT2022. However, atmospheric inversions and EC-based upscaling approaches do not necessarily represent identical carbon-flux components. Therefore, the manuscript should more carefully establish the comparability between MemoryFlux and atmospheric inversion products before interpreting their agreement as evidence of improved NEE estimation. Indeed, the authors acknowledge that harvest, river transport, fires, and other carbon fluxes may contribute to discrepancies between bottom-up and top-down estimates, and that atmospheric inversions themselves are subject to uncertainties associated with atmospheric transport and prior assumptions. Although fire emissions were removed from the OCO-2 v10 MIP comparison and are reported to be excluded from CT2022, the manuscript should more explicitly define the flux components included in each product and demonstrate that the comparisons are as consistent as possible. In addition, the mean annual NEE values reported for the different products are calculated over different temporal periods, which may influence the apparent similarity among products. I recommend either recalculating key budget comparisons over common overlapping periods or providing a sensitivity analysis to demonstrate that the main conclusions are not driven by differences in temporal coverage.
(3) The manuscript reports a mean annual NEE of approximately −1.27 ± 0.12 Pg C yr-1. However, the reported ±0.12 value appears to represent the temporal standard deviation of annual estimates rather than an uncertainty estimate for MemoryFlux itself. This distinction is important because the product is generated using PFT-specific neural networks trained with a spatially uneven and, for some PFTs, limited number of EC sites. The manuscript explicitly recognizes substantial observational limitations in boreal and tropical regions and for sparsely sampled PFTs. For a data product, users need to understand not only interannual variability but also the uncertainty associated with model structure and spatial extrapolation. If possible, explicitly reporting such uncertainty information would substantially improve the usability of MemoryFlux.
Minor comments:
Line 40: The phrase “all the recent extreme drought and flood events” is overly broad. I suggest revising this expression to “caused by the six selected severe drought and flood events during the study period.
Please explicitly define the NEE sign convention at an appropriate point in the manuscript. Although the manuscript implies that negative NEE values indicate carbon uptake, this convention should be clearly stated for users of the dataset.
Line 133: The manuscript states that 84 flux sites were used, whereas the Code and data availability section indicates that 35 FLUXNET sites and 47 AmeriFlux sites were included, resulting in a total of 82 sites. This inconsistency is confusing and should be carefully checked and corrected.
Table 1 should provide the units for all predictors, and the specific ERA5-Land soil moisture layer used in this analysis should be clearly specified.
Lines 189-192: I agree with the approach used to generate the static PFT map. However, the potential impacts of neglecting land-cover changes should at least be briefly discussed, particularly with respect to croplands, forest disturbances, and fire-affected regions.
Line 192: The phrase “major interpolation method” used to describe the land-cover resampling procedure is unclear. I assume that the authors refer to a majority aggregation method. Please specify the exact resampling algorithm used.
Line 297: The seasonal correlation coefficients (r = 0.96 and 0.97) are high; however, correlations calculated from a 12-month climatological cycle can be strongly influenced by the shared seasonal pattern. I therefore suggest complementing these correlation coefficients with one or two magnitude-based metrics, such as RMSE, mean bias, or differences in seasonal amplitude.
Lines 334-338: The IAV correlations with CT2022 (r = 0.41, p = 0.07) and OCO-2 v10 MIP (r = 0.70, p = 0.12) are not statistically significant at the p < 0.05 level. Therefore, the statement that MemoryFlux “effectively captured” IAVs should be moderated, particularly for these two comparisons.
Citation: https://doi.org/10.5194/essd-2026-418-RC2
Data sets
Improved estimation of net ecosystem CO2 exchange over North America using LSTM-based flux upscaling (2001–2021) Chengcheng Huang and Wei He https://doi.org/10.5281/zenodo.20482274
Model code and software
Improved estimation of net ecosystem CO2 exchange over North America using LSTM-based flux upscaling (2001–2021) Chengcheng Huang and Wei He https://doi.org/10.5281/zenodo.20482274
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 138 | 41 | 9 | 188 | 27 | 16 | 14 |
- HTML: 138
- PDF: 41
- XML: 9
- Total: 188
- Supplement: 27
- BibTeX: 16
- EndNote: 14
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This study introduces MemoryFlux, an LSTM-based monthly NEE dataset for North America at 0.1° spatial resolution covering the period from 2001 to 2021. By incorporating multiple vegetation and environmental variables with historical information from a six-month look-back window, the authors aim to improve the representation of ecosystem temporal memory in regional carbon flux upscaling. The resulting product reproduces several well-known spatial and seasonal features, including strong growing season carbon uptake in the U.S. Midwest, and shows greater agreement with selected atmospheric inversion products than several existing empirical upscaling products in terms of seasonal amplitude and continental carbon sink magnitude. The comparison between MemoryFlux and nonMemoryFlux, together with the analyses of major drought and flood events, provides an interesting demonstration of the potential importance of temporal information in data-driven carbon cycle modelling. I consider the dataset suitable for publication in ESSD after minor revision; however, several aspects related tothe attribution of improvements to “memory effects” and the benchmarking strategy should be strengthened and interpreted more cautiously.
The details are as follows:
The central scientific claim of this study is that incorporating memory effects improves the representation of NEE IAVs and climate-extreme-related anomalies. The manuscript states that MemoryFluxincorporates information from previous time steps, whereas nonMemoryFlux relies only on synchronous explanatory variables. However, the definition and implementation of nonMemoryFlux are not described in sufficient detail. Is the temporal look-back structure the only difference between the two model configurations? In addition, variables such as NDVI, LAI, FAPAR, and SIF inherently contain temporal persistence and may themselves reflect legacy effects. Therefore, the implication that one model incorporates ecosystem memory while the other contains no memory may be overly simplistic or potentially misleading. Please clarify this issue and more clearly demonstrate that the differences between MemoryFlux and nonMemoryFluxcan be specifically attributed to the representation of temporal memory effects. It would also be valuable to report quantitative differences in model performance between MemoryFlux and nonMemoryFlux at the site or PFT level, rather than relying primarily on representative time-series examples shown in Fig. S6.
A major narrative throughout the paper is that MemoryFlux “narrows the gap” between bottom-up estimates and atmospheric inversions products. While this result is interesting, numerical agreement with an inversion product should not automatically be interpreted as evidence of higher accuracy. EC-based upscaling does not explicitly account for harvest, some disturbance emissions, and other processes that may be included in top-down net land–atmosphere carbon flux estimates.Therefore, the difference between the approximately −1.27 Pg C yr-1 estimated by MemoryFlux and the approximately −0.7 to −0.8 Pg C yr-1 estimated by atmospheric inversion products should not be interpreted solely as model error; part of this discrepancy may arise from differences in flux definitions and system boundaries. I recommend that the authors precisely define the flux components extracted from CT2022 and OCO-2 MIP and describe in sufficient detail how fire emissions are removed. Furthermore, statements such as “more accurate” or “overestimate” should be replaced with more appropriate terms, such as “closer to,” “larger uptake than,” or “more consistent with,”.
The evaluation of extreme event performance is largely qualitative and should be supported by quantitative metrics. Figure 7 is visually compelling; however, the conclusion that MemoryFlux provides “more accurate constraints” on drought- and flood-related anomalies appears to rely primarily on visual agreement among spatial patterns and predefined event regions. Given the strength of this conclusion, additional quantitative evaluation would be preferable. Such analyses would transform the extreme-event section from a predominantly qualitative demonstration into a more robust assessment of the product’s performance. In addition, anomaly baselines are calculated using the individual temporal coverage of each product. Because the products have different temporal periods, this choice may influence anomaly magnitudes and complicate inter-product comparisons. A common reference period should be adopted for products with overlapping temporal coverage.
Specific comments
L136: “terrestrial vegetation biomes”? These classes are actually IGBP land-cover/PFT categories. Please consider using more precise terminology throughout the manuscript.
L268-270: The manuscript repeatedly emphasizes that MemoryFlux captures the strongest carbon uptake in the Midwest Corn Belt during the growing season. However, on an annual basis, MemoryFlux identifies the strongestcarbon sink in the southeastern U.S., whereas the atmospheric inversion products still indicate stronger uptake in the Midwest. Please explain more clearly how the successful reproduction of the peak growing-season Corn Belt signal can be reconciled with the disagreement in annual spatial patterns relative to the inversions.
L499: Several PFT-specific models are based on a very limited number of sites. Table S1 includes only two sites for SAV, with similarly sparse representation for MF, WSA, and CSH. Although the manuscript already discusses SAV as an example, I think it is necessary to report either the number of sites or the number of monthly NEE observations for each PFT, as this information is critical for assessing the reliability of the regional upscaling results.
L250-252: Please clarify whether the gridded SAV model uses a different set of predictors from the other PFT-specific models and, if so, whether this difference affects the comparability of model performance among PFTs.
Please avoid interpreting standard deviations across years as uncertainty intervals. In Fig. 4 and Tables S2–S3, please clearly specify whether the shading and error bars represent interannual variability, model ensemble spread, or another source of variation.
L549-550: The increasing carbon sink trend is not statistically significant and should therefore not be described as an established long-term increase.
The authors acknowledge in the Discussion that atmospheric inversions are not direct measurements of NEE. In this context, statements such as “overestimate JJA carbon uptake” and “overestimate carbon uptake in recent years” may be overly strong, as they implicitly treat the reference estimates as ground truth. Please revise similar wording throughout the manuscript.Furthermore, the phrase “monitoring climate extremes” may be somewhat too strong for these model-generated NEE products. Terms such as “representing” or “capturing” may be more appropriate.
The manuscript refers to the “2020–2021 southwestern North American drought,” whereas the event analysis presented in the figures appears to focus on a specific summer period. Please clarify which summer or summers were used to calculate the anomaly and provide the rationale for this choice.