the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A 1-km dataset of crop residue production and usage pathways in the conterminous U.S. from 2001 to 2021
Abstract. Crop residues represent an important biomass resource that supports soil organic carbon maintenance, livestock production, and emerging bioeconomy sectors. In the United States, crop residue production is concentrated in a few dominant cropping systems, whereas residue demand for livestock and other off-field uses is often geographically separated from production regions. However, spatially explicit datasets that jointly quantify crop residue production, allocation pathways, and spatial imbalances between production and consumption remain limited. Here, we developed a 1 km × 1 km gridded dataset of crop residue production and usage pathways across the conterminous United States for 2001–2021, covering nine major crops. A mass-balance framework was applied to reconcile residue production and consumption, allocating residues into four pathways: left on field, animal-use, off-field use, and burnt. An implied domestic transfer layer was also derived as an indicator of spatial mismatches between residue production and consumption. Results indicate that total U.S. residue production averaged 4.91×1011 kg/yr, with corn residue consistently contributing over 60 % of the total biomass. Residues left on field dominated nationally, accounting for 86.4 % of total residue production. Livestock use and off-field uses represented smaller but spatially heterogeneous pathways (13.4 % combined), while burnt residue accounted for less than 0.2 %. Residue production was concentrated in the Midwest, whereas higher consumption demand occurred in the Southeast, West Coast, and the Southern Great Plains. The national production-consumption mismatch ratio increased from 7.6 % to 8.4 % over the study period, highlighting a growing spatial imbalance between residue availability and consumption demand. By providing 1-km gridded, mass-balanced estimates of residue production, allocation pathways, and regional production-consumption mismatches, this dataset offers a spatially explicit foundation for quantifying crop residue flows across U.S. agricultural landscapes and supports improved representation of residue management in terrestrial biosphere models, soil carbon dynamics assessments, and sustainable residue biomass utilization strategies. The dataset is available at https://doi.org/10.5281/zenodo.18453064 (Zhang et al., 2026).
- Preprint
(2807 KB) - Metadata XML
-
Supplement
(4107 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on essd-2026-237', Andrew Smerald, 11 Jun 2026
-
AC1: 'Reply on RC1', Yongfa You, 11 Sep 2026
Comment [1]: Does national level data allow for crop-specific residue management to be determined? For example soybean and maize cultivation produce residues in very different amounts and with different nutritional qualities, presumably leading to different management choices. Knowing crop-specific management would, for example, be very useful for driving biogeochemical models. If the available data doesn't allow for separating crop-specific management, it may be worth adding a sentence or two to the discussion about why this is not currently possible and what data would be required.
Response: Thank you for raising this important point. In our dataset, residue production is estimated separately for each of the nine crops. However, the available national datasets used to estimate livestock feeding/bedding and other off-field pathways do not provide information on the crop origin of the residues consumed or removed. Therefore, these pathways cannot be partitioned at the crop level with the available data. We have added this limitation to the Discussion and clarified that crop-specific pathway estimates would require spatially explicit information linking residue removal and use to individual crops. (Please see Section 4.3, Lines 384-388)
Comment [2]: Does the USA have data about where maize is grown for silage as opposed to for grain? When maize is harvested for silage the majority of the above ground plant is normally removed. Maybe this could help further improve the dataset or give an additional way to validate it (especially what is fed to animals).
Response: Thank you for this suggestion. USDA-NASS provides separate statistics for corn harvested for grain and for silage. In our dataset, corn residue production is estimated from grain-harvest statistics only because silage harvest removes most aboveground biomass for livestock feed, leaving little residue available for the usage pathways represented in this dataset. We have clarified this point in the Methods. (Please see Section 2.2.2, Lines 117-119)
Comment [3]: In my understanding the redistribution of crop residues between surplus and deficit grid cells does not take the distance between the grid cells into account. Since residue transport is expensive, I assume there is a strong incentive to keep the transport distances as small as possible. I don't see this as a very critical point, since only quite a small amount of residue (7-8%) is moved between grid cells. However, it may be worth adding a sentence or two about this to the discussion/limitations section.
Response: Thank you for pointing this out. We agree that actual residue redistribution would be strongly constrained by transport distance and cost. In our framework, the required domestic inflow is allocated proportionally across surplus grid cells to satisfy the national mass balance. The resulting transfer layer therefore represents the amount of spatial reallocation required to maintain mass balance rather than actual transportation routes or source-destination flows. We have clarified this assumption in the Methods and added this limitation to the Discussion. (Please see Section 2.2.4, Lines 208-209; and Section 4.3, Lines 381-382)
Comment [4]: Is there any way to quantify the uncertainty for residue production?
Response: Thank you for this comment. We have now quantified the sensitivity of residue-production estimates to harvest index (HI), which is a key parameter used to convert grain yield to residue biomass. Specifically, we recalculated residue production using literature-derived constant HI values representing alternative plausible assumptions and compared these estimates with those obtained using the crop-specific, time-dependent HI adopted in our current analysis. The resulting national residue-production estimates were generally comparable across HI assumptions, although regional and event-specific uncertainty in HI remains. The sensitivity analysis has been added as Table S3 and Figure S10 and is discussed in Section 4.3 (Please see Lines 401-404).
Comment [5]: I assume all values for biomass weights are given in dry matter as opposed to fresh matter. It is probably worth stating this somewhere to avoid possible confusion. If you needed to convert e.g. grain yields from fresh to dry matter using standard crop-specific conversion factors it is probably worth documenting these values in the supplementary material.
Response: Thank you for pointing this out. All biomass quantities in the dataset are reported on a dry-matter basis. The USDA-NASS grain-yield data used in our calculations were already provided as dry-matter yields, so no additional moisture conversion was needed. We have clarified this point in the revised Methods to avoid confusion. (Please see Section 2.2.2, Line 117)
Citation: https://doi.org/10.5194/essd-2026-237-AC1
-
AC1: 'Reply on RC1', Yongfa You, 11 Sep 2026
-
RC2: 'Comment on essd-2026-237', Anonymous Referee #2, 02 Aug 2026
This manuscript presents a timely, high-resolution dataset on crop residue production and usage pathways across the conterminous U.S. Generating a spatially explicit, mass-balanced framework at this scale is a formidable and commendable effort. This dataset will undoubtedly serve as a vital resource for agricultural landscape planning and terrestrial biosphere modeling.
The methodology is generally robust, and the downscaling approach is well-conceived. Nevertheless, to further strengthen the manuscript, I have outlined a few minor suggestions regarding the underlying assumptions, error propagation, and validation strategies. My overall recommendation is that this manuscript is highly suitable for publication after minor revisions.
1. In the mass-balance framework, the left on field amount is calculated as the final residual. While mathematically necessary, this means any uncertainties or systemic biases in calculating total production or consumption pathways will accumulate entirely within the final residual estimates. It would be favored to discuss this error propagation explicitly. If possible, a brief comparison with independent survey or ground-truth data could greatly enhance confidence. That said, it is sometimes difficult to do so, as the USDA survey data often suffer from spatial and temporal mismatches. If a pixel-by-pixel ground-truthing is unfeasible due to these data constraints, the authors are encouraged at least include a dedicated qualitative or semi-quantitative discussion addressing how this scale mismatch and residual error accumulation might affect misestimation.
2. In Section 4.1, the author presents the spatial consistency check with Smerald et al. (2023) as a useful benchmark. Yet, strictly, it is not independent validation. Relying on this comparison without qualification risks misleading readers due to homologous bias. As I found both datasets likely share homologous Primary Data Sources (e.g., the GFED4s product used by Smerald et al. for on-field burning is built upon the same MODIS MCD64A1 dataset used in this study). Not all primary data sources have been checked, but I may like to suggest checking and acknowledging this distinction in the discussion section to avoid overstating the independence of this benchmark.3. Build on the above one: the figure 6 presents side-by-side visual comparisons against the benchmark dataset by Smerald et al. (2023). It can be subjective and often masks localized spatial discrepancies or variance dampening resulting from resolution mismatches. I may like to suggest merging a1 and a2 into a combined panel a by calculating and presenting quantitative spatial metrics (e.g., pixel-by-pixel RMSE or a spatial correlation coefficient) rather than relying solely on visual comparisons, which should apply to all subplots in Figure 4.
4. The use of harvest index may ignore the bias during climate extremes such as the 2012 north american drought and the 2019 flood event in the Midwest. Using a trend-based harvest index during such years might misestimate residue production. If a correction is not possible for this dataset, a brief caveat in the discussion regarding the dataset's performance during extreme weather years would be favored for modelers analyzing temporal dynamics.
Citation: https://doi.org/10.5194/essd-2026-237-RC2 -
AC2: 'Reply on RC2', Yongfa You, 11 Sep 2026
Comment [1]: In the mass-balance framework, the left on field amount is calculated as the final residual. While mathematically necessary, this means any uncertainties or systemic biases in calculating total production or consumption pathways will accumulate entirely within the final residual estimates. It would be favored to discuss this error propagation explicitly. If possible, a brief comparison with independent survey or ground-truth data could greatly enhance confidence. That said, it is sometimes difficult to do so, as the USDA survey data often suffer from spatial and temporal mismatches. If a pixel-by-pixel ground-truthing is unfeasible due to these data constraints, the authors are encouraged at least include a dedicated qualitative or semi-quantitative discussion addressing how this scale mismatch and residual error accumulation might affect misestimation.
Response: We thank the reviewer for this thoughtful comment. We agree that, because residue left on field is calculated as the residual, uncertainty in total residue production and the removal pathways directly propagate into the left-on-field (L) estimate. Specifically, given , overestimation of total residue production (P) increases the residual , whereas overestimation of burning (B), animal use (A), or off-field removal (O) decreases . We have added a detailed discussion of this error propagation in the revised manuscript (Please see Section 4.3, Lines 363-367).
To further characterize this uncertainty, we added a scenario-based sensitivity analysis to examine how livestock feed demand assumptions affect the left-on-field estimate, as livestock demand is a major source of uncertainty due to the wide range of published intake coefficients. Across the S0–SH scenarios, the fraction of residues allocated to animal and off-field uses ranged from 5.8% to 27.0%, while residue left on field ranged from 72.9% to 94.1% (Table S1). Figure S12 further shows that the sensitivity of the left-on-field fraction varies across years and regions, with the difference between the highest and lowest national estimates averaging 21.2 percentage points over 2001-2021 and larger pixel-level differences occurring in livestock-intensive regions such as Texas. (Table S1, Figure S12, Lines 367-372).
We also examined several potential external datasets, including USDA tillage statistics and aggregated state-level estimates of corn residue use by livestock (Schmer et al., 2017). However, differences in variable definitions, crop coverage, and spatial resolution prevent a direct comparison with our 1-km pathway estimates. We therefore did not use these datasets for direct validation.
Reference:
Schmer, M.R., Brown, R.M., Jin, V.L., Mitchell, R.B. and Redfearn, D.D., 2017. Corn residue use by livestock in the United States. Agricultural & Environmental Letters, 2(1), p.160043.
Comment [2]: In Section 4.1, the author presents the spatial consistency check with Smerald et al. (2023) as a useful benchmark. Yet, strictly, it is not independent validation. Relying on this comparison without qualification risks misleading readers due to homologous bias. As I found both datasets likely share homologous Primary Data Sources (e.g., the GFED4s product used by Smerald et al. for on-field burning is built upon the same MODIS MCD64A1 dataset used in this study). Not all primary data sources have been checked, but I may like to suggest checking and acknowledging this distinction in the discussion section to avoid overstating the independence of this benchmark.
Response: We agree with the reviewer that the comparison with Smerald et al. (2023) should not be interpreted as an independent validation. The two datasets share some underlying data sources and methodological assumptions; such overlap may introduce correlated biases and reduces the independence of the comparison. We have therefore revised the manuscript to present this analysis as a spatial consistency assessment rather than an independent validation, and we have updated the corresponding section title and discussion accordingly. (Please see Section 4.1, Lines 310-311)
Comment [3]: Build on the above one: the figure 6 presents side-by-side visual comparisons against the benchmark dataset by Smerald et al. (2023). It can be subjective and often masks localized spatial discrepancies or variance dampening resulting from resolution mismatches. I may like to suggest merging a1 and a2 into a combined panel a by calculating and presenting quantitative spatial metrics (e.g., pixel-by-pixel RMSE or a spatial correlation coefficient) rather than relying solely on visual comparisons, which should apply to all subplots in Figure 4.
Response: We agree that visual comparison alone may obscure localized spatial differences. Following the reviewer’s suggestion, we aggregated our 1-km estimates to the native 0.5° grid of Smerald et al. (2023), restricted the analysis to common valid grid cells within the conterminous U.S., and performed a pixel-level quantitative comparison for each residue pathway (Figure S11). The analysis provides a quantitative measure of where and by how much the two datasets differ spatially. Based on these results, we revised our previous statement of ‘high geospatial consistency’ and now describe the two products as showing similar overall spatial patterns with regional differences for some pathways. We also clarify that this comparison is intended as a consistency assessment rather than an independent validation, given differences in crop coverage, spatial resolution, and pathway-allocation assumptions. (Please see Section 4.1, Lines 311-312)
Comment [4]: The use of harvest index may ignore the bias during climate extremes such as the 2012 north american drought and the 2019 flood event in the Midwest. Using a trend-based harvest index during such years might misestimate residue production. If a correction is not possible for this dataset, a brief caveat in the discussion regarding the dataset's performance during extreme weather years would be favored for modelers analyzing temporal dynamics.
Response: We thank the reviewer for raising this point. We agree that harvest index (HI) can deviate from its prescribed temporal trend during extreme climate events because environmental stress can alter biomass partitioning between grain and vegetative tissues. Annual USDA-NASS grain yield records capture year-to-year variation associated with climate variability and extremes such as the 2012 drought and the 2019 Midwest flooding. However, the prescribed HI trend used in our framework does not explicitly represent event-specific changes in biomass partitioning. We have therefore added this limitation to the Discussion and noted that residue production estimates may be more uncertain during extreme weather years. We also report a sensitivity analysis using alternative literature-derived HI values, which produced national residue-production estimates comparable to those from the prescribed time-dependent HI approach (Table S3; Figure S10). Relevant literature has also been added to support this discussion (Gebre et al., 2022; Ludemann et al., 2022). (Please see Section 4.3, Lines 396-402)
References:
Gebre, Michael Gebretsadik, Istvan Rajcan, and Hugh James Earl. "Genetic variation for effects of drought stress on yield formation traits among commercial soybean [Glycine max (L.) Merr.] cultivars adapted to Ontario, Canada." Frontiers in Plant Science 13 (2022): 1020944.
Ludemann, Cameron I., et al. "Estimating maize harvest index and nitrogen concentrations in grain and residue using globally available data." Field Crops Research 284 (2022): 108578.
Citation: https://doi.org/10.5194/essd-2026-237-AC2 -
AC3: 'Reply on AC2', Yongfa You, 11 Sep 2026
A quick clarification: the equation in our previous response to AC2-Comment [1] did not display correctly in the system. The rest of the response was displayed as intended. The corrected response to Comment [1] is provided below. We apologize for the formatting issue.
Comment [1]: In the mass-balance framework, the left on field amount is calculated as the final residual. While mathematically necessary, this means any uncertainties or systemic biases in calculating total production or consumption pathways will accumulate entirely within the final residual estimates. It would be favored to discuss this error propagation explicitly. If possible, a brief comparison with independent survey or ground-truth data could greatly enhance confidence. That said, it is sometimes difficult to do so, as the USDA survey data often suffer from spatial and temporal mismatches. If a pixel-by-pixel ground-truthing is unfeasible due to these data constraints, the authors are encouraged at least include a dedicated qualitative or semi-quantitative discussion addressing how this scale mismatch and residual error accumulation might affect misestimation.
Response: We thank the reviewer for this thoughtful comment. We agree that, because residue left on field is calculated as the residual, uncertainty in total residue production and the removal pathways directly propagate into the left-on-field (L) estimate. Specifically, given L=P-(B+A+O), overestimation of total residue production (P) increases the residual L, whereas overestimation of burning (B), animal use (A), or off-field removal (O) decreases L. We have added a detailed discussion of this error propagation in the revised manuscript (Please see Section 4.3, Lines 363-367).
To further characterize this uncertainty, we added a scenario-based sensitivity analysis to examine how livestock feed demand assumptions affect the left-on-field estimate, as livestock demand is a major source of uncertainty due to the wide range of published intake coefficients. Across the S0–SH scenarios, the fraction of residues allocated to animal and off-field uses ranged from 5.8% to 27.0%, while residue left on field ranged from 72.9% to 94.1% (Table S1). Figure S12 further shows that the sensitivity of the left-on-field fraction varies across years and regions, with the difference between the highest and lowest national estimates averaging 21.2 percentage points over 2001-2021 and larger pixel-level differences occurring in livestock-intensive regions such as Texas. (Table S1, Figure S12, Lines 367-372).
We also examined several potential external datasets, including USDA tillage statistics and aggregated state-level estimates of corn residue use by livestock (Schmer et al., 2017). However, differences in variable definitions, crop coverage, and spatial resolution prevent a direct comparison with our 1-km pathway estimates. We therefore did not use these datasets for direct validation.
Reference:
Schmer, M.R., Brown, R.M., Jin, V.L., Mitchell, R.B. and Redfearn, D.D., 2017. Corn residue use by livestock in the United States. Agricultural & Environmental Letters, 2(1), p.160043.
Citation: https://doi.org/10.5194/essd-2026-237-AC3
-
AC3: 'Reply on AC2', Yongfa You, 11 Sep 2026
-
AC2: 'Reply on RC2', Yongfa You, 11 Sep 2026
Data sets
A 1-km dataset of crop residue production and usage pathways in the conterminous U.S. from 2001 to 2021 Yikun Zhang, Hua Yan, Wenzhe Jiao, Yanghui Kang, Marty R. Schmer, Ryan D. Stewart, Benjamin F. Tracy, and Yongfa You https://doi.org/10.5281/zenodo.18453064
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 542 | 87 | 93 | 722 | 53 | 100 | 81 |
- HTML: 542
- PDF: 87
- XML: 93
- Total: 722
- Supplement: 53
- BibTeX: 100
- EndNote: 81
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript reports a dataset for the production and usage of crop residues in the USA. There is a clear need for high quality national level datasets on residue production and usage, for example as a key input to model changes in soil organic carbon stocks. The methodology is quite similar to that used to produce a similar dataset on a global scale by Smerald et al. However, the use of more detailed, national-level input data and assumptions and a considerably finer spatial resolution means that, for the USA, this dataset is almost certainly superior to the global dataset.
In my opinion the methodology is reasonable and clearly described. I have a few relatively minor points that the authors may like to include, but overall my recommendation is to publish this manuscript.
1. Does national level data allow for crop-specific residue management to be determined? For example soybean and maize cultivation produce residues in very different amounts and with different nutritional qualities, presumably leading to different management choices. Knowing crop-specific management would, for example, be very useful for driving biogeochemical models. If the available data doesn't allow for separating crop-specific management, it may be worth adding a sentence or two to the discussion about why this is not currently possible and what data would be required.
2. Does the USA have data about where maize is grown for silage as opposed to for grain? When maize is harvested for silage the majority of the above ground plant is normally removed. Maybe this could help further improve the dataset or give an additional way to validate it (especially what is fed to animals).
3. In my understanding the redistribution of crop residues between surplus and deficit grid cells does not take the distance between the grid cells into account. Since residue transport is expensive, I assume there is a strong incentive to keep the transport distances as small as possible. I don't see this as a very critical point, since only quite a small amount of residue (7-8%) is moved between grid cells. However, it may be worth adding a sentence or two about this to the discussion/limitations section.
4. Is there any way to quantify the uncertainty for residue production?
5. I assume all values for biomass weights are given in dry matter as opposed to fresh matter. It is probably worth stating this somewhere to avoid possible confusion. If you needed to convert e.g. grain yields from fresh to dry matter using standard crop-specific conversion factors it is probably worth documenting these values in the supplementary material.