the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Agricultural Land Management Practices in the Conterminous United States from 1980–2023
Abstract. Agricultural land management practices affect environmental variables including greenhouse gas emissions, carbon and nitrogen cycles, soil profiles, water quality, and air pollution. To better understand agricultural land management practices, we compiled a land management history for the conterminous United States (CONUS) from 1980–2023. We used the National Resources Inventory as the basis for our sample-based approach to impute planting and harvest dates, synthetic N fertilizer and manure N and C application rates and timing, tillage systems and intensity, and cover crop adoption. We aggregated the imputations to a 0.25-degree grid and compiled a comprehensive dataset detailing the management practices used on agricultural lands across CONUS. From 1980–2023, we found trends towards later planting dates for cotton and spring grains and trends towards earlier planting dates for soybeans. Synthetic N fertilizer rates increased steadily from 1980–2000 and then stabilized from 2000–2023, while manure N amendments were low between 1980 and 2000 and then increased rapidly from 2000–2023. Generally, there were increases in no-till and reduced-till systems across CONUS, with more notable increases in central and eastern regions. Cover crop adoption increased across CONUS from 1980–2023, with the highest level of adoption occurring in the Northeast. The results from this product align with previously published histories, although our work provides a more comprehensive representation of cropland management practices than previous works. Our dataset is available on Dryad under public domain license (Hoskovec et al., 2025) and can be used to inform studies of agricultural lands that evaluate processes such as greenhouse gas emissions, environmental impacts, and food production.
- Preprint
(13330 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2025-842', Yushu Xia, 16 Feb 2026
-
RC2: 'Comment on essd-2025-842', Anonymous Referee #2, 08 May 2026
This paper describes a method that creates spatiotemporal datasets across CONUS (1980-2023) of agricultural land management practices including planting and harvest dates for major crops, nutrient management (fertilizer/manure application rates, timing), tillage, and cover crops. It uses a machine learning approach to impute estimates based primarily on point-based NRI surveys that are then adjusted based on state-level inventory datasets and then estimated spatially (wall-to-wall) across CONUS using weighted averaging techniques.
The method is novel and the dataset created will be a useful resource that leverages the rich survey datasets and the use of Gradient Boosted Modeling appears appropriate and well justified. The writing is high quality and the manuscript is well-organized. My main concerns are related to 1) the presentation of the methods and tracing the various datasets and their ultimate raw sources and derivatives and 2) the conceptual underpinning of the method, which does not seem to adequately leverage datasets that capture more detailed spatial variability (e.g., county-scale Ag Census data) and, thus, inherently smoothes away a lot of spatial variability in an increasingly concentrated (livestock) and intensifying agricultural landscape. I recommend reconsider with major revisions.
Regarding the first major concern, it would be really helpful to include a table of all of the major datasets used for each variable, what spatial and temporal scales they cover, etc. You rely on several foundational datasets that often are not adequately explained and you force the reader to search for the most salient points from the cited sources. Obviously, you don’t need to have an extensive description of each dataset (although the supplemental could be expanded) but describing the basic underpinning of each of the major ones used is essential. For instance, the EPA (2024) dataset is used throughout the method and the reader is not told that it relies on other datasets like the USDA-NASS annual livestock surveys to estimate things like manure N. As a data journal article, more time needs to be spent describing the underlying input datasets to the method.
For the second point, I am concerned about the utility of the spatial estimates and their inability to capture real and important spatial variability that is clearly shown by other datasets. For example, a reader hypothesizes that concentrated animal feeding operations (CAFOs) (or at least counties within a state that have relatively high densities of livestock) play an outsized role in certain environmental outcomes. Would your dataset be adequate to assess this question when it relies so heavily on state-level livestock estimates? I strongly recommend that the authors consider incorporating county-scale (and finer, if available) datasets to better constrain the spatial estimates based on the GBMs. If they are unable to incorporate such datasets because that is too out of scope, then at the very least they need to communicate the limitations of the produced dataset in terms of estimating within-state spatial variability.
I expand on the major comments below with reference to specific locations in the manuscript.
Specific comments:
Line 2: What is meant by “soil profiles”? This term is only used once here in the abstract. Please be more specific.
Lines 8-11: Several variables are described as increasing through time. Could you please put in percentages to provide more detail on how large the increases are?
Line 17: This list is much shorter than the one in the first sentence of the abstract. I would be consistent with the framing. It doesn’t have to be an exhaustive list but more than GHG/carbon here makes sense.
Line 24: Again, same comment as above. You had a much broader list of potential applications for this dataset but in this paragraph you are choosing to only focus on GHG/carbon. Why? This also happens again in Line 47.
Figures 1 and 2: These two figures should be re-created using publicly available shapefiles from the USDA (e.g., https://nrcs.maps.arcgis.com/home/item.html?id=58c18a7690fa4b2c86c5a9a069e0457b) I don’t know what the journal policy is on reproducing images from other sources but I think it should be avoided whenever possible.
Lines 66-77: How many points are included in the NRI sample? How much does it change through time? Is there a map that shows their distribution that could be included here or in the supplemental?
Line 102: One overarching question that comes up for me for each of the output variables from your method is why do you often aggregate spatially explicit data to the broader state-level and then use a different technique to distribute the data in a spatially explicit manner? Why not take advantage of the spatial variability present in some of the underlying datasets (e.g. OpTIS is available at the sub-state level, Ag Census has many variables available at the county level). I understand that your methodology, which relies heavily on the NRI survey sample points, would have to change substantially. But from a purely conceptual standpoint, why start with higher resolution data, upscale it, and then downscale it? Are there ways to incorporate county-scale datasets into your method?
Line 118: Based on this description of the use of the EPA (2024) dataset (i.e., manure N available to soils), does this mean that your dataset is estimating N “downstream” of the losses associated with collection and storage of manure? If so, what would you tell a researcher who is interested in determining the total impact of agricultural production on atmospheric nitrogen losses (ammonia and N2O) and wants to include the losses at CAFOs where collection and storage losses can be quite high? If your dataset has this limitation, then it needs to be directly communicated somewhere in the methods and discussion, especially because you state in the abstract that air pollution is a potential application.
Lines 124-125: Was there a reason that alfalfa was excluded from your analysis? Area harvested nationally for alfalfa is much larger compared to barley, oats, and sorghum combined.
Line 182: Could you please explain in more detail how manure N availability was estimated in EPA (2024)? I am roughly familiar with that dataset but looking at the link in the references does not provide any immediate indication of how manure N availability was determined. I would strongly recommend that at least a sentence or two is devoted to explaining how EPA (2024) came up with these estimates and what underlying datasets they relied upon (USDA-NASS survey livestock data)
Line 185: Just to note, this is where I really started to wonder why county-scale Ag Census data wasn’t considered at all (see major comment above)
Lines 186-195: Related to the major comment above about capturing spatial variability, at some point (either in methods or discussion) it would be very useful to have more discussion on the conceptual idea behind using a quite spatially-limited NRI point-scale survey data (and associated soil and geographic characteristics) along with GBMs to characterize the spatial variability of nutrient application rates and management. It’s just a fundamentally different way of understanding the drivers of spatial variability than I am used to. While far from perfect, at least county-scale datasets (Falcone, 2021; https://doi.org/10.3133/ofr20201153) give an indication of things like where livestock are concentrated within a state. This county-scale spatial variability could be a function of a lot of different things including historical roots of livestock production and recent development of livestock production infrastructure. Your conceptual model seems to suggest that spatial variability is solely a function of soil and geographic conditions that is mediated by a very limited spatial dataset of NRI point surveys. I don’t mean to dismiss this method out-of-hand but it just seems to have a lot of limitations that are not adequately explained.
Lines 210-213: Again, this is where I became increasingly confused as to why you would spend considerable effort to break apart estimates into each animal type at the state scale, while not incorporating animal type specific inventories at the county scale (Falcone, 2021)
Lines 226-227: What basis do you have for assuming that manure is only applied in the fall or the spring? Very often in livestock intensive places, manure is applied both in the fall and the spring (see this example from a survey of farmers in Wisconsin: https://dnr.wisconsin.gov/sites/default/files/topic/TMDLs/NEL/appendix_e_ag_survey.pdf)
Line 273: I think I know what you mean by “donor pools” but it is not immediately apparent. Please explain further.
Line 289-292: Again, is there any way to incorporate county-scale estimates of cover crop adoption from Ag Census data? Seems to be a real missed opportunity.
Lines 330-331: Isn’t this just circular logic? You assumed that manure was only applied in the spring or the fall in the methods and now it is being described as a result?
Figures 6 and 7: I would strongly consider changing these line charts to stacked area charts. It would perhaps be harder to see trends for individual crop types but then the reader could easily visualize how the total is changing through time.
Line 391: I tried to track down the USDA-ERS (2020) reference from the link in the references but it only went to a report from 2000. Please update.
Line 394-395: Is this finding surprising, though? If I understood the methods correctly, didn’t you adjust your manure rates based on EPA (2024) which also uses the USDA-NASS livestock numbers? This is where it occurred to me that there definitely needs to be more description of the different datasets used and what underlying datasets are used to inform them. I realize that it’s a lot to keep track of. But these interdependencies in the agricultural data ecosystem are incredibly important to understand and communicate.
Citation: https://doi.org/10.5194/essd-2025-842-RC2 -
AC1: 'AC Comment on essd-2025-842', Lauren Hoskovec, 26 Jun 2026
Author Comments:
We thank the reviewers for their thoughtful and constructive comments on our manuscript. Their feedback has been instrumental in strengthening our methodological justifications, structuring our data inputs, and sharpening the overall clarity of our analysis. We have addressed each comment specifically below.
Responses to RC 1 Comments:
The manuscript “Agricultural Land Management Practices in the Conterminous United States from 1980–2023” presents a grid-based crop management dataset derived from a data fusion and modeling approach. The authors provide visualizations that illustrate spatial and temporal patterns across multiple management variables produced in this study. The manuscript also describes the novelty of the dataset and includes comparisons with existing products. Overall, this work represents a timely contribution with clear potential value for modelers and a broad range of stakeholders. I offer several minor suggestions below that may further improve the clarity and accessibility of the manuscript for a wider audience.
L2: The manuscript did mention applications for GHG and C accounting quite a few times, but there was not much mention of water quality or air pollution, despite their inclusion in the Abstract. I suggest that the authors mention these applications in both the Introduction and the Discussion sections. It may be particularly helpful to add a short paragraph that lists example use cases of this dataset, which would clarify its broader applicability and help a wider audience appreciate its full value.
Yes, we agree that additional references can be added for other types of applications. We have added references to and more details regarding water quality studies in the Introduction. We extended the Conclusions to elaborate on example use cases of this product.
L49-50. The historical agricultural management data is very important for model spinup and making historical model simulations for counterfactual analysis. Consider mentioning these applications in the Introduction and/or Discussion sections.
Yes, these data are also important, particularly for applications that require a longer spin-up period, but it was not in the scope of this study to develop the longer time series of historical management prior to 1980. We have added a sentence in the conclusions acknowledging that these data can be combined with other datasets to model a longer time series for model application with the goal of evaluating patterns further in the past, or models that require a longer period to initialize state variables such as soil organic carbon pools.
L57-60. The language in this section gives the impression that the work primarily focuses on gap filling. However, the study also includes evaluation results, as described in the Appendix, and quality control steps are probably included as well. I suggest revising the language to better reflect the full range of activities conducted in this study. It may also be beneficial to highlight elements of novelty here rather than waiting until the Discussion section.
Thank you. We have revised this section to focus on creating the field-level imputation product, evaluating the imputations, ensuring quality control, and aggregating the field-level imputations to the 0.25-degree grid. We also highlight the elements of novelty including the broad range of management activities, detailed tillage information, and dissemination of a publicly available data set based on the NRI survey.
L65-77. Based on my understanding, the original NRI data are not openly accessible to the broader research community. If this is correct, an important contribution of this work lies in making information derived from this valuable dataset available through a gridded product. I suggest emphasizing this aspect.
Thank you. We added text to both the Introduction and Discussion to emphasize the novelty of our approach being based on the NRI sample and how aggregating the imputations to the grid allows for dissemination of data derived from this survey. We also added a paragraph to the Discussion regarding the major sources of innovation, which include using both the NRI survey as the basis for our sample-based approach and the detailed CEAP data to inform field-level management practices.
Figure 1. Is it necessary to include both Figures 1 and 2 in the main manuscript? They are not mentioned frequently in the manuscript, and the results are not often aggregated to these units. Consider moving to SI.
We have moved both maps to supplement.
L85. For audience less familiar with NRI and various agricultural management data product, it might be helpful to mention whether the following data products are completely independent from NRI or has partial overlaps. There are many products discussed in this section and a flowchart or a table (mentioning existing products and what novelty this study brings) might be helpful to illustrate the methodology.
We added a flow chart to describe our data assimilation process (now Figure 1), which shows that CEAP I and II, CDL, and NRI overlap at the field level, while other data sets are used to inform state-level trends or as top-down constraints. In response to other comments, we also added a table describing the data sources used to impute each management practice.
L169-170. Later on, the manuscript also mentioned manure animal type, which is important information and should probably be mentioned here as well. What about fertilizer types? Is it possible to include such information in the dataset as well and if not, what is the main reason for the exclusion/ challenge?
We added that we imputed manure animal type to the beginning of Section 3.1.2. We did not include information on fertilizer types because these data were not currently available in a usable format at the time of manuscript submission. In addition, there are clear dependencies between fertilizer rate, type, placement, and timing that need to be considered before predicting these activities accurately at the field-level. Hence, we are working to develop a more comprehensive fertilizer application data set before publishing those results.
L175-177. With this calculation method, is it possible to encounter area double counting when both manure and synthetic fertilizers? The ARMS data does include %area receiving both (even though it’s a small value).
It is possible that NRI survey locations were treated with both synthetic N fertilizer and manure amendment, and the stochasticity of our approach permitted this overlap. We hope to improve on modeling the correlation between fertilizer and manure applications in area treated, rates, and timing in future developments of this product.
L183. Assume the N availability data is assigned with animal type and manure form for the subsequent calculation?
We did not assign the N availability data with animal type and manure form when calculating the relative manure N available for application to soils because the EPA (2024) manure N availability data were aggregated to the state level. The “relative manure N available” predictor variable was meant to be a measure of the amount of manure N available to help establish baseline state-level average manure amendment rates where we had little data. We tried other derivations of manure N availability variables and found this derived variable to be the best predictor.
L200. Would it be accurate to say that the dataset of this study is more focused on crops, and therefore is not suited for the modeling of pastureland or vegetables?
Yes, the data set of this study focused on the major crops: corn, cotton, soybeans, sorghum, wheat, barley, and oats. It is possible to predict the management practices mentioned here on minor crops, hay, and pasture using some alternative data sources and modeling techniques. Our data set focused on the major crops due to the consistency in data and methods used to model these crops.
L220-221. The availability of crop-specific dataset from ARMS is highly dependent upon the time period. Is it possible to show some statistics about how much gap filling needs to be done for each time period? The gap-filling aspect can also influence conclusions drawn on temporal trends.
Yes, the ARMS data for fertilizer timing was only available every 4-5 years between 2001-2021. We have added this information to the text to clarify. We also clarified that results on fertilizer application timing before 2001 reflect the 2001 data.
Table 1. Is it possible to connect these levels with typical tillage practice names or expand the “Example” column for the broader audience who is interested in using this dataset?
We added additional examples of tillage implements that fall within each intensity class for use by the broader audience.
L283-292. Would something like CDL be helpful since crop types are identified in time? Or is CDL already embedded in one of the source datasets here? If not, I wonder if that data could be used for some sort of validation.
We did incorporate some of the Cropland Data Layer (CDL) data directly into our product. Specifically, the Cropland Data Layer has been incorporated into the analysis for filling gaps in crop types from 2018 to 2023. This information is provided in the methods section. In response to another reviewer comment, we also created a table (now Table 1) describing the data sources used to inform each management activity, and list CDL as one of the sources for informing crop type.
L294. Explain why 0.25 degree is selected as the final grid size.
The NRI program agreed to allow us to share the data product, which has underlying confidential information, at the 0.25-degree resolution. At the end of the Introduction, we added a sentence explaining that the 0.25-degree grid was selected to maintain confidentiality.
L296-298. It is not fully clear what the six imputations correspond to – is it for calculating the quantiles mentioned in the next paragraph?
We produced six imputations of management activities in our analysis to capture the variability in predicted management activities at the field level. This variability is based on random sampling inherent in the statistical algorithms used to impute each management practice. Measures of variation in the gridded dataset include both variability among and within the NRI survey locations in each grid cell. We have added this description to Section 3.2.
L305. Explain more about the “data suppression” procedure and whether there is follow up interpolations.
To preserve the confidentiality of the NRI survey locations and associated field information, gridded data were suppressed if there were fewer than three NRI survey locations within the grid cell. Notably, grid cells with fewer than three NRI survey locations were areas with little cropland and contribute minimally to larger regional patterns in cropland management. We did not provide follow-up interpolations for the suppressed data. We revised the text to include this information.
L306. Point out corresponding “Appendix” for the Method section, and consider adding some languages related to validation/ evaluation, as some of the results in the Appendix seems quite important.
Thank you. We added some language to Section 4 summarizing the evaluation of the imputations and referencing the associated Appendix.
Briefly, our survey location-level imputations closely match predictions from statistical models and align with the observed data when moderate to high amounts of data are available. Where data are scarce, the differences between predicted values (e.g. predictions from linear models, GBMs or Markov transition models) and observed values are large, leading to larger differences between imputed values and observed values. Hence, heightening the scale of temporal and spatial data availability will lead to more accurate inference on field-level management activities in future products.
Figure 3. The right edge of the (f) figure was cut off.
Thank you. This has been fixed.
Figure 8. Could this temporal trend be explained?
These values reflect patterns in the manure available for application to soils, and the amount of municipal waste in biosolids. The dynamics are presumably related to the size of livestock populations and changes in efficiency of production over the time series. We did not study those drivers directly in our development of this dataset. However, other research has shown a clear pattern of increasing proportion of smaller operations over the time period and there have been improvements in efficiency of production with more production per animal. Fancher et al. (2025) for example discuss shifts in beef cattle operations and efficiencies. Another study by Njuki showed substantial shifts in herd size and efficiencies in the United States (Njuki, E. 2022). While we are not aware of any study evaluating the drivers for manure production directly, these structural changes are likely driving the shifts in proportion of manure among the livestock categories. Future research is needed to further address this question. We included a discussion of this trend alongside the referenced articles in the Discussion.
Figure 10. Could there be a bit more explanation in the caption to make it stand alone. For example, consider mentioning that the tillage intensity increases from A to K.
We added additional text to the caption to better explain the A to K classifications and make the figure stand alone.
Figure 11. Consider modifying the color scheme or value range associated with the legend – the distinction of color in these four figures is pretty minor.
We have changed the color scheme to improve distinction.
L374-377. It would be great if the authors could also mention main differences in methodology in addition to the similarity. Otherwise, it would be expected that the results are similar given the underlying datasets and methodology.
Thank you for this suggestion. We included information on differences in methodology and data sets used that could be driving some of the differences in results of synthetic N fertilizer use between our product and that described in Cao et al. (2018).
Responses to RC2 Comments:
This paper describes a method that creates spatiotemporal datasets across CONUS (1980-2023) of agricultural land management practices including planting and harvest dates for major crops, nutrient management (fertilizer/manure application rates, timing), tillage, and cover crops. It uses a machine learning approach to impute estimates based primarily on point-based NRI surveys that are then adjusted based on state-level inventory datasets and then estimated spatially (wall-to-wall) across CONUS using weighted averaging techniques.
The method is novel and the dataset created will be a useful resource that leverages the rich survey datasets and the use of Gradient Boosted Modeling appears appropriate and well justified. The writing is high quality and the manuscript is well-organized. My main concerns are related to 1) the presentation of the methods and tracing the various datasets and their ultimate raw sources and derivatives and 2) the conceptual underpinning of the method, which does not seem to adequately leverage datasets that capture more detailed spatial variability (e.g., county-scale Ag Census data) and, thus, inherently smoothes away a lot of spatial variability in an increasingly concentrated (livestock) and intensifying agricultural landscape. I recommend reconsider with major revisions.
Regarding the first major concern, it would be really helpful to include a table of all of the major datasets used for each variable, what spatial and temporal scales they cover, etc. You rely on several foundational datasets that often are not adequately explained and you force the reader to search for the most salient points from the cited sources. Obviously, you don’t need to have an extensive description of each dataset (although the supplemental could be expanded) but describing the basic underpinning of each of the major ones used is essential. For instance, the EPA (2024) dataset is used throughout the method and the reader is not told that it relies on other datasets like the USDA-NASS annual livestock surveys to estimate things like manure N. As a data journal article, more time needs to be spent describing the underlying input datasets to the method.
Thank you. We agree that a table of the major datasets is needed to clarify how the numerous data sources are used to create our product. We have created this table, now Table 1, which describes the data sets used to inform each management practice as well as the spatial and temporal scales used in our product. In addition, we describe in Section 2 that the EPA (2024) data set relies on USDA NASS and US Agricultural Census data sets.
For the second point, I am concerned about the utility of the spatial estimates and their inability to capture real and important spatial variability that is clearly shown by other datasets. For example, a reader hypothesizes that concentrated animal feeding operations (CAFOs) (or at least counties within a state that have relatively high densities of livestock) play an outsized role in certain environmental outcomes. Would your dataset be adequate to assess this question when it relies so heavily on state-level livestock estimates? I strongly recommend that the authors consider incorporating county-scale (and finer, if available) datasets to better constrain the spatial estimates based on the GBMs. If they are unable to incorporate such datasets because that is too out of scope, then at the very least they need to communicate the limitations of the produced dataset in terms of estimating within-state spatial variability.
Thank you for this suggestion. First, no, it was not within the scope of this work to assess the impact of CAFOs. The manure management data we use from EPA (2024) provides the amount of manure N available for application to soils and is downstream from manure management in CAFOs. The EPA (2024) data was provided at the state-level and was not available at county-level at the time of manuscript submission. County-level estimates of kg N from manure, for example from Falcone (2021), are not necessarily representative of what is available to be applied to the croplands in our analysis. These county-level estimates include total manure N kg produced, including the amount applied to pasturelands. Hence, we retain our current state-level analysis of manure animal types, while noting these limitations in the Discussion.
We are revising our analysis of cover crops using county-level data from US Agricultural Census to better capture the spatial variability of this management practice within a state. We are testing the new approach and expect to incorporate these results in the revised manuscript.
Regarding the other management practices, we emphasize that the field-level CEAP surveys include important spatial drivers (MLRA, latitude, and longitude) and represent a substantial subset (>21,000 locations) of the NRI survey data set. Hence, local spatial trends in manure amendment rates and fertilization rates can be identified in the CEAP data through machine learning algorithms. Providing multiple imputations of the management data allows us to capture some of the spatial variability regarding the locations managed with these nutrients and the rates of nutrient application. We clarified our methodological approach in Section 3 and added text to the Discussion regarding the strengths and limitations of using the field-level CEAP data.
We address the specific comments regarding this major point below.
I expand on the major comments below with reference to specific locations in the manuscript.
Specific comments:
Line 2: What is meant by “soil profiles”? This term is only used once here in the abstract. Please be more specific.
This term has been removed from the Abstract since it is not used later.
Lines 8-11: Several variables are described as increasing through time. Could you please put in percentages to provide more detail on how large the increases are?
Thank you for this comment. We have provided specific information in the Abstract regarding the percentages and/or rates to provide more detail on how large the trends are over time. In addition, we provided more detail regarding these trends in the Results section.
Line 17: This list is much shorter than the one in the first sentence of the abstract. I would be consistent with the framing. It doesn’t have to be an exhaustive list but more than GHG/carbon here makes sense.
We have added water quality and nitrogen cycles to this list and added references describing the effects of agricultural management activities on water quality.
Line 24: Again, same comment as above. You had a much broader list of potential applications for this dataset but in this paragraph you are choosing to only focus on GHG/carbon. Why? This also happens again in Line 47.
Thank you for this comment. We have updated the text to expand the applications of our product in both in the Introduction and Conclusions. Specifically, we added text and references regarding the impact of agricultural land management practices on water quality in addition to the existing references on soil conservation, carbon and nitrogen cycles, and GHG emissions.
Figures 1 and 2: These two figures should be re-created using publicly available shapefiles from the USDA (e.g., https://nrcs.maps.arcgis.com/home/item.html?id=58c18a7690fa4b2c86c5a9a069e0457b) I don’t know what the journal policy is on reproducing images from other sources but I think it should be avoided whenever possible.
Thank you. We recreated these figures using shapefiles.
Lines 66-77: How many points are included in the NRI sample? How much does it change through time? Is there a map that shows their distribution that could be included here or in the supplemental?
We added text describing the number of points in the NRI sample. Specifically, in our analysis, we used a total of 299,570 distinct NRI survey locations. The number of locations per year varied from 216,349 locations in 1980 to 168,409 locations in 2023 due to variations in land use over time. We cannot include a map of the NRI locations because they are confidential and the locations cannot be publicly disclosed.
Line 102: One overarching question that comes up for me for each of the output variables from your method is why do you often aggregate spatially explicit data to the broader state-level and then use a different technique to distribute the data in a spatially explicit manner? Why not take advantage of the spatial variability present in some of the underlying datasets (e.g. OpTIS is available at the sub-state level, Ag Census has many variables available at the county level). I understand that your methodology, which relies heavily on the NRI survey sample points, would have to change substantially. But from a purely conceptual standpoint, why start with higher resolution data, upscale it, and then downscale it? Are there ways to incorporate county-scale datasets into your method?
Thank you for this comment. We believe clarifying our data sources through the table as you suggested helps address this comment. Specifically, the spatially and temporally rich NRI data set and field-level CEAP operations data provide high resolution information regarding environmental variables and management practices. We supplement these data sources with coarser data sets from NASS, ARMS, ERS, EPA, USGS, US Agricultural Census, OpTIS, and CTIC to estimate long-term baseline trends at the state-crop-year level. Then we use statistical and machine learning models to identify patterns in the field-level data, which, combined with the broad trends, are used to impute management practices on the NRI survey locations. Hence, our analysis is conducted at the field level, using both field-specific surveys and long-term spatial and temporal trends from other surveys and data sources. It is a novel methodology for this type of work, and we show that our results are similar to previously published data products using other data sets, including county-level inventories. Our work leverages our access to the NRI and CEAP surveys, which are not publicly available. Aggregating to a 0.25-degree grid, we permit public access to the gridded results derived from these surveys.
In response to your remark about OpTIS, we took advice from the USDA, who funded this work, and who recommended we use the OpTIS product at the state-scale rather than original product scale.
We have added text to clarify our approach in Section 3.
Line 118: Based on this description of the use of the EPA (2024) dataset (i.e., manure N available to soils), does this mean that your dataset is estimating N “downstream” of the losses associated with collection and storage of manure? If so, what would you tell a researcher who is interested in determining the total impact of agricultural production on atmospheric nitrogen losses (ammonia and N2O) and wants to include the losses at CAFOs where collection and storage losses can be quite high? If your dataset has this limitation, then it needs to be directly communicated somewhere in the methods and discussion, especially because you state in the abstract that air pollution is a potential application.
Our dataset is specific to soil management and is downstream from manure management in CAFOs. We will clarify in the methods that our dataset is restricted to soil management, and does not provide data on the collection, storage and handling of manure.
Lines 124-125: Was there a reason that alfalfa was excluded from your analysis? Area harvested nationally for alfalfa is much larger compared to barley, oats, and sorghum combined.
Our product was focused on major annual commodity crops rather than perennial crops. We hope to extend the product in the future to include perennial crops and grasslands.
Line 182: Could you please explain in more detail how manure N availability was estimated in EPA (2024)? I am roughly familiar with that dataset but looking at the link in the references does not provide any immediate indication of how manure N availability was determined. I would strongly recommend that at least a sentence or two is devoted to explaining how EPA (2024) came up with these estimates and what underlying datasets they relied upon (USDA-NASS survey livestock data).
There is a detailed description of the methods in Annex 3.11 of EPA (2024). EPA (2024) compiled population data on livestock from the USDA NASS and Census datasets. Manure production was estimated from the population data using Intergovernmental Panel on Climate Change (IPCC) methods, along with emissions and losses from manure management systems (IPCC 2006, 2019). The remaining manure is available for application to soils and used in our analysis of manure amendments as described in Section 3.1.2 in the manuscript.
We have added this information to the data descriptions in Section 2 in the manuscript.
Line 185: Just to note, this is where I really started to wonder why county-scale Ag Census data wasn’t considered at all (see major comment above)
Thank you for this note. We address this comment below.
Lines 186-195: Related to the major comment above about capturing spatial variability, at some point (either in methods or discussion) it would be very useful to have more discussion on the conceptual idea behind using a quite spatially-limited NRI point-scale survey data (and associated soil and geographic characteristics) along with GBMs to characterize the spatial variability of nutrient application rates and management. It’s just a fundamentally different way of understanding the drivers of spatial variability than I am used to. While far from perfect, at least county-scale datasets (Falcone, 2021; https://doi.org/10.3133/ofr20201153) give an indication of things like where livestock are concentrated within a state. This county-scale spatial variability could be a function of a lot of different things including historical roots of livestock production and recent development of livestock production infrastructure. Your conceptual model seems to suggest that spatial variability is solely a function of soil and geographic conditions that is mediated by a very limited spatial dataset of NRI point surveys. I don’t mean to dismiss this method out-of-hand but it just seems to have a lot of limitations that are not adequately explained.
Thank you for this point. We have updated and rephrased the referenced section (Section 3.1.2). We now clarify that we conducted the analysis of fertilizer and manure amendment rates at the field level using CEAP data, and used ARMS, ERS, and EPA data to establish baseline trends in crop-specific state-level average rates over the time series. The field-level analysis utilized the baseline trends as a mean, modeling individual CEAP field data as residuals to ensure consistency across the time series. We did not use county-level Agricultural Census data because the CEAP data are at an even finer scale (field-level) and provide important spatial drivers of application rates (e.g. MLRA, latitude, longitude). Patterns in these spatial drivers are captured through the machine learning model. County-level data may provide an explanation of these spatial patterns (e.g. livestock concentrations) and can supplement the CEAP data. We acknowledge that the CEAP data is limited temporally and discuss the limitations of this data set in the Discussion. We also discuss the possibility of supplementing the CEAP data set with county-level Agricultural Census data in future versions of this product. However, the county level data are not crop-specific, as in ARMS and CEAP, and often not measuring precisely what we are modeling. While adapting the county-level datasets to match our specific application is a valuable effort, doing so requires alternative data assimilation and statistical methods that we defer to a dedicated future study.
The NRI data set, with nearly 300,000 survey locations, provides a spatially robust representation of CONUS. In addition, the CEAP data set contains a subset of over 21,000 NRI survey locations. We added these details to the manuscript to justify our approach of conducting a bottoms-up analysis using field-level data, with constraints imposed at the state-level. Our product offers a new method for estimating fertilizer and manure amendment rates using field-level management data, and our results are similar to other products that use different data sources and methods.
We recognize that the fine-scale CEAP data was not used to estimate cover crop adoption rates. In response to this major comment, we are modifying our analysis of cover crops to include county-scale data since other fine-scale data sources were not used. We expect to include these results in the revised manuscript.
Lines 210-213: Again, this is where I became increasingly confused as to why you would spend considerable effort to break apart estimates into each animal type at the state scale, while not incorporating animal type specific inventories at the county scale (Falcone, 2021)
Thank you for this comment. The EPA (2024) data on livestock animal type fractions for manure N available for application to soils were only available on the state scale. Unlike county-level inventories, which aggregate total livestock populations without separating unapplied or pasture-deposited waste, the state-level EPA dataset isolates manure explicitly available for soil application, ensuring a more accurate input for our cropland management product.
We mention the limitation of this approach in the Discussion, as well as future work that could combine our approach with the information in the county-level data sets.
Lines 226-227: What basis do you have for assuming that manure is only applied in the fall or the spring? Very often in livestock intensive places, manure is applied both in the fall and the spring (see this example from a survey of farmers in Wisconsin: https://dnr.wisconsin.gov/sites/default/files/topic/TMDLs/NEL/appendix_e_ag_survey.pdf)
Based on manure amendment timing percentages reported in the ARMS data, we assumed manure was applied in fall, spring, or after planting. Since there was no information on split applications in the data set, we assumed a single application time at each NRI survey location.
Line 273: I think I know what you mean by “donor pools” but it is not immediately apparent. Please explain further.
We have restructured this paragraph to clarify what is meant by CEAP donor pools. Specifically, we used a hot-deck imputation approach stratified by CEAP donor pools to impute annual tillage intensity patterns on the NRI sample. The CEAP donor pools were constructed by combining the CEAP data within years (CEAP survey 1 was used for 1980-2008 and CEAP survey 2 for 2009-2023), CEAP region, crop (or crop group), and five-year tillage system type. We matched each NRI survey location to a CEAP donor pool, then randomly selected a tillage intensity pattern from the donor pool to represent the annual tillage intensity pattern at that NRI survey location.
Line 289-292: Again, is there any way to incorporate county-scale estimates of cover crop adoption from Ag Census data? Seems to be a real missed opportunity.
Thank you for this suggestion. We are updating our methods for imputing cover crop adoption to include county-level data from the US Agricultural Census. We expect to include these results in the revised manuscript.
Lines 330-331: Isn’t this just circular logic? You assumed that manure was only applied in the spring or the fall in the methods and now it is being described as a result?
To clarify, based on the ARMS data, manure could be applied in fall before planting, spring before planting, or after planting, and we report that most manure was applied in fall before planting or spring before planting. That is, little manure was applied after planting. We added a sentence to clarify this.
Figures 6 and 7: I would strongly consider changing these line charts to stacked area charts. It would perhaps be harder to see trends for individual crop types but then the reader could easily visualize how the total is changing through time.
We respectfully disagree. The first figure in the panel presents the annual total for each crop. Our analysis was conducted at the crop-level, so showing crop-specific trends over time remains the primary focus of our results. In addition, because not all crops were modeled in our analysis, we believe showing a total rate would be misleading, as it represents only the total of the crops modeled, and should not be interpreted as total application to cropland.
Line 391: I tried to track down the USDA-ERS (2020) reference from the link in the references but it only went to a report from 2000. Please update.
Thank you for catching this. The reference has been updated.
Line 394-395: Is this finding surprising, though? If I understood the methods correctly, didn’t you adjust your manure rates based on EPA (2024) which also uses the USDA-NASS livestock numbers? This is where it occurred to me that there definitely needs to be more description of the different datasets used and what underlying datasets are used to inform them. I realize that it’s a lot to keep track of. But these interdependencies in the agricultural data ecosystem are incredibly important to understand and communicate.
Thank you for this point. We have updated this paragraph to focus on the methodological differences between the two data products. Specifically, we note that we coupled machine learning models with field-level application data from CEAP to impute manure amendment rates, while Bian et al. (2021) calculated applicated rates using livestock numbers and manure recoverability rates. Despite the differences in methods, we arrived at similar conclusions.
Additional References
- EPA: Inventory of U.S. Greenhouse Gas Emissions and Sinks: 1990-2022., https://www.epa.gov/ghgemissions/inventory-us-greenhouse-gas-emissions-and-sinks-1990-2022, 2024.
- Fancher, H., Nagler, A., Ritten, J., and Wulfhorst, J. D.: US Beef Cattle Inventory Trends With Implications for Land Use and Rangelands, Rangeland Ecology and Management, 103, 545–553, https://doi.org/10.1016/j.rama.2025.01.007, 2025.
- Falcone, J. A.: Estimates of county-level nitrogen and phosphorus from fertilizer and manure from 1950 through 2017 in the conterminous United States, Open-File Report 2020–1153, U.S. Geological Survey, ISSN 2331-1258, https://doi.org/10.3133/ofr20201153, 2021.
- IPCC: 2006 IPCC Guidelines for National Greenhouse Gas Inventories, Institute for Global Environmental Strategies (IGES), Hayama, Japan, https://www.ipcc-nggip.iges.or.jp/public/2006gl/index.html, 2006.
- IPCC: 2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories, IPCC, Geneva, Switzerland, https://www.ipcc-nggip.iges.or.jp/public/2019rf/index.html, 2019
- Njuki, E.: Sources, Trends, and Drivers of U.S. Dairy Productivity and Efficiency, www.ers.usda.gov, 2022.
Citation: https://doi.org/10.5194/essd-2025-842-AC1
Data sets
Agricultural land management activities for conterminous US from 1980-2023 Lauren Hoskovec et al. https://datadryad.org/share/LINK_NOT_FOR_PUBLICATION/WmO29tLOABEqXMRp0H-NZpr1ddUotryamH2b78bK42g
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 345 | 281 | 26 | 652 | 42 | 68 |
- HTML: 345
- PDF: 281
- XML: 26
- Total: 652
- BibTeX: 42
- EndNote: 68
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General Comments:
The manuscript “Agricultural Land Management Practices in the Conterminous United States from 1980–2023” presents a grid-based crop management dataset derived from a data fusion and modeling approach. The authors provide visualizations that illustrate spatial and temporal patterns across multiple management variables produced in this study. The manuscript also describes the novelty of the dataset and includes comparisons with existing products. Overall, this work represents a timely contribution with clear potential value for modelers and a broad range of stakeholders. I offer several minor suggestions below that may further improve the clarity and accessibility of the manuscript for a wider audience.
Specific Comments:
L2: The manuscript did mention applications for GHG and C accounting quite a few times, but there was not much mention of water quality or air pollution, despite their inclusion in the Abstract. I suggest that the authors mention these applications in both the Introduction and the Discussion sections. It may be particularly helpful to add a short paragraph that lists example use cases of this dataset, which would clarify its broader applicability and help a wider audience appreciate its full value.
L49-50. The historical agricultural management data is very important for model spinup and making historical model simulations for counterfactual analysis. Consider mentioning these applications in the Introduction and/or Discussion sections.
L57-60. The language in this section gives the impression that the work primarily focuses on gap filling. However, the study also includes evaluation results, as described in the Appendix, and quality control steps are probably included as well. I suggest revising the language to better reflect the full range of activities conducted in this study. It may also be beneficial to highlight elements of novelty here rather than waiting until the Discussion section.
L65-77. Based on my understanding, the original NRI data are not openly accessible to the broader research community. If this is correct, an important contribution of this work lies in making information derived from this valuable dataset available through a gridded product. I suggest emphasizing this aspect.
Figure 1. Is it necessary to include both Figures 1 and 2 in the main manuscript? They are not mentioned frequently in the manuscript, and the results are not often aggregated to these units. Consider moving to SI.
L85. For audience less familiar with NRI and various agricultural management data product, it might be helpful to mention whether the following data products are completely independent from NRI or has partial overlaps. There are many products discussed in this section and a flowchart or a table (mentioning existing products and what novelty this study brings) might be helpful to illustrate the methodology.
L169-170. Later on, the manuscript also mentioned manure animal type, which is important information and should probably be mentioned here as well. What about fertilizer types? Is it possible to include such information in the dataset as well and if not, what is the main reason for the exclusion/ challenge?
L175-177. With this calculation method, is it possible to encounter area double counting when both manure and synthetic fertilizers? The ARMS data does include %area receiving both (even though it’s a small value).
L183. Assume the N availability data is assigned with animal type and manure form for the subsequent calculation?
L200. Would it be accurate to say that the dataset of this study is more focused on crops, and therefore is not suited for the modeling of pastureland or vegetables?
L220-221. The availability of crop-specific dataset from ARMS is highly dependent upon the time period. Is it possible to show some statistics about how much gap filling needs to be done for each time period? The gap-filling aspect can also influence conclusions drawn on temporal trends.
Table 1. Is it possible to connect these levels with typical tillage practice names or expand the “Example” column for the broader audience who is interested in using this dataset?
L283-292. Would something like CDL be helpful since crop types are identified in time? Or is CDL already embedded in one of the source datasets here? If not, I wonder if that data could be used for some sort of validation.
L294. Explain why 0.25 degree is selected as the final grid size.
L296-298. It is not fully clear what the six imputations correspond to – is it for calculating the quantiles mentioned in the next paragraph?
L305. Explain more about the “data suppression” procedure and whether there is follow up interpolations.
L306. Point out corresponding “Appendix” for the Method section, and consider adding some languages related to validation/ evaluation, as some of the results in the Appendix seems quite important.
Figure 3. The right edge of the (f) figure was cut off.
Figure 8. Could this temporal trend be explained?
Figure 10. Could there be a bit more explanation in the caption to make it stand alone. For example, consider mentioning that the tillage intensity increases from A to K.
Figure 11. Consider modifying the color scheme or value range associated with the legend – the distinction of color in these four figures is pretty minor.
L374-377. It would be great if the authors could also mention main differences in methodology in addition to the similarity. Otherwise, it would be expected that the results are similar given the underlying datasets and methodology.