the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Statewide Forest Monitoring Data for California
Abstract. High severity wildfires, driven by a complex interplay between climate change, vegetation growth, and human activity, present a critical threat to human safety and to the ecosystems of California. Forecasting and managing these fire systems is hindered in part by a lack of comprehensive, long-term, and regularly updated data regarding the vegetation fuel loads and forest structural conditions that drive fire behavior. To address this data need, we developed and launched the California Forest Observatory in the fall of 2020, which is a system designed to map forest structure and fuels at fine spatial and temporal scales. An extensive airborne LiDAR training dataset was integrated with multi-sensor satellite data from Sentinel-1, Sentinel-2, and PlanetScope constellations. LiDAR-derived canopy fuel metrics — including canopy height, cover, base height, and ladder fuel density — were modeled using convolutional neural networks, which leveraged spatial context and sensor fusion to generate predictive models at 10-meter and 3-meter resolutions. Canopy bulk density and surface fuel data were estimated from process-based models. The canopy fuel models demonstrated strong performance: 10-meter model predictions ranged from r2 = 0.51 to r2 = 0.79, 3-meter performance metrics ranged from r2 = 0.72 to r2 = 0.79, with low net bias across all variables, outperforming comparable, recently-released data products. The output data have successfully supported wildfire hazard mapping, utility risk mitigation, and carbon neutrality planning across the state, and are available via PANGAEA (https://doi.org/10.1594/PANGAEA.989871). By providing a data-driven evidence base for monitoring vegetation change, the California Forest Observatory offers a blueprint for the next generation of satellite-based conservation technologies aiming to inform sustainable land management strategies.
- Preprint
(6837 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-255', Anonymous Referee #1, 19 Aug 2026
-
RC2: 'Comment on essd-2026-255', Anonymous Referee #2, 21 Aug 2026
This manuscript presents a statewide forest monitoring dataset for California that integrates airborne LiDAR, multi-sensor satellite observations, and different models to map forest structure and vegetation fuels. My main concern is the lack of independent validation of the time-series changes. The manuscript emphasizes the Forest Observatory as a system for monitoring changes in forest structure and fuels over time. However, the validation mainly evaluates individual predictions rather than changes between years. The authors also acknowledge that year-to-year variability can result from changing observation conditions and model uncertainty. I suggest that the manuscript more clearly discuss how reliably temporal changes can be interpreted.
Specific comments:
- Spatial independence of model evaluation. The 64 × 64 pixel tiles were randomly divided into 80/10/10 training, validation, and testing sets. Could neighboring tiles or tiles from the same LiDAR acquisition occur in both the training and testing sets? If so, spatial autocorrelation could lead to optimistic performance estimates.
-
Canopy-height saturation. Figure 7 clearly shows canopy-height saturation, particularly above about 30 m. Were any correction or post-processing approaches applied, or are the saturated predictions retained in the final products? Please discuss whether this bias may affect derived variables that depend on canopy height.
-
Zero-value patterns in Figs. 5 and 6. The canopy-cover 2-D histograms show many observations with substantial canopy cover but predicted values near zero. This pattern is especially clear for the 3 m model. Could the authors explain the source of this pattern?
-
Generation of 3 m and 10 m response variables. Section 2.2 states that canopy height and canopy cover were first rasterized at 1 m resolution, while the models operate at 3 m and 10 m resolutions. Please clarify how the 1 m LiDAR response variables were aggregated or resampled to these resolutions.
-
Figure 8 and canopy bulk density. Figure 8A shows substantially lower canopy bulk density estimates from the Forest Observatory than from LANDFIRE. The manuscript explains why the distributions have different shapes, but it does not fully explain the difference in magnitude. Please discuss the causes and effects.
-
Uncertainty in derived fuel products. Canopy bulk density combines several modeled or externally derived quantities, while the surface fuel product is generated using a rules-based classification. Please provide more discussion of uncertainty in these derived products and how errors in the inputs may propagate. It may also be helpful to clarify the use of the term “process-based models” for these different approaches.
Citation: https://doi.org/10.5194/essd-2026-255-RC2 -
RC3: 'Comment on essd-2026-255', Anonymous Referee #3, 02 Sep 2026
The paper “Statewide Forest Monitoring Data for California” by Anderson and Marvin summarizes and demonstrates the potential for the California Forest Observatory dataset to be a tool for users to understand forest structure and composition via remotely sensed data. While the potential for this research is very exciting and novel, the presentation in the manuscript presents some major concerns about how these estimates are truly feasible when much of the predicted metrics are below the canopy and there are little to no ground-truthed validations of predictions. From the applications described in the discussion and conclusion, there appear to be a lot of valuable and tractable uses for these data, however, as presented, it is hard to interpret the development and validation of these predictors. My major concerns are:
- It is hard to understand how you can effectively estimate canopy bulk densities and surface fuels from remotely sensed data based on the content of the paper. For example, the framework outlined in figure 3 is very cool, in theory, but how does LiDAR detect difference between ladder fuels that are separate from or part of the same main canopy to then separate it into the appropriate surface fuels category. While the Scott and Burgan (2005) categories are fairly standard, it is unclear how to map these on surface fuel categories that are collected in-situ making ground truthing / validation extremely difficult. Finally, using the scale of all of CA to test these estimates seems like a challenging place to start. CA is big and full of diverse ecosystems.
- The scope of predicting canopy height and cover at 3 m pixels does not feel very functional or tractable. My main concerns stem from the fact that (1) the training LiDAR misses much of the Sierra Nevadas (figure 2), (2) there is little to no ground-truthing for these variables, so how do you ground-truth these predictions and what in-situ data does it track with?
Minor comments:
- Define decisions and content more clearly throughout the text
- Surface fuels max height is not defined anywhere
- 135: why did you go with > 3% change?
- 147: canopy bulk density units not defined until line 163
- 153: LC
- 241: U-nets not until 263
- 278: how can you jump from lower to higher resolution model by concatenating the predictions?
- Why did you choose to compare with LANDFIRE? (figure 8)
- I agree with reviewer 1’s comment about section 4&5.
Citation: https://doi.org/10.5194/essd-2026-255-RC3
Data sets
California Forest Observatory: Statewide Forest Monitoring Data Christopher Anderson and David Marvin https://doi.org/10.1594/PANGAEA.989871
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 191 | 69 | 86 | 346 | 117 | 82 |
- HTML: 191
- PDF: 69
- XML: 86
- Total: 346
- BibTeX: 117
- EndNote: 82
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript titled “Statewide Forest Monitoring Data for California” attempts to develop a modeling framework leveraging Sentinel-1/2 and high-resolution PlanetScope data to map multiple aspects of fuel structure across California. The study itself provides an interesting and potentially useful exploration of the extent to which satellite products can inform our understanding of understory fuel structure. However, given that none of these satellite predictors directly observes through the canopy and that no independent validation is provided, the reliability of the resulting fuel maps when extrapolated across the highly heterogeneous landscapes of the entire state is unlikely to meet the operational goals outlined or intended by the study. Despite the claimed operational value, I view the current exercise more as a modeling experiment, as it lacks the robust and comprehensive assessment needed to demonstrate that these remotely sensed estimates can provide information comparable to intensive field-based measurements. Specifically, I have the following major concerns, followed by more detailed line-by-line comments.
Line-by-line comments: