the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
The global forest diameter spectrum using a machine learning approach
Abstract. Global forest assessments assist climate policy development, ecosystem science, and conservation planning, yet they rely on biomass and canopy data that do not explicitly represent the stand structural attributes derived from tree diameter measurements. This limits the ability to compare size-related structure and within-stand heterogeneity at large spatial scales. Here we present a global, spatially explicit dataset of stand-level tree diameter structure for forest cover in 2020 at 0.027° (~3 km) resolution, based on 1,203,524 georeferenced forest inventory plots comprising 54.6 million trees (≥10 cm DBH) integrated with more than 50 environmental and satellite-derived covariates into machine learning models. The dataset provides the first globally consistent maps of three complementary diameter-based metrics: arithmetic mean diameter (Dmean), quadratic mean diameter (Dqm), and the coefficient of variation of diameter (Dcv), representing average tree size, large-tree dominance, and within-stand size variability, respectively. Model performance of the ecozone-specific Random Forest framework ranged from R² = 0.41–0.82 (RMSE = 3.91–4.63 cm) for Dmean, R² = 0.43–0.83 (RMSE = 4.38–5.27 cm) for Dqm, and R² = 0.47–0.62 with (RMSE = 0.10–0.13) for Dcv across different forest ecozones. By jointly quantifying central tendency and variability in tree size, the dataset revealed spatial patterns of forest structural organization not captured by existing biomass or canopy-height products. It provides a consistent baseline for cross-biome comparison of forest structure, supporting parameterization and evaluation of vegetation and Earth system models, while offering an independent benchmark for remotely sensed structural proxies. Furthermore, it enables spatial assessment of stand structural attributes, including large-tree dominance and structural complexity, facilitating integration of diameter-based structure into global analyses of carbon dynamics and ecosystem functioning.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Earth System Science Data.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(1173 KB) - Metadata XML
-
Supplement
(814 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on essd-2026-226', Jonas Lembrechts, 27 Jun 2026
-
AC1: 'Reply on RC1', Ankita Mitra, 26 Jul 2026
We sincerely thank the reviewer for the positive evaluation of our work and for the thoughtful, constructive comments and suggestions regarding data quality, model robustness, and validation. We greatly appreciate the time and effort devoted to reviewing our manuscript, and the feedback has helped us improve its clarity, rigor, and overall quality. Detailed point-by-point responses are provided in the accompanying file, "Comments_Responseletter_R1.pdf." In this document, each comment is addressed individually, with the reviewer's comments shown in black, our responses shown in green, and the corresponding revisions to the manuscript also highlighted in green to facilitate clarity and ease of review.
Regards.
-
AC1: 'Reply on RC1', Ankita Mitra, 26 Jul 2026
-
RC2: 'Comment on essd-2026-226', Anonymous Referee #2, 06 Oct 2026
This data paper develops three global forest diameter products by combining forest inventory data with several geospatially explicit layers representing vegetation characteristics, topography, soils, climate, etc., using a machine learning approach. The products also provide spatially explicit uncertainty estimates. The manuscript is well written and concise, and the methodology is straightforward to follow. I expect the resulting datasets to be useful to a broad community.
I have two thoughts that the authors may wish to consider:
1. Training models by ecozone versus ecozone × continent
My understanding is that the machine learning models are trained separately for each ecozone. These ecozones can span multiple continents. For example, boreal forests occur across both North America and Eurasia. However, forest composition can differ substantially among continents within the same ecozone. For example, spruce may be more dominant in parts of North America, whereas pine and larch may be more important in Eurasia. These differences in species composition could translate into substantial differences in forest structural characteristics, including diameter distributions, as well as differences in the environmental variables that best predict them.
The same issue may potentially apply to temperate and tropical forests. I therefore wonder whether training separate models for each ecozone–continent combination could improve predictive performance compared with the current approach. Is this something the authors considered? If so, could the authors briefly discuss what they found? If not, do the authors think that this would be a potentially useful sensitivity analysis?
2. Number and importance of predictor variables
I understand that, for an ESSD paper, the primary emphasis is on producing spatially explicit datasets rather than on explaining the mechanisms underlying spatial patterns. Nevertheless, I wonder whether the authors could provide some additional information on the number and importance of the predictor variables.
In particular, how does model performance change as the number of predictors is reduced? For example, do the best 10 or 20 predictors achieve performance close to that obtained using the full set of 50+ predictors? Conversely, is there a point at which adding additional predictors provides little further improvement in predictive performance? It would also be interesting to know whether this relationship varies spatially or among ecozones.
Relatedly, which predictor variables are most influential, and how does their importance vary spatially? I appreciate that correlated predictors can make variable-importance measures difficult to interpret, but even a broad assessment of predictor importance could provide useful context for interpreting the resulting products, particularly when considered alongside the uncertainty layers.
Finally, retaining only the most influential predictors could potentially reduce computational costs and make the models easier to interpret. Could the authors comment on how they considered the trade-off between model performance, computational complexity, and interpretability when selecting the predictor set?
Citation: https://doi.org/10.5194/essd-2026-226-RC2
Data sets
Rasters- Global prediction of tree diameter across forest cover Ankita Mitra https://figshare.com/s/689e4be80a63d05d5189
CSV-Grid predictions of forest diameter structure Ankita Mitra https://figshare.com/s/91c6cc4f2b9757d92da8
Metadata associated with the study Ankita Mitra https://figshare.com/s/34c21c03883a01ef708f
Global dataset used to train the models for tree diameter study Ankita Mitra https://figshare.com/s/c795146ea9e8aa42f81f
Model code and software
RF models for the tree diameter study Ankita Mitra https://figshare.com/s/60eac94e4c64a95c6d13
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 792 | 325 | 128 | 1,245 | 109 | 151 | 130 |
- HTML: 792
- PDF: 325
- XML: 128
- Total: 1,245
- Supplement: 109
- BibTeX: 151
- EndNote: 130
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This is a nice global collaborative study bringing together a mindboggling number of tree DBH-measurements from across the globe, and using these to make 3 km-resolution global maps for a few key metrics.
This will be a highly valuable product for many as the need for more forest structural data increases. It’s a bit of a pity that the resolution is not higher though (1 km would have brought it in line with so many other datasets). Perhaps I’m missing a bit of discussion in the paper why a higher resolution was not feasible, despite the number of observations. I assume it has to do with the relatively low R2-values, which indicate the presence of a lot of noise and unexplained variance. Perhaps a discussion of that noise would also be helpful: what makes the data so noisy? Is it disturbance or land use history?
L1: I think ‘using a machine learning approach’ would read better
Abstract: I wonder if it makes sense to highlight some results as well. For example, do the findings confirm existing literature that temperature forests have the largest mean diameters (Dmean), or is that new information? In both cases it could be a nice thing to mention in the abstract so it gets very concrete what kind of information one can get out of this dataset. Related, wouldn’t it make sense to have one paragraph at the end of the paper to discuss if the observed patterns are at least roughly in line with the literature (and thus cite some references), or if these patterns are novel? That would also help readers to find ecological underpinnings of the observed patterns whenever they want to use the maps.
L514-515: it’s not intuitive to me what the difference between a ‘fixed-area plot’ and an ‘inventory unit’ is (besides the size). Why are they named such? Does an inventory unit not have a fixed-area? Is it not a plot?
L516: how exactly is data expanded? If you have DBHs for a number of trees in a plot, how do you get the DBHs from the trees in the surrounding area from there? Are the parameters (e.g., mean) sensitive to plot size and, if so, how does the scaling solve that if you don’t have information on the surrounding area?
L533: add an example of such a name, perhaps?
L549: I assume the exclusion based on the 10 cm threshold does not apply to the summing of multi-stem diameters? The order of sentences now makes it slightly misleading.
L551: how often was disturbance information available? Is there a risk that many observations influenced by disturbance are still included, because the information was not available? If this information was rare, satellite-based proxies of disturbance may have to be used.
L618: how did you screen for multicollinearity?
L639: values of what? Of all predictor layers?
L639: I’m a bit worried about the idea of extracting values only at the centroid of the pixel. Wouldn’t the mean of the pixel be more representative? What if now for example the centroid of the pixel just has a forest gap – would then the forest height (a 30 m resolution variable) be entirely unrepresentative for the pixel?
Table 2: could you make a clearer distinction between the categories? Now the covariates of a certain category are listed above the category-name (top-aligning might help, or adding a horizontal line between categories).
L677: do you have references that argue for modelling separately per ecozone? Intuitively, I would say that by reducing the environmental range in your dataset (which you do by splitting in ecozones) and the number of datapoints, you reduce the explanatory power of your model. Also, these machine learning models deal well with non-linearity, so why is the splitting necessary?
L693-695: separately across each model technique, or averaged across techniques?
L727: and which function for the OLS?
Figure 2: I’m surprised to see such high values – a lot of green colours e.g. in Europe, with values far above the mean and standard deviation as reported in Table 1.
Figure 2: do we see abrupt changes between ecozones, stemming from the running of separate models?
Finally, I would love to see a map of data coverage (i.e. observation points) in the main text.
Overall, a great dataset and a massive effort, congratulations!
Kind regards,
Jonas Lembrechts