the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
ChinaTCC30: An annual 30 m tree canopy cover dataset for China from the 1970s to 2024
Abstract. Tree canopy cover (TCC) is a key biophysical indicator of forest structure and function, and long-term fine-resolution TCC data are essential for quantifying carbon cycling, monitoring forest succession, and supporting forest management under ongoing climate and land-use change. However, accurate long-term national-scale TCC mapping remains challenging because strong spatial heterogeneity requires large and representative reference samples that are difficult to obtain consistently across space and time. This challenge is particularly evident in China, where early historical TCC baselines and long-term annual TCC records remain insufficient. To address this gap, we developed an integrated framework that combines active learning, Random Forest, and stand-growth-based historical backtracking to reconstruct China TCC at 30 m (ChinaTCC30) for 1975, 1980, and annually from 1985 to 2024. Specifically, the historical TCC for 1975 and 1980 was reconstructed from MSS imagery using pseudo-samples generated under limited reference conditions, while annual wall-to-wall TCC maps for 1985–2024 were produced from Landsat imagery using visually interpreted TCC samples. Pixel-level uncertainty was further quantified using multi-model predictions. Validation against multiple independent reference datasets demonstrated the robustness and reliability of the resulting dataset, including visually interpreted samples (R² = 0.83, RMSE = 15.14 %) and multi-period National Forest Inventory data (6th – 9th NFI; R² = 0.66 – 0.82, RMSE = 14.68 % – 18.81 %). Compared with existing products, the resulting ChinaTCC30 dataset shows good national-scale spatial consistency while better capturing fine-scale spatial heterogeneity and long-term temporal continuity. As a spatially explicit TCC dataset spanning nearly five decades, it provides valuable support for forest monitoring, climate change research, and forest conservation and management, but also offers a cost-effective reference for TCC mapping at national and regional scales. The ChinaTCC30 dataset is available via an open-data repository: Part 1 (1975, 1980, and 1985–1988; DOI: https://doi.org/10.5281/zenodo.21194015), Part 2 (1989–1994; DOI: https://doi.org/10.5281/zenodo.21195911), Part 3 (1995–2000; DOI: https://doi.org/10.5281/zenodo.21199497), Part 4 (2001–2006; DOI: https://doi.org/10.5281/zenodo.21245669), Part 5 (2007–2012; DOI: https://doi.org/10.5281/zenodo.21262725), Part 6 (2013–2018; DOI: https://doi.org/10.5281/zenodo.21263677), and Part 7 (2019–2024; DOI: https://doi.org/10.5281/zenodo.21265832) (Yu et al., 2026a, b, c, d, e, f, g).
Competing interests: The authors declare no competing interests. Zhen Yu is a member of the Editorial Board for Earth System Science Data but had no role in the peer review, editorial management, or final decision regarding this manuscript.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(23557 KB) - Metadata XML
-
Supplement
(11495 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-546', Anonymous Referee #1, 20 Aug 2026
-
RC2: 'Comment on essd-2026-546', Anonymous Referee #2, 20 Aug 2026
This manuscript presents an annual 30 m tree canopy cover dataset for China from the 1970s to 2024. The methodology is systematic and the dataset is of high value. However, please address the following issues before publication.
1. The interpretation samples are mainly derived from VHR imagery of 2015–2024, whereas the product provides annual mapping for 1985–2024. Please clarify whether all annual models share the same sample set, and discuss the validity of the underlying assumption that the spectral–TCC relationship is temporally stable and transferable across years, particularly given the radiometric differences between early TM and recent OLI imagery.
2. The GEDI footprint is 25 m, while the product is 30 m. Please explain the spatial matching/aggregation method between the two, and whether the GEDI geolocation error was considered.
3. The pseudo-samples for the 1975/1980 baselines are based only on natural even-aged pure stands, whereas many of China's forests are dominated by plantations and mixed stands. Please discuss in the limitations the potential systematic bias this may introduce in major afforestation regions (e.g., the Three-North and Grain-for-Green areas).
4. Section 2.3 does not describe the trend analysis method, yet Section 3.5 and Fig. 10 use Sen's slope and significance testing (p < 0.05). Please add the trend detection method used in the Methods.
5. The compositing windows are 1972–1977 for the 1975 baseline and 1978–1984 for the 1980 baseline, which are relatively wide. Please discuss how these windows affect the temporal representativeness of the baseline years.
6. Fig. 10 is titled 1985–2024, while Section 3.5 discusses 1975–2024. Please clarify whether the pixel-level trend analysis includes the 1975/1980 baselines, and keep this consistent throughout.
7. Section 3.5 states a "westward shift of the center-of-gravity trajectory," while the Conclusion states a "southwestward shift." Please make these consistent.
8. The title of panel e is missing in Figure 12.
9. Figure numbers are missing in Section 4.2 (Line 670-671); "Fig. a9–10," etc. should read "Fig. 16 a9–10."?
10. The caption of Fig. 8, "across nation four inventory periods," should be revised to "across four national inventory periods."
11. Typo: "Radom Forest" in Section 3.3 (Line 340) should be "Random Forest".
Citation: https://doi.org/10.5194/essd-2026-546-RC2 -
RC3: 'Comment on essd-2026-546', Anonymous Referee #3, 20 Aug 2026
General comment
The dataset has potential value for long-term tree canopy cover monitoring in China. However, several methodological and data-consistency issues should be clarified before the reliability and originality of ChinaTCC30 can be fully assessed. My main concerns relate to the distinction from the existing CATCD framework, the historical reconstruction, temporal transferability of the model, auxiliary-data consistency, and temporal breakpoint correction.
Major Comments1. The methodological novelty relative to CATCD should be clarified more explicitly.
The 1985–2024 mapping framework shares several key components with Cai et al. (2024), including nationwide 30 m VHR interpretation plots, Random Forest with 500 trees, five-fold ensemble prediction, averaging of the five prediction layers, and the use of their standard deviation as an uncertainty layer. CATCD had already implemented this RF ensemble framework. ChinaTCC30 introduces additional components, including the 10 × 10 interpretation grid, KDAL, historical backcasting, ALCC predictors, and temporal breakpoint correction.
I suggest that the authors explicitly distinguish which methodological components are inherited or adapted from CATCD and which are newly developed in this study. A concise comparison with Cai et al. (2024) would help readers assess the methodological contribution of ChinaTCC30.
Relevant manuscript locations: Introduction, Lines ~95–105; Sect. 2.3.1–2.3.2; Sect. 2.3.6.
2. The reconstruction of the 1975 and 1980 historical baselines requires additional methodological detail.The manuscript states that historical TCC pseudo-samples were generated by back-calculating stand age and stem density and combining these variables with climate and topography. However, several important details remain unclear, including how “stable natural even-aged pure stands” were identified, the source and reference year of the stand-age data.
In addition, 1,700 MSS scenes acquired during 1972–1977 and 1978–1984 were composited to represent the nominal 1975 and 1980 baselines, respectively. Please clarify the compositing procedure. For example, were observations from all years weighted equally, how were multiple observations combined, and what fraction of the final mosaics was contributed by each acquisition year? This information is important because the “1975” and “1980” maps represent multi-year epochs rather than observations from single years.
Relevant manuscript locations: Sect. 2.3.3, Lines ~220–240; Sect. 2.3.6, Lines ~263–273.
3. More information is needed on the fitting of the stand age–stem density relationships.Supplementary Tables S4–S5 provide empirical age–stem-density relationships for nine species or genera, but the sample size used to fit each relationship, the optimization method, and the parameter ranges or boundary constraints are not reported.
Several fitted coefficients are exactly 100000 or 15000. For example, several power models have a coefficient of 100000, while the exponential models for Pinus densata and Betula platyphylla have coefficients of exactly 15000. These repeated exact values raise the possibility that some fitted parameters reached imposed optimization bounds.
Please report the sample sizes, fitting algorithm, initial values where relevant, and parameter bounds, and confirm whether any fitted parameters reached the prescribed limits. This is particularly important because these relationships are subsequently used to reconstruct historical stand conditions.
Relevant location: Supplementary Tables S4–S5, pp. 3–4.
4. The temporal coverage of the auxiliary datasets is inconsistent with their stated use.Table S1 reports ALCC for 1985–2022 and TerraClimate for 1985–2024. However, the manuscript states that historical climate variables were used for 1975 and 1980, and ALCC is also used to generate zero-TCC pseudo-samples for the historical reconstruction. ALCC is additionally described as an auxiliary predictor in the annual mapping framework.
Please clarify:
- the actual climate dataset and years used for the 1975 and 1980 reconstruction;
- how ALCC information beginning in 1985 was used to identify persistent bare-land/water areas for 1975 and 1980;
- how the land-cover predictor was handled for 2023 and 2024, when ALCC is unavailable.In particular, the use of post-1985 land-cover information for constructing historical zero-TCC pseudo-samples should be discussed because it may introduce information from later periods into the historical reconstruction.
Relevant manuscript locations: Sect. 2.1.2, Lines ~125–135; Sect. 3.1, around Line 305; Supplement Table S1.
5. Temporal transfer of the model from recent reference labels to 1985–2014 has not been sufficiently evaluated.The 10,000 visually interpreted plots were derived predominantly from VHR imagery acquired during 2015–2024, and Landsat predictors were extracted from the corresponding label years. The resulting model is nevertheless used to generate annual TCC maps back to 1985.
This implies substantial temporal transfer from predominantly recent training labels to earlier Landsat periods, particularly 1985–2014. Although sensor harmonization may reduce spectral differences, it does not by itself demonstrate that the TCC–predictor relationship remains transferable across several decades.
Please provide a more direct evaluation of model performance outside the primary reference-label period, preferably using temporally independent observations from earlier years. At minimum, the manuscript should discuss this temporal extrapolation more explicitly and quantify performance for earlier versus later periods where reference data are available. I also suggest conducting a Landsat spectral-consistency assessment across the major sensor periods, using temporally stable reference targets or other suitable invariant samples. Such an analysis could help demonstrate whether the harmonized spectral feature distributions remain sufficiently consistent for transferring a model trained predominantly on recent observations to earlier decades.
Relevant manuscript locations: Sect. 2.2.1, Lines ~145–150; Sect. 2.3.6, Lines ~263–280.
6. The temporal breakpoint correction is not sufficiently described to be reproduced.The manuscript states that stable ALCC pixels and local temporal trajectories were used to identify and correct non-physical breakpoints. However, the breakpoint-detection algorithm, temporal window, significance threshold, definition of stable pixels, and correction equation are not provided.
This distinction is important because genuine forest disturbances, such as harvesting, fire, typhoons, or land conversion, can also produce abrupt changes in TCC.
Please provide the full correction procedure in the Methods and report, at least, the number or proportion of pixels corrected, the years most affected, and the typical correction magnitude. It would also be useful to demonstrate that genuine abrupt canopy disturbances are retained by the correction procedure.
Relevant manuscript location: Discussion, Lines ~630–655. This procedure would be more appropriate as part of the Methods.
Specific Comments1. Title and spatial/temporal resolution
The title describes ChinaTCC30 as an “annual 30 m” dataset from the 1970s to 2024. However, the Methods describe 60 m historical TCC for 1975 and 1980 and annual 30 m TCC only from 1985 onward. The Abstract also describes ChinaTCC30 as a 30 m product for 1975, 1980, and annually from 1985–2024.
Please make the title, abstract, Methods, and released-product description consistent and explicitly state whether the released 1975/1980 rasters were resampled from their native 60 m MSS resolution.
2. The trend and center-of-gravity analyses need methodological descriptions.The Results report pixel-wise significance at p < 0.05 and use mean Sen’s slopes, but I could not find a corresponding statistical-method description in the Methods section. Please specify the significance test used, and indicate whether temporal autocorrelation and the large number of simultaneous pixel-level tests were considered.
The center-of-gravity analysis is also presented without an equation. Please specify whether the centroid was weighted by TCC, forest area, or area-weighted TCC, and state the coordinate system used for the calculation.
There is also an internal inconsistency: the Results describe a westward shift, whereas the Conclusion reports a southwestward shift. Please verify the calculation and make the wording consistent.
Relevant manuscript locations: Results, Lines ~405–435; Fig. 9–10; Conclusion, Lines ~715–720.
3. Please clarify the numerical representation of the released uncertainty rasters.In targeted checks of the released uncertainty rasters for 1975, 1980, 1985, and 2024, the valid uncertainty values were stored as integers, with the inspected range approximately 0–8, and a large number of pixels had an uncertainty value of exactly 0.
The manuscript states that uncertainty is calculated as the standard deviation among five model predictions. Please clarify whether the standard deviations were rounded or quantized before storage, the units and scale factor of the released uncertainty rasters, and whether a value of zero truly indicates zero inter-model variability. A short numerical description of the uncertainty files in the Data Availability section or repository documentation would be useful.
Relevant manuscript locations: Sect. 2.3.6, Lines ~275–280; Data Availability, Lines ~700 onward.
4. The spatial dimensions of the historical and annual GeoTIFF files are not fully consistent.In the released data, the inspected 1975/1980 rasters have dimensions of:
237482 × 148426
whereas the inspected 1985 and 2024 rasters have dimensions of:
237483 × 148427.
The files have the same inspected grid origin and pixel size, but the historical files contain one fewer column on the eastern edge and one fewer row on the southern edge.
For a dataset intended for long-term pixel-wise analysis, it would be preferable for all rasters to use exactly the same grid extent and array dimensions. Please check and, if possible, harmonize the historical and annual grids. If the difference is intentional, it should be documented clearly in the repository metadata.
Relevant manuscript location: Data Availability, Lines ~700 onward.
5. Typographical error at Line 340“K-means-based diversity active learning-Radom Forest” should be corrected to “Random Forest”.
6. Number of accuracy metricsThe manuscript states that “five standard statistical metrics” were calculated, but only r, R², RMSE, and Bias are subsequently listed. Please revise “five” to “four” or add the missing metric.
Relevant manuscript location: Sect. 2.3.4, Lines ~243–247.
7. Please moderate causal interpretations of the observed TCC trends.The Results state that the positive TCC trajectory reflects the “notable effectiveness of afforestation and ecological restoration efforts.” However, the present analysis maps TCC changes rather than explicitly attributing them to individual conservation or restoration programs.
I suggest using more cautious wording such as “consistent with” or “may be associated with” ecological restoration unless a dedicated attribution analysis is provided.
Relevant manuscript location: Results, around Lines ~420–435.
8. Please justify the use and resampling of the climate data.The manuscript states that temperature and precipitation from TerraClimate were bilinearly resampled to 30 m and then used for plot-scale pseudo-sample construction and prediction. Because the native spatial resolution of TerraClimate is approximately 4.6 km, bilinear resampling to 30 m does not introduce new fine-scale climatic information.
Please explain why TerraClimate was selected instead of a higher-spatial-resolution climate dataset, if suitable alternatives are available, and clarify the intended role of these resampled climate variables in the 30 m modeling framework. It would also be useful to discuss whether the coarse native resolution of the climate data may limit the representation of local climatic gradients.
Relevant manuscript location: Sect. 2.1.2, Lines ~130–135.
Citation: https://doi.org/10.5194/essd-2026-546-RC3
Data sets
ChinaTCC30: An annual 30 m tree canopy cover dataset for China from the 1970s to 2024 (Part 1: 1975, 1980, and 1985-1988) Jinge Yu et al. https://doi.org/10.5281/zenodo.21194015
ChinaTCC30: An annual 30 m tree canopy cover dataset for China from the 1970s to 2024 (Part 2: 1989-1994) Jinge Yu et al. https://doi.org/10.5281/zenodo.21195911
ChinaTCC30: An annual 30 m tree canopy cover dataset for China from the 1970s to 2024 (Part 3: 1995-2000) Jinge Yu et al. https://doi.org/10.5281/zenodo.21199497
ChinaTCC30: An annual 30 m tree canopy cover dataset for China from the 1970s to 2024 (Part 4: 2001-2006) Jinge Yu et al. https://doi.org/10.5281/zenodo.21245669
ChinaTCC30: An annual 30 m tree canopy cover dataset for China from the 1970s to 2024 (Part 5: 2007-2012) Jinge Yu et al. https://doi.org/10.5281/zenodo.21262725
ChinaTCC30: An annual 30 m tree canopy cover dataset for China from the 1970s to 2024 (Part 6: 2013-2018) Jinge Yu et al. https://doi.org/10.5281/zenodo.21263677
ChinaTCC30: An annual 30 m tree canopy cover dataset for China from the 1970s to 2024 (Part 7: 2019-2024) Jinge Yu et al. https://doi.org/10.5281/zenodo.21265832
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 182 | 84 | 11 | 277 | 33 | 14 | 13 |
- HTML: 182
- PDF: 84
- XML: 11
- Total: 277
- Supplement: 33
- BibTeX: 14
- EndNote: 13
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This study develops ChinaTCC30, a long-term tree canopy cover dataset for China spanning 1975-2024 by integrating Landsat observations, visually interpreted samples, machine learning, active learning, and stand-growth-based historical reconstruction. Overall, the manuscript is well structured and addresses an important need for long-term, spatially explicit tree canopy cover information in China. The model shows generally good performance based on the reported validation results, and the comparisons with independent datasets and existing TCC products demonstrate the potential value of ChinaTCC30 for studying forest dynamics and ecological change. However, several concerns remain regarding the validation of the long-term time series, the reliability of the 1975 and 1980 historical reconstructions, the design and independence of the training and testing samples, and the treatment of temporal consistency and uncertainty. My major and specific comments are provided below for clarity.
Major comments:
Specific comments: