the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Dense labelled time-series for mapping European forest disturbance agents
Abstract. Attributing disturbance agents to canopy mortality in European forests remains difficult due to sparse, heterogeneous, and often single-agent reference data. We present DISFOR (Viehweger et al., 2026), a uniformly re-interpreted ground-truth dataset of forest disturbance agents, designed to serve as training data for multi-temporal classification and analysis of disturbance agents with Sentinel-2 time-series. The dataset comprises 3,822 unique sample points, each defined at the 10×10 m pixel level, labelled by disturbance event and agent and fully temporally segmented into consecutive events and forest states for the years 2015–2024. Labels follow a three-level hierarchical scheme that supports analyses from broad "alive vs. disturbed" partitions to specific agents such as bark beetle, windthrow, wildfire, and salvage logging. Samples were drawn from multi-source ancillary data (e.g. EFFIS, FORWIND, Copernicus EMS, and regional forestry datasets) and consistently re-interpreted using an open-source interface and generally available primary data like Sentinel-2 and very high resolution imagery. For each sample we provide interpreter confidence, cluster identifiers to capture spatio-temporal autocorrelation, and additional metadata. Alongside interpreted samples, we release two Sentinel-2 data products tailored to complementary use cases: (i) a tabular single pixel reflectance time series between 2015–2024 and (ii) georeferenced 32×32-pixel image chips centred on the sampled points suitable for computer vision applications, with Python utilities for reproducible data loading and filtering. The dataset is suited for training and calibration of both change detection and agent attribution algorithms at sub-annual resolution, it supports training on single timestamps as well as on time-series, and facilitates studies that integrate spectral dynamics with spatial context. By harmonizing multiple first-level disturbance agent products and providing dense, temporally explicit labels, this resource lowers a major barrier to developing European forest disturbance and recovery monitoring.
- Preprint
(2496 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-185', Loïc Dutrieux, 29 May 2026
-
RC2: 'Comment on essd-2026-185', Anonymous Referee #2, 14 Aug 2026
The manuscript and dataset by Viehweger et al. presents a novel reference dataset on tree cover changes for Europe that includes timing and causal agents. As such, it presents an important addition to existing datasets, and I applauded the authors for investing the time to create this dataset. The novelty and quality certainly justify consideration for publications in ESSD. That said, I have several major and minor points that need to be addressed before publication.
Major points:
- First: What is the definition of a disturbance? A disturbance – in the classical sense – describes a relatively abrupt loss of biomass that allows for the establishment of new vegetation from freed resources (mostly light). This is often also described as canopy mortality or tree cover loss. As thinning, drought and defoliation are included, I wonder whether the dataset mixes tree loss events with changes in leaf area that do not lead to the death of trees. A clear definition (at best ecologically motivated) is needed.
- Second: How did you interpret routine versus non-routine management interventions? How did you make sure that management interventions were not triggered by bark beetle and thus should be classified as salvage logging? Already a single bark beetle tree can force managers to cut a whole stand to prevent spread (sanitation cutting). I doubt that this can be seen in the S2 time series.
- Third: Related to my first two points, the definition and inclusion of segments and events is not properly described, nor how they were separated during interpretation. This relates mainly to Table 1, for which I have several remarks:
- Re-vegetation is an odd name. There will always be vegetation establishing on disturbed sites. Maybe „recovery“ or „post-disturbance“ is a more suitable term?
- When does revegetation end? Once the canopy is closed? How is canopy closing defined?
- How did you check whether a harvest was planned or not? There could have been a single bark beetle tree that triggered harvest.
- Why not use unplanned harvest, but salvage logging as name? There is more than salvage cutting, sometimes foresters do sanitation cutting to prevent spread, even though the trees are not infested yet. This is also unplanned but not salvage logging.
- Bark beetle can be relatively short (1-2 years), but decline is often used for slower processes. I would reconsider this naming choice.
- Gypsy moth often does not kill trees, as trees can resprout. Did you check this?
- Is drought a disturbance agent? How did you verify this? Often drought incites biotic disturbances, which cannot be ruled out and would overlap with the biotic category.
- Overall, a bit more description would help to understand the different categories. Maybe add a description column to the table with example figures?
- Fourth: Disturbance timing relates to the first time the disturbance was evident in sentinel-2, but this might be much later than the actual disturbance event. The authors acknowledge this at some point in the manuscript, but this could be highlighted a bit more.
Minor points:
L. 45: The „isn’t“ is colloquial, please use “is not”.
L. 51: Rephrase to: „Because of this, there is…“
L. 62: Maybe don’t name the state of vegetation here already, as it can be confusing without clear definitions.
L. 66: What are short-duration disturbances and how are they characteristic? Also unclear in the second research gap stated below. Short duration can be anything from a few seconds to 1-2 years. In the US-American literature, short duration disturbances characterized disturbances <5 years, but I believe the authors refer to shorter time frames. Clear definitions are needed.
L. 67ff: Please write out the last paragraph, instead of using bullet points.
L. 94: Consider rewording into high-severity and low-severity disturbances, as stand-replacing refers to a spatial unit (a stand), but your interpretation is at the pixel level. Also, please give clear definitions of what constitutes a stand-replacing/non-stand-replacing disturbance.
L. 124: Were the tiles selected randomly or subjectively? Any inclusion criteria (i.e. forest cover) used? From eyeballing I would think they’re a bit too regularely spaced for random selection.
L. 126: Did you sample equalized by disturbance/non-disturbance stratum? Can you please provide the stratum weights? Those are needed for estimation. Generally, a bit more context on the sampling design would be warranted.
Section 2.2.: Many line breaks and one-sentence paragraphs. Don’t break text into too many paragraphs but find a good flow within paragraphs, with paragraphs guiding the reader from one message to the next.
L. 160: Bark beetles can also be relatively fast. The distinction between short and long disturbances is a bit arbitrary and never really defined, nor supported by literature. Also, if the event changes the canopy, it should be considered a change event, even though it changes it over several years (like in unmanaged bark beetle).
L. 206: Please rephrase „used first level data“ into „the first level data used“; otherwise data was bought second-hand ;-)
L. 231ff: There is some confusion between model accuracy and map accuracy. You will need an independent sample following a proper probabilistic sample design for assessing map accuracy (i.e. estimating the probability of any a pixel being correct). I suggest using the terms model and map accuracy specifically.
Figure 6: The increase in planned harvests after 2018 makes we wonder whether this class actually includes a lot of sanitation (and potentially salvage) logging. What other reason would result in an increase in harvest after one of the biggest waves of bark beetle in recent history.
Citation: https://doi.org/10.5194/essd-2026-185-RC2
Data sets
DISFOR Jonas Viehweger et al. https://doi.org/10.57967/hf/7983
Model code and software
DISFOR – Python library Jonas Viehweger https://github.com/JR-DIGITAL/DISFOR
DISFOR – Python library documentation Jonas Viehweger https://jr-digital.github.io/DISFOR/
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 303 | 75 | 20 | 398 | 21 | 22 |
- HTML: 303
- PDF: 75
- XML: 20
- Total: 398
- BibTeX: 21
- EndNote: 22
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Very interesting data compilation and re-interpretation effort that will certainly greatly benefit the community as data with this type of thematic details are either scarce or highly heterogeneous in quality. I have a few overall comments and some more minor remarks linked to sections or statements made in the manuscript.
Coverage:
It could have included at least two more datasets:
Details on labeling process:
The manuscript includes little details on the labeling process. I would have like to see illustrations and textual descriptions of spectro-spatio-temporal aspects of the various dynamics considered in the dataset. Some kind of feedback on the ideal visual interpretation parameters (color composites, vegetation indices) would be a nice addition too. Finally further discussion of some specific labeling challenges would make the data creation process fully transparent.
Temporal conceptualization of classification framework:
The manuscript discusses how various dynamics are either mapped to temporal segments of abrupt events, but it's not clear how this conceptualization is later handled in the dataset. I see only temporal segment labels in the dataset fields description.
Minor remarks:
Line 30: The community often makes bark beetle its own class and consider it to be the only type of biotic disturbance. It surely is by far the largest biotic agent but it's not the only one.
Line 41: metadata format heterogeneity is true and datasets cannot be ingested in their raw form in machine learning pipeline, but given data scarcity, this is hardly a blocking issue
Line 42: DEFID2 is in theory broader in scope than just insects; it includes all pests and diseases
Line 61-66: I believe the process you described can be characterized as augmentation
Line 69: You use "training data"; good to make it clear whether it's indeed intended for training, or if it is a more generic "reference dataset"
Line 96: It would be good to have some idea of the amount of sample for which the visual re-interpretation did not confirm the agent or even the disturbance. In my experience compiled datasets are widely heterogeneous in quality
Lines 104-110: I'm not entirely sure it is relevant to name the projects at this stage
Line 134: How were local news reports discovered?
Line 138: FORWIND has not been updated in a long time and marginally overlap with the sentinel era
Line 147: In my opinion you're missing an introductory paragraph here detailing the general idea and key assumption behind this app (visual interpretation potential, simultaneous spatial and temporal visualization, ergonomic, etc).
Figure 2: The right most explanation bubble explains that the sample footprint is not the same. Why?
Line 160: Dieback is a common term used instead of decline
Line 161: Stretching the concept of segments/events, not only disturbances would fit into events, but any transition between two segments, such as a stabilization at the end of a regrowth period.
Line 164: I do not understand why you'd need two labels per segment
Perhaps conceptual figure illustrating segments/labels/events would help here
Line 210: I think this was already mentioned a few paragraphs earlier
Line 235: Fully agree, I would even further insist on that. Because of unknown or non-zero inclusion probabilities, the dataset cannot be used for assessment of mapping accuracy. It can be used for model performance assessment though, but that does not necessarily translate to mapping accuracy.
Line 247: OK, but which baseline was used. Differences are marginal between baselines, but still good to be exhaustive
Line 248: How do you justify applying the offset but not the scaling factor?
Line 259: Not sure it makes sense to use COG for a 32 pixels chip that is probably made of a single data block. That's fine, the greater includes the lesser.
Lines 303-309: These are new elements; it's uncommon to mention things that have not been discussed previously in a conclusion
Loïc Dutrieux