the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Constructing nationally comprehensive annual rice paddy maps for Madagascar from 2017 to 2025
Abstract. Reliable information on rice paddy cultivation is essential for food security planning in Madagascar, where rice accounts for more than half of the national caloric intake. However, the scarcity of ground reference data, high prevalence of small cropping areas, and the heterogeneity of agroecological zones have limited the production of high-resolution annual rice paddy maps across the country. Here, we present the first country-wide, annual rice paddy maps of Madagascar at 10 m resolution from 2017 to 2025. Our framework integrates three components: (1) phenology-based pseudo-label generation from 30 m merged Harmonized Landsat Sentinel-2 (HLS) time series, exploiting the flooding-to-greenup signal characteristic of transplanted rice paddy; (2) 10 m two-stage Random Forest classification on Google Satellite Embedding (GSE) annual features, refined through targeted augmentation with a small set of manually labeled samples from low-confidence regions; and (3) harmonic NDVI fitting to characterize annual cropping intensity. The pseudo-labels formed compact and clearly separated clusters in the GSE feature space across all nine years, and only five GSE dimensions consistently contributed to rice discrimination, indicating that pre-trained embeddings encoded phenologically meaningful information for rice paddy. The two-stage classifier achieved an overall accuracy (OA) of 91.2 %, a precision of 99.0 %, a recall of 83.2 %, and an F1-score of 0.904 on independent validation samples, outperforming the SAR-based benchmark product by a wide margin. Our maps indicated an increase in mapped rice paddy extent from 886,112 ha in 2017 to 1,195,766 ha in 2025 (+34.9 %), with most gains occurring along the margins of existing paddies rather than in new frontiers. This study demonstrates that combining phenology-based pseudo-labels with pre-trained satellite embeddings provides a scalable, near label-free approach for rice paddy mapping in data-scarce regions. The data are publicly available at https://doi.org/10.5281/zenodo.20654510 (Kwon et al., 2026).
- Preprint
(4671 KB) - Metadata XML
-
Supplement
(2641 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- CC1: 'Comment on essd-2026-518', Koen De Vos, 11 Sep 2026
-
RC1: 'Comment on essd-2026-518', Anonymous Referee #1, 14 Sep 2026
Summary.
This study develops a 10-m annual rice paddy mapping framework for Madagascar from 2017 to 2025. The method combines phenology-guided pseudo-label generation from HLS time series with Google Satellite Embedding (GSE) features and a two-stage random forest classifier, in which manually interpreted hard samples are further used to refine low-confidence areas. The study addresses an important challenge in large-scale crop mapping under limited training data and provides a potentially valuable long-term rice dataset for Madagascar. However, several issues regarding the near-label-free claim, training–validation independence, and benchmark comparability should be further clarified before the robustness and generalizability of the framework can be fully assessed.
Major Comment 1.
The manuscript repeatedly characterizes the proposed workflow as “near label-free” and emphasizes that it avoids costly field reference data. However, the final second-stage classifier is refined by adding 1,000 manually labeled hard samples per class per year from low-confidence regions. Given three target classes and nine years (2017–2025), this design may involve up to 27,000 manually interpreted samples. More importantly, the reported F1-score increases markedly from 0.728 for the first-stage pseudo-label-only classifier to 0.904 for the final second-stage classifier. Therefore, the strongest reported accuracy is not achieved by the near-label-free component alone, but only after substantial targeted manual intervention. This raises a fundamental question about whether “near label-free” accurately characterizes the final mapping framework and how much of the reported performance can actually be attributed to the pseudo-label-based component.
Major Comment 2.
The manuscript states that the independent validation samples were manually interpreted using high-resolution Sentinel-2 RGB imagery and HLS-derived NDVI/LSWI time series, supported by local expertise. The second-stage hard samples were also manually interpreted using high-resolution imagery and NDVI/LSWI time-series information, and they were specifically selected from low-confidence or error-prone areas identified by the first-stage model. Because the training refinement and validation procedures rely on highly similar interpretation sources, strict separation between training and evaluation is essential. At present, the manuscript does not provide enough information to determine whether the hard samples and validation points are spatially and temporally independent, whether the same interpreters or reference information were used, or whether validation results influenced hard-sample selection. Consequently, it is difficult to assess whether the reported OA of 91.2%, precision of 99.0%, recall of 83.2%, and F1-score of 0.904 represent a fully independent estimate of generalization performance.
Major Comment 3.
The manuscript compares the proposed 10 m annual rice maps with the 20 m Africa rice distribution map for 2023 and repeatedly refers to this dataset as a “SAR-based benchmark.” However, according to the manuscript’s own description, the benchmark product was generated using both Sentinel-1 SAR and Sentinel-2 optical data, together with object-based segmentation and Random Forest classification. Referring to it simply as “SAR-based” may therefore be misleading. In addition, the proposed product covers 2017–2025 at 10 m, whereas the benchmark is a 2023 product at 20 m. The manuscript does not sufficiently explain how spatial resolution, temporal scope, validation sampling, and target-class definitions were harmonized before calculating the very large performance difference. The subsequent attribution of the benchmark’s poorer performance mainly to SAR viewing geometry, terrain shadow, overlapping orbits, and backscatter confusion is also difficult to isolate from other differences in training data, algorithm design, spatial resolution, and class definition. These factors make the fairness of the comparison and the causal interpretation of the benchmark’s lower accuracy uncertain.
Citation: https://doi.org/10.5194/essd-2026-518-RC1 -
RC2: 'Comment on essd-2026-518', Anonymous Referee #2, 15 Sep 2026
The major contribution of this study is the development of annual rice paddy maps for Madagascar, an important rice-producing country in Africa where high-resolution rice datasets remain relatively limited. In my view, any effort to improve spatially explicit rice information in Africa deserves attention, particularly when the resulting products are made publicly available. However, I still have several concerns regarding the method used to estimate cropping intensity and several other methodological aspects. My specific comments are as follows.
1. My major concern is the definition and estimation of cropping intensity. If I understand the method correctly, the authors first identify rice pixels and then estimate cropping intensity by fitting the HLS NDVI time series and counting vegetation peaks. Pixels with one NDVI peak are classified as single-cropped rice, whereas those with two peaks are classified as double-cropped rice. However, what is actually detected here is the number of vegetation growth cycles within pixels classified as rice, rather than the number of rice cultivation cycles themselves. These two concepts are not necessarily equivalent. Even if rice-wheat rotation or rice-other crop rotations are uncommon in Madagascar, I do not think that rice cropping intensity can be defined directly from the number of NDVI peaks. Each identified rice cycle should be supported by rice-specific phenological evidence rather than vegetation peaks alone. I therefore suggest that the authors reconsider and improve the current method used to estimate cropping intensity.
2.The validation dataset described in Section 2.3.4 requires substantially more detail. For a data product paper, the construction of the reference dataset is a central part of the study and should be sufficiently transparent for readers to judge its reliability and reproducibility. The manuscript currently states that samples were manually interpreted using Sentinel-2 RGB imagery, HLS-derived NDVI and LSWI time series, and local expertise, but the actual labeling protocol remains unclear. For example, what specific visual or temporal criteria were used to label a sample as rice or non-rice? How were NDVI and LSWI used together during the interpretation process? These details are necessary, particularly because the reported accuracy is one of the major strengths claimed for the dataset.
3. I have a significant concern regarding the initial candidate-area extraction step. The workflow first trains a land-cover classifier using Dynamic World-derived samples and GSE features, and only pixels classified as cropland or flooded vegetation are retained for subsequent rice mapping. This creates a potentially important source of irreversible omission error. If a true rice pixel is classified as grass, shrub, water, or another land-cover class during this first step, it can never be recovered by the subsequent rice classifier. Although Section 2.3.4 states that the land-cover classification was evaluated using independent reference samples, what matters for this masking step is specifically its recall for true rice pixels. I therefore suggest reporting the proportion of independent rice reference samples that fall within the cropland/flooded-vegetation candidate mask. This would directly quantify the omission error introduced before the rice classifier is applied.
4. The first-stage classifier is indeed trained using automatically generated pseudo-labels, but the final product relies on manually interpreted hard samples from low-confidence areas.Therefore, I have some concerns about repeatedly describing the framework as “near label-free.” In my view, this wording somewhat overstates the degree to which the final mapping process is independent of manually labeled data.
5. I suggest revising Section 2.3.5, currently titled “Benchmark rice mapping product.” The Africa-wide 20 m rice product is useful for comparison, but I do not think it needs to be framed as a “benchmark” dataset, particularly when the manuscript later emphasizes that the proposed product substantially outperforms it on the authors’ validation samples.
6. Some parts of the Methods contain unnecessary textbook-style explanations that interrupt the scientific flow. For example, the sentence stating that “The Random Forest algorithm is an ensemble technique that produces predictions by aggregating the outputs of multiple decision trees for classification or regression tasks” is not necessary for an ESSD readership. Random Forest is a well-established method, and a citation is sufficient.
Citation: https://doi.org/10.5194/essd-2026-518-RC2
Data sets
Annual rice paddy maps of Madagascar at 10 m resolution (2017-2025) Ryoungseob Kwon, Youngryel Ryu, Giacomo De Nicola, Hervet Randriamady, Oladimeji Ezekiel Mudele, Tinashe Tapera, Hyeyoung Jo, and Christopher D. Golden https://zenodo.org/records/20654510
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 285 | 141 | 106 | 532 | 67 | 140 | 105 |
- HTML: 285
- PDF: 141
- XML: 106
- Total: 532
- Supplement: 67
- BibTeX: 140
- EndNote: 105
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
An interesting work that nicely reflects the recent developments in national crop type/rice mapping research. While I agree with most of the findings of the authors, it would be interesting to further investigate how the authors methodological choices indeed improve on existing national-scale maps.
A comparison with the continental rice distribution map shows clear improvement, but this is expected upfront because the difference in scale (continental vs. national) and thereby representation of data used for training. The authors identify other existing national-scale maps such as the one we made a couple of years ago (De Vos et al. (2023)). The latter map is definitely biased towards very clear rice paddies and is missing out on smaller and further-from-the-road fields because of parts of its sampling design (particularly the ones only using crowdsourced data). It would be very informative to understand how this mapping model compares to our earlier maps and whether this study could highlight for which areas/regions the improvements are most visible. It would therefore be interesting to see the training samples and/or national-scale maps to be used as additional benchmark. Feel free to reach out to gain access to our training samples and maps.
Other than that, I would advise the authors to consider uncertainty estimates on the regional area calculations. Particularly the best practices outlined by Olofsson et al. (2014) are very informative on this and could aid framing the observed changes in a correct and uncertainty-aware way.