the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
PlanktoShare: A large (50k+) and FAIR learning set for the Plankton Imager (Pi-10) for the Greater North Sea and NE Atlantic, based on a new flexible classification protocol
Abstract. The use of imaging techniques for the study of particles and plankton is a rapidly advancing field in marine sciences. The data the tools produce require automated classification solutions that are trained on learning sets of manually labelled images. In this study we present PlanktoShare, a comprehensive (50k+ images) database with manually labelled images captured by the vessel-mounted Plankton Imager (Pi-10), in the 200 – 2,000 μm size range and including phytoplankton, holoplankton, meroplankton and various gelatinous taxa. The Pi-10 images particles continuously in a flow-through mode and can operate alongside research operations and during transits making it a popular choice for plankton monitoring. PlanktoShare provides a robust resource for training classifiers as an open resource. A key challenge in developing classifiers such as these is that commonly arises when merging learning sets from different sources, because images are often organized in a folder-like structure with incompatible or inconsistent nomenclature. To address this, we propose a database approach which separates the taxonomic information from descriptive attributes. Each image is assigned to one of the classes ‘Organism’ (whole organism), ‘Taxo_particle‘ (particle with taxonomic information, such as exuvia) and ‘Non_taxo_particle’ (particle without taxonomic information, such as marine snow aggregates). Taxonomic information is standardised using the aphiaID system from the WOrld Register of Marine Species while additional descriptive information (e.g. Life_stage) is stored as attributes. This database approach ensures full interoperability across learning sets from research groups allowing rapid expansion of geographical coverage and improved classification performance. Finally, we provide open-source code to apply pre-trained classifiers for users and outline future directions for collaborative plankton imaging.
- Preprint
(2413 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-215', Anonymous Referee #1, 04 Jul 2026
-
RC2: 'Comment on essd-2026-215', Jonas Mortelmans, 24 Jul 2026
First of all, i have to excuse myself for the brief comments, often prone to spelling errors. My initial revision was accidentally deleted so i had to redo it all (and to save time i have quite some spelling errors).
The manuscript then! It provides a good contribution useful to many (including myself). On one end, the provided training set is open-access and allows users to start building their own classifiers quickly. On the other end, it provides a good scheme to work on image datasets.
The contribution is very much needed, especially when working on high-throughput imaging sensors. Despite the nice contribution, the textual and graphical aspect remains quite poor.
- all figures need (major) updates
- the need for interoperability - perhaps mentions how the scheme can work for other sensors? We should indeed go beyound the mere 'taxonomic' scope.
- there is a strong need for additional tables and documentation.
- there is a strong mismatch between a training set on on end (relatively well tackled); but on modeling, CNN, classificiatin tools on the other end which is not tackled at all and just comes in like that.
- i have strong issues with the current architecture of the paper - it reads quite difficult and lacks a more conventient style (m&m, results and discussion, conclusion).
- the numbers, figures, and actual images in the training set are not matching at all - or it is poorly described; or there are several versions of PlanktoShare which are not published as yet? This needs to be fixed. Similarly - a canonical record manifest, exact file formats, folder structure, identifiers, metadata files, checksums, licences and a version-specific DOI is entirely missing.
- i have strong issues for the lack of scientific backgorund. To elevate the paper to a worthy release, much more scientific background is needed (read and catch up on recent literature).
- there are many typographic errors and remnant of draft versions throughout the manuscript (i stopped commenting them after a while i have to admit). I hope authors were equally involved in proof-reading as obvious typos are still there.
- consistent terminology. The manuscript alternates between database, learning set, training set, library and folder. These terms should be clearly defined and used consistently.
- the annotation quality controll eneds to be addressed in detail. Who, how, more then one expert? CNN only? I strongly believe the scheme proposed needs to be revised. On one end, to include such information (essential as in time we evolve to a non-human validated dataset); but also to avoid the ZERO/NA thing which is poorly explained?
- The confusion matrices are too small and insufficient by themselves. Per-class precision, recall, F1 and support, as well as macro and weighted F1, balanced accuracy, confidence intervals and major confusion patterns...
- audit the bibliography - quite some missing.
- explain way more thoroughly the applicability for users; and similarly (as said above) make sure to explain what is in the scheme and in what format (table!).
- you speak of a database, but i see nothing hierarchical? how are thing coupled? how are images saved - are these in a noSQL server database? I have the feeling I "miss" something?
- the model software, Plankton Imager Classifier by geojoost; is not documented at all here.
Many of these, are elaborated in comment, throughout the manuscript.
I advice a major revision.
-
RC3: 'Reply on RC2', Jonas Mortelmans, 24 Jul 2026
So just to be sure: please, download the attachment to see all comments.
Citation: https://doi.org/10.5194/essd-2026-215-RC3
-
AC1: 'Comment on essd-2026-215', Lodewijk van Walraven, 04 Sep 2026
We have read the reviews of our submitted ESSD manuscript “PlanktoShare: A large (50k+) and FAIR learning set for the Plankton Imager (Pi-10) for the Greater North Sea and NE Atlantic, based on a new flexible classification protocol”. We thank the reviewers for their constructive and thorough comments. We think we can address the reviewers’ comments and would like to submit a revised version. Below we present an overview of how we aim to address the reviewers’ comments and suggestions.
Overall changes to structure
Both reviewers are positive about our paper, note its value as a data paper and support publication after appropriate revision. Upon reading the reviewers’ comments we noticed that a substantial part of the comments is related to the example application of our learning set and classification protocol by training two ResNet 50 image classification CNN models. As we mentioned on line 85, our main aim is to present a flexible, standardise annotation scheme for plankton imaging data, provide a training set of Plankton Imager PI-10 images and describe a framework to support the long term management and development of this training set. In section 5 we added the CNN models as an example application of this framework and learning set. However, upon reading the reviewers’ comments we think the CNN examples should be removed from the manuscript as the trained CNN models detract from the main focus of the paper. Removal of the CNN models is also suggested by reviewer 2 in a comment on line 85.
In the revision we will therefore remove the section and figures concerning the CNN models. Instead we will refer the readers to a Github repository where these models, including later improved versions, will be available. With this change we hope to resolve the comments of the reviewers related to the CNN models.
Please find attachedf a detailed response to the comments by reviewer 1 and 2.
Data sets
PlanktoShare: A large (50k+) and FAIR learning set for the Plankton Imager (Pi-10) for the Greater North Sea and NE Atlantic, based on a new flexible classification protocol Lodewijk van Walraven et al. https://doi.org/10.5281/zenodo.19119233
Model code and software
Plankton Imager Classifier Joost van Dalen https://github.com/geoJoost/planktoshare
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 530 | 137 | 43 | 710 | 45 | 38 |
- HTML: 530
- PDF: 137
- XML: 43
- Total: 710
- BibTeX: 45
- EndNote: 38
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript presents PlanktoShare, an open plankton-image learning set for the Plankton Imager, comprising more than 50,000 manually labelled images from the Greater North Sea and the NE Atlantic. It also proposes a flexible classification and database framework. Overall, the manuscript has clear value as a data paper. The dataset is timely and useful for promoting Pi-10 data sharing, classifier training, and regional plankton monitoring. I support publication after appropriate revision.
Major comments
Minor comments