the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
PlanktoShare: A large (50k+) and FAIR learning set for the Plankton Imager (Pi-10) for the Greater North Sea and NE Atlantic, based on a new flexible classification protocol
Abstract. The use of imaging techniques for the study of particles and plankton is a rapidly advancing field in marine sciences. The data the tools produce require automated classification solutions that are trained on learning sets of manually labelled images. In this study we present PlanktoShare, a comprehensive (50k+ images) database with manually labelled images captured by the vessel-mounted Plankton Imager (Pi-10), in the 200 – 2,000 μm size range and including phytoplankton, holoplankton, meroplankton and various gelatinous taxa. The Pi-10 images particles continuously in a flow-through mode and can operate alongside research operations and during transits making it a popular choice for plankton monitoring. PlanktoShare provides a robust resource for training classifiers as an open resource. A key challenge in developing classifiers such as these is that commonly arises when merging learning sets from different sources, because images are often organized in a folder-like structure with incompatible or inconsistent nomenclature. To address this, we propose a database approach which separates the taxonomic information from descriptive attributes. Each image is assigned to one of the classes ‘Organism’ (whole organism), ‘Taxo_particle‘ (particle with taxonomic information, such as exuvia) and ‘Non_taxo_particle’ (particle without taxonomic information, such as marine snow aggregates). Taxonomic information is standardised using the aphiaID system from the WOrld Register of Marine Species while additional descriptive information (e.g. Life_stage) is stored as attributes. This database approach ensures full interoperability across learning sets from research groups allowing rapid expansion of geographical coverage and improved classification performance. Finally, we provide open-source code to apply pre-trained classifiers for users and outline future directions for collaborative plankton imaging.
- Preprint
(2413 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-215', Anonymous Referee #1, 04 Jul 2026
-
RC2: 'Comment on essd-2026-215', Jonas Mortelmans, 24 Jul 2026
First of all, i have to excuse myself for the brief comments, often prone to spelling errors. My initial revision was accidentally deleted so i had to redo it all (and to save time i have quite some spelling errors).
The manuscript then! It provides a good contribution useful to many (including myself). On one end, the provided training set is open-access and allows users to start building their own classifiers quickly. On the other end, it provides a good scheme to work on image datasets.
The contribution is very much needed, especially when working on high-throughput imaging sensors. Despite the nice contribution, the textual and graphical aspect remains quite poor.
- all figures need (major) updates
- the need for interoperability - perhaps mentions how the scheme can work for other sensors? We should indeed go beyound the mere 'taxonomic' scope.
- there is a strong need for additional tables and documentation.
- there is a strong mismatch between a training set on on end (relatively well tackled); but on modeling, CNN, classificiatin tools on the other end which is not tackled at all and just comes in like that.
- i have strong issues with the current architecture of the paper - it reads quite difficult and lacks a more conventient style (m&m, results and discussion, conclusion).
- the numbers, figures, and actual images in the training set are not matching at all - or it is poorly described; or there are several versions of PlanktoShare which are not published as yet? This needs to be fixed. Similarly - a canonical record manifest, exact file formats, folder structure, identifiers, metadata files, checksums, licences and a version-specific DOI is entirely missing.
- i have strong issues for the lack of scientific backgorund. To elevate the paper to a worthy release, much more scientific background is needed (read and catch up on recent literature).
- there are many typographic errors and remnant of draft versions throughout the manuscript (i stopped commenting them after a while i have to admit). I hope authors were equally involved in proof-reading as obvious typos are still there.
- consistent terminology. The manuscript alternates between database, learning set, training set, library and folder. These terms should be clearly defined and used consistently.
- the annotation quality controll eneds to be addressed in detail. Who, how, more then one expert? CNN only? I strongly believe the scheme proposed needs to be revised. On one end, to include such information (essential as in time we evolve to a non-human validated dataset); but also to avoid the ZERO/NA thing which is poorly explained?
- The confusion matrices are too small and insufficient by themselves. Per-class precision, recall, F1 and support, as well as macro and weighted F1, balanced accuracy, confidence intervals and major confusion patterns...
- audit the bibliography - quite some missing.
- explain way more thoroughly the applicability for users; and similarly (as said above) make sure to explain what is in the scheme and in what format (table!).
- you speak of a database, but i see nothing hierarchical? how are thing coupled? how are images saved - are these in a noSQL server database? I have the feeling I "miss" something?
- the model software, Plankton Imager Classifier by geojoost; is not documented at all here.
Many of these, are elaborated in comment, throughout the manuscript.
I advice a major revision.
-
RC3: 'Reply on RC2', Jonas Mortelmans, 24 Jul 2026
So just to be sure: please, download the attachment to see all comments.
Citation: https://doi.org/10.5194/essd-2026-215-RC3
Data sets
PlanktoShare: A large (50k+) and FAIR learning set for the Plankton Imager (Pi-10) for the Greater North Sea and NE Atlantic, based on a new flexible classification protocol Lodewijk van Walraven et al. https://doi.org/10.5281/zenodo.19119233
Model code and software
Plankton Imager Classifier Joost van Dalen https://github.com/geoJoost/planktoshare
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 499 | 127 | 35 | 661 | 34 | 29 |
- HTML: 499
- PDF: 127
- XML: 35
- Total: 661
- BibTeX: 34
- EndNote: 29
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript presents PlanktoShare, an open plankton-image learning set for the Plankton Imager, comprising more than 50,000 manually labelled images from the Greater North Sea and the NE Atlantic. It also proposes a flexible classification and database framework. Overall, the manuscript has clear value as a data paper. The dataset is timely and useful for promoting Pi-10 data sharing, classifier training, and regional plankton monitoring. I support publication after appropriate revision.
Major comments
Minor comments