the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
OceanTACO: A Multi-Sensor Global Ocean Sea Surface State Dataset
Abstract. We present OceanTACO, a harmonised global collection of sea surface state datasets designed to support reproducible Earth system research. The collection integrates satellite altimetry, sea surface temperature, salinity, surface winds, reanalysis fields, and Argo in situ observations within a unified cloud-optimised specification based on Transparent Access to Cloud-optimised datasets (TACO). It includes Level-3 observations, Level-4 gap-filled products, and reanalysis outputs while preserving native spatial and temporal resolution. The core dataset spans 29 March 2023 to 1 August 2025, covering the Surface Water and Ocean Topography (SWOT) mission, with an extended record from 1 January 2015 until 29 March 2023 for non-SWOT sources.
Datasets are harmonised through standardised metadata, spatial referencing, and temporal indexing, enabling consistent spatiotemporal queries across sensors and processing levels. A uniform internal structure reduces product-specific preprocessing and allows the same data-access routines to be applied across regions, sensors, and studies. This supports Earth systems analyses workflows such as validation against in situ observations, comparisons between observation and mapped products, observation system experiments, and multivariate sensor analyses.
Example applications demonstrate cross-product collocation with Argo, analysis of sea surface height variability during extreme events, and relationships between surface variables relevant for data-driven reconstruction. OceanTACO improves accessibility to coordinated multi-source analyses while preserving data provenance and native observation characteristics, and can be extended with new missions without restructuring the dataset. The core and extended dataset are available at https://doi.org/10.57967/hf/8171 (Lehmann and Aybar, 2026a) and https://doi.org/10.57967/hf/8172 (Lehmann and Aybar, 2026b) respectively.
- Preprint
(10906 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 14 Aug 2026)
- CC1: 'Comment on essd-2026-232', Giuseppe M.R. Manzella, 01 Jun 2026 reply
-
RC1: 'Comment on essd-2026-232', Anonymous Referee #1, 15 Jun 2026
reply
The manuscript describes a collection of ocean surface fields from observations and one ocean and one atmospheric reanalysis, which already existed on, or were brought to, regular grids close to their original resolution. Apart from the simple interpolation/aggregation, the product compresses the data by using int16 representation together with a standard compression routine. As regridding and lossy compressions (Reichelt et al. 2026) are widely available standard procedures, the main merit of the product is mostly the provision of a unique STAC catalog that unifies all products. This lowers the barrier to data use, particularly for disciplines unfamiliar with ocean and atmosphere data.
The manuscript could be better organized. The most important information about the temporal and spatial resolution of the source data and of the final product is not available for all products. It is also hard to find since listed at diverse places. Also, information about depth for GLORYS-12 and Argo is missing, suggesting that some data include vertical information.
Details:
What is the reason to limit the range to 2015-2025? Most data is available for decades before. Altimeter and GLORYS-12 are from 1993 and Argo from early 2000s. I think the short range limits the applications because the collection will be most useful for statistical analyses.
Section 2.1 Information about temporal and spatial resolution should be provided along with the data information.
L 150-151 It would be useful to know the temporal and spatial resolution in the OceanTACO product for the other data as well.L158-159 Are the original time and location stamps are preserved or is the data interpolated in any way?
L 163-164 Why are products listed that do not fall within the OceanTACO period?
L189-190 Are the complete profiles provided or only surface data. If only near surface data, which depth?
L221 Not up to 4? For instance, if the region covers a corner where 4 regions meet.
2.3.2/2.3.3 Strange way of numbering the sections. 2.3.2 seems to contain 2.3.3 otherwise it is and empty section. Consequently 2.3.3 should named 2.3.2.1.
Equation 2: My understanding is that the along-track data has a resolution of about 7 km, such that the number of SLA observations per cell is about one. Do these metrics make sense with such low number of samples or is the input data of much higher resolution?
Compression: Is it possible to compare your compression with other methods discussed in (https://doi.org/10.5194/egusphere-2026-60) and maybe evaluate by ClimateBenchPress?
Fig.4 What is shown is obviously a rather trivial consequence of how many significant digits int16 can represent. At least add a word about that.
The scatter plot is not useful since only pretty large errors will be visible there. I could be interested in a relative error normalized by the std of the data. For regions with low variability of a few cm, the error may reach 1% or so.L324-325 Maybe mention other eddy rich regions, e.g Gulf Stream or in the Southern Ocean as well.
L 332-336 Not so clear what the advantage is here. It seems much harder to do along-track spectra with the gridded fields, since it is not so easy anymore to find the points belonging to one track. Also it seems that the analysis of along-track SSH and SWOT is performed over different points. I guess the whole regions shown in the Figure were considered. That would explain the different PSD levels for the longer wavelengths where the difference in resolution should not be an issue. Seems this example is more a demonstration of pitfalls if analysis becomes to easy, people will not care about important details anymore.
Table A1 It would be good to add the temporal resolution and for GLORYS-12 that this is only surface data.
Table B1 It would be useful to report the error as a noise-to-signal error to provide a better idea if the resulting product is still useful. Or for which applications it may still be usable. For instance, the salinity signal is often less than 0.1. The RMSE value suggests that in some regions the error reaches 1% leaving only 2 digits significant. For SSS better use g/Kg, if this is meant.
Citation: https://doi.org/10.5194/essd-2026-232-RC1 -
AC2: 'Reply on RC1', Nils Lehmann, 27 Jul 2026
reply
We thank the reviewer for a careful report that improved both the manuscript and the released dataset. References are to the revised manuscript.General comments- Novelty / "mainly a STAC catalog." The reviewer is right that the component procedures (regridding, metadata standardisation, `int16`+`zlib`) are not inherently novel. We have more clearly reframed the contribution throughout as the harmonised, provenance-preserving integration of SWOT-era L3, L4, reanalysis, and Argo into one sample-invariant, queryable specification, and we draw an explicit line between this internal sample organisation and STAC-style asset discovery. (Abstract, Introduction, framework-contrast, STAC discussion, Conclusions.)- Organisation / scattered resolution and depth. We have consolidated temporal cadence and depth/coverage into Table 1 and the appendix variable dictionary (Table A1), with a pointer in Section 2.1, so that both source and stored resolution are given for every product (GLORYS surface subset with 15 m currents; Argo full 0 to 2000 m; L4 wind stored daily from hourly).Detailed comments1. Why 2015 to 2025? The sources reach back to 1993 or the early 2000s, so 2015 is not a data-availability limit. The start date follows the salinity record: the multi-observation salinity product we use combines SMAP, SMOS, satellite SST, and in situ salinity through multivariate optimal interpolation, and reaches this configuration in 2015 when SMAP becomes available. From 2015 onward the sensor constellation is constant apart from the L3 SSH missions and SWOT from 2023, which add sampling without removing any modality. Earlier years would append data from a thinner observing system rather than extend this one. Storage supports the choice (581 GB for the extended period alone; reaching back to 1993 would roughly triple it), and the horizon matches recent work such as [Martin et al. 2025](https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2025MS005063). The choice is not a hard boundary: all of the processing code is open source, so users with more storage and processing power can append earlier years for back-compatible sources through the public code and TACO folder-mode without restructuring the archive. (Changed in the Dataset description; Conclusions.)2./3. Resolution with the data (Sec. 2.1; L150 to 151). This is covered by the consolidation above; source and stored resolution are now given for all modalities in Tables 1 and A1.4. Stamps preserved or interpolated? (L158 to 159). We clarified this: L4 and reanalysis products keep their native grids and are only encoded, L3 nadir and SWOT observations are conservatively binned rather than interpolated (with `obs_mean_lon/lat`, `n_obs`, and `n_tracks` retained), and Argo keeps its native coordinates and timestamps.5. Products outside the period (L163 to 164). The list describes the DUACS product's mission heritage, not the missions active during 2015 to 2025; we corrected the wording.6. Full profiles or surface? (L189 to 190). Argo profiles are kept in full over 0 to 2000 m, and GLORYS is included as a surface subset (15 m currents), following [Martin et al. 2025](https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2025MS005063). This is now stated in the depth column of the table.7. "Up to 4" tiles (L221). The reviewer is correct: a query at a shared corner can overlap four tiles, and it is served transparently across all of them. We fixed the wording.8. Section numbering 2.3.2/2.3.3. We resolved this structurally by converting the per-variable descriptions into run-in labels under the Level 4 and Level 3 data subsections.9. Eq. 2, SEM at about one observation per cell. The reviewer's concern was well founded. The previous within-track quantity was not meaningful at this sampling, so we replaced it with an inter-track dispersion measure, σ_track (the observation-weighted standard deviation of per-track means), reported only for crossover cells where two or more passes overlap, and added `n_tracks` as a provenance variable. It is left undefined (NaN) elsewhere. Where it is defined (2.26 % of cells for L3 SSH and 4.04 % for SWOT), σ_track is physically meaningful but modest, with a median of 0.65 to 0.79 cm, about 7 % of the field standard deviation. We recomputed and re-uploaded the L3 nadir and SWOT collections. (Regridding section: the inter-track dispersion equation and surrounding prose; variable dictionary in Table A1.)10. Compression vs Reichelt / ClimateBenchPress. We carried out the suggested benchmark on the five gridded products, using a ClimateBenchPress-style evaluation with SZ3, SPERR, ZFP, EBCC, and bitround variants. Under matched strict bounds these codecs reach higher ratios than the deployed encoding (for example L4 SSH 5.16x against 16.11x, and L4 SST 7.00x against 17.27x), and we report this in a new benchmark table. We have chosen to keep the deployed `int16`+`zlib` scheme, however, for a practical reason: these error-bounded codecs are typically deployed as array-store codecs for Zarr, where the decoder is applied transparently on read, and they are not directly compatible with the TACO framework, which stores self-contained files that any standard reader can open without an additional codec dependency. The deployed scheme keeps the archive simple, transparent, and directly readable, and preserves each field at full numerical resolution rather than degrading it to any single user's error tolerance, so we present the benchmark as a separate, additive result rather than as a replacement.We also took the opportunity to separate two things that the original text conflated: the lossy-encoding error of the `int16` scheme, and the disk-space savings reported in the storage table. The two measure different quantities. The storage-table savings also include product restructuring rather than encoding alone, for example the consolidation of native hourly L4 wind fields into daily OceanTACO variables and of many individual L3 track files into a single regridded product, which is why their ratios are larger than the encoding step by itself. We revised the table captions and footnotes accordingly. We also corrected the L4 SST benchmark setup, since the source is natively `int16` and error bounds below that quantisation step are not meaningful. (Compression section; appendix Table B2; storage-table caption and footnotes.)11. Fig. 4, int16 significant digits. We added this: the caption now states that the absolute error is bounded by the `int16` quantisation step.12. Scatter / relative error. We added an "RMSE / field std." column to the appendix compression table and a second figure panel mapping the relative error (`|orig-comp|/σ_local`). As the reviewer anticipated, this shows the largest relative error in low-variability regions while the absolute error stays at the floor.13. Other eddy-rich regions (L324 to 325). We added the Southern Ocean and Antarctic Circumpolar Current alongside the Gulf Stream and Kuroshio.14. PSD example (L332 to 336). The reviewer's comment was correct: the original comparison did not enforce identical sampling, so the long-wavelength offset was a sampling artifact rather than a resolution difference. We replaced it with a matched L4-DUACS-against-SWOT comparison, in which the L4 field is sampled at the exact SWOT positions. The spectra now coincide at large scales (above roughly 100 to 150 km) and diverge only at short wavelengths, which isolates the genuine resolution difference, and constructing this matched comparison is itself straightforward with OceanTACO. To avoid over-interpreting the short-wavelength tail, we use the denoised SWOT L3 field and frame the divergence as a genuine resolution difference only over the roughly 30 to 150 km band, noting in the caption and text that the L3 spectrum flattens at the shortest scales near the effective-resolution limit of the product rather than resolving indefinitely.15. Table A1, temporal resolution and GLORYS surface. Addressed in the updated table.16. Table B1, noise-to-signal and SSS units. We added the noise-to-signal column (see comment 12). On units, we have kept practical salinity but conceded the labelling: the product is practical salinity (PSS-78, dimensionless), inherited from CMEMS, whereas g/kg is absolute salinity (TEOS-10), a different quantity. We relabelled salinity consistently as PSS-78 throughout.We are grateful for the reviewer's comment and believe they have improved the manuscript.Citation: https://doi.org/
10.5194/essd-2026-232-AC2
-
AC2: 'Reply on RC1', Nils Lehmann, 27 Jul 2026
reply
-
RC2: 'Comment on essd-2026-232', Giuseppe M.R. Manzella, 08 Jul 2026
reply
The paper presents a "data system" composed of data from various sources and therefore subject to potentially different processing methods and/or algorithms. Highlighting these possibilities and providing information on the challenges encountered when using data from various sources is a plus.
The data system can certainly be useful for many purposes, and therefore the paper, while not original, is valuable.
Some information should be added for further useful references. Any data set contains noise and spurious data. Have these been eliminated from the data residing in the various repositories? Was a check performed within the "OceanTACO" system?
A weakness of the article is represented by the usage examples, which are rather superficial and do not provide sufficient detail, as in 3.3 and 3.5. In the case of Observation System Experiments and Mission Impact Studies, how much of an impact might the definition of very large regions (perhaps unable to highlight subregional phenomena) have on (e.g.) impact studies? Even more vague is the paragraph regarding machine learning, which could also be considered data mining. There are many machine learning models, including supervised, unsupervised, semisupervised, and reinforcement learning. Some details on what has been used by the cited authors could be very useful.
Citation: https://doi.org/10.5194/essd-2026-232-RC2 -
AC1: 'Reply on RC2', Nils Lehmann, 27 Jul 2026
reply
We thank the reviewer for a positive and constructive report. We are glad the reviewer finds the data system useful for many purposes, and we agree that its value lies not innew algorithms but in the harmonised, provenance-preserving organisation of heterogeneous sources. We have taken the opportunity to sharpen the usage examples and to make thedataset's quality-control stance explicit. References are to the revised manuscript.General comments- Challenges of combining data from various sources: We fully agree, and this friction is central to the contribution. Harmonising these products required reconciling heterogeneousgrids and geometries onto a common reference, aligning differing temporal cadences, and unifying incompatible fill-value, metadata, and licensing conventions across ten sources.These steps are documented in the Introduction and the Processing methods section. OceanTACO removes this friction once, under a single sample-invariant, provenance-preserving structure.Detailed comments- Noise and spurious data, was a check performed within OceanTACO? Each product is ingested at the published, quality-controlled processing level provided by its authoritativesource. We rely on this provider-level quality control and deliberately do not layer additional processing on top of it. Different downstream tasks or use cases impose conflicting requirements. Some analyses call for aggressive outlier rejection, while data assimilation and machine-learning reconstruction often need the raw observations and their error characteristics preserved, so asingle fixed filtering step would suit some applications and silently degrade others. We have added a sentence to Section 2.4 making this stance explicit.- Usage examples in 3.3 and 3.5 are superficial. We agree the examples could be more concrete, and we have expanded both sections with specific pointers and a fuller account of the relevantmethods (detailed below). We note one boundary: the ESSD data-description manuscript type scopes out "extensive analysis and interpretation of data," which "should seek publication in anappropriate regular journal." This applies to running a complete observation-system experiment or a full model-development study, not to describing how the dataset supports them or to situating them in the literature, both of which we have expanded here.- Impact of very large regions on subregional phenomena (3.3). This is an important point, and we realise the text was unclear: the eight "regions" are purely a storage-and-retrieval efficiencymechanism (regional tiling for selective I/O), not analysis regions. Any arbitrarily sized bounding box, or any user-defined subregion, can be retrieved, served transparently across up to four tileswhere a box straddles tile boundaries (Section 2.4.1), and each product is preserved at its native resolution, down to about 2 km for SWOT. Subregional and fine-scale phenomena therefore remain fully accessible. The Hurricane Milton case study (Section 3.4) already demonstrates a localised subregional analysis at full resolution. We have added a clarifying sentence to Section 3.3 to state this explicitly.- Observation system experiments (3.3) A genuine OSE requires running a data-assimilation system with observing systems selectively withheld, which is a research study in its own right and falls under the "extensive analysis" that ESSD places out of scope for this manuscript type. What the dataset provides is the machinery that makes such experiments configurable, and we have made this concrete in Section 3.3: OceanTACO preserves per-mission identifiers and sensor-level indexing; the per-mission temporal coverage that determines which sensors are available for a given experiment is tabulated in the Appendix (Table B1); and the matched SWOT-versus-mapped spectral comparison in Section 3.2 alreadyillustrates the kind of resolved-scale, mission-impact diagnostic these queries support.- Machine learning methods and specifics (3.5). We have expanded Section 3.5 to describe, at a high level, what each cited method does rather than referring to them collectively: supervised networks that fuse multiple satellite variables into gridded SSH, end-to-end learned variational assimilation schemes for nadir and wide-swath altimetry, dynamically constrained joint reconstruction of SSH and temperature from multi-sensor observations, generative assimilation of multi-modal satellite observations, and unsupervised spatio-temporal interpolation of altimetry maps. The sea surface temperature and salinity work is predominantly supervised super-resolution and gap filling. Their common requirement isconsistently collocated multi-variable, multi-sensor inputs, which is exactly the harmonisation OceanTACO supplies, and the accompanying documentation shows how to build ML-ready PyTorch datasets from the same spatiotemporal queries.We thank the reviewer again for the helpful and encouraging assessment.Citation: https://doi.org/
10.5194/essd-2026-232-AC1
-
AC1: 'Reply on RC2', Nils Lehmann, 27 Jul 2026
reply
Data sets
OceanTACO Core Dataset Nils Lehmann https://doi.org/10.57967/hf/8171
Model code and software
Data Generation Code Nils Lehmann and Cesar Aybar https://github.com/nilsleh/oceanTACO
Interactive computing environment
ReadTheDocs Documentation Page Nils Lehmann https://oceantaco.readthedocs.io/en/latest/index.html
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 452 | 180 | 33 | 665 | 26 | 27 |
- HTML: 452
- PDF: 180
- XML: 33
- Total: 665
- BibTeX: 26
- EndNote: 27
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The paper presents a "data system" composed of data from various sources and therefore subject to potentially different processing methods and/or algorithms. Highlighting these possibilities and providing information on the challenges encountered when using data from various sources is a plus.
The data system can certainly be useful for many purposes, and therefore the paper, while not original, is valuable.
Some information should be added for further useful references. Any data set contains noise and spurious data. Have these been eliminated from the data residing in the various repositories? Was a check performed within the "OceanTACO" system?
A weakness of the article is represented by the usage examples, which are rather superficial and do not provide sufficient detail, as in 3.3 and 3.5. In the case of Observation System Experiments and Mission Impact Studies, how much of an impact might the definition of very large regions (perhaps unable to highlight subregional phenomena) have on (e.g.) impact studies? Even more vague is the paragraph regarding machine learning, which could also be considered data mining. There are many machine learning models, including supervised, unsupervised, semisupervised, and reinforcement learning. Some details on what has been used by the cited authors could be very useful.