the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Updating the ESA Earth System Model for Future Gravity Mission Simulation Studies: ESA ESM 3.0
Abstract. The ESA Earth System Model (ESA ESM) provides a synthetic data set of the time-variable global gravity field that includes realistic mass variations in the atmosphere, oceans, terrestrial water storage, continental ice sheets, and the solid Earth on a wide range of spatial and temporal frequencies.
We here present the details of the newly released version 3.0 of the ESA ESM which has been updated to reflect the increased capabilities of future satellite gravimetry missions and is designed to be used as a source input model in closed loop simulation studies. The changes to the previous ESA ESM 2.0 include the utilization of a small ensemble of co- and post-seismic earthquake signals, an updated GIA model, additional mass balance signals from previously not considered Arctic glaciers, sub-monthly surface-mass balance changes and a more realistic representation of ice sheet dynamics. Extreme hydrometeorological events as well as climate-driven and anthropogenic impacts on continental water storage are represented through an update of the hydrological component. Additionally, the ESM separately includes ocean bottom pressure variations along the western slope of the Atlantic, representing variations in the meridional overturning circulation as a critically important component of the interactively coupled global climate system. ESA ESM 3.0 is available with 6-hour resolution from January 2007 until December 2020 as separate layers representing the respective dynamic processes considered. It is augmented with synthetic error time series for atmosphere and ocean as well as hydrology to facilitate stochastical modelling of residual background model errors.
The updated ESA ESM 3.0 is available under:
Shihora, Linus; Klemann, Volker; Jensen, Laura; Dill, Robert; Sasgen, Ingo; Wouters, Bert; Sauber, Jeanne; Han, Shin-Chan; Tanaka, Yoshiyuki; Braitenberg, Carla; Javed, Muhammad Tahir; Maurizio, Gerardo; Lecomte, Hugo; Babeyko, Andrey; Hovius, Niels; Dobslaw, Henryk (2026): The ESA Earth System Model 3.0. GFZ Data Services. https://doi.org/10.5880/GFZ.UEOS.2026.001
- Preprint
(6120 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on essd-2026-576', Jennifer Bonin, 22 Sep 2026
-
RC2: 'Comment on essd-2026-576', Anonymous Referee #2, 24 Sep 2026
Shihora et al. present version 3.0 of the ESA Earth System Model (ESM) for gravity mission simulation studies. Compared to its predecessor, the model benefits from a more comprehensive representation of mass changes in the different components of the climate system and inclusion of a several interesting target signals (AMOC-related boundary pressures, marine deposition in the northern Bay of Bengal). The manuscript is already in good shape, with only minor issues scattered here and there; see my comments below. In general, the authors could exert some more effort to better connect the different pieces and openly discuss product limitations and inconsistencies among the underlying datasets.
Detailed comments:
# As far as I can see, none of the figures is referenced (let alone discussed) in the main text. This creates a major disconnect that needs to be resolved. I have looked at the journal’s submission guidelines, and they don’t incentivize such approach either.
# L78-82: I found this section somewhat hard to follow and imprecise. Please ensure that you define ‘OBP anomalies’ and explain how you handled spatial mean OBP variability, which may have physical and unphysical causes in an ocean model (e.g, Boussinesq approximation, as mentioned). The explanations later in Section 2.6 provide some additional clues, but the matter should preferably be clarified right away in Section 2.4.
# L101-103: Far too concise for a description of the ice layer. The authors should provide further specifications and references, as well as insights into how modeled mass signals were calibrated against satellite gravimetry. Moreover, the text suggests that SMB variability is sampled daily, while ice dynamics (D) are modeled as trends plus accelerations. Does the exclusion of short-period variability in D compromise the realism of ESA ESM 3.0?
# L111: The five-year sampling of glacier changes in Alaska, Europe, and Central Asia contrasts with the temporal resolution adopted for other layers. The authors should address this limitation explicitly and attempt to quantify the signal lost due to under-sampling (similar to the previous point).
# Section 2.8, D-layer: Both the sediment deposition and AMOC-related OBP signals are small compared to the oceanic background variability. Exploration of those signals in closed-loop simulations therefore requires idealized settings, such as exclusion of the AO component (BTW: is that also recommended for studies of sediment mass change rates?). I’m wondering whether more could be done to strengthen the case for the D-layer and work toward a more realistic representation … Perhaps the authors can discuss options to exploit the spatial connectivity of the boundary pressures and river delta OBP anomalies? In addition, would it make sense to represent all other oceanic mass changes as uncorrelated time-varying stochastic field that could be switched on or off for closed-loop simulations?
Technical corrections:
Check all abbreviations – should be introduced upon first mentioning and then be used throughout the manuscript (e.g., L65 introduces ‘OBP’ a second time, ISSM and ISMIP7 on L101 are undefined)
L7: ‘previously not considered’ à ‘previously unconsidered’
L11: ‘critically important component of the interactively coupled global climate system’ – overly verbose phrasing
L49: ‘… representing a sea-level variations’ à ‘… representing manometric sea-level variations’
L50: Start new sentence after ‘(D)’
L58-59: This lead-in feels unnecessary. Consider moving it to the short paragraph immediately preceding Section 2.
L81: ‘In contrast to …, we explicitly note …’ – awkward sentence structure, rephrase for clarity
L99: ‘Arctic permafrost’
L132: Quoting the exact time spans of the LIA loading variations would help delineate these calculations from what was done for glaciers in the ice layer
L138: ‘coefficient level’
L142: ‘This layer balances mass changes …’
L147-148: Reformulate
L200: ‘changes in meridional overturning’
L201: ‘the two signal sources are essentially unrelated’
L204: ‘given by’ à ‘manifested in’
L229: ‘… regardless of their surface exposure’ – clumsy, reformulate
L279: MAGIC (fully capitalized)
L300: ‘sub-annual band’ instead of ‘sub-yearly band’?
L311: ‘… to be representative of a real-world mission’
L316: ‘… we provide, unlike in ESM 2.0, an error …’
L342: ‘spatially variable sea-level changes’ – imprecise, please reformulate and refer to manometric sea level changes
L367: ‘… in line with the greater variation’ - unclear reasoning
Citation: https://doi.org/10.5194/essd-2026-576-RC2
Data sets
The ESA Earth System Model 3.0 Linus Shihora et al. https://doi.org/10.5880/GFZ.UEOS.2026.001
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 327 | 105 | 74 | 506 | 81 | 75 |
- HTML: 327
- PDF: 105
- XML: 74
- Total: 506
- BibTeX: 81
- EndNote: 75
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review of “Updating the ESA Earth System Model…” by Shihora et al (2026)
This paper describes a new simulation set that is underway in preparation for the GRACE-C and NGGM missions, something that is highly worthwhile. The model components described are the logical ones, and the general methods make sense. However, I have some questions concerning things, which I’m hoping are just a matter of my own confusion, and which can mainly be resolved by fairly simple clarifications.
1) First, it is unclear to me what’s going on with the atmosphere and especially the ocean components. Section 2.1 and 2.2 seemed clear. But then, in Sections 3 and (I believe?) 4, the A and O terms are omitted from the simulation comparisons. I do not understand why. Have the A/O simulated series actually been tested, or not? Also, how were the error estimates for the A and O terms created, and where are those being used? And finally, why do you need to add a separate estimate of the AMOC effects, when you already have an ocean model involved? Does MPIOM not estimate the AMOC transport? Please explain.
2) Second, I believe some sort of comparison over the oceans and over the ice sheets only makes sense. Showing me only limited hydrologic comparisons omits 2/3+ the world’s surface, and both the “O” and “I” models (plus whatever “A” terms there are). How accurately does the simulated GRACE-C (etc) “mission” pick up the simulated ocean and ice signals? How does the groundtrack impact them, etc? I realize that these questions can’t all be answered in one paper, but showing us SOMETHING about these major regions only makes sense.
I also noted some smaller concerns:
3) L65: Why is only a static atmosphere applied over the ocean?
4) Section 2.7: While I support the testing of earthquake effects, I am concerned that by applying 15 major earthquakes over a single year, you’re going to completely mess up the view of non-earthquake ocean signals. Have you tested that? Additionally, since GRACE is estimated in global harmonics, having multiple major quakes occur at nearly the same time means that every harmonic is going to be rattling around unpredictably for all of 2008 – globally. I have no idea what that’s going to look like, even in “quiet” areas far from any quake, but I’d expect it to be visible in the ocean and deserts at least. Moreover, in the real world, the effect of nearby earthquakes are layered on top of each other, interfering both constructively and destructively – which means that their relative timing is critical. You recognized that in Japan. I see that in the Andaman area, you separated the Sumatra and Nias quakes by the 4 months they really were separated, but you then put the Indian Ocean quake (very nearby) only a month after Nias, rather than 7 years later. Not only is that not realistic, but it’s actually going to make it next to impossible for people using this simulation to see if those 3 quakes can be separated (a known issue with the Sumatra and Nias ones, of course). Does lining this many earthquake up so closely in time really make any sense? Or would it make more sense to plop them down at the times they actually occurred (or the same number of years apart), instead?
5) Why does the unconserved mass from the sediment and AMOC tests not get added back in the “B” layer? I’m sure it’s a small effect, but since you’ve got a GRD-based mass conservation tool, why not use it on everything?
6) L225: This explanation for why you add an extra AMOC signal confused me, rather than clarifying anything. Again, please make your goals clear here. Why add an extra signal, rather than use the AMOC signal which (I presume) is inside MPIOM?
7) L266: What do you mean by “target layers” here? Why is the simulation not based on all the modeled input that you’ve previously described? Why leave out the A and O components, particularly? If I’m going to use your simulated product to determine, say, whether I will be able to resolve some ocean signal in NGGM, I need an ocean model in the simulated NGGM output. Is the A/O data in there? Or not? This is incredibly unclear to me.
8) L268: What sampling distribution (groundtracks) are used? Realistic ones that shift over time as the altitude decreases? Or some idealized groundtrack with perfect, uniform spacing throughout?
9) L278: Wait. You’re saying that the ESM 3.0 improvements stem from the fact that you’ve artificially applied different error characteristics in the AOerr inputs? Why? How are those error characteristics defined, and does the change make any sense? Please back up and give an explanation of the AO error terms in the relevant sections of Section 2.
10) L282: I do not know if you can say that ESM3.0 is a “suitable model” if you’ve avoided testing the A/O portions that cover most of the world, honestly.
11) Fig 6: While the general shape of this plot is as expected, I noticed a few oddities that I don’t understand. First, why do the resonance spikes in the “GRACE-like” curves every 15-degrees stop at deg 75? Where’s the resonance at deg 90, etc? Second, why is there a sudden shift in trend in the MAGIC-like case at deg 80? Third, I grant it’s been a while since I’ve made one of these figures, but in the real GRACE/FO cases, I don’t ever remember seeing the long parabolic(ish) arcs between each resonance harmonic, or the smaller spikes/waves every few degrees that I see in this figure. What’s causing these unexpected features?
12) L289: What is a Vader filter? Please add a citation. Actually, you should probably add a citation for the DDK3 filter too, despite it being more well-known.
13) L290: This is the first time you’ve mentioned the terms L2a and L2b. Please define them and explain what the difference is between the sub-figures in Fig 7 because of it.
14) Actually… maybe you’d best explain what message we’re supposed to be getting out of Fig 7, regardless. Because the only message I’m getting is “all these subplots look about the same”. Is there another message intended? The colorbar on the figure makes it impossible to get much info out of this plot. For instance, anything from 0-20mm looks the same shade. And since all four sub-figures are similar, it’s even harder. Can you change the colorbar to include more hues, or something? Also, have you considered plotting one map, and then showing the DIFFERENCES in RMS relative to that, for the other plots? Show us the change in TWS variability caused by switching from one method to the next? Oh – and this is a plot of the variability after the trend and annual is removed. Are those trends and (especially) annuals impacted at all by which version you’re looking at?
15) L300: Just be warned: “broadly consistent results” just means that your models/simulations are all the same. It doesn’t mean they’re “right” or useful for what people need out of a simulation.
16) L305 and Fig 8: Please explain what a z-score is computed, in regards to the TWS signals. I’d assumed it would be a “z=2 means it’s outside the 2-sigma range for the timeseries” or something, but my eyeball doesn’t get that out of Fig 8. It is unclear to me how the z-score connects to the EWH at any time. Please add details.
17) Fig 8: I am concerned that this might be an example of “unfounded wiggle matching”, which (alas) is very easy to do with hydrology signals. You note that the 2010 year was extra snow-rich, okay. But multiple other years reach the same height of EWH and nearly the same z-score height; were those years also pretty snowy/rainy? The year 2009 has a z-score of ~0 for the whole year – was it atypically dry? Basically, a single point doesn’t give me much confidence, in a highly variable series like this.
18) Fig 8: Please draw vertical/horizontal bars at the tic marks, so we can more easily determine what times things are happening at in these plots. And please make extra vertical bars at the start/end of 2010, since you’re using it as an example case. Also, in the smaller subplot, the L3 from L2b is invisible on my copy. Is it on top of the L2a case? Can you alter the colors/dashes to make this more visible?
19) What am I supposed to understand, by noting that the “L3 from L2a” and “L3 from L2b” lines are identical? Please spell it out for me, since I’m missing it.
20) L313: This is where the AOe07 first comes up, I believe. I do not understand why an error series is being used for de-aliasing. Why not just put in one model as a “truth” series, and then use a second (probably MPIOM) to ACTUALLY de-alias it? Actually test the entire, real process. If that’s what you’re doing, I apologize. In any case, the phrasing needs to be adjusted to make it more clear what you’re really doing and why.
21) L315: Where is this HIe07 coming from? Also, there has been no testing/demonstration of either the A, O, or I components.
22) L321: Typo: You put two “two”s into this line. Oops.
Thanks for all your hard work on this needed simulation. Hopefully most of my confusions can be rapidly overcome simply by more detailed explanations. If I’ve harped excessively on the topic of including a full and proper ocean model, well, it’s because the last NGGM simulation I looked at didn’t have one, and that kind of stinks if you’re an oceanographer! This is a valuable product for preparing for the upcoming gravity mission launches, and I appreciate your efforts towards it.
-- Jennifer Bonin