Hi, I'm hoping to get time in the call to touch on this a little. Similar to previous posts about worse nightmares I would like us to think about questions that we want to answer with the analog-error-models. The idea would be to pick 8 of these to run with for each 3 year benchmarking cycle - one for each world. If these start from something simple to something horrible this will give us a chance to see where algorithms begin to struggle. Lets focus here on monthly means to make things a little simpler for the time being - please comment with your ideas. For example:
1) Do homogenisation algorithms detect discontinuities when none are present?
Analog-error-world 1 = A historical forcing model analog-known-world with no-errors added
2) Do homogenisation algorithms cope with discontinuities that affect the variance?
Analog-error-world 2 = A historical forcing model analog-known-world with seasonally constant changes applied
Analog-error-world 3 = A historical forcing model analog-known-world with seasonally varying changes applied at the same location and approximate magnitude as analog-error-world 2
2) Can homogenisation algorithms cope with non-stationary worlds/ where there is a background trend?
Analog-error-world 4 = A control forcing (constant pre-industrial emissions) model analog-known-world with mixed error structure applied
Analog-error-world 5 = An A1B (high emissions) forcing model analog-known-world basis with identical error structure to World 4
3) Can homogenisation algorithms cope when discontinuities are small and frequent?
Analog-error-world 6 = A historical forcing model analog-known-world with many small discontinuities added of various sign biases - (seasonally varying to be realistic?).
4) Can homogenisation algorithms cope with layered gradual and abrupt discontinuities (i.e., urban warming + instrument shelter change)
Analog-error-world 7 = A historical forcing model analog-known-world with either gradual or abrupt discontinuities applied to a station (seasonally varying to be realistic?)
Analog-error-world 8 = A historical forcing model analog-known-world with both gradual and abrupt discontinuities applied to a station (seasonally varying to be realistic?) using the initial error structure from analog-error-world 7 with other errors added.
At present its probably useful just to come up with as many plausible questions as possible and examples of error world structures to explore these.
There is an argument for including a really nasty one that is perhaps unplausible - so feel free to be creative.
4 comments:
ref. 2. I would think it worth assessing whether the phasing of natural variability and its nature mattered. This would call for the #2 to be (at least) four analogs all with the same breakpoints applied. A control run, perhaps two c20c ensemble members from the same model and a c20c run from a different model (same forcing but presumably different estimates of e.g. ENSO signal as attested to in the literature).
I like the idea of putting in a nightmare world (=something worse than even the more skeptical fringes think plausible) as I think its important to ascertain breaking point for any algorithm. To the same extent is it worth throwing in one or more error models somewhat simpler than the real world on the basis that if an algorithm can't even get e.g solely large (>1K) step-like breaks with no overall sign bias + perfect metadata adequately they should be regarded with real suspicion?
I don't have the homogenisation background of most of the group, so can't suggest realistic perturbations in the same way you can. But since we're trying to cover lots of possible worlds, let's put forward something from a statisticians point of view based on mixture distributions, which may be completely unrealistic. Suppose that a measuring instrument can be in two or more states and it switches between them according to some random mechanism - I don't know a physical justification, a loose connection perhaps. In the different states the probability distributions of measurements are different, so we have a mixture of distributions. The distributions could differ in mean, in variance or in other ways. This is a bit like changes in mean, variance etc caused by a change of instrument, but the mechanism is different, things can switch back again and there are almost certainly no metadata. An added complication is having done some statistical analysis to identify which measurements come from each distribution in the mixture, how can it be known which distribution is the 'correct' one, if any?
Ian, your suggestion is good. There is some indication in the data, see the work of Peter Domonkos on platform-like pairs of breaks, that after one break, the next break restores the old situation. Especially if the homogeneous subperiod between these two breaks is short, just a few years. This would be logical if the first break is because of a measurement problem, which is removed as soon as it is discovered.
@Kate's post.
> Do homogenisation algorithms detect
> discontinuities when none are present?
Yes they do. In HOME we had one regional network without breaks, which was not known to the homogenizers. (I am sorry for being such a mean guy, i.e. a scientist. ;-) Most algorithms put in breaks (the main exception is the USHCN algorithm from NOAA). Inserting some breaks were there are non does not have to be too bad, if the corrections inserted to not change the data much. Thus rather than looking only at detection, it is important to actually look at the quality of the data after homogenization. Such a world without breaks would best be part of the validation part of our study and not the blind benchmarking part. If you know there is one world without breaks, it is easy to find out which one it is.
> Do homogenisation algorithms cope with
> discontinuities that affect the variance?
Not many algorithms explicitly correct for changes in the variance. The main exception is CLIMATOL by José A. Guijarro. Most will thus simply keep the variability as it is (except if the added variance is implemented as a change in the magnitude of the the annual cycle.) Still it would be good to include such a feature to see how the algorithms react to it; a difficulty may be that we might not know what are typical sizes for changes in the variance. We would first have to study some data sets.
> Can homogenisation algorithms cope with
> non-stationary worlds/ where there is a
> background trend?
Except for one contribution with a programming error, all relative homogenization algorithm passed this test without problem for the HOME benchmark. Which was to be expected as the design principle of relative homogenization is to remove trends and climate variability by only considering the difference between one station and its neighbours. Only absolute homogenization will have problems with such tasks, the more difficult (nonlinear) you make the trend, the more problem such algorithms have.
> Can homogenisation algorithms cope when
> discontinuities are small and frequent?
If these small and frequent breaks are random, the homogenization algorithms will not improve the data much. On the other hand, there would also hardly be any error in the trends and only small errors in the decadal variability due to such breaks. If the small breaks would be temporally correlated, i.e. if you would insert multiple breaks in the same direction, that would be very similar to a scenario in which you insert trends to single stations. The HOME project did not see any indication that this creates problems for homogenization algorithms.
> Can homogenisation algorithms cope with
> layered gradual and abrupt discontinuities
In HOME the breaks and local (station) trends were inserted independently. Thus this layered situation was typical. I would expect it to be realistic as well and thus something we should include. This question could thus already be answered by analysing the HOME benchmark. I expect no problems.
Post a Comment