We should state here what properties an ideal benchmark data set should have, right?
My wish would be for a data set that is as close as possible to real world problems (in addition to data sets that allow testing your methods until they break, which is of course very important).
As close a possible to real world problems could mean: Use physics-based error models (to simulate instrumental errors), simulate typical reporting errors (there should be plenty of experience around what can be wrong), simulate typical processing errors, etc. We will still not get around adding also simply perturbations in a statistical sense, but I think we can be more realistic than that.
Such a data set necessarily is a subdaily data set, and the monthly benchmark data set would simply be an average of the subdaily data (with an additional simulations of errors that can occur during the production of monthly means). Such a data set would necessarily be based on some sort of climate model or reanalysis data because other variables than temperature would be used, and they would be used in a high resolution.
I volunteer to produce such a data set if requested, but lacking experience on homogenizing data outside Europe, I would have to team up with more experienced people telling me what possibly can go wrong in Africa or the Arctic.
Purpose - To facilitate use of a robust, independent and global common benchmarking and assessment system for temperature data-product creation methodologies to aid product intercomparison and uncertainty quantification:
http://www.surfacetemperatures.org/benchmarking-and-assessment-working-group
Posting and comments are open for constructive advice and ideas.
Tuesday, 26 July 2011
Monday, 25 July 2011
Another radiosonde benchmarking paper
This is a little self serving but a further twist on the radiosonde benchmarking has just become available at JGR. A link to this paper is here . This uses the same benchmarks as used in Titchner et al. but looks at multiple additional impacts and also starts to try to address how you could use results from multiple different estimators to yield a grand unified estimate of the truth. There are, of course, many ways one could go about this step, but clearly such an effort may well be something the group would wish to consider how to approach ...
Thursday, 14 July 2011
Generating inhomogeneous worlds
The manuscript on the monthly temperature and precipitation benchmark of the COST Action HOME is now finished: manuscript. The inhomogeneities applied in this study, together with ideas for improvements in the discussion and outlook, may be a good basis for the benchmarking of the surface temperature initiative as well. Thus here I will only mention where I would suggest to do thinks differently for the surface temperatures initiative (STI). Next to improvements due to lessons learned from this study, the STI differs in two important aspects. 1) HOME considered regional networks, STI global datasets. 2) HOME focused on intercomparison of the homogenization algorithms, STI has the additional ambition that the benchmarking leads error estimates for the STI database.
I agree with Kate Willett that it would be valuable to have both data for which the truth is known as well as data for which the truth will be revealed after all homogenized contributions have been submitted. The latter leads to more reliable results because the homogenization algorithms cannot be tuned to solve the benchmark data well (and then possibly perform better on the benchmark than real data).
Data for which the truth is known would be a more classical validation study, which has the advantage that the homogenizers can learn during the exercise and also find (programming) errors; see outlook of our paper. In case of a benchmark one would have to wait another 3 year cycle to implement bug fixes. In HOME we had a number of such bugs, often in the parts newly written, e.g. to be able to handle multiple networks (which in not needed in daily work).
In HOME, we had random and clustered breaks. The clustered breaks occurred with a probability of 30% and if they occurred they affected 30% of the stations. As HOME generated regional networks, we could not insert breaks that occur in all stations, because they would not be detectable by relative homogenization. As the STI will generate global data, it is possible to insert breaks that occur simultaneously in one entire regional network, one would see them at the border of the countries. It would be interesting to see if the homogenization algorithms are able to move the information at the border into the heart of the network. However, also in this case, always inserting the clustered breaks in all stations would not be realistic. One country often has multiple networks (synoptic, volunteer networks or measurements by multiple organizations). Thus partial clustered breaks are also needed. In HOME the clustered breaks were perfectly simultaneous, which mimics a change in observational rules. Breaks clustered over a period, which mimics a change in the instrumentation or screens, would also be valuable.
The clustered breaks are not only correlated in time, but also in size. This is needed to generated biases in the continental and global averages due to inhomogeneities. This is seen in real data, especially in the early instrumental period and now again in the transition to automatic weather stations. These two periods may warrant special rules, i.e. maybe the timing of the clustered should not be fully random.
Another aspect that was not touched in the intercomparison study HOME, but needs to be treated in the STI are multiple elements. For many stations more than one temperature dataset will be available: Tmin, Tmax, Tmean, DTR (diurnal temperature range), or temperature observations for specific hours. Methods using multiple elements simultaneously likely perform better and if they are applied to the real data, they should also be validated in the benchmarking. Then we would need to know how the size of the inhomogeneities correlates between the various temperature variables. A break in Tmin, does not imply a (statistically significant) break in Tmax, but probably does make it more likely. Similar relations will likely hold among all temperature measures. I am not sure if there is data from literature on such cross-relations. On the other hand, you may also want to have some worlds with univariate validation data for the intercomparison of algorithms that do not use this information.
In the HOME benchmark, some stations had a local trend. The statistical properties of these trends were idealised, which was acceptable for an intercomparison, but to obtain realistic error estimates it should be studied in more detail how often such local trends occur in real datasets.
In HOME the perturbation were a constant number for every month. In the STI we have the opportunity to make the perturbations a function of insolation, wind and precipitation. This would make them partially stochastic, which is more realistic. In this case we would have to decide whether to make these covariates available to the homogenizers (potentially better homogenization) or not (most algorithms will currently not be able to use this information; making intercomparison more difficult). I could imagine that this is especially important for breaks during the early instrumental period, when measurement methods were not yet fully optimized to handle radiation and wetting problems.
It is important to add a stochastic and nonlinear large-scale trend to the data. It should be stochastic so that the homogenizer cannot see how well he did by computing the trend. And it should be nonlinear and contain decadal variability because homogenization algorithms always work and we should not mix our theoretical ideas with the data. Because the STI not only wants to make an intercomparison of homogenization algorithms, but also aims to produce representative errors, this trend will have to be modelling in a more realistic way as HOME did. I would suggest modelling the trend and the decadal variability separately and to vary it gradually on continental scales.
One of the outcomes of the study was that outliers are not important for the quality of the homogenization. Thus it may be best not to insert outliers in the benchmark. That would save work in inserting them and in removing them again later in the analysis of the homogenized data and homogenizers would not feel they have to program additional processing for outliers.
If I understood the science plan of the STI right, there will be multiple data levels (images of log books, digitized data, data in SI units, merged data, quality controlled data, homogenized data and gridded data). And everyone is invited to produce data or implement algorithms for the various levels. That would mean that there could be multiple input dataset for our benchmarking exercise. For homogenization, especially various ways of merging the data could be important. Merging methods that do not merge stations that move (to produce data that is suited to study local effects and relations with other (surface) parameters and variables) would produce much shorter time series than merging methods that aim to produce long time series (to make it easier to study secular trends due to climate change). The former case may be more difficult to homogenize because the time series are shorter, but may also be easier because overlapping data is not removed.
Many algorithms will likely not be able to automatically use the available metadata. For the intercomparison component of the project, it may thus be worthwhile to also have a dataset without metadata.
I agree with Kate Willett that it would be valuable to have both data for which the truth is known as well as data for which the truth will be revealed after all homogenized contributions have been submitted. The latter leads to more reliable results because the homogenization algorithms cannot be tuned to solve the benchmark data well (and then possibly perform better on the benchmark than real data).
Data for which the truth is known would be a more classical validation study, which has the advantage that the homogenizers can learn during the exercise and also find (programming) errors; see outlook of our paper. In case of a benchmark one would have to wait another 3 year cycle to implement bug fixes. In HOME we had a number of such bugs, often in the parts newly written, e.g. to be able to handle multiple networks (which in not needed in daily work).
In HOME, we had random and clustered breaks. The clustered breaks occurred with a probability of 30% and if they occurred they affected 30% of the stations. As HOME generated regional networks, we could not insert breaks that occur in all stations, because they would not be detectable by relative homogenization. As the STI will generate global data, it is possible to insert breaks that occur simultaneously in one entire regional network, one would see them at the border of the countries. It would be interesting to see if the homogenization algorithms are able to move the information at the border into the heart of the network. However, also in this case, always inserting the clustered breaks in all stations would not be realistic. One country often has multiple networks (synoptic, volunteer networks or measurements by multiple organizations). Thus partial clustered breaks are also needed. In HOME the clustered breaks were perfectly simultaneous, which mimics a change in observational rules. Breaks clustered over a period, which mimics a change in the instrumentation or screens, would also be valuable.
The clustered breaks are not only correlated in time, but also in size. This is needed to generated biases in the continental and global averages due to inhomogeneities. This is seen in real data, especially in the early instrumental period and now again in the transition to automatic weather stations. These two periods may warrant special rules, i.e. maybe the timing of the clustered should not be fully random.
Another aspect that was not touched in the intercomparison study HOME, but needs to be treated in the STI are multiple elements. For many stations more than one temperature dataset will be available: Tmin, Tmax, Tmean, DTR (diurnal temperature range), or temperature observations for specific hours. Methods using multiple elements simultaneously likely perform better and if they are applied to the real data, they should also be validated in the benchmarking. Then we would need to know how the size of the inhomogeneities correlates between the various temperature variables. A break in Tmin, does not imply a (statistically significant) break in Tmax, but probably does make it more likely. Similar relations will likely hold among all temperature measures. I am not sure if there is data from literature on such cross-relations. On the other hand, you may also want to have some worlds with univariate validation data for the intercomparison of algorithms that do not use this information.
In the HOME benchmark, some stations had a local trend. The statistical properties of these trends were idealised, which was acceptable for an intercomparison, but to obtain realistic error estimates it should be studied in more detail how often such local trends occur in real datasets.
In HOME the perturbation were a constant number for every month. In the STI we have the opportunity to make the perturbations a function of insolation, wind and precipitation. This would make them partially stochastic, which is more realistic. In this case we would have to decide whether to make these covariates available to the homogenizers (potentially better homogenization) or not (most algorithms will currently not be able to use this information; making intercomparison more difficult). I could imagine that this is especially important for breaks during the early instrumental period, when measurement methods were not yet fully optimized to handle radiation and wetting problems.
It is important to add a stochastic and nonlinear large-scale trend to the data. It should be stochastic so that the homogenizer cannot see how well he did by computing the trend. And it should be nonlinear and contain decadal variability because homogenization algorithms always work and we should not mix our theoretical ideas with the data. Because the STI not only wants to make an intercomparison of homogenization algorithms, but also aims to produce representative errors, this trend will have to be modelling in a more realistic way as HOME did. I would suggest modelling the trend and the decadal variability separately and to vary it gradually on continental scales.
One of the outcomes of the study was that outliers are not important for the quality of the homogenization. Thus it may be best not to insert outliers in the benchmark. That would save work in inserting them and in removing them again later in the analysis of the homogenized data and homogenizers would not feel they have to program additional processing for outliers.
If I understood the science plan of the STI right, there will be multiple data levels (images of log books, digitized data, data in SI units, merged data, quality controlled data, homogenized data and gridded data). And everyone is invited to produce data or implement algorithms for the various levels. That would mean that there could be multiple input dataset for our benchmarking exercise. For homogenization, especially various ways of merging the data could be important. Merging methods that do not merge stations that move (to produce data that is suited to study local effects and relations with other (surface) parameters and variables) would produce much shorter time series than merging methods that aim to produce long time series (to make it easier to study secular trends due to climate change). The former case may be more difficult to homogenize because the time series are shorter, but may also be easier because overlapping data is not removed.
Many algorithms will likely not be able to automatically use the available metadata. For the intercomparison component of the project, it may thus be worthwhile to also have a dataset without metadata.
Tuesday, 5 July 2011
Benchmarking temperature networks
The COST Action HOME has just finished a manuscript on benchmarking monthly of homogenization algorithms for regional monthly temperature and precipitation networks. I think it turned out quite interesting and provides a good base for discussions on the work of the surface temperatures benchmarking group; if only to avoid making the same mistakes.
Monday, 20 June 2011
Homogenization seminar
People interested in this blog are likely also interested in the upcoming
Seventh seminar for homogenization and quality control in climatological databases and COST ES-0601 “HOME” action management committee and final meeting in Budapest, Hungary on the 24 – 28 October 2011. Abstract deadline is the 16th of
September 2011. This is the main event in the homogenization community, in my opinion.
Seventh seminar for homogenization and quality control in climatological databases and COST ES-0601 “HOME” action management committee and final meeting in Budapest, Hungary on the 24 – 28 October 2011. Abstract deadline is the 16th of
September 2011. This is the main event in the homogenization community, in my opinion.
Thursday, 16 June 2011
If I had but one analog I could create ...
I think I would look to create something that tested in some sense the limits of homogenization methods whilst still retaining some realism.
I would take a forced component run such as c20c and then look to add change points that in the net removed that trend. This would penalize any algorithm that tended to introduce adjustments with a preferential zero bias.
The breaks I would add would be a mix of step like and slope like and a large number would have changes in seasonality and timeseries variance associated.
I'd have limited metadata and what metadata there was would be poor quality.
I would assume that most breaks were small (sigma <1K, in some cases perhaps <<1K) and that they happened fairly frequently (once every 5 to ten years say on average).
A number of breaks would be quasi-contemperaneous over countries and these would have very similar characteristics to each other.
There is documented evidence that these issues all to some extent pervade the network (e.g. US network move from stevenson screen to automated sensors happened largely within 5 years over 70% of the network).
So, whilst at the outer bounds of plausibility it would not be an entirely implausible error structure.
I would take a forced component run such as c20c and then look to add change points that in the net removed that trend. This would penalize any algorithm that tended to introduce adjustments with a preferential zero bias.
The breaks I would add would be a mix of step like and slope like and a large number would have changes in seasonality and timeseries variance associated.
I'd have limited metadata and what metadata there was would be poor quality.
I would assume that most breaks were small (sigma <1K, in some cases perhaps <<1K) and that they happened fairly frequently (once every 5 to ten years say on average).
A number of breaks would be quasi-contemperaneous over countries and these would have very similar characteristics to each other.
There is documented evidence that these issues all to some extent pervade the network (e.g. US network move from stevenson screen to automated sensors happened largely within 5 years over 70% of the network).
So, whilst at the outer bounds of plausibility it would not be an entirely implausible error structure.
Monday, 13 June 2011
Big questions with which to test homogenisation algorithms
Hi, I'm hoping to get time in the call to touch on this a little. Similar to previous posts about worse nightmares I would like us to think about questions that we want to answer with the analog-error-models. The idea would be to pick 8 of these to run with for each 3 year benchmarking cycle - one for each world. If these start from something simple to something horrible this will give us a chance to see where algorithms begin to struggle. Lets focus here on monthly means to make things a little simpler for the time being - please comment with your ideas. For example:
1) Do homogenisation algorithms detect discontinuities when none are present?
Analog-error-world 1 = A historical forcing model analog-known-world with no-errors added
2) Do homogenisation algorithms cope with discontinuities that affect the variance?
Analog-error-world 2 = A historical forcing model analog-known-world with seasonally constant changes applied
Analog-error-world 3 = A historical forcing model analog-known-world with seasonally varying changes applied at the same location and approximate magnitude as analog-error-world 2
2) Can homogenisation algorithms cope with non-stationary worlds/ where there is a background trend?
Analog-error-world 4 = A control forcing (constant pre-industrial emissions) model analog-known-world with mixed error structure applied
Analog-error-world 5 = An A1B (high emissions) forcing model analog-known-world basis with identical error structure to World 4
3) Can homogenisation algorithms cope when discontinuities are small and frequent?
Analog-error-world 6 = A historical forcing model analog-known-world with many small discontinuities added of various sign biases - (seasonally varying to be realistic?).
4) Can homogenisation algorithms cope with layered gradual and abrupt discontinuities (i.e., urban warming + instrument shelter change)
Analog-error-world 7 = A historical forcing model analog-known-world with either gradual or abrupt discontinuities applied to a station (seasonally varying to be realistic?)
Analog-error-world 8 = A historical forcing model analog-known-world with both gradual and abrupt discontinuities applied to a station (seasonally varying to be realistic?) using the initial error structure from analog-error-world 7 with other errors added.
At present its probably useful just to come up with as many plausible questions as possible and examples of error world structures to explore these.
There is an argument for including a really nasty one that is perhaps unplausible - so feel free to be creative.
1) Do homogenisation algorithms detect discontinuities when none are present?
Analog-error-world 1 = A historical forcing model analog-known-world with no-errors added
2) Do homogenisation algorithms cope with discontinuities that affect the variance?
Analog-error-world 2 = A historical forcing model analog-known-world with seasonally constant changes applied
Analog-error-world 3 = A historical forcing model analog-known-world with seasonally varying changes applied at the same location and approximate magnitude as analog-error-world 2
2) Can homogenisation algorithms cope with non-stationary worlds/ where there is a background trend?
Analog-error-world 4 = A control forcing (constant pre-industrial emissions) model analog-known-world with mixed error structure applied
Analog-error-world 5 = An A1B (high emissions) forcing model analog-known-world basis with identical error structure to World 4
3) Can homogenisation algorithms cope when discontinuities are small and frequent?
Analog-error-world 6 = A historical forcing model analog-known-world with many small discontinuities added of various sign biases - (seasonally varying to be realistic?).
4) Can homogenisation algorithms cope with layered gradual and abrupt discontinuities (i.e., urban warming + instrument shelter change)
Analog-error-world 7 = A historical forcing model analog-known-world with either gradual or abrupt discontinuities applied to a station (seasonally varying to be realistic?)
Analog-error-world 8 = A historical forcing model analog-known-world with both gradual and abrupt discontinuities applied to a station (seasonally varying to be realistic?) using the initial error structure from analog-error-world 7 with other errors added.
At present its probably useful just to come up with as many plausible questions as possible and examples of error world structures to explore these.
There is an argument for including a really nasty one that is perhaps unplausible - so feel free to be creative.