Tuesday, 26 July 2011

Benchmark for real-world problems

We should state here what properties an ideal benchmark data set should have, right?
My wish would be for a data set that is as close as possible to real world problems (in addition to data sets that allow testing your methods until they break, which is of course very important).
As close a possible to real world problems could mean: Use physics-based error models (to simulate instrumental errors), simulate typical reporting errors (there should be plenty of experience around what can be wrong), simulate typical processing errors, etc. We will still not get around adding also simply perturbations in a statistical sense, but I think we can be more realistic than that.
Such a data set necessarily is a subdaily data set, and the monthly benchmark data set would simply be an average of the subdaily data (with an additional simulations of errors that can occur during the production of monthly means). Such a data set would necessarily be based on some sort of climate model or reanalysis data because other variables than temperature would be used, and they would be used in a high resolution.
I volunteer to produce such a data set if requested, but lacking experience on homogenizing data outside Europe, I would have to team up with more experienced people telling me what possibly can go wrong in Africa or the Arctic.

Monday, 25 July 2011

Another radiosonde benchmarking paper

This is a little self serving but a further twist on the radiosonde benchmarking has just become available at JGR. A link to this paper is here . This uses the same benchmarks as used in Titchner et al. but looks at multiple additional impacts and also starts to try to address how you could use results from multiple different estimators to yield a grand unified estimate of the truth. There are, of course, many ways one could go about this step, but clearly such an effort may well be something the group would wish to consider how to approach ...

Thursday, 14 July 2011

Generating inhomogeneous worlds

The manuscript on the monthly temperature and precipitation benchmark of the COST Action HOME is now finished: manuscript. The inhomogeneities applied in this study, together with ideas for improvements in the discussion and outlook, may be a good basis for the benchmarking of the surface temperature initiative as well. Thus here I will only mention where I would suggest to do thinks differently for the surface temperatures initiative (STI). Next to improvements due to lessons learned from this study, the STI differs in two important aspects. 1) HOME considered regional networks, STI global datasets. 2) HOME focused on intercomparison of the homogenization algorithms, STI has the additional ambition that the benchmarking leads error estimates for the STI database.

I agree with Kate Willett that it would be valuable to have both data for which the truth is known as well as data for which the truth will be revealed after all homogenized contributions have been submitted. The latter leads to more reliable results because the homogenization algorithms cannot be tuned to solve the benchmark data well (and then possibly perform better on the benchmark than real data).
Data for which the truth is known would be a more classical validation study, which has the advantage that the homogenizers can learn during the exercise and also find (programming) errors; see outlook of our paper. In case of a benchmark one would have to wait another 3 year cycle to implement bug fixes. In HOME we had a number of such bugs, often in the parts newly written, e.g. to be able to handle multiple networks (which in not needed in daily work).

In HOME, we had random and clustered breaks. The clustered breaks occurred with a probability of 30% and if they occurred they affected 30% of the stations. As HOME generated regional networks, we could not insert breaks that occur in all stations, because they would not be detectable by relative homogenization. As the STI will generate global data, it is possible to insert breaks that occur simultaneously in one entire regional network, one would see them at the border of the countries. It would be interesting to see if the homogenization algorithms are able to move the information at the border into the heart of the network. However, also in this case, always inserting the clustered breaks in all stations would not be realistic. One country often has multiple networks (synoptic, volunteer networks or measurements by multiple organizations). Thus partial clustered breaks are also needed. In HOME the clustered breaks were perfectly simultaneous, which mimics a change in observational rules. Breaks clustered over a period, which mimics a change in the instrumentation or screens, would also be valuable.
The clustered breaks are not only correlated in time, but also in size. This is needed to generated biases in the continental and global averages due to inhomogeneities. This is seen in real data, especially in the early instrumental period and now again in the transition to automatic weather stations. These two periods may warrant special rules, i.e. maybe the timing of the clustered should not be fully random.

Another aspect that was not touched in the intercomparison study HOME, but needs to be treated in the STI are multiple elements. For many stations more than one temperature dataset will be available: Tmin, Tmax, Tmean, DTR (diurnal temperature range), or temperature observations for specific hours. Methods using multiple elements simultaneously likely perform better and if they are applied to the real data, they should also be validated in the benchmarking. Then we would need to know how the size of the inhomogeneities correlates between the various temperature variables. A break in Tmin, does not imply a (statistically significant) break in Tmax, but probably does make it more likely. Similar relations will likely hold among all temperature measures. I am not sure if there is data from literature on such cross-relations. On the other hand, you may also want to have some worlds with univariate validation data for the intercomparison of algorithms that do not use this information.

In the HOME benchmark, some stations had a local trend. The statistical properties of these trends were idealised, which was acceptable for an intercomparison, but to obtain realistic error estimates it should be studied in more detail how often such local trends occur in real datasets.

In HOME the perturbation were a constant number for every month. In the STI we have the opportunity to make the perturbations a function of insolation, wind and precipitation. This would make them partially stochastic, which is more realistic. In this case we would have to decide whether to make these covariates available to the homogenizers (potentially better homogenization) or not (most algorithms will currently not be able to use this information; making intercomparison more difficult). I could imagine that this is especially important for breaks during the early instrumental period, when measurement methods were not yet fully optimized to handle radiation and wetting problems.

It is important to add a stochastic and nonlinear large-scale trend to the data. It should be stochastic so that the homogenizer cannot see how well he did by computing the trend. And it should be nonlinear and contain decadal variability because homogenization algorithms always work and we should not mix our theoretical ideas with the data. Because the STI not only wants to make an intercomparison of homogenization algorithms, but also aims to produce representative errors, this trend will have to be modelling in a more realistic way as HOME did. I would suggest modelling the trend and the decadal variability separately and to vary it gradually on continental scales.

One of the outcomes of the study was that outliers are not important for the quality of the homogenization. Thus it may be best not to insert outliers in the benchmark. That would save work in inserting them and in removing them again later in the analysis of the homogenized data and homogenizers would not feel they have to program additional processing for outliers.

If I understood the science plan of the STI right, there will be multiple data levels (images of log books, digitized data, data in SI units, merged data, quality controlled data, homogenized data and gridded data). And everyone is invited to produce data or implement algorithms for the various levels. That would mean that there could be multiple input dataset for our benchmarking exercise. For homogenization, especially various ways of merging the data could be important. Merging methods that do not merge stations that move (to produce data that is suited to study local effects and relations with other (surface) parameters and variables) would produce much shorter time series than merging methods that aim to produce long time series (to make it easier to study secular trends due to climate change). The former case may be more difficult to homogenize because the time series are shorter, but may also be easier because overlapping data is not removed.

Many algorithms will likely not be able to automatically use the available metadata. For the intercomparison component of the project, it may thus be worthwhile to also have a dataset without metadata.

Tuesday, 5 July 2011

Benchmarking temperature networks

The COST Action HOME has just finished a manuscript on benchmarking monthly of homogenization algorithms for regional monthly temperature and precipitation networks. I think it turned out quite interesting and provides a good base for discussions on the work of the surface temperatures benchmarking group; if only to avoid making the same mistakes.