Thursday, 14 July 2011

Generating inhomogeneous worlds

The manuscript on the monthly temperature and precipitation benchmark of the COST Action HOME is now finished: manuscript. The inhomogeneities applied in this study, together with ideas for improvements in the discussion and outlook, may be a good basis for the benchmarking of the surface temperature initiative as well. Thus here I will only mention where I would suggest to do thinks differently for the surface temperatures initiative (STI). Next to improvements due to lessons learned from this study, the STI differs in two important aspects. 1) HOME considered regional networks, STI global datasets. 2) HOME focused on intercomparison of the homogenization algorithms, STI has the additional ambition that the benchmarking leads error estimates for the STI database.

I agree with Kate Willett that it would be valuable to have both data for which the truth is known as well as data for which the truth will be revealed after all homogenized contributions have been submitted. The latter leads to more reliable results because the homogenization algorithms cannot be tuned to solve the benchmark data well (and then possibly perform better on the benchmark than real data).
Data for which the truth is known would be a more classical validation study, which has the advantage that the homogenizers can learn during the exercise and also find (programming) errors; see outlook of our paper. In case of a benchmark one would have to wait another 3 year cycle to implement bug fixes. In HOME we had a number of such bugs, often in the parts newly written, e.g. to be able to handle multiple networks (which in not needed in daily work).

In HOME, we had random and clustered breaks. The clustered breaks occurred with a probability of 30% and if they occurred they affected 30% of the stations. As HOME generated regional networks, we could not insert breaks that occur in all stations, because they would not be detectable by relative homogenization. As the STI will generate global data, it is possible to insert breaks that occur simultaneously in one entire regional network, one would see them at the border of the countries. It would be interesting to see if the homogenization algorithms are able to move the information at the border into the heart of the network. However, also in this case, always inserting the clustered breaks in all stations would not be realistic. One country often has multiple networks (synoptic, volunteer networks or measurements by multiple organizations). Thus partial clustered breaks are also needed. In HOME the clustered breaks were perfectly simultaneous, which mimics a change in observational rules. Breaks clustered over a period, which mimics a change in the instrumentation or screens, would also be valuable.
The clustered breaks are not only correlated in time, but also in size. This is needed to generated biases in the continental and global averages due to inhomogeneities. This is seen in real data, especially in the early instrumental period and now again in the transition to automatic weather stations. These two periods may warrant special rules, i.e. maybe the timing of the clustered should not be fully random.

Another aspect that was not touched in the intercomparison study HOME, but needs to be treated in the STI are multiple elements. For many stations more than one temperature dataset will be available: Tmin, Tmax, Tmean, DTR (diurnal temperature range), or temperature observations for specific hours. Methods using multiple elements simultaneously likely perform better and if they are applied to the real data, they should also be validated in the benchmarking. Then we would need to know how the size of the inhomogeneities correlates between the various temperature variables. A break in Tmin, does not imply a (statistically significant) break in Tmax, but probably does make it more likely. Similar relations will likely hold among all temperature measures. I am not sure if there is data from literature on such cross-relations. On the other hand, you may also want to have some worlds with univariate validation data for the intercomparison of algorithms that do not use this information.

In the HOME benchmark, some stations had a local trend. The statistical properties of these trends were idealised, which was acceptable for an intercomparison, but to obtain realistic error estimates it should be studied in more detail how often such local trends occur in real datasets.

In HOME the perturbation were a constant number for every month. In the STI we have the opportunity to make the perturbations a function of insolation, wind and precipitation. This would make them partially stochastic, which is more realistic. In this case we would have to decide whether to make these covariates available to the homogenizers (potentially better homogenization) or not (most algorithms will currently not be able to use this information; making intercomparison more difficult). I could imagine that this is especially important for breaks during the early instrumental period, when measurement methods were not yet fully optimized to handle radiation and wetting problems.

It is important to add a stochastic and nonlinear large-scale trend to the data. It should be stochastic so that the homogenizer cannot see how well he did by computing the trend. And it should be nonlinear and contain decadal variability because homogenization algorithms always work and we should not mix our theoretical ideas with the data. Because the STI not only wants to make an intercomparison of homogenization algorithms, but also aims to produce representative errors, this trend will have to be modelling in a more realistic way as HOME did. I would suggest modelling the trend and the decadal variability separately and to vary it gradually on continental scales.

One of the outcomes of the study was that outliers are not important for the quality of the homogenization. Thus it may be best not to insert outliers in the benchmark. That would save work in inserting them and in removing them again later in the analysis of the homogenized data and homogenizers would not feel they have to program additional processing for outliers.

If I understood the science plan of the STI right, there will be multiple data levels (images of log books, digitized data, data in SI units, merged data, quality controlled data, homogenized data and gridded data). And everyone is invited to produce data or implement algorithms for the various levels. That would mean that there could be multiple input dataset for our benchmarking exercise. For homogenization, especially various ways of merging the data could be important. Merging methods that do not merge stations that move (to produce data that is suited to study local effects and relations with other (surface) parameters and variables) would produce much shorter time series than merging methods that aim to produce long time series (to make it easier to study secular trends due to climate change). The former case may be more difficult to homogenize because the time series are shorter, but may also be easier because overlapping data is not removed.

Many algorithms will likely not be able to automatically use the available metadata. For the intercomparison component of the project, it may thus be worthwhile to also have a dataset without metadata.

3 comments:

Kate Willett said...

Thanks Victor - some comments from me:

This is all really valuable - next call is time to iron out what we're going to put in these analog-error-worlds. The work done during COST HOME will make an excellent starting point.

Using spatially clustered breaks that where some are very close (temporally) and some are identical will be useful to approximate nationwide changes.

In the first instance we are using temperature alone but including multiple variables with physically and temporally consistent breaks should be considered for later versions.

Local trends will be part of the GCM base for the analog-known-worlds. These will differ between choice of forcing run and also regionally. They will be non-linear and unknown to users for the official benchmark release.

Ideally I would like to add breaks that are physically based on changes in insolation, wind and perhaps precip. In practise this could be done using simultaneous information from the GCM or by using approximations of these measures.

Useful info on outliers.

Data-product creators would be encouraged to use stage 3 (consolidated common format) data for QC and homogenisation so our benchmarks would aim to replicate this level of data but include metadata in some cases.

Victor Venema said...

With multiple variables, I was thinking of Tmean, Tmax and Tmin. Some datasets will have only Tmean, others Tmax and Tmin and yet others will have all three. For an intercomparison study it would be sufficient to just use Tmean, but if our worlds should be an image of the real data, then we will have to generate a multivariate benchmark, including sensible cross relations between the breaks in all these temperature measures.

Will there be only one stage 3 temperature product? Or will there the multiple ones because various methods to merge station data to longer time series are applied?

Kate Willett said...

I think there will be one stage 3 dataset but that this may contain multiple versions of the same station. These should be clearly identifiable as 'duplicate' versions of a station. A user will have to make a decision which one to use.

My aim for the first round of benchmarks was to limit this to Tmean for simplicity but if all goes well then including Tmax and Tmin would have high value and is arguably essential in order to ensure physically consistent breakpoint detection and adjustment.