Showing posts with label HOME. Show all posts
Showing posts with label HOME. Show all posts

Tuesday, 10 January 2012

New article: Benchmarking homogenization algorithms for monthly data

The main paper of the COST Action HOME on homogenization of climate data has been published today in Climate of the Past. This paper could be seen as a pre-study for the benchmarking in the international surface temperature initiative (ISTI)

Inhomogeneities

To study climatic variability the original observations are indispensable, but not directly usable. Next to real climate signals they may also contain non-climatic changes. Corrections to the data are needed to remove these non-climatic influences, this is called homogenisation. The best known non-climatic change is the urban heat island effect. The temperature in cities can be warmer than on the surrounding country side, especially at night. Thus as cities grow, one may expect that temperatures measured in cities become higher. On the other hand, many stations have been relocated from cities to nearby, typically cooler, airports. Other non-climatic changes can be caused by changes in measurement methods. Meteorological instruments are typically installed in a screen to protect them from direct sun and wetting. In the 19th century it was common to use a metal screen on a North facing wall. However, the building may warm the screen leading to higher temperature measurements. When this problem was realised the so-called Stevenson screen was introduced, typically installed in gardens, away from buildings. This is still the most typical weather screen with its typical double-louvre door and walls. Nowadays automatic weather stations, which reduce labor costs, are becoming more common; they protect the thermometer by a number of white plastic cones. This necessitated changes from manually recorded liquid and glass thermometers to automated electrical resistance thermometers, which reduces the recorded temperature values.



One way to study the influence of changes in measurement techniques is by making simultaneous measurements with historical and current instruments, procedures or screens. This picture shows three meteorological shelters next to each other in Murcia (Spain). The rightmost shelter is a replica of the Montsouri screen, in use in Spain and many European countries in the late 19th century and early 20th century. In the middle, Stevenson screen equipped with automatic sensors. Leftmost, Stevenson screen equipped with conventional meteorological instruments.
Picture: Project SCREEN, Center for Climate Change, Universitat Rovira i Virgili, Spain.


A further example for a change in the measurement method is that the precipitation amounts observed in the early instrumental period (about before 1900) are biased and are 10% lower than nowadays because the measurements were often made on a roof. At the time, instruments were installed on rooftops to ensure that the instrument is never shielded from the rain, but it was found later that due to the turbulent flow of the wind on roofs, some rain droplets and especially snow flakes did not fall into the opening. Consequently measurements are nowadays performed closer to the ground.

Homogenization

To reliably study the real development of the climate, non-climatic changes have to be removed. For this the small difference of one station to its direct neighbours are utilized. In this way non-climatic changes (shelter and instrument changes or station moves usually) in a single stations can be more clearly seen as in the record of one station by itself due to the strong natural climatic variability. This method does not work when changes are applied to a whole country’s network. Such extensive changes are less problematic, however, because are typically well documented.




Meteorological window suggested by Italian Central Office for Meteorology and Climate in 1879 (Tacchini, 1879). In the last decades of the 19th century most of Italian observations were performed in urban environments, in screens located outside a north-facing window of the highest floor of a “meteorological tower”. The purpose of using such towers was to perform observations above the level of the roofs of the surrounding buildings. Picture: Michele Brunetti, ISAC-CNR, Bologna, Italy.

Benchmarking

To study the performance of the various homogenisation methods, the COST Action HOME has performed a test with artificial climate data. The advantage of artificial data is that the non-climatic changes are known to those who created the data. The artificial data used mimics climatic networks and their data problems with unprecedented realism. For this we have used the IAAFT algorithm, which can generate non-Gaussian data with arbitrary temporal variability and cross-correlations between the stations. The artificial data may have a warming, a cooling or no trend, to ensure objective testing of the methods. The main novelty is that the test was blind. In other words, while homogenising the data the scientists did not know which station contained which non-climatic problem. The artificial data were generated and the analysis of results was performed by independent researchers, who did not homogenise the data themselves. Consequently, the COST Action is sure that the results are an honest appraisal of the true power of homogenisation algorithms.

Some people remaining sceptical of climate change claim that adjustments applied to the data by climatologists, to correct for the issues described above, lead to overestimates of global warming. The results clearly show that homogenisation improves the quality of temperature records and makes the estimate of climatic trends more accurate.



The photo on the right shows an open shelter for meteorological instruments at the edge of the school square of the primary school of La Rochelle, in 1910. La Rochelle is a coastal city in western France and a seaport on the Bay of Biscay. On the left one sees the current situation, a Stevenson-like screen located closer to the ocean, along the Atlantic shore, in place named "Le bout blanc". Behind the fence you see the water of the port. Picture: Olivier Mestre, Meteo France, Toulouse, France.

Methodological advances in homogenization

In the past it was customary in homogenisation to compare a station with its neighbours by creating a reference time series from averaging over multiple neighbouring stations. Due to the averaging the influence of random non-climatic factors is strongly reduced. Thus if a jump was found in the difference time series of a station with its reference, the jump was assumed to be in the station, not in the reference, which was assumed to be homogeneous. In recent years climatologists and statisticians have worked on advanced statistical methods that do not need a homogeneous reference. The traditional methods reduced the influence of non-climatic factors on the temperature measurements, but the complex modern methods clearly improved the data much more. This finding could only be reached using the benchmark data simulating complete networks with realistic non-climatic problems. Thus now we can recommend with confidence that climatologists should use the new methods. These recommendations are, of course, not only based on the numerical results, but also on our mathematical understanding of the algorithms.

Open-access publishing

The scientific article with 31 authors describing this study has been published today in the journal Climate of the Past. This international journal is an open-access and an open-review journal of the European Geosciences Union. The articles of open-access journals can be freely read by anyone; the costs of publication are born by the authors. Open-access publishing makes it easier for researchers, also from poorer countries, to stay up to date and to participate in science. Also the general public can profit from open-access publishing as the access to the primary source can make the public debate on current scientific issues in newspapers and blogs more informed. Especially for this topic, we felt it was important that everyone can read the article. Next to many EGU journals, last week also the meteorological journal Tellus joined the open access movement.

Climate of the Past is also an open-review journal. This new way of reviewing scientific articles is public, everyone has the possibility to respond to the initial draft of the paper and everyone can read these comments as well as the comments of the official peer reviewers of the manuscript.

For more information

Venema, V., O. Mestre, E. Aguilar, I. Auer, J.A. Guijarro, P. Domonkos, G. Vertacnik, T. Szentimrey, P. Stepanek, P. Zahradnicek, J. Viarre, G. Müller-Westermeier, M. Lakatos, C.N. Williams, M. Menne, R. Lindau, D. Rasol, E. Rustemeier, K. Kolokythas, T. Marinova, L. Andresen, F. Acquaotta, S. Fratianni, S. Cheval, M. Klancar, M. Brunetti, Ch. Gruber, M. Prohom Duran, T. Likso, P. Esteban, Th. Brandsma. Benchmarking homogenization algorithms for monthly data. , Climate of the Past, 8, pp. 89-115, 2012.

If you would like to analyse the data used in this study, please go to this page for a link to the data as well as to documents that describe the dataset and data formats in detail.

The homepage of the Action HOME with amongst others a bibliography with most if not all articles on homogenization of climate networks.

If you are interested in homogenization, please send me an e-mail and I will put you on our email distribution list.

Thursday, 14 July 2011

Generating inhomogeneous worlds

The manuscript on the monthly temperature and precipitation benchmark of the COST Action HOME is now finished: manuscript. The inhomogeneities applied in this study, together with ideas for improvements in the discussion and outlook, may be a good basis for the benchmarking of the surface temperature initiative as well. Thus here I will only mention where I would suggest to do thinks differently for the surface temperatures initiative (STI). Next to improvements due to lessons learned from this study, the STI differs in two important aspects. 1) HOME considered regional networks, STI global datasets. 2) HOME focused on intercomparison of the homogenization algorithms, STI has the additional ambition that the benchmarking leads error estimates for the STI database.

I agree with Kate Willett that it would be valuable to have both data for which the truth is known as well as data for which the truth will be revealed after all homogenized contributions have been submitted. The latter leads to more reliable results because the homogenization algorithms cannot be tuned to solve the benchmark data well (and then possibly perform better on the benchmark than real data).
Data for which the truth is known would be a more classical validation study, which has the advantage that the homogenizers can learn during the exercise and also find (programming) errors; see outlook of our paper. In case of a benchmark one would have to wait another 3 year cycle to implement bug fixes. In HOME we had a number of such bugs, often in the parts newly written, e.g. to be able to handle multiple networks (which in not needed in daily work).

In HOME, we had random and clustered breaks. The clustered breaks occurred with a probability of 30% and if they occurred they affected 30% of the stations. As HOME generated regional networks, we could not insert breaks that occur in all stations, because they would not be detectable by relative homogenization. As the STI will generate global data, it is possible to insert breaks that occur simultaneously in one entire regional network, one would see them at the border of the countries. It would be interesting to see if the homogenization algorithms are able to move the information at the border into the heart of the network. However, also in this case, always inserting the clustered breaks in all stations would not be realistic. One country often has multiple networks (synoptic, volunteer networks or measurements by multiple organizations). Thus partial clustered breaks are also needed. In HOME the clustered breaks were perfectly simultaneous, which mimics a change in observational rules. Breaks clustered over a period, which mimics a change in the instrumentation or screens, would also be valuable.
The clustered breaks are not only correlated in time, but also in size. This is needed to generated biases in the continental and global averages due to inhomogeneities. This is seen in real data, especially in the early instrumental period and now again in the transition to automatic weather stations. These two periods may warrant special rules, i.e. maybe the timing of the clustered should not be fully random.

Another aspect that was not touched in the intercomparison study HOME, but needs to be treated in the STI are multiple elements. For many stations more than one temperature dataset will be available: Tmin, Tmax, Tmean, DTR (diurnal temperature range), or temperature observations for specific hours. Methods using multiple elements simultaneously likely perform better and if they are applied to the real data, they should also be validated in the benchmarking. Then we would need to know how the size of the inhomogeneities correlates between the various temperature variables. A break in Tmin, does not imply a (statistically significant) break in Tmax, but probably does make it more likely. Similar relations will likely hold among all temperature measures. I am not sure if there is data from literature on such cross-relations. On the other hand, you may also want to have some worlds with univariate validation data for the intercomparison of algorithms that do not use this information.

In the HOME benchmark, some stations had a local trend. The statistical properties of these trends were idealised, which was acceptable for an intercomparison, but to obtain realistic error estimates it should be studied in more detail how often such local trends occur in real datasets.

In HOME the perturbation were a constant number for every month. In the STI we have the opportunity to make the perturbations a function of insolation, wind and precipitation. This would make them partially stochastic, which is more realistic. In this case we would have to decide whether to make these covariates available to the homogenizers (potentially better homogenization) or not (most algorithms will currently not be able to use this information; making intercomparison more difficult). I could imagine that this is especially important for breaks during the early instrumental period, when measurement methods were not yet fully optimized to handle radiation and wetting problems.

It is important to add a stochastic and nonlinear large-scale trend to the data. It should be stochastic so that the homogenizer cannot see how well he did by computing the trend. And it should be nonlinear and contain decadal variability because homogenization algorithms always work and we should not mix our theoretical ideas with the data. Because the STI not only wants to make an intercomparison of homogenization algorithms, but also aims to produce representative errors, this trend will have to be modelling in a more realistic way as HOME did. I would suggest modelling the trend and the decadal variability separately and to vary it gradually on continental scales.

One of the outcomes of the study was that outliers are not important for the quality of the homogenization. Thus it may be best not to insert outliers in the benchmark. That would save work in inserting them and in removing them again later in the analysis of the homogenized data and homogenizers would not feel they have to program additional processing for outliers.

If I understood the science plan of the STI right, there will be multiple data levels (images of log books, digitized data, data in SI units, merged data, quality controlled data, homogenized data and gridded data). And everyone is invited to produce data or implement algorithms for the various levels. That would mean that there could be multiple input dataset for our benchmarking exercise. For homogenization, especially various ways of merging the data could be important. Merging methods that do not merge stations that move (to produce data that is suited to study local effects and relations with other (surface) parameters and variables) would produce much shorter time series than merging methods that aim to produce long time series (to make it easier to study secular trends due to climate change). The former case may be more difficult to homogenize because the time series are shorter, but may also be easier because overlapping data is not removed.

Many algorithms will likely not be able to automatically use the available metadata. For the intercomparison component of the project, it may thus be worthwhile to also have a dataset without metadata.

Tuesday, 5 July 2011

Benchmarking temperature networks

The COST Action HOME has just finished a manuscript on benchmarking monthly of homogenization algorithms for regional monthly temperature and precipitation networks. I think it turned out quite interesting and provides a good base for discussions on the work of the surface temperatures benchmarking group; if only to avoid making the same mistakes.