Wednesday, 5 November 2014

Release of a daily benchmark dataset - version 1

Kate Willett's blog post from 6th October gives a detailed over-view of the benchmarking process that forms part of the ISTI's aims. It is hoped that in the long term these benchmarks will not only be produced at the monthly level, but also for daily data.

This post announces the release of a smaller daily benchmark dataset focusing on four regions in North America. These regions can be seen in Figure 1.


Figure 1 Station locations of the four benchmark regions. Blue stations are in all worlds. Red stations only appear in worlds 2 and 3.

These benchmarks have similar aims to the global benchmarks that are currently being produced by the ISTI working group, namely to:
 

  1. Assess the performance of current homogenisation algorithms and provide feedback to allow for their improvement 
  2. Assess how realistic the created benchmarks are, to allow for improvements in future iterations 
  3. Quantify the uncertainty that is present in data due to inhomogeneities both before and after homogenisation algorithms have been run on them

A perfect algorithm would return the inhomogeneous data to their clean form – correctly identifying the size and location of the inhomogeneities and adjusting the series accordingly. The inhomogeneities that have been added will not be made known to the testers until the completion of the assessment cycle – mid 2015. This is to ensure that the study is as fair as possible with no testers having prior knowledge of the added inhomogeneities.

The data are formed into three worlds, each consisting of the four regions shown in Figure 1. World 1 is the smallest and contains only those stations shown in blue in Figure 1, Worlds 2 and 3 are the same size as each other and contain all the stations shown.

Homogenisers are requested to prioritise running their algorithms on a single region across worlds instead of on all regions in a single world. This will hopefully maximise the usefulness of this study in assessing the strengths and weaknesses of the process. The order of prioritisation for the regions is Wyoming, South East, North East and finally the South West.

This study will be more effective the more participants it has and if you are interested in participating please contact Rachel Warren (rw307 AT exeter.ac.uk). The results will form part of a PhD thesis and therefore it is requested that they are returned no later than Friday 12th December 2014. However, interested parties who are unable to meet this deadline are also encouraged to contact Rachel.

There will be a further smaller release in the next week that is just focussed on Wyoming and will explore climate characteristics of data instead of just focusing on inhomogeneity characteristics.

Monday, 6 October 2014

A framework for benchmarking of homogenisation algorithm performance on the global scale - Paper now published

We have just had our first benchmarking paper accepted at Geoscientific Instrumentation, Methods and Data Systems:

Willett, K., Williams, C., Jolliffe, I. T., Lund, R., Alexander, L. V., Brönnimann, S., Vincent, L. A., Easterbrook, S., Venema, V. K. C., Berry, D., Warren, R. E., Lopardo, G., Auchmann, R., Aguilar, E., Menne, M. J., Gallagher, C., Hausfather, Z., Thorarinsdottir, T., and Thorne, P. W.: A framework for benchmarking of homogenisation algorithm performance on the global scale, Geosci. Instrum. Method. Data Syst., 3, 187-200, doi:10.5194/gi-3-187-2014, 2014.

Benchmarking, in this context, is the assessment of homogenisation algorithm performance against a set of realistic synthetic worlds of station data where the locations and size/shape of inhomogeneities are known a priori. Crucially, these inhomogeneities are not known to those performing the homogenisation, only those performing the assessment. Assessment of both the ability of algorithms to find changepoints and accurately return the synthetic data to its clean form (prior to addition of inhomogeneity) has three main purposes:

      1) quantification of uncertainty remaining in the data due to inhomogeneity
      2) inter-comparison of climate data products in terms of fitness for a specified purpose
      3) providing a tool for further improvement in homogenisation algorithms

Here we describe what we believe would be a good approach to a comprehensive homogenisation algorithm benchmarking system. This includes an overarching cycle of: benchmark development; release of formal benchmarks; assessment of homogenised benchmarks and an overview of where we can improve for next time around (Figure 1).

Figure 1 Overview the ISTI comprehensive benchmarking system for assessing performance of homogenisation algorithms. (Fig. 3 of Willett et al., 2014)

There are four components to creating this benchmarking system. 

Creation of realistic clean synthetic station data

Firstly, we must be able to synthetically recreate the 30000+ ISTI stations such that they have the correct variability, auto-correlation and interstation cross-correlations as the real data but are free from systematic error. In other words, they must contain a realistic seasonal cycle and features of natural variability (e.g., ENSO, volcanic eruptions etc.). There must be a realistic persistence month-to-month in each station and geographically across nearby stations. 

Creation of realistic error models to add to the clean station data

The added inhomogeneities should cover all known types of inhomogeneity in terms of their frequency, magnitude and seasonal behaviour. For example, inhomogeneities could be any or a combination of the following:

     -  geographically or temporally clustered due to events which affect entire networks or regions (e.g. change in observation time);
     -  close to end points of time series;
     -  gradual or sudden;
     -  variance-altering;
     -  combined with the presence of a long-term background trend;
     - small or large;

     - frequent;
     - seasonally varying.

Design of an assessment system

Assessment of the homogenised benchmarks should be designed with the three purposes of benchmarking in mind. Both the ability to correctly locate changepoints and to adjust the data back to its homogeneous state are important. It can be split into four different levels:

     - Level 1: The ability of the algorithm to restore an inhomogeneous world to its clean world state in terms of climatology, variance and trends.

     - Level 2: The ability of the algorithm to accurately locate changepoints and detect their size/shape.

     - Level 3: The strengths and weaknesses of an algorithm against specific types of inhomogeneity and observing system issues.

     - Level 4: A comparison of the benchmarks with the real world in terms of detected inhomogeneity both to measure algorithm performance in the real world and to enable future improvement to the benchmarks.

The benchmark cycle

This should all take place within a well laid out framework to encourage people to take part and make the results as useful as possible. Timing is important. Too long a cycle will mean that the benchmarks become outdated. Too short a cycle will reduce the number of groups able to participate.

Producing the clean synthetic station data on the global scale is a complicated task that has now taken several years but we are close to completion of a version 1. We have collected together a list of known regionwide inhomogeneities and a comprehensive understanding of the many many different types of inhomogeneities that can affect station data. We have also considered a number of assessment options and decided to focus on levels 1 and 2 for assessment within the benchmark cycle. Our benchmarking working group is aiming for release of the first benchmarks by January 2015.

Thursday, 7 November 2013

Summary of Regional Inhomogeneities

A few months ago the Benchmarking group made a call to the homogenisation community to submit any times/characters of known inhomogeneities occurring in different regions.

Thanks to the many contributors we now have an overview for a number of countries:
Spain
Poland
Canada
Switzerland
Australia
Tropics
Italy
Central Europe
Russia
Slovenia
Netherlands
UK
Austria
USA

There is always room for more. If you come across any useful information or would just like to see what is there so far please go to:

https://docs.google.com/spreadsheet/ccc?key=0Al6ocsUAaINSdHpTREJzVkRZUTdfVjNPRlh0Q1V3WUE&usp=sharing#gid=0

This is an online editable document so please do add info. A static version will be scraped once a month.




Kate










ISTI/Benchmarking talk at Reading University Meteorology Department lunchtime seminar

I was invited to present the work of the ISTI at Reading University Meteorology Department.

The talk covered the aims and workings of the ISTI: to facilitate robust climate analysis of land surface temperature.

1) Provide a version controlled, comprehensive, traceable to known origin, openly shared databank of surface temperature data with digitally attached metadata.

2) Set up an international standard for benchmark testing homogenisation algorithms and other methodological choices for building climate data-products to help quantify homogenisation and methodological uncertainty and aid product selection and methodological advancement.

3) Provide a portal for all ISTI related products with user tools to help inform choice of product, visualisation and intercomparison of products.

I then focussed on the benchmarking side of things to explain why this is needed and how we are going about it.

The slides for the talk are available here: The ISTI: Dragging the land surface temperature data kicking and screaming into the 21st Century

A few of the interesting questions that I remember:

Will the ISTI databank include non-standard station data?
Yes, there are plans to incorporate data from amateur stations. This will be flagged and prioritised appropriately.

How do you deal with uncertainties between the true shaded air temperature verses the biases created by poor ventilation or insolation issues in screen temperatures?
This is a big problem as we need to quantify changes/impacts on the smaller, more local scale. By making as much data available as possible, in addition to metadata, the ISTI can help address this problem. Having a higher density of stations increases the statistical power in detection of errors. Making it easier to build surface temperature products enables more people to tackle the problem and apply a wider spread of methodological choices. This may help to constrain the uncertainty to come extent.

Can you provide a list of stations used to make up each gridbox in gridded products?
This is something that we would encourage anyone building data-products to do. Ideally they would list the stations and which stage and version they come from. The data-portal will have to be able to cope with hosting ancillary information pertaining to the data-products in an easily searchable and usable manner.

Can other people make their own error worlds/synthetic data?
Yes - all of our code will be written in R and published on the website. The idea is that others can use the code and tweak various parameters to make their own clean synthetic stations and add errors as they wish.

Monday, 19 August 2013

Benchmarking Workshop Agenda and Report

The workshop agenda and full report can be found here: https://sites.google.com/a/surfacetemperatures.org/home/benchmarking-and-assessment-working-group#Minutes

Below is the executive summary.

1st – 3rd July 2013 Benchmarking Working Group Workshop Report Executive Summary National Climatic Data Center (NCDC) of the National Oceanic and Atmospheric Administration (NOAA), Asheville, NC, USA

Attended in person:
Kate Willett (UK), Matt Menne (USA), Claude Williams (USA), Robert Lund (USA), Enric Aguilar (Spain), Colin Gallagher (USA), Zeke Hausfather (USA), Peter Thorne (USA), Jared Rennie (USA)
 

Attended by phone:
Ian Jolliffe (UK), Lisa Alexander (Australia), Stefan Brönniman (Switzerland), Lucie A. Vincent (Canada), Victor Venema (Germany), Renate Auchmann (Switzerland), Thordis Thorarinsdottir (Norway), Robert Dunn (UK), David Parker (UK)


A three day workshop was held to bring together some members of the ISTI Benchmarking working group with the aim of making significant progress towards the creation and dissemination of a homogenisation algorithm benchmark system. Specifically, we hoped to have: the method for creating the analog-clean-worlds finalised; the error-model worlds defined and a plan of how to develop these; and the concepts for assessment finalised including a decision on what data/statistics to ask users to return. This was an ambitious plan for three days with numerous issues and big decisions still to be tackled.

The complexity of much of the discussion throughout the three days really highlighted the value of this face-to-face meeting. It was important to take time to ensure that everyone understood and had come to the same conclusion. This was aided by whiteboard illustrations and software exploration, which would not have been possible over a teleconference.

In overview, we made significant progress in terms of developing and converging on concepts and important decisions. We did not complete the work of Team Creation as hoped, but necessary exploration of the existing methods was undertaken revealing significant weaknesses and ideas for new avenues to explore have been found.

The blind and open error-worlds concepts are 95% complete and progress was made on the specifics of the changepoint statistics for each world. Important decisions were also made regarding missing data, length of record and changepoint location frequency. Seasonal cycles were discussed at length and more research has been actioned. A significant first go was made at designing a build methodology for the error-models with some coding examples worked through and different probability distributions explored.

We converged on what we would like to receive from benchmark users for the assessment and worked through some examples of aggregating station results over regions. We will assess both retrieval of trends and climate characteristics in addition to ability to detect changepoints. Contingency tables of some form will also be used. We also hope to have some online or assessment software available so that users can make their own assessment of the open worlds and past versions of benchmarks. We plan to collaborate with the VALUE downscaling validation project where possible.

From an intense three days all participants and teleconference participants gained a better understanding of what we're trying to achieve and how we are going to get there. This was a highly valuable three days, not least through its effect of focussing our attention prior to the meeting and motivating further collaborative work after the meeting. Two new members have agreed to join the effort and their expertise is a fantastic contribution to the project.

Specifically, Kate and Robert are to work on their respective methods for Team Creation, utilising GCM data and the vector autoregressive method. This will result in a publication describing the methodology. We aim to finalise this work in August.

Follow on teleconferences, Team Corruption will focus on completing the distribution specifications and building the probability model to allocate station changepoints. This work is planned for completion by October 2013. Release of the benchmarks is scheduled for November 2013.

Team Validation will continue to develop the specific assessment tests and work these into a software package that can be easily implemented. This work is hoped to be completed by December 2013, but there is more time available as assessment will take place at least 1 year after benchmark release.


Monday, 13 May 2013

Call for regional inhomogeneity info

To create realistic benchmarks we would like to reproduce times and locations of known sources of inhomogeneity as best we can. Please can you help us. If you know of any regional/countrywide changes to the observing system over time please can you list them here or point us to some documentation/reference. Any information is valuable - even if its quite vague.

Ideally we'd like to know:

WHEN - specific date or month or year or even decade etc.
WHERE - a region, a country, an international GTS/WMO change etc.
WHAT - a change in shelter, thermometer type, automation, observing time/practice etc.
HOW - are there any estimates of the size/direction/nature of the effect of this change?

Please post here and encourage others to do so. We then hope to reward you with some realistic error-worlds to play with.

Kate

Sunday, 15 January 2012

Mailing list on homogenisation of climate data

There is now a homogenisation mailing list, which aims to strengthen the communication and co-operation in the scientific community on homogenenisation. Suggested uses of the list are announcements on:

  • Conferences, workshops, etc.
  • Important papers
  • Discussions
  • Job opportunities
  • New projects started
  • Requests for data, information, and cooperation partners
  • ...
For more information (on how to subscribe), please to go my page on the homogenisation mailing list.

Tuesday, 10 January 2012

New article: Benchmarking homogenization algorithms for monthly data

The main paper of the COST Action HOME on homogenization of climate data has been published today in Climate of the Past. This paper could be seen as a pre-study for the benchmarking in the international surface temperature initiative (ISTI)

Inhomogeneities

To study climatic variability the original observations are indispensable, but not directly usable. Next to real climate signals they may also contain non-climatic changes. Corrections to the data are needed to remove these non-climatic influences, this is called homogenisation. The best known non-climatic change is the urban heat island effect. The temperature in cities can be warmer than on the surrounding country side, especially at night. Thus as cities grow, one may expect that temperatures measured in cities become higher. On the other hand, many stations have been relocated from cities to nearby, typically cooler, airports. Other non-climatic changes can be caused by changes in measurement methods. Meteorological instruments are typically installed in a screen to protect them from direct sun and wetting. In the 19th century it was common to use a metal screen on a North facing wall. However, the building may warm the screen leading to higher temperature measurements. When this problem was realised the so-called Stevenson screen was introduced, typically installed in gardens, away from buildings. This is still the most typical weather screen with its typical double-louvre door and walls. Nowadays automatic weather stations, which reduce labor costs, are becoming more common; they protect the thermometer by a number of white plastic cones. This necessitated changes from manually recorded liquid and glass thermometers to automated electrical resistance thermometers, which reduces the recorded temperature values.



One way to study the influence of changes in measurement techniques is by making simultaneous measurements with historical and current instruments, procedures or screens. This picture shows three meteorological shelters next to each other in Murcia (Spain). The rightmost shelter is a replica of the Montsouri screen, in use in Spain and many European countries in the late 19th century and early 20th century. In the middle, Stevenson screen equipped with automatic sensors. Leftmost, Stevenson screen equipped with conventional meteorological instruments.
Picture: Project SCREEN, Center for Climate Change, Universitat Rovira i Virgili, Spain.


A further example for a change in the measurement method is that the precipitation amounts observed in the early instrumental period (about before 1900) are biased and are 10% lower than nowadays because the measurements were often made on a roof. At the time, instruments were installed on rooftops to ensure that the instrument is never shielded from the rain, but it was found later that due to the turbulent flow of the wind on roofs, some rain droplets and especially snow flakes did not fall into the opening. Consequently measurements are nowadays performed closer to the ground.

Homogenization

To reliably study the real development of the climate, non-climatic changes have to be removed. For this the small difference of one station to its direct neighbours are utilized. In this way non-climatic changes (shelter and instrument changes or station moves usually) in a single stations can be more clearly seen as in the record of one station by itself due to the strong natural climatic variability. This method does not work when changes are applied to a whole country’s network. Such extensive changes are less problematic, however, because are typically well documented.




Meteorological window suggested by Italian Central Office for Meteorology and Climate in 1879 (Tacchini, 1879). In the last decades of the 19th century most of Italian observations were performed in urban environments, in screens located outside a north-facing window of the highest floor of a “meteorological tower”. The purpose of using such towers was to perform observations above the level of the roofs of the surrounding buildings. Picture: Michele Brunetti, ISAC-CNR, Bologna, Italy.

Benchmarking

To study the performance of the various homogenisation methods, the COST Action HOME has performed a test with artificial climate data. The advantage of artificial data is that the non-climatic changes are known to those who created the data. The artificial data used mimics climatic networks and their data problems with unprecedented realism. For this we have used the IAAFT algorithm, which can generate non-Gaussian data with arbitrary temporal variability and cross-correlations between the stations. The artificial data may have a warming, a cooling or no trend, to ensure objective testing of the methods. The main novelty is that the test was blind. In other words, while homogenising the data the scientists did not know which station contained which non-climatic problem. The artificial data were generated and the analysis of results was performed by independent researchers, who did not homogenise the data themselves. Consequently, the COST Action is sure that the results are an honest appraisal of the true power of homogenisation algorithms.

Some people remaining sceptical of climate change claim that adjustments applied to the data by climatologists, to correct for the issues described above, lead to overestimates of global warming. The results clearly show that homogenisation improves the quality of temperature records and makes the estimate of climatic trends more accurate.



The photo on the right shows an open shelter for meteorological instruments at the edge of the school square of the primary school of La Rochelle, in 1910. La Rochelle is a coastal city in western France and a seaport on the Bay of Biscay. On the left one sees the current situation, a Stevenson-like screen located closer to the ocean, along the Atlantic shore, in place named "Le bout blanc". Behind the fence you see the water of the port. Picture: Olivier Mestre, Meteo France, Toulouse, France.

Methodological advances in homogenization

In the past it was customary in homogenisation to compare a station with its neighbours by creating a reference time series from averaging over multiple neighbouring stations. Due to the averaging the influence of random non-climatic factors is strongly reduced. Thus if a jump was found in the difference time series of a station with its reference, the jump was assumed to be in the station, not in the reference, which was assumed to be homogeneous. In recent years climatologists and statisticians have worked on advanced statistical methods that do not need a homogeneous reference. The traditional methods reduced the influence of non-climatic factors on the temperature measurements, but the complex modern methods clearly improved the data much more. This finding could only be reached using the benchmark data simulating complete networks with realistic non-climatic problems. Thus now we can recommend with confidence that climatologists should use the new methods. These recommendations are, of course, not only based on the numerical results, but also on our mathematical understanding of the algorithms.

Open-access publishing

The scientific article with 31 authors describing this study has been published today in the journal Climate of the Past. This international journal is an open-access and an open-review journal of the European Geosciences Union. The articles of open-access journals can be freely read by anyone; the costs of publication are born by the authors. Open-access publishing makes it easier for researchers, also from poorer countries, to stay up to date and to participate in science. Also the general public can profit from open-access publishing as the access to the primary source can make the public debate on current scientific issues in newspapers and blogs more informed. Especially for this topic, we felt it was important that everyone can read the article. Next to many EGU journals, last week also the meteorological journal Tellus joined the open access movement.

Climate of the Past is also an open-review journal. This new way of reviewing scientific articles is public, everyone has the possibility to respond to the initial draft of the paper and everyone can read these comments as well as the comments of the official peer reviewers of the manuscript.

For more information

Venema, V., O. Mestre, E. Aguilar, I. Auer, J.A. Guijarro, P. Domonkos, G. Vertacnik, T. Szentimrey, P. Stepanek, P. Zahradnicek, J. Viarre, G. MĂ¼ller-Westermeier, M. Lakatos, C.N. Williams, M. Menne, R. Lindau, D. Rasol, E. Rustemeier, K. Kolokythas, T. Marinova, L. Andresen, F. Acquaotta, S. Fratianni, S. Cheval, M. Klancar, M. Brunetti, Ch. Gruber, M. Prohom Duran, T. Likso, P. Esteban, Th. Brandsma. Benchmarking homogenization algorithms for monthly data. , Climate of the Past, 8, pp. 89-115, 2012.

If you would like to analyse the data used in this study, please go to this page for a link to the data as well as to documents that describe the dataset and data formats in detail.

The homepage of the Action HOME with amongst others a bibliography with most if not all articles on homogenization of climate networks.

If you are interested in homogenization, please send me an e-mail and I will put you on our email distribution list.

Friday, 6 January 2012

Benchmarking of USHCN

Cross-posted from the main blog. Please submit comments there.

Firstly, an up-front caveat, I am third (last) author on this paper.

Today the Journal of Geophysical Research has published a paper that applies the benchmarking and assessment principles of the surface temperature initiative to the USHCN dataset of land surface air temperatures.

Williams, C. N., Jr., M. J. Menne, and P. Thorne
Benchmarking the performance of pairwise homogenization of surface temperatures in the United States
J. Geophys. Res., doi:10.1029/2011JD016761, in press.  (behind a paywall - sorry)

The analysis takes the pairwise homogenization algorithm used to create the GHCN and USHCN products and does two things.

Firstly, it identifies a large number of decision points within the algorithm that do not have an absolute basis and allows these to vary. A good climate science analogy here is the perturbed climate model runs of the climateprediction.net project and other similar projects. These decision points were varied by random seeding of values to create a 100 member ensemble of solutions. This at least starts to explore the parametric uncertainties within the algorithm (and any interdependency's) and their implications for our understanding of the observed temperature record evolution. However, what it does not necessarily do is give us any better an idea as to what the true climate evolution may have been. Which brings us on to the second innovation and the focus of this post ...

Secondly, in addition to running on the observations these ensembles were run on a set of eight analogs to the USHCN network. These consisted of sets of data which directly mimicked the observational availability of the USHCN network itself through time. They were based upon climate model runs from a range of models and a range of forcing scenarios. This ensures some 'plausible' spatio-temporal coherency to the large-scale temperature fields. On top of these were super-imposed additional differences to mimic potential random and systematic influences of instrumental and operational artifacts. A set of distinct possibilities were explored across the eight worlds ranging from no systematic biases at all (highly improbable but a useful 'algorithm does no harm' test) through to a scenario where the network was bedevilled with very many largely small breaks with a sign-bias tendency - a situation which any algorithm would find hard to cope with. Unlike the real-world these analogs afford a luxury of knowing the true answer so that it is possible to actually benchmark and understand the fundamental algorithm performance and any limitations. Then it is possible to re-evaluate the real-world results afresh with these new insights gleaned from such realistic test-beds.

The results from the analogs were broadly encouraging. First and foremost when applied to the data with no breaks added virtually no adjustments were made and the impact on large-scale averaged timeseries and trends was so minuscule a magnifying glass would be required to tell the difference. So, in the implausible eventuality that the raw data are bias free the algorithm really would do no harm. For the other analogs the performance was mixed. Easier cases where breaks were bigger and metadata better it did better. Harder cases it fared worse. Where there was no overall sign bias in the applied breaks the ensemble was spread relatively evenly around the starting data. But where there was an overall bias in the raw data presented to the algorithm it consistently moved the overall data in the right direction but rarely far enough. The implication being that this uncertainty is effectively one-tailed. The chances of over-shooting the adjustments is substantially smaller than the chances of under-shooting the required adjustments.

So, what does the real-world look like when reassessed through this new understanding?

For minimum temperatures over the longest timescales the trends are spread around the raw data. But in the periods 1951 onwards and 1979 onwards it is spread distinctly either side of the raw data. Post 1951 trends in the raw data exhibit too little warming, and post-1979 too much. This is consistent with prior understanding of the biases of the change of time of observation (largely 1950s-1970s; spurious cooling) and move to MMTS sensors (early to mid-1980s and often associated with a microclimate relocation; spurious warming) on Tmin.

For maximum temperatures the ensemble consistently precludes the raw observations over all considered timescales. The raw data are almost certainly biased and show too little warming. Over all periods the ensemble of solutions show more warming than the raw data. Again, this is consistent with current understanding of the impact of time of observation biases (again a spurious cooling) and the transition to MMTS which unlike for minimum temperatures imparts a spurious cooling effect. The less than encouraging implication from the analogs, however, is that in the operational USHCN algorithm we are substantially more likely to be under-estimating the required adjustments, and hence rate of warming in maximum temperatures, than over-estimating it.

Is this the last word on the issue? Certainly not. As this was a first step along this path the analogs were perhaps not as sophisticated as would ideally be the case. And here we hope the benchmarking and assessment group can provide more realistic (and global) analogs later this year. Further, this considered solely parametric uncertainty - varying choices within one algorithm. The larger and more difficult uncertainty to understand is the structural uncertainty that would result from applying fundamentally distinct approaches to the same problem and allow a better exploration of the possible solution space. And here we have to look to you, readers, to develop new and novel approaches to homogenizing the data and submit them to the same raw data and analogs to allow consistent benchmarking and better understanding.

Finally, over coming weeks we will be hosting code, data, and metadata (including the analogs) online. I'll provide an update when it is all up there but given that its several Gb and requires fitting around other duties its not going to be instantaneous by any stretch.

Comments are welcome, but please remember that this is a strictly moderated blog and to follow the house rules.

Monday, 19 December 2011

Metadata

The quality of homogenized data does not only depend on the performance of the homogenization algorithm, but also on the metadata, documentary evidence of possible break points. Therefore, we (benchmarking working group of the international surface temperature initiative) want to include metadata in our benchmark.

The amount and quality of metadata is thus likely important to obtain realistic estimates of the uncertainty in climate variability and trends due to remaining inhomogeneities. For the quantity of metadata we can just mimic the metadata in the ISTI database.

I am wondering whether we have enough information on the quality of the metadata. Especially as it will be difficult to tell whether the meta data was right. That there was no jump in the (monthly mean) data, does not mean that nothing has changed. Do we have any idea about the accuracy of the available (machine-readable) metadata? That the distribution of break sizes at dates with a known break (from metadata) was normally distributed (Menne and Williams, 2005) gives some hope that the quality is good enough.

Friday, 11 November 2011

Team Validation - thoughts from the Homogenisation Meeting

Notes for Team Validation:

Use the existing benchmarks and validation to look at which methods are more or less useful. Importance of looking at both ability to recreate the 'truth' and to detect the different types of breaks so that algorithm creators can get something positive about this – what exactly is causing problems for the algorithms? (station density, break frequency, break magnitude, background trend, seasonal cycle, natural variability, missing data, breaks near end-points etc.).

Reference period – this should be the most recent homogeneous subperiod. This is a problem, especially for algorithms doing seasonal shifts, when the last breakpoint is very close to the end of the record. However, this could be a real break location and so should not deliberately be avoided. Assessment should be aware of this though – algorithms could be penalised by this because they would not be able to model the seasonality effectively but assessments may look like the algorithm is failing because of the types of breaks or another complicating feature that was added – importance of useful assessment.

I think that false alarm rates are very important. I would rather a conservative and low false alarm rate than one that gets a higher number of breaks but adds a lot of error too. The false alarm rate should take into account the impact of incorrectly detecting a break given the adjustment applied. A detected break with a negligible adjustment applied is not so bad.

RMSE error seems to be a simple and useful metric – to root or not to root though?
analog-error-worlds minus analog-known-worlds = FULLRMSE
adjusted analog-error-worlds minus analog-known-worlds = REDUCEDRMSE
REDUCEDRMSE should be less than FULLRMSE if the algorithm is improving the network

Watch out for temporal variation in contingency scores – fewer breaks detected near the end of series?

Validating on annual verses monthly (or daily) – should validate on the highest resolution that will be used – so monthly I would say. This will penalise against flat adjustments but we know that inhomogeneities are not flat changes – this means that the errors added MUST be as realistic as possible.

How to calculate True negatives for the contingency scores?

Ensemble approach – hopefully this approach will be growing in popularity and so we need to be able to cope with this. An argument for keeping validation simple.

Some algorithms are trying to adjust more than just the mean, some of the higher order moments. Do we know enough to be able to add in errors in this way? Can we assess this fairly?

Team Corruption - Thoughts from the Homogenisation Meeting

Notes for Team Corruption:

Need to add in realistic inhomogeneities that do not reward specific algorithms by being too obvious/exaggerated. For example, having an over exaggerated seasonally dependent shift will penalise algorithms with a flat detection/adjustment more than necessary and reward algorithms detecting/adjusting based on strong seasonal shifts. This is a difficult balance to achieve but having final errors added by those not building the algorithms and keeping the benchmarks blind will help.

Specific types of inhomogeneity:
- Add in station moves by cutting a pasting a nearby station series. May have to tweak a little to avoid exact duplication though – could create 'duplicate' stations by using the average of 2-3 neighbouring stations to downscale the GCM gridbox therefore creating a unique but realistic station. These 'duplicate' stations will differ slightly and can be substituted for part of a station series to mimic a station move.
- Instrument change/calibration error – this could be a flatter change but could also be a change to the variance on hourly timescales (not necessarily monthly). Instrument sensitivity may change.
- Shelter change – cotton region to stevenson screen – would be a seasonally varying change
- Manual to automated – more missing data, more repeated data (QC), fewer outliers? (QC), more or less sensitivity?
- Changes in observation times – how will this be manifested in monthly data?
- Significant changes to network density – a very real problem that may be reflected in the analogs anyway as they follow the real station drop-in/out – although do we want 100+ years of benchmarks? If we're shortening the record we need to ensure a similar station fall out in at least one of the worlds. When validating we need to be clear on the reasons why algorithms are failing if possible. 1972 seems to be an important year in ISD (NCDC's global sub-daily data) where vast numbers of digitised records drop out and then come back in in 1973.
- Changes in observation frequency and reporting resolution. Increases in reporting frequency from 6 hourly to hourly may mean that lower minimums/higher maximums are now recorded – and vice versa. Rounding procedures may lead to changes from resolution changes – do they truncate or round?

Have a few established break characteristics to input but make them not too predictable or people will know what to look for.

Reference period – this should be the most recent homogeneous subperiod. This is a problem, especially for algorithms doing seasonal shifts, when the last breakpoint is very close to the end of the record. However, this could be a real break location and so should not deliberately be avoided. Assessment should be aware of this though – algorithms could be penalised by this because they would not be able to model the seasonality effectively but assessments may look like the algorithm is failing because of the types of breaks or another complicating feature that was added – importance of useful assessment.

Future benchmarks:
- should be realistic
- Correlations in perturbations within a network – geographical clusters
- study seasonal cycle
- Provide metadata – some good, some bad, some incomplete, some negligible

Include other key climate features – solar radiation/sunshine duration affects the break characteristics, wind, ENSO etc. Largest effects in clear skies – full solar radiation. This info can be stored from the climate model data when creating the analog-known-worlds for later use by team creation.

Be realistic but also have ability to isolate certain break types/questions to make analysis useful – need for a series of worlds with well posed questions.

Regional knowledge is valuable – how to obtain this?
- Norway: Most breaks due to relocation (55%), screen changes (14%), instrument change (15%), other (15%) - very little effect of changing observer – NOT QUITE SURE HOW THAT ADDS UP TO 100%? SIMULTANEOUS CHANGES?
- France/Germany found most changes due to changes in shelters. Norway may have less changes with shelters because of radiation? Or many changes happen at the same time so difficult to distinguish.
- Norwegian data are composites of multiple nearby stations – not official station moves but later station mergers! Similarly in Czech Republic.

Proportion of known to unknown breaks – I would expect that for most countries there are more 'unknown' breaks than 'known' breaks – Czech has 50% backed up by metadata.

Some algorithms are trying to adjust more than just the mean, some of the higher order moments. Do we know enough to be able to add in errors in this way? Can we assess this fairly?

Team Creation - thoughts from the Homgenisation Meeting

Notes for Team Creation:

Future benchmarks:
should be realistic
realistic outliers/random errors - assume a good QC has been undertaken
insert random missing data (which we will have masked from the real stations anyway)
study frequency and size of local trends (which will come from the climate models)

Adding the noise term – some of this will be uncorrelated with other stations – simple random errors, some of this would be the weather term although how this would play out on monthly timescales is unclear – persistent cold or hot events – these would be correlated across networks. Some kind of simple weather generator? Could this sort of thing be modelled from the real stations? Study periodicities in common or something like that? Could use geospatial statistics to get at spatial covariance and add 'weather' based on these underlying relationships?

May be worth storing some other information from the models to be used by team corruption – incoming solar radiation, windspeed? This wouldn't be public info but could help with 'realistic' error input.

2011 Progress Report Now Published

The 2011 Progress Report has just been accepted by the Steering Committee and is now available on out website: http://www.surfacetemperatures.org/benchmarking-and-assessment-working-group#Working%20Group%20Documents.

Thanks for all the work from the group so far! There's been a lot of discussion of novel concepts. The next phase, arguably the hardest, is to get something up and running by November 2012. One year to go!

Kate

Friday, 26 August 2011

Team Validation

Here are my first set of thoughts on what our team could and should
be doing. There may be things that I've completely overlooked. Please
send any comments you have on omissions, or on any of my thoughts, as
soon as you think of them.


I see three things that fall within our remit:

1. Identify any experts that we would like to join us and issue invitations.

2. Identify which validation/verification techniques we should use.

3. Find or write software/code to implement the chosen techniques.

Let's say a bit more on each of these:

1. Until we have made progress on 2., it is difficult to decide who
best to invite. We could, of course, ask someone with general
verification knowledge rather than someone specialising in the types
of data format we identify in 2. Please email any suggestions to me.
I can think of several, but no one of them stands out as first choice.
At this stage there may not be much for them to get their teeth into.

2. We will not be certain what the data format will be until Team
Corruption have made decisions. I guess this will not be finalised
until sometime in 2012, so we could argue that we can do nothing
until then. However, I'm sure we can make some educated guesses
as to what will need to be validated. If there are some things we can
be fairly certain of, we can make decisions for those formats, but
not waste time considering scenarios that might not be used.

A quick look at some possibilities/questions:

was a changepoint found Yes/No?

was the nature of the changepoint correctly identified - this could
again be Yes/No or it could quantify how closely the magnitude
of a change was estimated. Different types of change would need
different validation methodology.

a key question is whether validation will be done station-by-station
and the results simply added, or whether an attempt will be made
to assess how well the spatial pattern of the 'corrected' data match
the 'true' data? The answer to this would determine whether we
want to bring on board an expert in spatial verification.

will we want to compare the distribution of 'corrected' data with
that of the true 'data', as well as looking at how well individual
corrected and true data sets match?

some ideas are given in Section 2.5 of
whitepaper_Benchmarking_Jun2011_v2.pdf
(available on the group website) and also in Section 5 of the COST
(HOME) paper circulated by Victor on July 5th.

3. This stage needs some serious investment of time, and computing
expertise, and is not something I can contribute much to. We could
certainly do with an expert on this side of things. It can't be started
until we are well advanced with 2. but we could start thinking about
who/how will do this.

Ian

Tuesday, 26 July 2011

Benchmark for real-world problems

We should state here what properties an ideal benchmark data set should have, right?
My wish would be for a data set that is as close as possible to real world problems (in addition to data sets that allow testing your methods until they break, which is of course very important).
As close a possible to real world problems could mean: Use physics-based error models (to simulate instrumental errors), simulate typical reporting errors (there should be plenty of experience around what can be wrong), simulate typical processing errors, etc. We will still not get around adding also simply perturbations in a statistical sense, but I think we can be more realistic than that.
Such a data set necessarily is a subdaily data set, and the monthly benchmark data set would simply be an average of the subdaily data (with an additional simulations of errors that can occur during the production of monthly means). Such a data set would necessarily be based on some sort of climate model or reanalysis data because other variables than temperature would be used, and they would be used in a high resolution.
I volunteer to produce such a data set if requested, but lacking experience on homogenizing data outside Europe, I would have to team up with more experienced people telling me what possibly can go wrong in Africa or the Arctic.

Monday, 25 July 2011

Another radiosonde benchmarking paper

This is a little self serving but a further twist on the radiosonde benchmarking has just become available at JGR. A link to this paper is here . This uses the same benchmarks as used in Titchner et al. but looks at multiple additional impacts and also starts to try to address how you could use results from multiple different estimators to yield a grand unified estimate of the truth. There are, of course, many ways one could go about this step, but clearly such an effort may well be something the group would wish to consider how to approach ...

Thursday, 14 July 2011

Generating inhomogeneous worlds

The manuscript on the monthly temperature and precipitation benchmark of the COST Action HOME is now finished: manuscript. The inhomogeneities applied in this study, together with ideas for improvements in the discussion and outlook, may be a good basis for the benchmarking of the surface temperature initiative as well. Thus here I will only mention where I would suggest to do thinks differently for the surface temperatures initiative (STI). Next to improvements due to lessons learned from this study, the STI differs in two important aspects. 1) HOME considered regional networks, STI global datasets. 2) HOME focused on intercomparison of the homogenization algorithms, STI has the additional ambition that the benchmarking leads error estimates for the STI database.

I agree with Kate Willett that it would be valuable to have both data for which the truth is known as well as data for which the truth will be revealed after all homogenized contributions have been submitted. The latter leads to more reliable results because the homogenization algorithms cannot be tuned to solve the benchmark data well (and then possibly perform better on the benchmark than real data).
Data for which the truth is known would be a more classical validation study, which has the advantage that the homogenizers can learn during the exercise and also find (programming) errors; see outlook of our paper. In case of a benchmark one would have to wait another 3 year cycle to implement bug fixes. In HOME we had a number of such bugs, often in the parts newly written, e.g. to be able to handle multiple networks (which in not needed in daily work).

In HOME, we had random and clustered breaks. The clustered breaks occurred with a probability of 30% and if they occurred they affected 30% of the stations. As HOME generated regional networks, we could not insert breaks that occur in all stations, because they would not be detectable by relative homogenization. As the STI will generate global data, it is possible to insert breaks that occur simultaneously in one entire regional network, one would see them at the border of the countries. It would be interesting to see if the homogenization algorithms are able to move the information at the border into the heart of the network. However, also in this case, always inserting the clustered breaks in all stations would not be realistic. One country often has multiple networks (synoptic, volunteer networks or measurements by multiple organizations). Thus partial clustered breaks are also needed. In HOME the clustered breaks were perfectly simultaneous, which mimics a change in observational rules. Breaks clustered over a period, which mimics a change in the instrumentation or screens, would also be valuable.
The clustered breaks are not only correlated in time, but also in size. This is needed to generated biases in the continental and global averages due to inhomogeneities. This is seen in real data, especially in the early instrumental period and now again in the transition to automatic weather stations. These two periods may warrant special rules, i.e. maybe the timing of the clustered should not be fully random.

Another aspect that was not touched in the intercomparison study HOME, but needs to be treated in the STI are multiple elements. For many stations more than one temperature dataset will be available: Tmin, Tmax, Tmean, DTR (diurnal temperature range), or temperature observations for specific hours. Methods using multiple elements simultaneously likely perform better and if they are applied to the real data, they should also be validated in the benchmarking. Then we would need to know how the size of the inhomogeneities correlates between the various temperature variables. A break in Tmin, does not imply a (statistically significant) break in Tmax, but probably does make it more likely. Similar relations will likely hold among all temperature measures. I am not sure if there is data from literature on such cross-relations. On the other hand, you may also want to have some worlds with univariate validation data for the intercomparison of algorithms that do not use this information.

In the HOME benchmark, some stations had a local trend. The statistical properties of these trends were idealised, which was acceptable for an intercomparison, but to obtain realistic error estimates it should be studied in more detail how often such local trends occur in real datasets.

In HOME the perturbation were a constant number for every month. In the STI we have the opportunity to make the perturbations a function of insolation, wind and precipitation. This would make them partially stochastic, which is more realistic. In this case we would have to decide whether to make these covariates available to the homogenizers (potentially better homogenization) or not (most algorithms will currently not be able to use this information; making intercomparison more difficult). I could imagine that this is especially important for breaks during the early instrumental period, when measurement methods were not yet fully optimized to handle radiation and wetting problems.

It is important to add a stochastic and nonlinear large-scale trend to the data. It should be stochastic so that the homogenizer cannot see how well he did by computing the trend. And it should be nonlinear and contain decadal variability because homogenization algorithms always work and we should not mix our theoretical ideas with the data. Because the STI not only wants to make an intercomparison of homogenization algorithms, but also aims to produce representative errors, this trend will have to be modelling in a more realistic way as HOME did. I would suggest modelling the trend and the decadal variability separately and to vary it gradually on continental scales.

One of the outcomes of the study was that outliers are not important for the quality of the homogenization. Thus it may be best not to insert outliers in the benchmark. That would save work in inserting them and in removing them again later in the analysis of the homogenized data and homogenizers would not feel they have to program additional processing for outliers.

If I understood the science plan of the STI right, there will be multiple data levels (images of log books, digitized data, data in SI units, merged data, quality controlled data, homogenized data and gridded data). And everyone is invited to produce data or implement algorithms for the various levels. That would mean that there could be multiple input dataset for our benchmarking exercise. For homogenization, especially various ways of merging the data could be important. Merging methods that do not merge stations that move (to produce data that is suited to study local effects and relations with other (surface) parameters and variables) would produce much shorter time series than merging methods that aim to produce long time series (to make it easier to study secular trends due to climate change). The former case may be more difficult to homogenize because the time series are shorter, but may also be easier because overlapping data is not removed.

Many algorithms will likely not be able to automatically use the available metadata. For the intercomparison component of the project, it may thus be worthwhile to also have a dataset without metadata.

Tuesday, 5 July 2011

Benchmarking temperature networks

The COST Action HOME has just finished a manuscript on benchmarking monthly of homogenization algorithms for regional monthly temperature and precipitation networks. I think it turned out quite interesting and provides a good base for discussions on the work of the surface temperatures benchmarking group; if only to avoid making the same mistakes.

Monday, 20 June 2011

Homogenization seminar

People interested in this blog are likely also interested in the upcoming
Seventh seminar for homogenization and quality control in climatological databases and COST ES-0601 “HOME” action management committee and final meeting in Budapest, Hungary on the 24 – 28 October 2011. Abstract deadline is the 16th of
September 2011. This is the main event in the homogenization community, in my opinion.