Tuesday, 26 July 2011

Benchmark for real-world problems

We should state here what properties an ideal benchmark data set should have, right?
My wish would be for a data set that is as close as possible to real world problems (in addition to data sets that allow testing your methods until they break, which is of course very important).
As close a possible to real world problems could mean: Use physics-based error models (to simulate instrumental errors), simulate typical reporting errors (there should be plenty of experience around what can be wrong), simulate typical processing errors, etc. We will still not get around adding also simply perturbations in a statistical sense, but I think we can be more realistic than that.
Such a data set necessarily is a subdaily data set, and the monthly benchmark data set would simply be an average of the subdaily data (with an additional simulations of errors that can occur during the production of monthly means). Such a data set would necessarily be based on some sort of climate model or reanalysis data because other variables than temperature would be used, and they would be used in a high resolution.
I volunteer to produce such a data set if requested, but lacking experience on homogenizing data outside Europe, I would have to team up with more experienced people telling me what possibly can go wrong in Africa or the Arctic.

3 comments:

Victor Venema said...

A good suggestion, which would use the strength of this study, the use of model data.

We may have to take care that not all models store their sub-daily data.

Kate Willett said...

I believe the CMIP5 set of model runs should include some long high resolution output. I will investigate. Failing that the 20th Century Reanalysis is free from most instrumental inhomogeneity as it only takes in surface pressure data and is available at 6 hourly resolution so may be useful.

Stefan if you can step forward to build this I will help in whatever way I can. Claude, Robert and Lucie should be able to provide some input into known North American inhomogeneities. Lisa and Blair Trewin may have some incite for Australia. Perhaps between us we can do some digging for information of other regions. We have some African scientists working here at the moment. I will see if they can offer any advice.

Blair Trewin said...

I've been doing a lot of work on homogenising Australian data (at the daily timescale). Some of the more 'complex' inhomogeneities that I've found, which might be worth thinking about for examples in a benchmark data set, include:

- no significant change in annual means but large changes of opposite sign in winter and summer. (This could happen when a site moves from a coastal location to a more continental one or vice versa, or, as in one example I found, when a site moved from an in-town location where local surface conditions were similar all year, to cropland where the site was surrounded by green vegetation in winter/spring but bare soil in summer/autumn).

- an inhomogeneity that has little effect on one end of the daily frequency distribution but a large effect on the other end - this can happen when a site moves from an urban location to a rural one or vice versa, with little difference in minimum temperatures on cloudy/windy nights but a large one on clear/calm nights.

One particularly complicated form of behaviour I've seen in some cases for warm-season maximum temperatures on coasts with a strong land-sea temperature contrast is where temperature differences between a coastal and inland site progressively increase with increasing temperature, then collapse to near zero on the very hottest days when strong offshore winds cause the collapse of the sea breeze. Australia has its share of sites like this but I imagine they would also be found in other places with a hot land/cold ocean contrast (e.g. the US west coast).