Friday, 11 March 2011

Creating the Benchmark 'Truths'

The Steering Committee is drafting an Implementation Plan to ensure success of the Surface Temperature Initiative. This will be posted on the website (www.surfacetemperatures.org) when finalised. Crucially it has a list of deadlines - some of which relate to our Benchmarking and Assessment Working group. These are our goals:

1. Defining methods to create the benchmark analog truth stations - to mirror the databank consolidated master database (these will not be made publicly available immediately)

2. Defining the spread of error models to be applied to the benchmark analog truths

3. Creating the benchmark analog truths and error worlds - these will be publicly available

4. Running some kind of review workshop (possibly online) at the end of the 3 year cycle to release the 'truths' and review benchmark production/implementation - can this be improved for the second cycle.

I would like to try and iron out goal 1. in the coming conference call (Wed March 30th 2pm GMT) GMT). I think we have two options - purely synthetic utilising statistical models or part synthetic using a combination of physical models (GCMs or reanalyses) and statistical models. Both will likely need to use information from the databank consolidated master database to govern individual station climate characteristics.

I think it is important to characterise true climate features of real data for each station - so its climatology, variance, background trend, natural variability, serial autocorrelation and its relationship to other stations within the spatial covariance structure.

Pure Synthetic

X(t,l,h) = S(t,l,h) + T(t,l,h) + RE(t,l,h)

Where X = station at time t, location l and height h, S = seasonal cycles, T = background features (e.g., trend, ENSO, volcanoes, solar cycles etc.) and RE = residual random error.

Could the actual stations be used to give basic climate characteristics? I'm not too worried about mimicking stations exactly but we do want to represent the spatial covariance between a global network of stations quite accurately (Tropics, mid-latitudes, coastal, mountainous, etc.) and the station 'noise' around the errors that we eventually apply which will be our signal.

Part Synthetic

This would take gridded fields of GCM or 20th Century Reanalyses data (no inhomogeneities due to station moves, instrument changes or data type ingestion changes) as the base for creating analog stations to mimic the consolidated master database. The grids would have to be downscaled, using statistical models and actual stations from the database to give the basic climate characteristics - similar to the above. Here realistic S and T are already there in the models - they just need tweaking to create individual stations with appropriate autocorrelation and within a realistic spatial covariance structure. RE would need to be added.

So there are some ideas to get us started and bash to pieces. I'm not a statistician, and so certainly need help!

1 comment:

Victor Venema said...

Both routes are good. Both "purely synthetic data" and "GCM plus downscaling" are able to produce data that is realistic enough to be able to benchmark homogenisation algorithms.

I have a slight preference for "GCM plus downscaling".

First of all simply because it is new and may thus produce new insights.

Second, it would allow for difference time series with some decadal variability. It would be interesting to see how state-of-the-art relative homogenisation methods respond to this, as these method assume that decadal is only in the station data itself, but none is present any more in the difference time series.