Monday, 19 December 2011

Metadata

The quality of homogenized data does not only depend on the performance of the homogenization algorithm, but also on the metadata, documentary evidence of possible break points. Therefore, we (benchmarking working group of the international surface temperature initiative) want to include metadata in our benchmark.

The amount and quality of metadata is thus likely important to obtain realistic estimates of the uncertainty in climate variability and trends due to remaining inhomogeneities. For the quantity of metadata we can just mimic the metadata in the ISTI database.

I am wondering whether we have enough information on the quality of the metadata. Especially as it will be difficult to tell whether the meta data was right. That there was no jump in the (monthly mean) data, does not mean that nothing has changed. Do we have any idea about the accuracy of the available (machine-readable) metadata? That the distribution of break sizes at dates with a known break (from metadata) was normally distributed (Menne and Williams, 2005) gives some hope that the quality is good enough.

Friday, 11 November 2011

Team Validation - thoughts from the Homogenisation Meeting

Notes for Team Validation:

Use the existing benchmarks and validation to look at which methods are more or less useful. Importance of looking at both ability to recreate the 'truth' and to detect the different types of breaks so that algorithm creators can get something positive about this – what exactly is causing problems for the algorithms? (station density, break frequency, break magnitude, background trend, seasonal cycle, natural variability, missing data, breaks near end-points etc.).

Reference period – this should be the most recent homogeneous subperiod. This is a problem, especially for algorithms doing seasonal shifts, when the last breakpoint is very close to the end of the record. However, this could be a real break location and so should not deliberately be avoided. Assessment should be aware of this though – algorithms could be penalised by this because they would not be able to model the seasonality effectively but assessments may look like the algorithm is failing because of the types of breaks or another complicating feature that was added – importance of useful assessment.

I think that false alarm rates are very important. I would rather a conservative and low false alarm rate than one that gets a higher number of breaks but adds a lot of error too. The false alarm rate should take into account the impact of incorrectly detecting a break given the adjustment applied. A detected break with a negligible adjustment applied is not so bad.

RMSE error seems to be a simple and useful metric – to root or not to root though?
analog-error-worlds minus analog-known-worlds = FULLRMSE
adjusted analog-error-worlds minus analog-known-worlds = REDUCEDRMSE
REDUCEDRMSE should be less than FULLRMSE if the algorithm is improving the network

Watch out for temporal variation in contingency scores – fewer breaks detected near the end of series?

Validating on annual verses monthly (or daily) – should validate on the highest resolution that will be used – so monthly I would say. This will penalise against flat adjustments but we know that inhomogeneities are not flat changes – this means that the errors added MUST be as realistic as possible.

How to calculate True negatives for the contingency scores?

Ensemble approach – hopefully this approach will be growing in popularity and so we need to be able to cope with this. An argument for keeping validation simple.

Some algorithms are trying to adjust more than just the mean, some of the higher order moments. Do we know enough to be able to add in errors in this way? Can we assess this fairly?

Team Corruption - Thoughts from the Homogenisation Meeting

Notes for Team Corruption:

Need to add in realistic inhomogeneities that do not reward specific algorithms by being too obvious/exaggerated. For example, having an over exaggerated seasonally dependent shift will penalise algorithms with a flat detection/adjustment more than necessary and reward algorithms detecting/adjusting based on strong seasonal shifts. This is a difficult balance to achieve but having final errors added by those not building the algorithms and keeping the benchmarks blind will help.

Specific types of inhomogeneity:
- Add in station moves by cutting a pasting a nearby station series. May have to tweak a little to avoid exact duplication though – could create 'duplicate' stations by using the average of 2-3 neighbouring stations to downscale the GCM gridbox therefore creating a unique but realistic station. These 'duplicate' stations will differ slightly and can be substituted for part of a station series to mimic a station move.
- Instrument change/calibration error – this could be a flatter change but could also be a change to the variance on hourly timescales (not necessarily monthly). Instrument sensitivity may change.
- Shelter change – cotton region to stevenson screen – would be a seasonally varying change
- Manual to automated – more missing data, more repeated data (QC), fewer outliers? (QC), more or less sensitivity?
- Changes in observation times – how will this be manifested in monthly data?
- Significant changes to network density – a very real problem that may be reflected in the analogs anyway as they follow the real station drop-in/out – although do we want 100+ years of benchmarks? If we're shortening the record we need to ensure a similar station fall out in at least one of the worlds. When validating we need to be clear on the reasons why algorithms are failing if possible. 1972 seems to be an important year in ISD (NCDC's global sub-daily data) where vast numbers of digitised records drop out and then come back in in 1973.
- Changes in observation frequency and reporting resolution. Increases in reporting frequency from 6 hourly to hourly may mean that lower minimums/higher maximums are now recorded – and vice versa. Rounding procedures may lead to changes from resolution changes – do they truncate or round?

Have a few established break characteristics to input but make them not too predictable or people will know what to look for.

Reference period – this should be the most recent homogeneous subperiod. This is a problem, especially for algorithms doing seasonal shifts, when the last breakpoint is very close to the end of the record. However, this could be a real break location and so should not deliberately be avoided. Assessment should be aware of this though – algorithms could be penalised by this because they would not be able to model the seasonality effectively but assessments may look like the algorithm is failing because of the types of breaks or another complicating feature that was added – importance of useful assessment.

Future benchmarks:
- should be realistic
- Correlations in perturbations within a network – geographical clusters
- study seasonal cycle
- Provide metadata – some good, some bad, some incomplete, some negligible

Include other key climate features – solar radiation/sunshine duration affects the break characteristics, wind, ENSO etc. Largest effects in clear skies – full solar radiation. This info can be stored from the climate model data when creating the analog-known-worlds for later use by team creation.

Be realistic but also have ability to isolate certain break types/questions to make analysis useful – need for a series of worlds with well posed questions.

Regional knowledge is valuable – how to obtain this?
- Norway: Most breaks due to relocation (55%), screen changes (14%), instrument change (15%), other (15%) - very little effect of changing observer – NOT QUITE SURE HOW THAT ADDS UP TO 100%? SIMULTANEOUS CHANGES?
- France/Germany found most changes due to changes in shelters. Norway may have less changes with shelters because of radiation? Or many changes happen at the same time so difficult to distinguish.
- Norwegian data are composites of multiple nearby stations – not official station moves but later station mergers! Similarly in Czech Republic.

Proportion of known to unknown breaks – I would expect that for most countries there are more 'unknown' breaks than 'known' breaks – Czech has 50% backed up by metadata.

Some algorithms are trying to adjust more than just the mean, some of the higher order moments. Do we know enough to be able to add in errors in this way? Can we assess this fairly?

Team Creation - thoughts from the Homgenisation Meeting

Notes for Team Creation:

Future benchmarks:
should be realistic
realistic outliers/random errors - assume a good QC has been undertaken
insert random missing data (which we will have masked from the real stations anyway)
study frequency and size of local trends (which will come from the climate models)

Adding the noise term – some of this will be uncorrelated with other stations – simple random errors, some of this would be the weather term although how this would play out on monthly timescales is unclear – persistent cold or hot events – these would be correlated across networks. Some kind of simple weather generator? Could this sort of thing be modelled from the real stations? Study periodicities in common or something like that? Could use geospatial statistics to get at spatial covariance and add 'weather' based on these underlying relationships?

May be worth storing some other information from the models to be used by team corruption – incoming solar radiation, windspeed? This wouldn't be public info but could help with 'realistic' error input.

2011 Progress Report Now Published

The 2011 Progress Report has just been accepted by the Steering Committee and is now available on out website: http://www.surfacetemperatures.org/benchmarking-and-assessment-working-group#Working%20Group%20Documents.

Thanks for all the work from the group so far! There's been a lot of discussion of novel concepts. The next phase, arguably the hardest, is to get something up and running by November 2012. One year to go!

Kate

Friday, 26 August 2011

Team Validation

Here are my first set of thoughts on what our team could and should
be doing. There may be things that I've completely overlooked. Please
send any comments you have on omissions, or on any of my thoughts, as
soon as you think of them.


I see three things that fall within our remit:

1. Identify any experts that we would like to join us and issue invitations.

2. Identify which validation/verification techniques we should use.

3. Find or write software/code to implement the chosen techniques.

Let's say a bit more on each of these:

1. Until we have made progress on 2., it is difficult to decide who
best to invite. We could, of course, ask someone with general
verification knowledge rather than someone specialising in the types
of data format we identify in 2. Please email any suggestions to me.
I can think of several, but no one of them stands out as first choice.
At this stage there may not be much for them to get their teeth into.

2. We will not be certain what the data format will be until Team
Corruption have made decisions. I guess this will not be finalised
until sometime in 2012, so we could argue that we can do nothing
until then. However, I'm sure we can make some educated guesses
as to what will need to be validated. If there are some things we can
be fairly certain of, we can make decisions for those formats, but
not waste time considering scenarios that might not be used.

A quick look at some possibilities/questions:

was a changepoint found Yes/No?

was the nature of the changepoint correctly identified - this could
again be Yes/No or it could quantify how closely the magnitude
of a change was estimated. Different types of change would need
different validation methodology.

a key question is whether validation will be done station-by-station
and the results simply added, or whether an attempt will be made
to assess how well the spatial pattern of the 'corrected' data match
the 'true' data? The answer to this would determine whether we
want to bring on board an expert in spatial verification.

will we want to compare the distribution of 'corrected' data with
that of the true 'data', as well as looking at how well individual
corrected and true data sets match?

some ideas are given in Section 2.5 of
whitepaper_Benchmarking_Jun2011_v2.pdf
(available on the group website) and also in Section 5 of the COST
(HOME) paper circulated by Victor on July 5th.

3. This stage needs some serious investment of time, and computing
expertise, and is not something I can contribute much to. We could
certainly do with an expert on this side of things. It can't be started
until we are well advanced with 2. but we could start thinking about
who/how will do this.

Ian

Tuesday, 26 July 2011

Benchmark for real-world problems

We should state here what properties an ideal benchmark data set should have, right?
My wish would be for a data set that is as close as possible to real world problems (in addition to data sets that allow testing your methods until they break, which is of course very important).
As close a possible to real world problems could mean: Use physics-based error models (to simulate instrumental errors), simulate typical reporting errors (there should be plenty of experience around what can be wrong), simulate typical processing errors, etc. We will still not get around adding also simply perturbations in a statistical sense, but I think we can be more realistic than that.
Such a data set necessarily is a subdaily data set, and the monthly benchmark data set would simply be an average of the subdaily data (with an additional simulations of errors that can occur during the production of monthly means). Such a data set would necessarily be based on some sort of climate model or reanalysis data because other variables than temperature would be used, and they would be used in a high resolution.
I volunteer to produce such a data set if requested, but lacking experience on homogenizing data outside Europe, I would have to team up with more experienced people telling me what possibly can go wrong in Africa or the Arctic.

Monday, 25 July 2011

Another radiosonde benchmarking paper

This is a little self serving but a further twist on the radiosonde benchmarking has just become available at JGR. A link to this paper is here . This uses the same benchmarks as used in Titchner et al. but looks at multiple additional impacts and also starts to try to address how you could use results from multiple different estimators to yield a grand unified estimate of the truth. There are, of course, many ways one could go about this step, but clearly such an effort may well be something the group would wish to consider how to approach ...

Thursday, 14 July 2011

Generating inhomogeneous worlds

The manuscript on the monthly temperature and precipitation benchmark of the COST Action HOME is now finished: manuscript. The inhomogeneities applied in this study, together with ideas for improvements in the discussion and outlook, may be a good basis for the benchmarking of the surface temperature initiative as well. Thus here I will only mention where I would suggest to do thinks differently for the surface temperatures initiative (STI). Next to improvements due to lessons learned from this study, the STI differs in two important aspects. 1) HOME considered regional networks, STI global datasets. 2) HOME focused on intercomparison of the homogenization algorithms, STI has the additional ambition that the benchmarking leads error estimates for the STI database.

I agree with Kate Willett that it would be valuable to have both data for which the truth is known as well as data for which the truth will be revealed after all homogenized contributions have been submitted. The latter leads to more reliable results because the homogenization algorithms cannot be tuned to solve the benchmark data well (and then possibly perform better on the benchmark than real data).
Data for which the truth is known would be a more classical validation study, which has the advantage that the homogenizers can learn during the exercise and also find (programming) errors; see outlook of our paper. In case of a benchmark one would have to wait another 3 year cycle to implement bug fixes. In HOME we had a number of such bugs, often in the parts newly written, e.g. to be able to handle multiple networks (which in not needed in daily work).

In HOME, we had random and clustered breaks. The clustered breaks occurred with a probability of 30% and if they occurred they affected 30% of the stations. As HOME generated regional networks, we could not insert breaks that occur in all stations, because they would not be detectable by relative homogenization. As the STI will generate global data, it is possible to insert breaks that occur simultaneously in one entire regional network, one would see them at the border of the countries. It would be interesting to see if the homogenization algorithms are able to move the information at the border into the heart of the network. However, also in this case, always inserting the clustered breaks in all stations would not be realistic. One country often has multiple networks (synoptic, volunteer networks or measurements by multiple organizations). Thus partial clustered breaks are also needed. In HOME the clustered breaks were perfectly simultaneous, which mimics a change in observational rules. Breaks clustered over a period, which mimics a change in the instrumentation or screens, would also be valuable.
The clustered breaks are not only correlated in time, but also in size. This is needed to generated biases in the continental and global averages due to inhomogeneities. This is seen in real data, especially in the early instrumental period and now again in the transition to automatic weather stations. These two periods may warrant special rules, i.e. maybe the timing of the clustered should not be fully random.

Another aspect that was not touched in the intercomparison study HOME, but needs to be treated in the STI are multiple elements. For many stations more than one temperature dataset will be available: Tmin, Tmax, Tmean, DTR (diurnal temperature range), or temperature observations for specific hours. Methods using multiple elements simultaneously likely perform better and if they are applied to the real data, they should also be validated in the benchmarking. Then we would need to know how the size of the inhomogeneities correlates between the various temperature variables. A break in Tmin, does not imply a (statistically significant) break in Tmax, but probably does make it more likely. Similar relations will likely hold among all temperature measures. I am not sure if there is data from literature on such cross-relations. On the other hand, you may also want to have some worlds with univariate validation data for the intercomparison of algorithms that do not use this information.

In the HOME benchmark, some stations had a local trend. The statistical properties of these trends were idealised, which was acceptable for an intercomparison, but to obtain realistic error estimates it should be studied in more detail how often such local trends occur in real datasets.

In HOME the perturbation were a constant number for every month. In the STI we have the opportunity to make the perturbations a function of insolation, wind and precipitation. This would make them partially stochastic, which is more realistic. In this case we would have to decide whether to make these covariates available to the homogenizers (potentially better homogenization) or not (most algorithms will currently not be able to use this information; making intercomparison more difficult). I could imagine that this is especially important for breaks during the early instrumental period, when measurement methods were not yet fully optimized to handle radiation and wetting problems.

It is important to add a stochastic and nonlinear large-scale trend to the data. It should be stochastic so that the homogenizer cannot see how well he did by computing the trend. And it should be nonlinear and contain decadal variability because homogenization algorithms always work and we should not mix our theoretical ideas with the data. Because the STI not only wants to make an intercomparison of homogenization algorithms, but also aims to produce representative errors, this trend will have to be modelling in a more realistic way as HOME did. I would suggest modelling the trend and the decadal variability separately and to vary it gradually on continental scales.

One of the outcomes of the study was that outliers are not important for the quality of the homogenization. Thus it may be best not to insert outliers in the benchmark. That would save work in inserting them and in removing them again later in the analysis of the homogenized data and homogenizers would not feel they have to program additional processing for outliers.

If I understood the science plan of the STI right, there will be multiple data levels (images of log books, digitized data, data in SI units, merged data, quality controlled data, homogenized data and gridded data). And everyone is invited to produce data or implement algorithms for the various levels. That would mean that there could be multiple input dataset for our benchmarking exercise. For homogenization, especially various ways of merging the data could be important. Merging methods that do not merge stations that move (to produce data that is suited to study local effects and relations with other (surface) parameters and variables) would produce much shorter time series than merging methods that aim to produce long time series (to make it easier to study secular trends due to climate change). The former case may be more difficult to homogenize because the time series are shorter, but may also be easier because overlapping data is not removed.

Many algorithms will likely not be able to automatically use the available metadata. For the intercomparison component of the project, it may thus be worthwhile to also have a dataset without metadata.

Tuesday, 5 July 2011

Benchmarking temperature networks

The COST Action HOME has just finished a manuscript on benchmarking monthly of homogenization algorithms for regional monthly temperature and precipitation networks. I think it turned out quite interesting and provides a good base for discussions on the work of the surface temperatures benchmarking group; if only to avoid making the same mistakes.

Monday, 20 June 2011

Homogenization seminar

People interested in this blog are likely also interested in the upcoming
Seventh seminar for homogenization and quality control in climatological databases and COST ES-0601 “HOME” action management committee and final meeting in Budapest, Hungary on the 24 – 28 October 2011. Abstract deadline is the 16th of
September 2011. This is the main event in the homogenization community, in my opinion.

Thursday, 16 June 2011

If I had but one analog I could create ...

I think I would look to create something that tested in some sense the limits of homogenization methods whilst still retaining some realism.

I would take a forced component run such as c20c and then look to add change points that in the net removed that trend. This would penalize any algorithm that tended to introduce adjustments with a preferential zero bias.

The breaks I would add would be a mix of step like and slope like and a large number would have changes in seasonality and timeseries variance associated.

I'd have limited metadata and what metadata there was would be poor quality.

I would assume that most breaks were small (sigma <1K, in some cases perhaps <<1K) and that they happened fairly frequently (once every 5 to ten years say on average).

A number of breaks would be quasi-contemperaneous over countries and these would have very similar characteristics to each other.

There is documented evidence that these issues all to some extent pervade the network (e.g. US network move from stevenson screen to automated sensors happened largely within 5 years over 70% of the network).

So, whilst at the outer bounds of plausibility it would not be an entirely implausible error structure.

Monday, 13 June 2011

Big questions with which to test homogenisation algorithms

Hi, I'm hoping to get time in the call to touch on this a little. Similar to previous posts about worse nightmares I would like us to think about questions that we want to answer with the analog-error-models. The idea would be to pick 8 of these to run with for each 3 year benchmarking cycle - one for each world. If these start from something simple to something horrible this will give us a chance to see where algorithms begin to struggle. Lets focus here on monthly means to make things a little simpler for the time being - please comment with your ideas. For example:

1) Do homogenisation algorithms detect discontinuities when none are present?
Analog-error-world 1 = A historical forcing model analog-known-world with no-errors added

2) Do homogenisation algorithms cope with discontinuities that affect the variance?
Analog-error-world 2 = A historical forcing model analog-known-world with seasonally constant changes applied
Analog-error-world 3 = A historical forcing model analog-known-world with seasonally varying changes applied at the same location and approximate magnitude as analog-error-world 2

2) Can homogenisation algorithms cope with non-stationary worlds/ where there is a background trend?
Analog-error-world 4 = A control forcing (constant pre-industrial emissions) model analog-known-world with mixed error structure applied
Analog-error-world 5 = An A1B (high emissions) forcing model analog-known-world basis with identical error structure to World 4

3) Can homogenisation algorithms cope when discontinuities are small and frequent?
Analog-error-world 6 = A historical forcing model analog-known-world with many small discontinuities added of various sign biases - (seasonally varying to be realistic?).

4) Can homogenisation algorithms cope with layered gradual and abrupt discontinuities (i.e., urban warming + instrument shelter change)
Analog-error-world 7 = A historical forcing model analog-known-world with either gradual or abrupt discontinuities applied to a station (seasonally varying to be realistic?)
Analog-error-world 8 = A historical forcing model analog-known-world with both gradual and abrupt discontinuities applied to a station (seasonally varying to be realistic?) using the initial error structure from analog-error-world 7 with other errors added.

At present its probably useful just to come up with as many plausible questions as possible and examples of error world structures to explore these.

There is an argument for including a really nasty one that is perhaps unplausible - so feel free to be creative.

Friday, 11 March 2011

Creating the Benchmark 'Truths'

The Steering Committee is drafting an Implementation Plan to ensure success of the Surface Temperature Initiative. This will be posted on the website (www.surfacetemperatures.org) when finalised. Crucially it has a list of deadlines - some of which relate to our Benchmarking and Assessment Working group. These are our goals:

1. Defining methods to create the benchmark analog truth stations - to mirror the databank consolidated master database (these will not be made publicly available immediately)

2. Defining the spread of error models to be applied to the benchmark analog truths

3. Creating the benchmark analog truths and error worlds - these will be publicly available

4. Running some kind of review workshop (possibly online) at the end of the 3 year cycle to release the 'truths' and review benchmark production/implementation - can this be improved for the second cycle.

I would like to try and iron out goal 1. in the coming conference call (Wed March 30th 2pm GMT) GMT). I think we have two options - purely synthetic utilising statistical models or part synthetic using a combination of physical models (GCMs or reanalyses) and statistical models. Both will likely need to use information from the databank consolidated master database to govern individual station climate characteristics.

I think it is important to characterise true climate features of real data for each station - so its climatology, variance, background trend, natural variability, serial autocorrelation and its relationship to other stations within the spatial covariance structure.

Pure Synthetic

X(t,l,h) = S(t,l,h) + T(t,l,h) + RE(t,l,h)

Where X = station at time t, location l and height h, S = seasonal cycles, T = background features (e.g., trend, ENSO, volcanoes, solar cycles etc.) and RE = residual random error.

Could the actual stations be used to give basic climate characteristics? I'm not too worried about mimicking stations exactly but we do want to represent the spatial covariance between a global network of stations quite accurately (Tropics, mid-latitudes, coastal, mountainous, etc.) and the station 'noise' around the errors that we eventually apply which will be our signal.

Part Synthetic

This would take gridded fields of GCM or 20th Century Reanalyses data (no inhomogeneities due to station moves, instrument changes or data type ingestion changes) as the base for creating analog stations to mimic the consolidated master database. The grids would have to be downscaled, using statistical models and actual stations from the database to give the basic climate characteristics - similar to the above. Here realistic S and T are already there in the models - they just need tweaking to create individual stations with appropriate autocorrelation and within a realistic spatial covariance structure. RE would need to be added.

So there are some ideas to get us started and bash to pieces. I'm not a statistician, and so certainly need help!

Thursday, 10 February 2011

Assessing the Benchmarks

When we eventually design the benchmarks we need to keep in mind our assessment aims and outputs. They need to be designed such that it is easy to obtain useful and meaningful measures of skill. I've been chatting to Ian Jolliffe about the use of simple forecast verification techniques (break present and break detected = very good, break present and break not detected = bad, no break and break detected = very bad, no break and no break detected = good) including ROC plots. I'll go into more detail on this at some point soon but just thought I would throw the thought out there in case anyone has any bright ideas/thoughts.

Review paper references

Hi All, all the comments so far have given me a lot to think about which is great. The references are really useful. I'm hoping to get chance to start the review paper this weekend so if anyone can think of anymore useful references for me to browse please can you list them here.

Tuesday, 1 February 2011

My first time using blog...

I had a look at a number of homogeneous stations (stations with no major problems) to retrieve some statistical properties. I used the monthly means of the daily maximum temperature for about 30 stations, covering as much as possible 1920-2009. I examined the monthly departures from their long-term means (12 monthly means) so each series had about 1080 data points. Overall, the best-fit linear trend ranged from 0.3 to 1.5°C for 1920-2009. Once the trend was removed, the standard deviation ranged from 2.5 to 3.5 and the autocorrelation from 0.15 to 0.25. The reverse process can be used to generate datasets. However, this approach does not capture the long-term variation, such as warming from the 1900s to the 1930s, followed by a cooling to the 1970s, and warming to the 2000s.

The types of “inhomogeneities” that I have seen in temperature series are:

1. change in annual means: 2 or 3 steps are common over 90 years; a more complex situation would be a step every 10 years
2. change in monthly means: here, the magnitude of the step would be different depending of the month (e.g. an instrument relocation could generate a positive step in the summer and a negative step in the winter)
3. change in mean and variance due to a change of instrument or observing practices
4. change in mean followed by a change in trend direction: this occurs rarely and I can’t think of a physical process that would generate this situation in temperature series.

These “inhomogeneities” could be detected using a network of neighbour stations which could be highly (or not) correlated with the tested series. In addition, the same (or other) “inhomogeneities” could be found in the neighbouring series.

Thursday, 27 January 2011

Homogenization aspects that scare me

I promised a list of what scares me (beyond spiders and ice climbing) when homogenizing temperature time series. As our focus is limited to temperatures, I will not worry with changes in variances, extremes, etc. The list:

1) Multiple changepoints
2) Seasonal features, especially those typically present in daily data
3) Autocorrelation
4) Missing and/or erroneous data
5) Mean misspecification, e.g., not accounting for a trend
6) Multiple reference series

Missing data is usually just a programming hassle, so maybe it shouldn't make this list. All of these issues have been tackled in the changepoint/homogenization literature to some degree, but I do not know of a reference where all issues are considered in tandem.

Kate's Pseudo-worlds work

I've uploaded a document describing the pseudo-worlds onto our website. I have created these as part of some work with Lisa Alexander, her post-doc Marcus and Peter Thorne to help us design/choose a suit of homogenisation algorithms for creating daily global datasets. I'm happy to provide this data to anyone that is interested and the 'solutions'. It could certainly be improved but may be a useful example/starting point for the benchmarks we're designing here. Hopefully, as it is used, its strengths/weaknesses will become apparent. I've started this thread as a place for people to comment their thoughts/advice on this work.

Tuesday, 18 January 2011

Benchmarking and Assessment Open Comment - January

Welcome to the Benchmarking and Assessment Working Group's work space. This post is open to everyone to comment freely about anything relating to benchmarking and assessment. Full details of who we are, our aims and our progress to date can be found here: http://www.surfacetemperatures.org/benchmarking-and-assessment-working-group.