Notes for Team Validation:
Use the existing benchmarks and validation to look at which methods are more or less useful. Importance of looking at both ability to recreate the 'truth' and to detect the different types of breaks so that algorithm creators can get something positive about this – what exactly is causing problems for the algorithms? (station density, break frequency, break magnitude, background trend, seasonal cycle, natural variability, missing data, breaks near end-points etc.).
Reference period – this should be the most recent homogeneous subperiod. This is a problem, especially for algorithms doing seasonal shifts, when the last breakpoint is very close to the end of the record. However, this could be a real break location and so should not deliberately be avoided. Assessment should be aware of this though – algorithms could be penalised by this because they would not be able to model the seasonality effectively but assessments may look like the algorithm is failing because of the types of breaks or another complicating feature that was added – importance of useful assessment.
I think that false alarm rates are very important. I would rather a conservative and low false alarm rate than one that gets a higher number of breaks but adds a lot of error too. The false alarm rate should take into account the impact of incorrectly detecting a break given the adjustment applied. A detected break with a negligible adjustment applied is not so bad.
RMSE error seems to be a simple and useful metric – to root or not to root though?
analog-error-worlds minus analog-known-worlds = FULLRMSE
adjusted analog-error-worlds minus analog-known-worlds = REDUCEDRMSE
REDUCEDRMSE should be less than FULLRMSE if the algorithm is improving the network
Watch out for temporal variation in contingency scores – fewer breaks detected near the end of series?
Validating on annual verses monthly (or daily) – should validate on the highest resolution that will be used – so monthly I would say. This will penalise against flat adjustments but we know that inhomogeneities are not flat changes – this means that the errors added MUST be as realistic as possible.
How to calculate True negatives for the contingency scores?
Ensemble approach – hopefully this approach will be growing in popularity and so we need to be able to cope with this. An argument for keeping validation simple.
Some algorithms are trying to adjust more than just the mean, some of the higher order moments. Do we know enough to be able to add in errors in this way? Can we assess this fairly?
Purpose - To facilitate use of a robust, independent and global common benchmarking and assessment system for temperature data-product creation methodologies to aid product intercomparison and uncertainty quantification:
http://www.surfacetemperatures.org/benchmarking-and-assessment-working-group
Posting and comments are open for constructive advice and ideas.
Friday, 11 November 2011
Team Corruption - Thoughts from the Homogenisation Meeting
Notes for Team Corruption:
Need to add in realistic inhomogeneities that do not reward specific algorithms by being too obvious/exaggerated. For example, having an over exaggerated seasonally dependent shift will penalise algorithms with a flat detection/adjustment more than necessary and reward algorithms detecting/adjusting based on strong seasonal shifts. This is a difficult balance to achieve but having final errors added by those not building the algorithms and keeping the benchmarks blind will help.
Specific types of inhomogeneity:
- Add in station moves by cutting a pasting a nearby station series. May have to tweak a little to avoid exact duplication though – could create 'duplicate' stations by using the average of 2-3 neighbouring stations to downscale the GCM gridbox therefore creating a unique but realistic station. These 'duplicate' stations will differ slightly and can be substituted for part of a station series to mimic a station move.
- Instrument change/calibration error – this could be a flatter change but could also be a change to the variance on hourly timescales (not necessarily monthly). Instrument sensitivity may change.
- Shelter change – cotton region to stevenson screen – would be a seasonally varying change
- Manual to automated – more missing data, more repeated data (QC), fewer outliers? (QC), more or less sensitivity?
- Changes in observation times – how will this be manifested in monthly data?
- Significant changes to network density – a very real problem that may be reflected in the analogs anyway as they follow the real station drop-in/out – although do we want 100+ years of benchmarks? If we're shortening the record we need to ensure a similar station fall out in at least one of the worlds. When validating we need to be clear on the reasons why algorithms are failing if possible. 1972 seems to be an important year in ISD (NCDC's global sub-daily data) where vast numbers of digitised records drop out and then come back in in 1973.
- Changes in observation frequency and reporting resolution. Increases in reporting frequency from 6 hourly to hourly may mean that lower minimums/higher maximums are now recorded – and vice versa. Rounding procedures may lead to changes from resolution changes – do they truncate or round?
Have a few established break characteristics to input but make them not too predictable or people will know what to look for.
Reference period – this should be the most recent homogeneous subperiod. This is a problem, especially for algorithms doing seasonal shifts, when the last breakpoint is very close to the end of the record. However, this could be a real break location and so should not deliberately be avoided. Assessment should be aware of this though – algorithms could be penalised by this because they would not be able to model the seasonality effectively but assessments may look like the algorithm is failing because of the types of breaks or another complicating feature that was added – importance of useful assessment.
Future benchmarks:
- should be realistic
- Correlations in perturbations within a network – geographical clusters
- study seasonal cycle
- Provide metadata – some good, some bad, some incomplete, some negligible
Include other key climate features – solar radiation/sunshine duration affects the break characteristics, wind, ENSO etc. Largest effects in clear skies – full solar radiation. This info can be stored from the climate model data when creating the analog-known-worlds for later use by team creation.
Be realistic but also have ability to isolate certain break types/questions to make analysis useful – need for a series of worlds with well posed questions.
Regional knowledge is valuable – how to obtain this?
- Norway: Most breaks due to relocation (55%), screen changes (14%), instrument change (15%), other (15%) - very little effect of changing observer – NOT QUITE SURE HOW THAT ADDS UP TO 100%? SIMULTANEOUS CHANGES?
- France/Germany found most changes due to changes in shelters. Norway may have less changes with shelters because of radiation? Or many changes happen at the same time so difficult to distinguish.
- Norwegian data are composites of multiple nearby stations – not official station moves but later station mergers! Similarly in Czech Republic.
Proportion of known to unknown breaks – I would expect that for most countries there are more 'unknown' breaks than 'known' breaks – Czech has 50% backed up by metadata.
Some algorithms are trying to adjust more than just the mean, some of the higher order moments. Do we know enough to be able to add in errors in this way? Can we assess this fairly?
Need to add in realistic inhomogeneities that do not reward specific algorithms by being too obvious/exaggerated. For example, having an over exaggerated seasonally dependent shift will penalise algorithms with a flat detection/adjustment more than necessary and reward algorithms detecting/adjusting based on strong seasonal shifts. This is a difficult balance to achieve but having final errors added by those not building the algorithms and keeping the benchmarks blind will help.
Specific types of inhomogeneity:
- Add in station moves by cutting a pasting a nearby station series. May have to tweak a little to avoid exact duplication though – could create 'duplicate' stations by using the average of 2-3 neighbouring stations to downscale the GCM gridbox therefore creating a unique but realistic station. These 'duplicate' stations will differ slightly and can be substituted for part of a station series to mimic a station move.
- Instrument change/calibration error – this could be a flatter change but could also be a change to the variance on hourly timescales (not necessarily monthly). Instrument sensitivity may change.
- Shelter change – cotton region to stevenson screen – would be a seasonally varying change
- Manual to automated – more missing data, more repeated data (QC), fewer outliers? (QC), more or less sensitivity?
- Changes in observation times – how will this be manifested in monthly data?
- Significant changes to network density – a very real problem that may be reflected in the analogs anyway as they follow the real station drop-in/out – although do we want 100+ years of benchmarks? If we're shortening the record we need to ensure a similar station fall out in at least one of the worlds. When validating we need to be clear on the reasons why algorithms are failing if possible. 1972 seems to be an important year in ISD (NCDC's global sub-daily data) where vast numbers of digitised records drop out and then come back in in 1973.
- Changes in observation frequency and reporting resolution. Increases in reporting frequency from 6 hourly to hourly may mean that lower minimums/higher maximums are now recorded – and vice versa. Rounding procedures may lead to changes from resolution changes – do they truncate or round?
Have a few established break characteristics to input but make them not too predictable or people will know what to look for.
Reference period – this should be the most recent homogeneous subperiod. This is a problem, especially for algorithms doing seasonal shifts, when the last breakpoint is very close to the end of the record. However, this could be a real break location and so should not deliberately be avoided. Assessment should be aware of this though – algorithms could be penalised by this because they would not be able to model the seasonality effectively but assessments may look like the algorithm is failing because of the types of breaks or another complicating feature that was added – importance of useful assessment.
Future benchmarks:
- should be realistic
- Correlations in perturbations within a network – geographical clusters
- study seasonal cycle
- Provide metadata – some good, some bad, some incomplete, some negligible
Include other key climate features – solar radiation/sunshine duration affects the break characteristics, wind, ENSO etc. Largest effects in clear skies – full solar radiation. This info can be stored from the climate model data when creating the analog-known-worlds for later use by team creation.
Be realistic but also have ability to isolate certain break types/questions to make analysis useful – need for a series of worlds with well posed questions.
Regional knowledge is valuable – how to obtain this?
- Norway: Most breaks due to relocation (55%), screen changes (14%), instrument change (15%), other (15%) - very little effect of changing observer – NOT QUITE SURE HOW THAT ADDS UP TO 100%? SIMULTANEOUS CHANGES?
- France/Germany found most changes due to changes in shelters. Norway may have less changes with shelters because of radiation? Or many changes happen at the same time so difficult to distinguish.
- Norwegian data are composites of multiple nearby stations – not official station moves but later station mergers! Similarly in Czech Republic.
Proportion of known to unknown breaks – I would expect that for most countries there are more 'unknown' breaks than 'known' breaks – Czech has 50% backed up by metadata.
Some algorithms are trying to adjust more than just the mean, some of the higher order moments. Do we know enough to be able to add in errors in this way? Can we assess this fairly?
Team Creation - thoughts from the Homgenisation Meeting
Notes for Team Creation:
Future benchmarks:
should be realistic
realistic outliers/random errors - assume a good QC has been undertaken
insert random missing data (which we will have masked from the real stations anyway)
study frequency and size of local trends (which will come from the climate models)
Adding the noise term – some of this will be uncorrelated with other stations – simple random errors, some of this would be the weather term although how this would play out on monthly timescales is unclear – persistent cold or hot events – these would be correlated across networks. Some kind of simple weather generator? Could this sort of thing be modelled from the real stations? Study periodicities in common or something like that? Could use geospatial statistics to get at spatial covariance and add 'weather' based on these underlying relationships?
May be worth storing some other information from the models to be used by team corruption – incoming solar radiation, windspeed? This wouldn't be public info but could help with 'realistic' error input.
Future benchmarks:
should be realistic
realistic outliers/random errors - assume a good QC has been undertaken
insert random missing data (which we will have masked from the real stations anyway)
study frequency and size of local trends (which will come from the climate models)
Adding the noise term – some of this will be uncorrelated with other stations – simple random errors, some of this would be the weather term although how this would play out on monthly timescales is unclear – persistent cold or hot events – these would be correlated across networks. Some kind of simple weather generator? Could this sort of thing be modelled from the real stations? Study periodicities in common or something like that? Could use geospatial statistics to get at spatial covariance and add 'weather' based on these underlying relationships?
May be worth storing some other information from the models to be used by team corruption – incoming solar radiation, windspeed? This wouldn't be public info but could help with 'realistic' error input.
2011 Progress Report Now Published
The 2011 Progress Report has just been accepted by the Steering Committee and is now available on out website: http://www.surfacetemperatures.org/benchmarking-and-assessment-working-group#Working%20Group%20Documents.
Thanks for all the work from the group so far! There's been a lot of discussion of novel concepts. The next phase, arguably the hardest, is to get something up and running by November 2012. One year to go!
Kate
Thanks for all the work from the group so far! There's been a lot of discussion of novel concepts. The next phase, arguably the hardest, is to get something up and running by November 2012. One year to go!
Kate
Friday, 26 August 2011
Team Validation
Here are my first set of thoughts on what our team could and should
be doing. There may be things that I've completely overlooked. Please
send any comments you have on omissions, or on any of my thoughts, as
soon as you think of them.
I see three things that fall within our remit:
1. Identify any experts that we would like to join us and issue invitations.
2. Identify which validation/verification techniques we should use.
3. Find or write software/code to implement the chosen techniques.
Let's say a bit more on each of these:
1. Until we have made progress on 2., it is difficult to decide who
best to invite. We could, of course, ask someone with general
verification knowledge rather than someone specialising in the types
of data format we identify in 2. Please email any suggestions to me.
I can think of several, but no one of them stands out as first choice.
At this stage there may not be much for them to get their teeth into.
2. We will not be certain what the data format will be until Team
Corruption have made decisions. I guess this will not be finalised
until sometime in 2012, so we could argue that we can do nothing
until then. However, I'm sure we can make some educated guesses
as to what will need to be validated. If there are some things we can
be fairly certain of, we can make decisions for those formats, but
not waste time considering scenarios that might not be used.
A quick look at some possibilities/questions:
was a changepoint found Yes/No?
was the nature of the changepoint correctly identified - this could
again be Yes/No or it could quantify how closely the magnitude
of a change was estimated. Different types of change would need
different validation methodology.
a key question is whether validation will be done station-by-station
and the results simply added, or whether an attempt will be made
to assess how well the spatial pattern of the 'corrected' data match
the 'true' data? The answer to this would determine whether we
want to bring on board an expert in spatial verification.
will we want to compare the distribution of 'corrected' data with
that of the true 'data', as well as looking at how well individual
corrected and true data sets match?
some ideas are given in Section 2.5 of
whitepaper_Benchmarking_Jun2011_v2.pdf
(available on the group website) and also in Section 5 of the COST
(HOME) paper circulated by Victor on July 5th.
3. This stage needs some serious investment of time, and computing
expertise, and is not something I can contribute much to. We could
certainly do with an expert on this side of things. It can't be started
until we are well advanced with 2. but we could start thinking about
who/how will do this.
Ian
be doing. There may be things that I've completely overlooked. Please
send any comments you have on omissions, or on any of my thoughts, as
soon as you think of them.
I see three things that fall within our remit:
1. Identify any experts that we would like to join us and issue invitations.
2. Identify which validation/verification techniques we should use.
3. Find or write software/code to implement the chosen techniques.
Let's say a bit more on each of these:
1. Until we have made progress on 2., it is difficult to decide who
best to invite. We could, of course, ask someone with general
verification knowledge rather than someone specialising in the types
of data format we identify in 2. Please email any suggestions to me.
I can think of several, but no one of them stands out as first choice.
At this stage there may not be much for them to get their teeth into.
2. We will not be certain what the data format will be until Team
Corruption have made decisions. I guess this will not be finalised
until sometime in 2012, so we could argue that we can do nothing
until then. However, I'm sure we can make some educated guesses
as to what will need to be validated. If there are some things we can
be fairly certain of, we can make decisions for those formats, but
not waste time considering scenarios that might not be used.
A quick look at some possibilities/questions:
was a changepoint found Yes/No?
was the nature of the changepoint correctly identified - this could
again be Yes/No or it could quantify how closely the magnitude
of a change was estimated. Different types of change would need
different validation methodology.
a key question is whether validation will be done station-by-station
and the results simply added, or whether an attempt will be made
to assess how well the spatial pattern of the 'corrected' data match
the 'true' data? The answer to this would determine whether we
want to bring on board an expert in spatial verification.
will we want to compare the distribution of 'corrected' data with
that of the true 'data', as well as looking at how well individual
corrected and true data sets match?
some ideas are given in Section 2.5 of
whitepaper_Benchmarking_Jun2011_v2.pdf
(available on the group website) and also in Section 5 of the COST
(HOME) paper circulated by Victor on July 5th.
3. This stage needs some serious investment of time, and computing
expertise, and is not something I can contribute much to. We could
certainly do with an expert on this side of things. It can't be started
until we are well advanced with 2. but we could start thinking about
who/how will do this.
Ian
Tuesday, 26 July 2011
Benchmark for real-world problems
We should state here what properties an ideal benchmark data set should have, right?
My wish would be for a data set that is as close as possible to real world problems (in addition to data sets that allow testing your methods until they break, which is of course very important).
As close a possible to real world problems could mean: Use physics-based error models (to simulate instrumental errors), simulate typical reporting errors (there should be plenty of experience around what can be wrong), simulate typical processing errors, etc. We will still not get around adding also simply perturbations in a statistical sense, but I think we can be more realistic than that.
Such a data set necessarily is a subdaily data set, and the monthly benchmark data set would simply be an average of the subdaily data (with an additional simulations of errors that can occur during the production of monthly means). Such a data set would necessarily be based on some sort of climate model or reanalysis data because other variables than temperature would be used, and they would be used in a high resolution.
I volunteer to produce such a data set if requested, but lacking experience on homogenizing data outside Europe, I would have to team up with more experienced people telling me what possibly can go wrong in Africa or the Arctic.
My wish would be for a data set that is as close as possible to real world problems (in addition to data sets that allow testing your methods until they break, which is of course very important).
As close a possible to real world problems could mean: Use physics-based error models (to simulate instrumental errors), simulate typical reporting errors (there should be plenty of experience around what can be wrong), simulate typical processing errors, etc. We will still not get around adding also simply perturbations in a statistical sense, but I think we can be more realistic than that.
Such a data set necessarily is a subdaily data set, and the monthly benchmark data set would simply be an average of the subdaily data (with an additional simulations of errors that can occur during the production of monthly means). Such a data set would necessarily be based on some sort of climate model or reanalysis data because other variables than temperature would be used, and they would be used in a high resolution.
I volunteer to produce such a data set if requested, but lacking experience on homogenizing data outside Europe, I would have to team up with more experienced people telling me what possibly can go wrong in Africa or the Arctic.
Monday, 25 July 2011
Another radiosonde benchmarking paper
This is a little self serving but a further twist on the radiosonde benchmarking has just become available at JGR. A link to this paper is here . This uses the same benchmarks as used in Titchner et al. but looks at multiple additional impacts and also starts to try to address how you could use results from multiple different estimators to yield a grand unified estimate of the truth. There are, of course, many ways one could go about this step, but clearly such an effort may well be something the group would wish to consider how to approach ...