Thursday, June 05, 2008

"A complete recording would make it difficult to establish the facts"

It's a funny world where politicians can say things like that without fear of public ridicule.

The politician in question is the Japanese justic minister, and the topic is the recording of interrogation sessions by the Japanese police in order that the courts can judge how "voluntary" the resulting confessions are. There is increasing pressure for this, as the public slowly wakes up to the standard operating practice of the judicial system here (such as 3 weeks in custody before any charges are laid, during which time the victim is treated to sleep deprivation and coercive interrogations aiming at a "confession" which - contrary to the letter of the law - is itself sufficient for a conviction even if subsequently withdrawn).

He is obviously worried by the plummeting conviction rate - down to 97.1% at the last count. This politician is the same moron who recently said he is in favour of the death penalty because the Japanese place much more importance on the value of life than the barbarian Westerners and that didn't get him laughed out of a job either. For that matter, neither did his boast that his friend knew a member of Al-Quaida who regularly visited Japan, and that he had advance knowledge of the Bali bombing but did nothing about it. [That latter story was easily debunked as the delusions of a fantasist, BTW.]

Monday, June 02, 2008

More on that SST change

As expected, there have been increasing volumes of hot air expended on the blogosphere about this. RC's post seems pretty reasonable as ever and so is this comment they point to. There is probably little point in speculating in too great detail about the implications, but then again, if one is not going to speculate pointlessly on a blog, there seems little point in having one.

Steve McIntyre was quick to present a hypothetical new surface temperature record, based on slowly phasing out buckets over the 2nd half of the century.

This gives a greatly reduced warming trend over the last 50 and even 30 years, which got Roger Pielke very excited. However, even though Steve may be right about the buckets themselves, his analysis (which, to be fair, he did not present as anything authoritative) ignores the fact that SST observations have increasingly come from other sources such as satellites and buoys. (Furthermore, that graph above seems to show a change of 0.3C for the global mean temperature, which is hard to justify since the adjustment under discussion is only 0.3C for the SST measurements made by some ships, albeit most of them.) I happen to have used some satellite SST data in previous research, and on checking the details I see that the Pathfinder AVHRR were flying from late 1981 (there may even have been older missions for all I know, but I would guess not). This fleet of satellites has provided hi-res global coverage on a regular basis from 1985 (and probably something from the few years prior), so although I do not know the details of the global SST analysis it seems inconceivable that bucket measurements from ships played much of a part subsequent to that date. Turning to buoys, Wikipedia tells me that the US National Data Buoy Development Program started in 1967 and the National Data Buoy Center was formed in 1970, so it seems that at least some data was coming from this source back then too (and note that even if the volume of data was quite small, its relative accuracy may make it outweigh a lot of ship measurements).

So I don't know exactly what the charges will be, but I suspect that the graphic here attributed to CRU is a reasonable first guess.


As you should be able to see at a glance (although Roger apparently cannot), the maximum change in the smoothed record shown there is a little short of 0.2C globally, which is just as expected given a maximum change of no more than 0.3C for the ocean (70% of the earth's surface).

So how does this affect the IPCC's latest report? Well, in the Technical Summary, there is a nice figure (Figure TS.6) which gives 25, 50, 100 and 150 year trends. So let's give this a little update shall we?

Actually, I cannot reproduce either that graph or the Independent's one precisely, as I do not know exactly how they did the smoothing, or (in the case of the IPCC) which data set they used. But this seems acceptably close to both (click for larger version):


The original data are the crosses and the blue line (5 year boxcar smooth), with the new smoothed data forming the dotted blue line which trends down more slowly from 1945-1960. The adjustment I used was a linear term which falls from 70%*0.3C in 1946, to 0 in 1960. Obviously I don't claim this is authoritative but I do believe it is more plausible than a number of other graphs I have seen on the intertubes...

The solid straight lines are the trends plotted by the IPCC. The dotted ones are equivalent trends of the adjusted data. With the help of a bit of rounding, the 50 year trend (1956-2005) changes from 0.13 to 0.12C/decade. By cherry-picking the start date to be 1946 I can get a change as large as 20%, from 0.11C/decade to 0.09C/decade. Woohoo. Let the blogorrhea continue...

Update. It certainly looks like the Independent graphic attributed to CRU is smoothed over rather more than 5 years. Roger says it is a 21y binomial filter, and if that is correct then the adjustment presented in that graph there must be just a hand-drawn guess rather than the results of any realistic calculation, since a 21y smooth would smudge out any adjustment over a rather longer interval than indicated. It's even possible that the attribution to CRU is just for the original data, and the adjustment was drawn on by a journalist. So don't take any of this too seriously for now. I would guess that the data analysts may take this opportunity to have another look at the various assumptions underlying the splicing together of different data sources, and by the time they are done there may be a whole bunch of minor adjustments throughout the time series.

Sunday, June 01, 2008

Corbyn's May forecast

After last month's success, I'm relieved to say that things are back to normal for May, so my hat is safe from the frying pan. Remember, Corbyn forecast a cool month (0.5 to 1C below normal) with average rainfall (90-115%). Well, it turned out rather warm at 13.6C (here and here), which is more than 2C above the normal for May (about 2 standard deviations), and the rain was also clearly above his range (currently 128% at Philip Eden's site, although there is one more day to update there, which cannot bring it into the forecast range now updated to 124% for the full month). The temperature looks like it should be among the 10 warmest on record (and it could possibly be the warmest May for 160 years), although we will have to wait for the official figures to be sure.

So now Corbyn is down to a 40% success rate for the year so far (4 hits from 10 forecasts). Even if he gets every remaining month for the rest of the year perfect for temperature and rainfall, he cannot climb back up to his claimed 80% success rate.

Friday, May 30, 2008

Coming out of the closet?

There is surely less to this story than meets the eye. I reckon she just took a wrong turning on the way home one day and didn't realise she had ended up in someone else's walk-in wardrobe.

Thursday, May 29, 2008

Oops

This seems pretty embarrassing for all concerned. Remember that mid-century cooling that people have been desperately fiddling their models to reproduce for years? It now turns out that it was just an artefact of dubious assumptions about measurement bias (at least, a large chunk of it - I haven't seen any official revised global surface temperature data).

It seems like it won't make much difference to climate predictions, although maybe one should expect it to reduce our estimates of both aerosol cooling and climate sensitivity marginally (I haven't read that linked commentary yet, so don't know how much detail they go into). It will also make it easier for the models to simulate the observed climate history. In fact one could almost portray this as another victory for modelling over observations, since the models have always struggled to reproduce this rather surprising dip in temperatures (eg SPM Fig 4). I've asked people about this problem myself in various seminars, and never got much of an answer. It's pretty shocking that such a problem could have been overlooked for so long.

It wasn't overlooked by everyone, actually. But I anticipate that plenty of people will try their best to avoid looking and linking in that particular direction...

Sunday, May 25, 2008

Once more unto the breach dear friends, once more...

...or fill up the bin with rejected manuscripts.

I wasn't going to bother blogging this, as there is really not that much to be said that has not already been covered at length. But I sent it to someone for comments, and (along with replying) he sent it to a bunch of other people, so there is little point in trying to pretend it doesn't exist.

It's a re-writing of the uniform-prior-doesn't-work stuff, of course. Although I had given up on trying to get that published some time ago, the topic still seems to have plenty of relevance, and no-one else has written anything about it in the meantime. I also have a 500 quid bet to win with jules over its next rejection. So we decided to warm it over and try again. The moderately new angle this time is to add some simple economic analysis to show that these things really matter. In principle it is obvious that by changing the prior, we change the posterior and this will change results of an economic calculation, but I was a little surprised to find that swapping between U[0,20C] and U[0,10C] (both of which priors have been used in the literature, even by the same authors in consecutive papers) can change the expected cost of climate change by a factor of more than 2!

We have also gone much further than before in looking at the robustness of results when any attempt at a reasonable prior is chosen. This was one of the more useful criticisms raised over the last iteration, and we no longer have the space limitations of previous manuscripts. The conclusion seems clear - one cannot generate these silly pdfs which assign high probability to very high sensitivity, other than by starting with strong (IMO ridiculous) prior belief in high sensitivity, and then ignoring almost all evidence to the contrary. Whether or not such a statement is publishable or not (at least, publishable by us), remains to be seen. I'm not exactly holding my breath, but would be very happy to have my pessimism proved wrong.

Monday, May 19, 2008

Question: when is 23% equal to 5%?

Answer: when the 23% refers to the proportion of models that are rejected by Roger Pielke's definition of "consistent with the models at the 95% level".

By the normal definition, this simple null hypothesis significance test should reject roughly 5% of the models. Eg Wilks Ch 5 again "the null hypothesis is rejected if the probability (as represented by the null distribution) of the test statistic, and all other results at least as unfavourable to the null hypothesis, is less than or equal to the test level. [...] Commonly the 5% level is chosen" (which we are using here).

I asked Roger what range of observed trends would pass his consistency test and he replied with -0.05 to 0.45C/decade. I then asked Roger how many of the models would pass his proposed test, and instead of answering, he ducked the question, sarcastically accusing me of a "nice switch" because I asked him about the finite sample rather than the fitted Gaussian. They are the same thing, Roger (to within sampling error). Indeed, the Gaussian was specifically constructed to agree with the sample data (in mean and variance, and it visually matches the whole shape pretty well).

The answer that Roger refused to provide is that 13/55 = 24% of the discrete sample lies outside his "consistent with the models at the 95% level" interval (from the graph, you can read off directly that it is at least 10, which is 18% of the sample, and at most 19, or 35%).

But that's only with my sneaky dishonest switch to the finite sample of models. If we use the fitted Gaussian instead, then roughly 23% of it lies outside Rogers proposed "consistent at the 95% level" interval. So that's entirely different from the 24% of models that are outside his range, and supports his claims...

I guess if you squint and wave your hands, 23% is pretty close to 5%. Close enough to justify a sequence of patronising posts accusing me and RC of being wrong, and all climate scientists of politicising climate science and trying to shut down the debate, anyway. These damn data and their pro-global-warming agenda. Life would be so much easier if we didn't have to do all these difficult sums.

I'm actually quite enjoying the debate over whether the temperature trend is inconsistent with the models at the "Pielke 5%" level :-) And so, apparently, are a large number of readers.

Sunday, May 18, 2008

Putting Roger out of his misery

OK, we've all had our fun, but perhaps it is time to put an end to it. There's obviously a simple conceptual misunderstanding underlying Roger's attempts at analysis, which some have spotted, but some others don't seem to have so I will try to make it as clear as possible.

The models provided a distribution of predictions about the real-world trend over the 8 years 2000-2007 inclusive. However, we have only one realisation of the real-world trend, even though there are various observational analyses of it. The spread of observational analyses is dependent on observational error and their distribution is (one hopes) roughly centred on the specific instance of the true temperature trend over that one interval, whereas the spread of forecasts depends on the (much larger) natural variability of the system and this distribution is centred on the models' estimate of the underlying forced response. Of course these distributions aren't the same, even in mean let alone width. There is no way they could possibly be expected to be the same (excepting some implausible coincidences). So of course when Roger asks Megan if these distributions differ, it is easy to see that they do. But what is that supposed to show?

People tend to get unreasonably hot under the collar in discussions about climate science, so let's change the scenario to a less charged situation. Roger, please riddle me this:

I have an apple sitting in front of me, mass unknown. I use some complex numerical models make a wild guess and estimate its mass at 100±50g (Gaussian, 2sd). I also have several weighing scales, all of which have independent Gaussian measuring errors of ±5g. I have two questions:

1. If I weight the apple once, what range of observed weights X are consistent with my estimate of 100±50g?

2. If I weigh the apple 100 times with 100 different sets of scales (each set of scales having independent errors of the same magnitude), what range of observed weight distributions are consistent with my estimate for the apple's mass of 100±50g. Hint: the distribution of observed weights can be approximated by the Gaussian form X±5g for some X. I am asking what values for X, the mean of the set of observations, would be consistent with (at the 95% level) my estimate for the true mass.

You can also ask Megan for help, if you like - but if so, please show her my exact words rather than trying to "interpret" them for her as you "interpreted" the question about climate models and observations. You can reassure her that I'm not looking for precise answers to N decimal places to a tricky mathematical problem so much as a understanding of the conceptual difference between the uncertainty in a prediction, and the uncertainty in the measurement of a single instance. It is not a trick question, merely a trivial one.

Or, dressing up the same issue in another format:

If the weather forecast for today says that the temperature should be 20±1C, and the thermometer in my garden says 19.4±0.1C, then I hope we would all agree that the observation is consistent with the forecast. Would that conclusion change if I had 10 thermometers, half of which said 19.4±0.1C and half 19.5±0.1C? Of course, in this case the distribution of observations is clearly seen to be markedly different from the distribution of the forecast. Nevertheless, the true temperature is just as predicted (within the forecast uncertainty). If there is anyone (not just Roger) who thinks that the mean observation of 19.45C is inconsistent with the forecast, please let me know what range of observed temperatures would be consistent.

Friday, May 16, 2008

Roger gets it right!

But only where he says "James is absolutely correct when he says that it would be incorrect to claim that the temperatures observed from 2000-2007 are inconsistent with the IPCC AR4 model predictions. In more direct language, any reasonable analysis would conclude that the observed and modeled temperature trends are consistent." (his bold)

Unfortunately, the bit where he tries cherry picking a shorter interval Jan 2001 - Mar 2008 and claims "there is a strong argument to be made that these distributions are inconsistent with one another" is just as wrong as the nonsense he came up with previously.

Really, I would have thought that if my previous post wasn't clear enough, he could have consulted a numerate undergraduate to explain it to him (or simply asked me about it) rather than just repeating the same stupid errors over and over and over and over again. This isn't politics where you can create your own reality, Roger.

So let's look at the interval Jan 2001-Mar 2008. I say (or rather, IDL's linfit procedure says) the trend for these monthly data from HadCRU is -0.1C/decade, which seems to agree with the value on Roger's graph.

The null distribution over this shorter interval of 7.25 years will be a little broader than the N(0.19,0.21) that I used previously, for exactly the same reason that the 20-year trend distribution is much tighter than the 8-year distribution (averaging out of short-term variability). I can't be bothered trying to calculate it from the data, but N(0.19,0.23) should be a reasonable estimate (based on an assumption of white noise spectrum, which isn't precisely correct but won't be horribly wrong). This adjustment doesn't actually matter for the overall conclusion, but it is important to be aware of it if Roger starts to cherry-pick even shorter intervals.

So, where does -0.1 lie in the null distribution? About 1.26 standard deviations from the mean, well within the 95% interval (which is numerically (-0.27,0.65) in this case). Even if the null hypothesis was true, there would be about a 21% probability of observing data this "extreme". There's nothing remotely unusual about it.

So no, Roger, you are wrong again.

Thursday, May 15, 2008

The consistently wrong chronicles

...or, how to perform the most elementary null hypothesis significance tests.

Roger Pielke has been saying some truly bizarre and nonsensical things recently. The pick of the bunch IMO is this post. The underlying question is: Are the models consistent with the observations over the last 8 years?

So Roger takes the ensemble of model outputs (8 year trend as analysed in this RC post), and then plots some observational estimates (about which more later), which clearly lie well inside the 95% range of the model predictions, and apparently without any shame or embarrassment adds the obviously wrong statement:
"one would conclude that UKMET, RSS, and UAH are inconsistent with the models".
Um....no one would not:


Update: OK, there are a number of things wrong with this picture. First, these "Observed temperature trends" stated on the left, calculated by Lucia, are actually per century not per decade, although I think they have been plotted in the right place. When OLS is used on the 8-year trends (to be consistent with the model analysis), the various obs give results of around 0.13 - 0.26C/decade, with my HadCRU analysis actually being at the lower end of this range. Second, the pale blue lines purporting to show "95% spread across model realizations" are in the wrong place. Roger seems to have done a 90% spread (5-95% coverage) which is about 20% too narrow, in terms of the range it implies.

I challenged this obvious absurdity and repeatedly asked him to back it up with a calculation. After a lot of ducking and weaving, about the 30th comment under the post, he eventually admits "I honestly don't know what the proper test is". Isn't thinking about the proper test a prerequisite for confidently asserting that the models fail it? Anyway, I'll walk through it here very slowly for the hard of understanding. I'll use Wilks "Statistical methods in the atmospheric sciences" (I have the 1st edition), and in particular Chapter 5: "Hypothesis testing". It opens:

Formal testing of hypotheses, also know as significance testing, is generally covered extensively in introductory courses in statistics. Accordingly, this chapter will review only the basic concepts behind formal hypothesis tests...[cut]

and then continues with:
5.1.3 The elements of any hypothesis test

Any hypothesis test proceeds according to the following five steps:

1. Identify a test statistic that is appropriate to the data and question at hand.

This is a gimme. Obviously, the question that Roger has posed is about the 8-year trend of observed mean surface temperature. I'm going to use an ordinary least squares (OLS) fit because that is what is already available for the models, and it is also by far the most commonly used method for trend estimation and has well understood properties. For some unstated reason, Roger chose to use Cochrane-Orcutt estimates for the observed data that he plotted in his picture, but I do not know how well that method performs for such a short time series or how it compares to OLS. Anyone who wishes to repeat the analysis using C-O should find it easy enough in principle, they will need to get the raw model output (freely available) and analyse it in that manner. I would bet a large sum of money that this will not change the results qualitatively.

2. Define a null hypothesis.

Easy enough, the null hypothesis H0 here that I wish to test is that the models correctly predict the planetary temperature trend over 2000-2007. If anyone has any other suggestion for what null hypothesis makes sense in this situation, I'm all ears.

3. Define an alternative hypothesis.

"H0 is false". This all seems too easy so far....there must be something scary around the corner.

4. Obtain the null distribution, which is simply the sampling distribution of the test statistic given that the null hypothesis is true.

OK, now the real work - such as it is - starts. First, we have the distribution of trends predicted by the models. As RC have shown, this is well approximated by a Gaussian N(0.19,0.21). (I am going to stick with decadal trends throughout rather than using a mix of time scales to give me less chance of embarassingly dropping factors of 10 as Roger has done in several places in his post. He has also plotted his blue "95%" lines in the wrong place too, but I've got bigger fish to fry.) There are firm theoretical reasons why we should expect a Gaussian to provide a good fit (basically the Central Limit Theorem). This distribution isn't quite what we need, however. The model output (as analysed) uses perfect knowledge of the model temperature, whereas the observed estimate for the planet is calculated from limited observational coverage. In fact, CRU estimate their observational errors at about 0.025 for each year's mean (at one standard deviation). This introduces a small additional uncertainty of about 0.04 on the decadal trend. That is, if the true planetary trend is X, say, then the observational analysis will give us a number in the range [X-0.08,X+0.08] with 95% probability.

Putting that together with the model output, we get the result that if the null hypothesis is true and the models' prediction of N(0.19,0.21) for the true planetary trend is correct, then the sampling distribution for the observed trend should also be N(0.19,0.21). I calculated 0.21 for the standard deviation there by adding the two uncertainties of 0.21 and 0.04 in quadrature (ie squaring, adding, taking the square root). This is the correct formula under the assumption that the observational error is independent of the true planetary temperature, which seems natural enough.

So, as I had guessed in my comments to Roger's post, considering observational uncertainty here has a negligible effect (is rounded off completely), so we could have simply used the existing model spread as the null distribution. Using this approach generally makes such tests stiffer than they should be, but it is often a small effect.

5. Compare the observed test statistic to the null distribution. If the test statistic falls in a sufficiently improbable region of the null distribution, H0 is rejected as too unlikely to have been true given the observed evidence. If the test statistic falls within the range of "ordinary" values described by the null distribution, the test statistic is seen as consistent with H0 which is then not rejected. [my emphasis]

OK, let's have a look at the test statistic. For HADCRU, the least squares trend is....0.11C/decade. That is a simple least squares to the last 8 full year values of these data. (I generally use the variance-adjusted version, on the ground that if they think there is a reason to adjust the variance, I see no reason to presume that this harms their analysis. It doesn't affect the conclusions of course.)

So, where does 0.11 lie in the null distribution N(0.19,0.21)? Just about slap bang in the middle, that's where. OK, it is marginally lower than the mean (by a whole 0.38 standard deviations), but actually closer to the mean than one could generally hope for, even if the null is true. In fact the probability of a sample statistic from the null distribution being worse than the observed test statistic is a whopping 70% (this value being 1 minus the integral of a Gaussian from -0.38 to +0.38 standard deviations)!

So what do we conclude from this?

First, that the data are obviously not inconsistent with the models at the 5% level.

Second...well I leave readers to draw their own conclusions about Roger "I honestly don't know what the proper test is" Pielke.