Friday, May 30, 2008

Coming out of the closet?

There is surely less to this story than meets the eye. I reckon she just took a wrong turning on the way home one day and didn't realise she had ended up in someone else's walk-in wardrobe.

Thursday, May 29, 2008

Oops

This seems pretty embarrassing for all concerned. Remember that mid-century cooling that people have been desperately fiddling their models to reproduce for years? It now turns out that it was just an artefact of dubious assumptions about measurement bias (at least, a large chunk of it - I haven't seen any official revised global surface temperature data).

It seems like it won't make much difference to climate predictions, although maybe one should expect it to reduce our estimates of both aerosol cooling and climate sensitivity marginally (I haven't read that linked commentary yet, so don't know how much detail they go into). It will also make it easier for the models to simulate the observed climate history. In fact one could almost portray this as another victory for modelling over observations, since the models have always struggled to reproduce this rather surprising dip in temperatures (eg SPM Fig 4). I've asked people about this problem myself in various seminars, and never got much of an answer. It's pretty shocking that such a problem could have been overlooked for so long.

It wasn't overlooked by everyone, actually. But I anticipate that plenty of people will try their best to avoid looking and linking in that particular direction...

Sunday, May 25, 2008

Once more unto the breach dear friends, once more...

...or fill up the bin with rejected manuscripts.

I wasn't going to bother blogging this, as there is really not that much to be said that has not already been covered at length. But I sent it to someone for comments, and (along with replying) he sent it to a bunch of other people, so there is little point in trying to pretend it doesn't exist.

It's a re-writing of the uniform-prior-doesn't-work stuff, of course. Although I had given up on trying to get that published some time ago, the topic still seems to have plenty of relevance, and no-one else has written anything about it in the meantime. I also have a 500 quid bet to win with jules over its next rejection. So we decided to warm it over and try again. The moderately new angle this time is to add some simple economic analysis to show that these things really matter. In principle it is obvious that by changing the prior, we change the posterior and this will change results of an economic calculation, but I was a little surprised to find that swapping between U[0,20C] and U[0,10C] (both of which priors have been used in the literature, even by the same authors in consecutive papers) can change the expected cost of climate change by a factor of more than 2!

We have also gone much further than before in looking at the robustness of results when any attempt at a reasonable prior is chosen. This was one of the more useful criticisms raised over the last iteration, and we no longer have the space limitations of previous manuscripts. The conclusion seems clear - one cannot generate these silly pdfs which assign high probability to very high sensitivity, other than by starting with strong (IMO ridiculous) prior belief in high sensitivity, and then ignoring almost all evidence to the contrary. Whether or not such a statement is publishable or not (at least, publishable by us), remains to be seen. I'm not exactly holding my breath, but would be very happy to have my pessimism proved wrong.

Monday, May 19, 2008

Question: when is 23% equal to 5%?

Answer: when the 23% refers to the proportion of models that are rejected by Roger Pielke's definition of "consistent with the models at the 95% level".

By the normal definition, this simple null hypothesis significance test should reject roughly 5% of the models. Eg Wilks Ch 5 again "the null hypothesis is rejected if the probability (as represented by the null distribution) of the test statistic, and all other results at least as unfavourable to the null hypothesis, is less than or equal to the test level. [...] Commonly the 5% level is chosen" (which we are using here).

I asked Roger what range of observed trends would pass his consistency test and he replied with -0.05 to 0.45C/decade. I then asked Roger how many of the models would pass his proposed test, and instead of answering, he ducked the question, sarcastically accusing me of a "nice switch" because I asked him about the finite sample rather than the fitted Gaussian. They are the same thing, Roger (to within sampling error). Indeed, the Gaussian was specifically constructed to agree with the sample data (in mean and variance, and it visually matches the whole shape pretty well).

The answer that Roger refused to provide is that 13/55 = 24% of the discrete sample lies outside his "consistent with the models at the 95% level" interval (from the graph, you can read off directly that it is at least 10, which is 18% of the sample, and at most 19, or 35%).

But that's only with my sneaky dishonest switch to the finite sample of models. If we use the fitted Gaussian instead, then roughly 23% of it lies outside Rogers proposed "consistent at the 95% level" interval. So that's entirely different from the 24% of models that are outside his range, and supports his claims...

I guess if you squint and wave your hands, 23% is pretty close to 5%. Close enough to justify a sequence of patronising posts accusing me and RC of being wrong, and all climate scientists of politicising climate science and trying to shut down the debate, anyway. These damn data and their pro-global-warming agenda. Life would be so much easier if we didn't have to do all these difficult sums.

I'm actually quite enjoying the debate over whether the temperature trend is inconsistent with the models at the "Pielke 5%" level :-) And so, apparently, are a large number of readers.

Sunday, May 18, 2008

Putting Roger out of his misery

OK, we've all had our fun, but perhaps it is time to put an end to it. There's obviously a simple conceptual misunderstanding underlying Roger's attempts at analysis, which some have spotted, but some others don't seem to have so I will try to make it as clear as possible.

The models provided a distribution of predictions about the real-world trend over the 8 years 2000-2007 inclusive. However, we have only one realisation of the real-world trend, even though there are various observational analyses of it. The spread of observational analyses is dependent on observational error and their distribution is (one hopes) roughly centred on the specific instance of the true temperature trend over that one interval, whereas the spread of forecasts depends on the (much larger) natural variability of the system and this distribution is centred on the models' estimate of the underlying forced response. Of course these distributions aren't the same, even in mean let alone width. There is no way they could possibly be expected to be the same (excepting some implausible coincidences). So of course when Roger asks Megan if these distributions differ, it is easy to see that they do. But what is that supposed to show?

People tend to get unreasonably hot under the collar in discussions about climate science, so let's change the scenario to a less charged situation. Roger, please riddle me this:

I have an apple sitting in front of me, mass unknown. I use some complex numerical models make a wild guess and estimate its mass at 100±50g (Gaussian, 2sd). I also have several weighing scales, all of which have independent Gaussian measuring errors of ±5g. I have two questions:

1. If I weight the apple once, what range of observed weights X are consistent with my estimate of 100±50g?

2. If I weigh the apple 100 times with 100 different sets of scales (each set of scales having independent errors of the same magnitude), what range of observed weight distributions are consistent with my estimate for the apple's mass of 100±50g. Hint: the distribution of observed weights can be approximated by the Gaussian form X±5g for some X. I am asking what values for X, the mean of the set of observations, would be consistent with (at the 95% level) my estimate for the true mass.

You can also ask Megan for help, if you like - but if so, please show her my exact words rather than trying to "interpret" them for her as you "interpreted" the question about climate models and observations. You can reassure her that I'm not looking for precise answers to N decimal places to a tricky mathematical problem so much as a understanding of the conceptual difference between the uncertainty in a prediction, and the uncertainty in the measurement of a single instance. It is not a trick question, merely a trivial one.

Or, dressing up the same issue in another format:

If the weather forecast for today says that the temperature should be 20±1C, and the thermometer in my garden says 19.4±0.1C, then I hope we would all agree that the observation is consistent with the forecast. Would that conclusion change if I had 10 thermometers, half of which said 19.4±0.1C and half 19.5±0.1C? Of course, in this case the distribution of observations is clearly seen to be markedly different from the distribution of the forecast. Nevertheless, the true temperature is just as predicted (within the forecast uncertainty). If there is anyone (not just Roger) who thinks that the mean observation of 19.45C is inconsistent with the forecast, please let me know what range of observed temperatures would be consistent.

Friday, May 16, 2008

Roger gets it right!

But only where he says "James is absolutely correct when he says that it would be incorrect to claim that the temperatures observed from 2000-2007 are inconsistent with the IPCC AR4 model predictions. In more direct language, any reasonable analysis would conclude that the observed and modeled temperature trends are consistent." (his bold)

Unfortunately, the bit where he tries cherry picking a shorter interval Jan 2001 - Mar 2008 and claims "there is a strong argument to be made that these distributions are inconsistent with one another" is just as wrong as the nonsense he came up with previously.

Really, I would have thought that if my previous post wasn't clear enough, he could have consulted a numerate undergraduate to explain it to him (or simply asked me about it) rather than just repeating the same stupid errors over and over and over and over again. This isn't politics where you can create your own reality, Roger.

So let's look at the interval Jan 2001-Mar 2008. I say (or rather, IDL's linfit procedure says) the trend for these monthly data from HadCRU is -0.1C/decade, which seems to agree with the value on Roger's graph.

The null distribution over this shorter interval of 7.25 years will be a little broader than the N(0.19,0.21) that I used previously, for exactly the same reason that the 20-year trend distribution is much tighter than the 8-year distribution (averaging out of short-term variability). I can't be bothered trying to calculate it from the data, but N(0.19,0.23) should be a reasonable estimate (based on an assumption of white noise spectrum, which isn't precisely correct but won't be horribly wrong). This adjustment doesn't actually matter for the overall conclusion, but it is important to be aware of it if Roger starts to cherry-pick even shorter intervals.

So, where does -0.1 lie in the null distribution? About 1.26 standard deviations from the mean, well within the 95% interval (which is numerically (-0.27,0.65) in this case). Even if the null hypothesis was true, there would be about a 21% probability of observing data this "extreme". There's nothing remotely unusual about it.

So no, Roger, you are wrong again.

Thursday, May 15, 2008

The consistently wrong chronicles

...or, how to perform the most elementary null hypothesis significance tests.

Roger Pielke has been saying some truly bizarre and nonsensical things recently. The pick of the bunch IMO is this post. The underlying question is: Are the models consistent with the observations over the last 8 years?

So Roger takes the ensemble of model outputs (8 year trend as analysed in this RC post), and then plots some observational estimates (about which more later), which clearly lie well inside the 95% range of the model predictions, and apparently without any shame or embarrassment adds the obviously wrong statement:
"one would conclude that UKMET, RSS, and UAH are inconsistent with the models".
Um....no one would not:


Update: OK, there are a number of things wrong with this picture. First, these "Observed temperature trends" stated on the left, calculated by Lucia, are actually per century not per decade, although I think they have been plotted in the right place. When OLS is used on the 8-year trends (to be consistent with the model analysis), the various obs give results of around 0.13 - 0.26C/decade, with my HadCRU analysis actually being at the lower end of this range. Second, the pale blue lines purporting to show "95% spread across model realizations" are in the wrong place. Roger seems to have done a 90% spread (5-95% coverage) which is about 20% too narrow, in terms of the range it implies.

I challenged this obvious absurdity and repeatedly asked him to back it up with a calculation. After a lot of ducking and weaving, about the 30th comment under the post, he eventually admits "I honestly don't know what the proper test is". Isn't thinking about the proper test a prerequisite for confidently asserting that the models fail it? Anyway, I'll walk through it here very slowly for the hard of understanding. I'll use Wilks "Statistical methods in the atmospheric sciences" (I have the 1st edition), and in particular Chapter 5: "Hypothesis testing". It opens:

Formal testing of hypotheses, also know as significance testing, is generally covered extensively in introductory courses in statistics. Accordingly, this chapter will review only the basic concepts behind formal hypothesis tests...[cut]

and then continues with:
5.1.3 The elements of any hypothesis test

Any hypothesis test proceeds according to the following five steps:

1. Identify a test statistic that is appropriate to the data and question at hand.

This is a gimme. Obviously, the question that Roger has posed is about the 8-year trend of observed mean surface temperature. I'm going to use an ordinary least squares (OLS) fit because that is what is already available for the models, and it is also by far the most commonly used method for trend estimation and has well understood properties. For some unstated reason, Roger chose to use Cochrane-Orcutt estimates for the observed data that he plotted in his picture, but I do not know how well that method performs for such a short time series or how it compares to OLS. Anyone who wishes to repeat the analysis using C-O should find it easy enough in principle, they will need to get the raw model output (freely available) and analyse it in that manner. I would bet a large sum of money that this will not change the results qualitatively.

2. Define a null hypothesis.

Easy enough, the null hypothesis H0 here that I wish to test is that the models correctly predict the planetary temperature trend over 2000-2007. If anyone has any other suggestion for what null hypothesis makes sense in this situation, I'm all ears.

3. Define an alternative hypothesis.

"H0 is false". This all seems too easy so far....there must be something scary around the corner.

4. Obtain the null distribution, which is simply the sampling distribution of the test statistic given that the null hypothesis is true.

OK, now the real work - such as it is - starts. First, we have the distribution of trends predicted by the models. As RC have shown, this is well approximated by a Gaussian N(0.19,0.21). (I am going to stick with decadal trends throughout rather than using a mix of time scales to give me less chance of embarassingly dropping factors of 10 as Roger has done in several places in his post. He has also plotted his blue "95%" lines in the wrong place too, but I've got bigger fish to fry.) There are firm theoretical reasons why we should expect a Gaussian to provide a good fit (basically the Central Limit Theorem). This distribution isn't quite what we need, however. The model output (as analysed) uses perfect knowledge of the model temperature, whereas the observed estimate for the planet is calculated from limited observational coverage. In fact, CRU estimate their observational errors at about 0.025 for each year's mean (at one standard deviation). This introduces a small additional uncertainty of about 0.04 on the decadal trend. That is, if the true planetary trend is X, say, then the observational analysis will give us a number in the range [X-0.08,X+0.08] with 95% probability.

Putting that together with the model output, we get the result that if the null hypothesis is true and the models' prediction of N(0.19,0.21) for the true planetary trend is correct, then the sampling distribution for the observed trend should also be N(0.19,0.21). I calculated 0.21 for the standard deviation there by adding the two uncertainties of 0.21 and 0.04 in quadrature (ie squaring, adding, taking the square root). This is the correct formula under the assumption that the observational error is independent of the true planetary temperature, which seems natural enough.

So, as I had guessed in my comments to Roger's post, considering observational uncertainty here has a negligible effect (is rounded off completely), so we could have simply used the existing model spread as the null distribution. Using this approach generally makes such tests stiffer than they should be, but it is often a small effect.

5. Compare the observed test statistic to the null distribution. If the test statistic falls in a sufficiently improbable region of the null distribution, H0 is rejected as too unlikely to have been true given the observed evidence. If the test statistic falls within the range of "ordinary" values described by the null distribution, the test statistic is seen as consistent with H0 which is then not rejected. [my emphasis]

OK, let's have a look at the test statistic. For HADCRU, the least squares trend is....0.11C/decade. That is a simple least squares to the last 8 full year values of these data. (I generally use the variance-adjusted version, on the ground that if they think there is a reason to adjust the variance, I see no reason to presume that this harms their analysis. It doesn't affect the conclusions of course.)

So, where does 0.11 lie in the null distribution N(0.19,0.21)? Just about slap bang in the middle, that's where. OK, it is marginally lower than the mean (by a whole 0.38 standard deviations), but actually closer to the mean than one could generally hope for, even if the null is true. In fact the probability of a sample statistic from the null distribution being worse than the observed test statistic is a whopping 70% (this value being 1 minus the integral of a Gaussian from -0.38 to +0.38 standard deviations)!

So what do we conclude from this?

First, that the data are obviously not inconsistent with the models at the 5% level.

Second...well I leave readers to draw their own conclusions about Roger "I honestly don't know what the proper test is" Pielke.

Monday, May 12, 2008

Are you avin a laff?

Round and round the mulberry bush...

Roger Pielke, 30 April:

there is in fact nothing that can be observed in the climate system that would be inconsistent with climate model predictions. If global cooling over the next few decades is consistent with model predictions, then so too is pretty much anything and everything under the sun.
Me (in the comments):
over the 30 year time frame there will be strong warming
Roger:

I see that you neglected to address my central question

Me:
I explicitly wrote "over the 30 year time frame there will be strong warming" - and actually 20y would be a safe bet too.

Roger:

I see that you have once again avoided addressing this question.

Me here in more detail:
Warming over 30 years is assured, 20 years must be "very likely", 10 years I would certainly say "likely" but that is a bit of a rough estimate.

I could do a detailed calculation about the probability of different trends over the next 30 years, but that's already been done.

Roger:

James, when you write, "Warming over 30 years is assured, 20 years must be "very likely", 10 years I would certainly say "likely" but that is a bit of a rough estimate" you are much closer to what I am looking for. I am asking for somewhat less "roughness" in these estimates, and grounding them more quantitatively than this sort of hand-waving which is a common response.
Me:
you asked for more quantitative estimates, but did you read the link I provided, where such quantitative estimates were explicitly presented 6 years ago?

If, after reading that (and the two papers it refers to) you still have a question then feel free to follow up.


Roger:

If you think that I'm focused on 2020-2030 (the subject of the essay in Nature that you linked to) then you are not really paying attention.


Me:

Roger, you started off with "if global cooling over the next few decades..." (my emphasis) which remains on your blog even after several people have pointed out that it is a gross mischaracterisation of the Keenlyside paper. So I pointed you to explicit probabilistic predictions about the next few decades which are as clear as day about the probability of cooling over that time frame.

Now you say you are not focussed on the next few decades...

If you want shorter term, I'm sure you have already seen the Smith et al Science paper, within which 50% of years post 2009 are predicted to beat the 1998 record. But as you can see, this is still a rather young area of science, and Keenlyside disagree to some extent (although not as strongly as some have portrayed it - I think their 10y mean forecast could still validate even if we see some new records).

Roger again (not in response to me, but bringing up the topic again on his blog):


You can just clear all of this up by answering my original question:

What observations of the climate system to 2020 would be inconsistent (lets say at the 95% level of certainty) with the climate model projections of the IPCC AR4? It is a simple question. use global average surface temps from UKMET as the variable of interest if you'd like, since that is what we've been discussing, or use a different one.


Me, hopefully for the last time:
Stott and Kettleborough estimate that the global mean temperature in the decade 2020–30 will be 0.3–1.3 K greater than in 1990–2000 (5–95% likelihood range)
and
Knutti et al. find that the projected distribution of likely surface warming is independent of the choice of emission scenario for the next several decades; that the probable warming for 2020–30 relative to 1990–2000 is about 0.5–1.1 K (5–95% likelihood range)
(these being direct quotes from the paper I linked to earlier).

As I mentioned back then, I think these forecasts do have some limitations, but since I pointed them out to Roger 10 days ago it is more than a little tendentious of him to repeatedly insist that they do not exist, and furthermore to pretend that he's been met with nothing but dodging and evasion in response to his question about which observations over the next few decades would be inconsistent (at the 5% level) with the model forecasts. The reason why that paper specifically looked at decadal averages over a 30 year interval is because on this time frame the GW signal is clearly visible above natural variability, but its magnitude is not very sensitive to emissions scenarios (within reason). But it is a simple matter of reading off the graphs for anyone who wants a different forecast interval. However, it seems quite clear that Roger is more interested in pretending that the answer has not been provided, than in actually looking at it. He's avin a laff.

Also relevant: Eli Rabett and RC.

Sunday, May 11, 2008

This is a local hospital for local people - there's nothing for you here.

This is an ugly story which I heard on the grapevine, and which is unlikely to feature in the Japanese press (unlikely to be made into a bizarre comedy either). Recently someone here in Japan needed some medical treatment, so they found an official web-site listing local hospitals with English-speaking staff, and when they tried to go there, they were refused on the grounds of nationality - the head doctor had simply decided they were going to stop taking any foreigners! It was not an urgent case, and they found treatment elsewhere, but it's still a rather shocking reminder of how this sort of bigotry is casually accepted at all levels in society here. They may be desperate for foreign tourists to come and spend their money here, and for "guest workers" to come to prop up their economy (so long as they don't get big ideas about settling here, and go home after a few years), but a large proportion don't actually think foreigners are human, and a "foreigners not welcome" attitude, although thankfully rare (except when renting accommodation, where it is the rule rather than the exception) is still considered quite acceptable.

Kerosene-soaked man catches fire after trying to smoke at Nagoya police station