Thursday, September 07, 2017
More on Bayesian approaches to detection and attribution
Saturday, April 25, 2015
BlueSkiesResearch.org.uk: The EGU review 2015
Tuesday, March 03, 2015
Climate change by numbers
Update: Oh, this is interesting. It's a blog post about the programme from the mathematician (Norman Fenton) who presented the 95% section. Turns out he is actually a Bayesian who clearly understands how that number is tarnished by prosecutor's fallacy, and he argues that the scientific debate would be improved by a greater use of Bayesian methods!
Sunday, March 31, 2013
Decadal prediction stuff part 2
They also compared their forecast to a couple of alternative approaches:
Wednesday, March 21, 2012
HadCRUT4: 1998 and all that
As was exclusively revealed on this blog some time ago, according to this updated analysis, 1998 is no longer regarded as the warmest year on record, having been overtaken first by 2005 and then again in 2010. Of course, this was not really controversial in scientific circles, as the small cold bias of HadCRUT3 (due to data voids in the most rapidly warming areas) had been well documented. Competing analyses (NCDC and GISTEMP) that smooth over the gaps already showed these results.
Some of the major differences between the data sets are illustrated by this figure from the paper:

I've added two green ovals to each map to highlight some differences - firstly, filling in the big gap in Siberia, where it was clear in HadCRUT3 that there had been lots of warming even though there were gaps in the coverage. The addition of more observations has merely confirmed what everyone knew, and the previous data has not been changed in any significant way. Moreover, the additional data in the relatively cool area of the south Indian ocean shows that they didn't just try to collect data in the hottest regions. So whatever desperate attempts the sceptics make to discredit this update, they simply don't have a leg to stand on.
The problem really is in the presentation of the data average as representing "global mean temperature" in the first place, when it doesn't, as the data are not missing at random. One good way to deal with missing data in this sort of situation is to fill the gaps with some sort of smoothing (there are a wide range of options of varying sophistication) to generate a global field before averaging, as NCDC and GISTEMP have always done. However, this doesn't matter in the context of the sort of detailed model-data comparisons that underpin detection and attribution, since they generally work at the level of the gridded data and voids are ignored. The mid-century changes to ocean data may however have a modest impact here.
David Whitehouse (or someone impersonating him) has been quick off the mark to bluster about how he would have really won the bet anyway. Unfortunately his comments are highly misleading.
Firstly, we never clearly specified HadCRUT3. I did try to get David to confirm exact details via email, but he refused, and I believe the precise phrase used by Tim Harford was "the Hadley Centre analysis". Maybe his behaviour should have been a red flag that he would resort to rewriting history whenever possible, but I thought it was unlikely to be ambiguous. Of course I didn't know the HC were planning to change their analysis, even though it is inevitable that these sort of changes do take place over time (hence the "3" in HadCRUT3). For how a bet of this type is more clearly specified, see here for example: "in the data set HadCRUTv. or successor data set. Successor data set is the data set used by the Hadley Centre to compare hottest years to the media at that time." Incidentally, Gabi has not won that bet yet, as it specifically refers to the setting of a new record, and 2010 was not.
Secondly, he claims that because 2010 is no hotter than 2005 in the new analysis, he wins anyway. This is more than a bit ridiculous, as the entire bet was predicated on the usually long interval after 1998 in which the record had (apparently) not been broken.
(Of course it will only be a matter of weeks before we start seeing sceptical arguments along the lines of "no global warming since 2005").
Monday, November 07, 2011
The null hypothesis in climate science
Trenberth argues that, since the null (that we have not changed the climate) is not true, we should try to test some other null hypothesis. He sounds like someone who has just discovered that the frequentist approach is actually pretty useless in principle (as I've said many times before, it is fundamentally incapable of even addressing the questions that people want answers to), but although he seems to be grasping towards a Bayesian approach, he hasn't really got there, at least not in a coherent and clear manner. Curry is just nonsense as usual, and beside noting that she has (1) grossly misrepresented the IAC report and (2) abjectly failed to back up the claims that Curry and Webster made in a previous paper, there isn't really anything meaningful to discuss in what she said.
Myles Allen's commentary is by some distance the best of the bunch, in fact I broadly agree (shock horror) with what he has said. If one is going to take a frequentist approach, the null hypothesis of no effect is often an entirely reasonable starting point. It is important to understand that rejecting the null does not simply mean learning that there has been some effect, but it also indicates that we know (at least at some level of confidence) the direction of the effect! That is, it is not only an effect of zero which is rejected, but all possible negative (say) effects of any magnitude too - this generalisation may not be strictly correct in all possible applications of this sort of methodology, but I'm pretty sure it is true in practice for the D&A field. Especially when we are talking about the local incidence of extreme weather, there really are many cases when we have little reason for a prior belief in an anthropogenically-forced increase versus a decrease in these events, so a reasonable Bayesian approach would also start from a prior which was basically symmetric around zero. The correct interpretation of a non-rejection of the null here is not "there has been no effect" but rather "we don't know if AGW is making these events more or less likely/large". Much of Trenberth's complaint could be more productively aimed at the routine misinterpretation of D&A results, rather than the method of their generation. Trenberth also sometimes sounds like he is arguing that we should always assume that every bad thing was caused by (or at least exacerbated by) AGW, but this simply isn't tenable. Even if storminess increases in general, changes in storm tracks might lead to reduction in events in some areas, with Zahn and von Storch's work on polar lows an obvious example of this. On the other hand, there are also some types of event where we may have decent prior belief in the nature of the anthropogenically-forced change (such as temperature extremes) and in these cases it would be reasonable for a Bayesian to use a prior that reflects this belief.
I can find one thing to object to in Myles' commentary though, and that's the manner in which he tries to pre-judge the "consensus" response to Trenberth's argument. Noting that he (Allen) is in fact a major figure in forming the "consensus" in these private meetings where the handful of IPCC authors decide what to say, it sounds to me rather like a pre-emptive strike against anyone who might be tempted to take the opposing view. I would prefer it if he restricted himself to arguing on the basis of the issues rather than that he holds/forms the majority view. His behaviour here is reminiscent of the way he (and others) tried to reject our arguments about uniform priors, on the basis that everyone had already agreed that his approach was the correct solution. All that achieved was to slow the progress of knowledge by a few years.
Tuesday, November 01, 2011
Judith Curry, Detection and Attribution and the IAC report
But while having a scan of the relevant blogs, I came across something else that struck me as worth mentioning.
It concerns La Curry's criticism of the IPCC, particularly the statement of WG1 that "Most of the observed increase in global average temperatures since the mid-20th century is very likely due to the observed increase in anthropogenic GHG concentrations". Curry (and Webster) apparently don't like this statement. If they left it at that, then fmdidgad might be the most natural response, but rather than leaving it at that, they double down on it by asserting here that the IAC supports their view.
The relevant section of the IAC report is here, and it can be readily checked that their criticism which C&W quote was in fact aimed at the far more vague statements frequently made in WG2, and was clearly not intended to be applied to the WG1 statement at all. In fact, the WG1 statement is not imprecise in this sense at all, it is simply a one-sided probability statement of the general form "The probability of x > y is p". In this case x is the warming caused by anthropogenic effects, y is ~0.3C (being half the observed trend of ~0.6C) and p is 90%. Another IPCC statement of logically equivalent form is that the equilibrium climate sensitivity is very likely to be greater than 1.5C. If the IAC had intended their criticism to apply to the huge number of similarly-structured probabilistic statements in WG1, it is surely inconceivable that they would not have mentioned this criticism anywhere in the section where they actually address the treatment of probability in WG1. Indeed, the only place in that chapter where the IAC mention this issue of "imprecise statements, made without reference to the time period under consideration or to a climate scenario under which the conclusions would be true" is in that one section and even page where they are quite explicitly and specifically addressing WG2. There is, of course, nothing significantly ambiguous about the time period under consideration or the scenario under consideration in the statement that C&W object to.
I invite either of Curry and Webster to explain why they believe that the IAC intended this criticism, so clearly aimed at WG2, to apply to that D&A-based statement in WG1. Or alternatively, they could abandon their patently untenable claim that the IAC "shares their concerns" over this statement. Of course, I've been looking forward to Curry explaining her "Italian flag" blether for a long time now, to no effect. So I'm not holding my breath.
Friday, August 12, 2011
How many of Roger's findings about probability manage to be wrong? Answer: he's more inventive than you might expect.
It turns out that 100%-28% = 72% is merely the average (lower bound) probability level associated with the statements they made. Such as "It is very likely that hot extremes, heat waves and heavy precipitation events will continue to become more frequent." Here "very likely" means greater than 90%. So, given 10 such statements, the IPCC is saying that they would expect the "very likely" outcome to occur about 9 times, and not occur about once. And similarly for "likely" (66%). Averaging over all the probabilistic statements, it should be expected that in about 28% of cases, the (probabilistically) preferred outcome will not actually happen.
And in Roger-world, this means that 28% of the statements are "incorrect". Note, however, that he does not make this silly claim in the paper itself, but only in his blog post.
To see why this interpretation is nonsensical, consider a single roll of a fair die. I state (accurately) that it is "likely" to lie in the range 1-5. If I roll a 6, then in Roger-world, my statement was incorrect. However, it was not incorrect, and Roger is simply wrong to claim so.
As you can see from the comments, I challenged Roger on this, and his response (entirely in character) is to duck and weave. In his comment #5, for example, he shamelessly misrepresents what I said, and brings up the red herring of a definitive prediction (when in fact I had clearly made a probabilistic one, and the distinction is of course absolutely fundamental to the point). The obvious elephant in the room that Roger cannot bring himself to acknowledge is that the statement is correct irrespective of the outcome of the roll. "Correctness" of a single statement simply isn't something that can be directly validated (or invalidated) by the outcome, and the accurate calibration of a probabilistic prediction system actually relies on having the appropriate number of "failures" for each level of probability.
I realise of course that having done some rather boring textual analysis that in his own words amounts to "Nothing too interesting, really", Roger is just rabble-rousing on his blog. I'm confident that any competent scientist will see straight though it, but that's hardly his target audience.
As for what the 72%/28% average actually does mean, it doesn't actually tell us anything except that the IPCC makes a lot of statements about things that it is only (by its own admission) moderately confident about. It might in principle be interesting to see how the confidence level changes over time, but only if the set of statements were to be held fixed from one assessment to the next. People have looked at climate sensitivity estimates (hardly changed) and detection and attribution (increased markedly in confidence) but not a lot else AIUI. I suppose we can anticipate Roger claiming that the next report is either more correct, or less, depending on what mix of statements they happen to include :-)
Incidentally, and although it's a minor point it is perhaps telling in terms of his overall level of competence, Roger is also wrong where he claims that if the statements are not independent, then the proportion of "incorrect" will be higher than 28%. Actually, if the statements are not independent (while still being correctly calibrated), then the proportion that do not come to pass would still be 28% in expectation, just with higher variance, meaning that either a larger or smaller proportion would not be surprising. Unlike the simple misinterpretation in his blog post title, this elementary error is actually made in the paper itself.
Wednesday, December 29, 2010
Snowfalls are now just a thing of the past
According to Dr David Viner, a senior research scientist at the climatic research unit (CRU) of the University of East Anglia,within a few years winter snowfall will become "a very rare and exciting event"."Children just aren't going to know what snow is," he said.
[...]
Professor Jarich Oosten, an anthropologist at the University of Leiden in the Netherlands, says that even if we no longer see snow, it will remain culturally important.
"We don't really have wolves in Europe any more, but they are still an important part of our culture and everyone knows what they look like," he said.
David Parker, at the Hadley Centre for Climate Prediction and Research in Berkshire, says ultimately, British children could have only virtual experience of snow. Via the internet, they might wonder at polar scenes - or eventually "feel" virtual cold.
Heavy snow will return occasionally, says Dr Viner, but when it does we will be unprepared. "We're really going to get caught out. Snow will probably cause chaos in 20 years time," he said.
Of course snow falling always has and always will cause chaos in Britain. Until it stops falling completely, which may now be a few years further away than was previously thought :-)
(for those who've been living in a cave for the last few years, or at least outside the UK, it appears that in fact reports of the demise of snowfall are greatly exaggerated, at least according to the winters of 2008, 2009 and 2010).
Actually, it's better than that, because the latest research is that all this snowfall is actually more proof of global warming after all! I await with amusement the reaction of the detection and attribution community to this proof that all their results are bogus, since they have already proved that the "warmer winters" are being caused by anthropogenic global warming (eg here, a paper I just saw today). Personally, while I accept it's theoretically possible for AGW to cause some localised cooling (at least on a temporary basis), my money is on the D&A results for the time being. Invest in Scottish skiing resorts (at least for the long term) at your peril...
(and as a footnote to the pedants, I know there doesn't necessarily have to be a contradiction between snowfall and warmth, but in fact this December has been not only snowy but also perhaps the coldest since records began in the UK).
Sunday, December 19, 2010
AGU part 2
Perhaps it's best to blog a little after the event, as it gives time for the boring and disappointing bits to fade away, leaving a more positive overall impression. And with an 11h flight I have ample time to wax lyrical about the good bits too, rather than struggling to scribble something down in a spare moment. Indeed I now see on checking through my notes that there were a couple of interesting talks on Tuesday, that I didn't mention before, in the section on uncertainty quantification. First Carol Snyder presented an unusually high estimate of climate sensitivity based on paleo data, which appears to hinge on a strong estimated LGM cooling. Not that this makes her wrong of course, but the previous week, Andreas Schmittner had presented a somewhat contrary result at the PMIP meeting, so I will email them to try to identify the discrepancy. Derek Lemoine also presented some evidence, probably compatible with Frank et al (who spoke in the morning), supporting a lowish (but positive) carbon cycle feedback at the bottom end of the model range. Jules went to a session on geoengineering where everyone seems to be trying to prove that even if we could restrain the rise in global mean temperature, it was only by screwing up all regional patterns of rainfall.
Wednesday
On Wednesday morning I started off at writers corner, where several authors of popular science books discussed their work and/or their lives. And then at the end of this session we had Greg Craven. It tended rather to the hysterical, and I don't mean that in a good way. My complete unabridged notes on his presentation read "I am insane. Apocalpyse soon." although in the interests of fairness I should point out that only the first sentence was a direct quote, the second was merely my personal summary of events. Perhaps the most I should say is that since Stephen Mosher wrote in detail about how much he didn't like the panel discussion later that day, I can only assume that he did not attend the prior presentations. I actually walked out of the panel discussion at the point that Greg started telling an uppity woman (president of some equal opportunities organisation, no less) that she had said her piece and could she please shut up and sit down. Pot, meet kettle. Oppenheimer, on the other hand, gave a sensible talk that I found interesting, if not too earth-shattering. He argued that scientists have a general duty to engage with the public, and even if we didn't want to do it individually, we can't necessarily avoid it in a world where merely being a climate scientist puts us in the firing line. However, under the rubric of "engaging" he discussed a wide range of options, and I was amused to see him list blogging as ranking higher than merely participating in assessments such as the IPCC and NAS :-) He also emphasised the importance of only speaking in areas where you had earnt credibility based on your published record, which formed an interesting backdrop to Judith Curry's talk later that day. She devoted her time to accusing the IPCC of ignoring the tails of the pdfs of climate sensitivity that were clearly presented in the very figure that she repeatedly referred to and explicitly emphasised in the summary ("values substantially higher than 4.5C cannot be excluded"), then read out a few cartoons and finally, literally out of nowhere, concluded that therefore they had underestimated the magnitude of decadal variability and that their detection and attribution results were unsound! Really, I'm not making this up, it was actually how it happened. These latter topics were first introduced on her concluding slide and there was no hint of supporting argument. She also talked about the "modal falsification" of Betz 2009, (which I haven't read but just googled now, is there a free version somewhere?) so I asked if and how this "falsification" (and she used the scare quotes herself) was distinct from assigning a low posterior probability in a Bayesian sense. She replied that it could be considered the same, at which point some of the audience were shaking their heads and others were nodding in agreement. From which I conclude that nobody, including Judith, knows what Judith means. Unfortunately, she didn't seem to be anywhere to be found at the end of the session and I didn't see her at any of the other relevant sessions where people actually dealing with these sorts of issues were actually presenting concrete results.
Thursday
There was more communication stuff on Thursday morning, much along the lines of wailing about how nasty everyone (well, Republicans and/or denialists, at least) was being to climate scientists, and how we need to educate everyone about the Truth of climate change. Which of course is true to some extent, but I'm not really convinced it is worth the time and energy that was devoted to it at the AGU. Tim Palmer gave a really good Bjerknes Lecture, which I nearly didn't go to as I've seen most of the content before (at the INI) but I'm glad I did as it seemed much better this time round. Of course, as he mentioned, it was a sort of anti-Bjerknes lecture in some ways, because the eponymous scientist was firmly rooted in the deterministic world and Tim is very much in the probabilistic/ensemble forecasting mould (as everyone in NWP has to be, of course, and I'm sure Bjerknes would agree were he around today). In fact Tim is a strong and convincing advocate of the use of stochastic parameterisations in climate models, which formed the main content of his talk. I agree they are a good idea and must remember to mention some time that Jim Hansen used a random number generator in the cloud scheme of his 1984 model :-) One of the numerous modelling groups here is using a more modern equivalent, so I don't think it's something that people are particularly hostile to. Where I part company with Tim is in his advocacy of a single coordinated model-building effort, which to be fair he hardly mentioned this time. I think it is clear that even with stochastic physics, we would still need a range of different models to investigate our uncertainty in long-term climate change meaningfully, and Suki Manabe made exactly this point in his inimitable style in the questions after the talk.
After lunch there was even more on communication. I'm not sure if it was deliberate or not, but the session discussing how to cope with this blogospheric "other" that scientists don't really understand was running in parallel with a panel of science bloggers offering advice to those who were prepared to join in. I only stayed for a little of the latter, it was pretty anodyne stuff. As jules noted, they all seemed to be "normal" geoscientists, none of them were involved in the post-normal world of climate science so their take on questions of debate, argument and abuse seemed somewhat rose-tinted to me. Not that that should put anyone off who is considering getting involved. I also thought their attitude towards discussing on-going and unpublished work was rather 20th century.
Then it was time for us to head off to defend our posters, which we had planned for beer o'clock to ease the pain. Unlike the EGU where the poster session was arranged every evening, here the poster session was scheduled for a full half-day (4h20m) out of which you were supposed to pick at least an hour that you would be there. Of course this is horribly inefficient and made it very hard to meet specific people, though perhaps there were enough random encounters that it didn't matter and a good poster should be pretty self-explanatory anyway. I didn't have that many visitors but made up for that by having long chats with several of them.
Friday
Friday was probably the most fun, with lots of sessions on using data together with models. First the reanalysis session, where various improvements and new analyses were presented. There are still lots of problems with inhomogeneities in observations which make the trends in precipitation extremely dubious - even at the global average scale, different analyses gave substantially different results. Jules heard about various attempts at carbon sequestration, but it wasn't clear if they would be practical and effective. Then in the afternoon there was a big session on using observations to test and validate models, which is of course right up our street. There is a big push for more of this in the context of the CMIP5 model runs, and it seems there will be some basket of comparisons proposed although it must be remembered that no-one actually knows what particular features are important in improving model predictions. As well as a number of fairly detailed and specific studies Wendy Parker spoke about the broader context of how we could think about model adequacy and performance. There had already been a whole session on more philosophical aspects of model interpretation which I did not attend but jules did, I think many people are generally aware of the issues, but struggling to find concrete solution. Anyway, I didn't think that anyone presented anything particularly earth-shattering in this session but unlike some of the other more disappointing parts earlier in the week, they were at least presenting some new ideas and results, rather than either merely re-reading an old long-published paper I already knew about (which are few had done) or else talking vaguely about what they were hoping to do in the next few years. I also stuck my oar in with a few questions and comments, which at least made it interesting to me :-) I also have a few more things to chase up with some of the speakers via emails as I didn't quite understand what they had done, or perhaps why they had done it.
So in the end I suppose the whole thing was quite fun and I'm certainly happier on the plane on the way home than I was a few days ago when I wrote the previous post. Given that we were tired before we even started (following on directly from the PMIP meeting in Kyoto last week) and that jules came down with a cold that hung around all week, we are glad to be heading home. In fact our devotion to duty is such that we turned down the offer of $800 each to get bumped off the plane which was over-booked. Maybe we should have taken the money and gone straight back to the Mac store but we are sufficiently worn out that another day of sightseeing in SF (especially given the rain) didn't really appeal. Jamstec would also have had a fit of course, which almost tipped the balance in favour of staying, but not quite. And anyway, they've already bought us everything that the Mac store sells :-)
Friday, August 06, 2010
Wiley Interdisciplinary Reviews: Climate Change
The Editors seem a bunch of slightly unconventional people, a little removed from the mainstream IPCC stalwarts though eminent enough and with some IPCC links: Hulme, Pielke, von Storch, Nicholls, Yohe are names that many will be familiar with. The others are probably all famous too, but I'm too ignorant to recognise them. I'm sure the journal is not intended as a direct rival to the IPCC, but it may turn out to provide an interesting and slightly alternative perspective.
The articles to date include a mix of authoritative reviews from leading experts - such as Parker on the urban heat island, Stott on detection and attribution, interspersed with perhaps more personal and less authoritative articles. I can safely say that without risk of criticism because one of them is mine - a review on Bayesian approaches to detection and attribution. This article had a rather difficult genesis. I was initially dubious about my suitability for the task and indeed the value of the article, but after declining once (and proposing another author, who also declined) I changed my mind and had a go. My basic difficulty with addressing the concept is that D&A has always seemed to me to be a rather limited and awkward approach to the question of estimating the effects of anthropogenic and natural forcing, which is tortured into a frequentist framework where it doesn't really fit. Eg, no sane person believes these forcings have zero effect, so what exactly is the purpose of a null hypothesis significance test in the first place? However, conventional D&A has such a stranglehold on the scientific conscious that most Bayesian approaches have actually mimicked this frequentist alternative of the Bayesian estimate that you really wanted in the first place. It all seems a bit tortured and long-winded to me.
Anyway, I eventually found some things to say, which hopefully aren't entirely stupid and help to show how a Bayesian approach might actually be useful in answering the questions that (sensible) people might want to know the answer to, rather than the relatively useless questions that frequentist methods can answer, which are then inevitably misinterpreted as answers to the questions that people wanted to answer in the first place (as I argue and document in the article).
Another of the personal and argumentative articles was contributed by Jules, who was invited to say something about skill and uncertainty in climate models. This was actually the article that sparked off our "Reliability" paper, as our discussions kept coming back to the odd inconsistency between the flat rank histogram evaluation that I know is standard in most ensemble prediction, versus the Pascal's triangle distribution that a truth-centred ensemble would generate (ie, if each model is independently and equiprobably greater than or less than the observations, then the obs should generally be very close to the ensemble median). Of course this problem didn't take long to solve once we had set out the issue clearly enough to recognise that there really were two incompatible paradigms in play, and Jules even ended up citing the GRL paper which overtook her WIREs one in the review process.
Perhaps of more widespread interest to other readers, is a simple analysis of the skill of Hansen's forecast which he made back in 1988 to the US Congress. We'd actually had lengthy discussions with several people (listed in the acknowledgements) a year or two ago, trying to resurrect the old model code that was used for this prediction in order to re-run it and analyse its outputs in more detail. But this proved to be impossible. (The code exists but has been updated and gives substantially different results. If only the code had been published in GMD!) Therefore we were left with nothing more than the single printed plot of global mean temperature to look at. This didn't seem much to base a proper peer-reviewed paper on, so the idea died a death. When this WIREs invitation came long it seemed like a good opportunity to publish the one usable result we had obtained, as an example of what skill means. The headline result is that under any reasonable definition of skill, the Hansen prediction was skillful. While no great surprise, I don't think it has been presented in quite those terms before. It's a shame that we weren't able to generate a more comprehensive set of outputs though which might have given a more robust result than this single statistic.
The null hypothesis of persistence (no change in temperature) was found to give best performance over the historical interval, compared to extrapolating a trend. So this is the appropriate zero-skill baseline for evaluating the forecast. Nowadays with the AGW trend well established, probably most would argue that a continuation of that trend is a good bet, though that still leaves open the question of how long a historical interval to fit the trend over. Anyway, the model forecast is clearly drifting on the high side by now - most likely due to some combination of high sensitivity, low thermal inertia and lack of tropospheric aerosols - but is still far closer to the observations than we would have achieved by assuming no change. Furthermore, the observed warming is also very close to the top end of the distribution of historical 20 year trends, meaning that the observed outcome would be very unlikely if the the climate was merely following some sort of random walk. This evidence for the power of climate models is obviously limited by lack of detailed outputs for validation, but what there is is clearly very strongly supportive.Friday, January 29, 2010
Adapt or die
Given the lengthy introduction citing numerous hyperbolic warnings about the primarily negative and potentially disastrous health effects anticipated due to climate change, it is mildly amusing (though hardly surprising to anyone familiar with the literature) that they observe that temperature-related mortality has decreased sharply over the interval studied. More interestingly, they find that a major cause of this decrease is not even the warmer winters the UK has seen, but rather better adaptation to cold weather. This of course directly refutes the "optimally adapted" meme, since if that was the case, we would expect to see a decrease rather than the observed increase in tolerance to cold extremes as the winters got warmer. But in reality, wealth and technology (eg insulation, heating and clothing) act to increase our tolerance whether or not the climate changes, although the latter may in theory act as an additional stimulus if it were sufficiently important (which it clearly is not, in the UK).
Of course they have to finish off by talking about increases in heatwaves, wondering "whether adaptation will manage to keep pace with such changes". I think it is patently obvious that it will, and would happily bet against any predictions of increasing temperature-related deaths in coming decades.
Saturday, September 08, 2007
Pile on!
I didn't want to distract myself before by fisking it in detail, but it seems largely misguided. His opening sentence "The latest report from the Intergovernmental Panel on Climate Change assesses the skill of climate models by their ability to reproduce warming over the twentieth century..." is obviously wrong and the article does not improve much from there. There is a whole chapter (number 8) explicitly dedicated to assessing model performance, there is another chapter (6) on paleoclimate, and large parts of chapters 9 and 10 are also concerned with how we gain confidence in projections/probabilistic predictions. While 20th century climate change detection and attribution does get (IMO) a slightly unhealthy prominence in the report as a whole, it is hardly the whole story - indeed it has long been known that merely matching the general temperature trend is a rather weak test of model performance at least for the longer term (one detail that is worth noting is that all plausible models indicate a roughly linear trend over say 1980-2030, which does contribute to our confidence in a continuation of the recent warming trend).
A response from several IPCC authors has just appeared here, together with a brief rebuttal from Schwartz. The IPCC authors point out a number of reasons why Schwartz is wrong, and in his final rebuttal all he does is repeat his original erroneous claim that "in assessing the skill of climate models by their ability to reproduce warming over the twentieth century, the latest report from the IPCC may give a false sense of their predictive capability". It's just not true.
Coming soon: more on Schwartz and that sensitivity estimate.
Wednesday, May 02, 2007
Why David Evans is wrong (along with all the other sceptics)
Usually, I don't bother addressing the sceptic stuff that can be found on a thousand blogs: people who want the facts can come looking for them, and I've got limited patience for wrestling with pigs (you both get mucky, but the pig enjoys it). However, David is a bit of a special case as he's actually been prepared to put a significant amount of his money where his mouth is. It turns out that his post is a fairly standard laundry list of excuses as to why he thinks climate science is a bit ropey. I could have some sympathy with some of his points, but there is one gaping hole in his analysis.
He starts off by acknowledging:
1. Carbon dioxide is a greenhouse gas. Proved in a laboratory a century ago.and then continues with a string of tenuous arguments as to how it is possible that other factors could be behind most of the recent warming of the Earth.
At no point does he actually produce any evidence for his implied belief that CO2 has a negligible effect.
This, to me, is the crux of the matter. With no feedbacks at all, the sensitivity to doubled CO2 is about 1C, based on the well-understood radiative physics. It is also obvious that a warmer atmosphere has the potential to hold more water vapour (itself a GHG of course), and although the magnitude of this effect isn't known with certainty, the most plausible first-order estimate (supported by models and data) would be that relative humidity will stay roughly constant. This gives another 1C, making 2C in total. [The numbers here are intended as ballpark estimates, please let's have no quibbles about the precision.] Clouds may have a significant effect to enhance or offset warming. We know the climate has varied plenty in the past (indeed this is generally one of the septics' favourite talking points), so it seems implausible that they are a very strong stabilising force. Almost all models, using a wide range of physical parameterisations, suggest a significant positive amplification, giving the typical range of 2-4.5C for sensitivity. All analyses of observational evidence also point towards a value of close to 3C (exactly how close is still subject to some debate).
So we are left with the question: why on Earth would anyone believe that CO2 has almost no effect?
The claim that a world without anthropogenic forcing could possibly have warmed this much, whether true or not, is almost entirely irrelevant to the question of what we expect the anthropogenic effect to be. It's true that in the event of stronger natural variability, that might suggest a slightly larger possibility that a future downturn in the natural component could exceed the anthropogenically-forced response in the short term. But as I mentioned above, more variability also implies smaller stabilising feedbacks, so we'd also expect to see a larger sensitivity to CO2 in this case. Are we really supposed to believe that the planet is highly sensitive to some speculative and unquantified mechanism such as cosmic rays, and simultaneously insensitive to an effect that's been reasonably well understood for over 100 years? Why?
Detection and attribution has a lot to answer for in respect of this confusion. D&A essentially addresses the question "could an unforced planet have warmed as much as the observations"? However, this (frequentist) question has only tangential relevance to a (Bayesian) estimate of future warming. IMO, the arguments over whether or not we have "detected" AGW, and at what level of confidence, entirely misses the point. Imagine that someone points a gun in roughly your direction, and pulls the trigger. According to D&A, nothing interesting is going on until the bullet hits you, but at that point it's too late. An intelligent Bayesian would believe that there was a significant probability of serious harm before the bullet arrived - hopefully even before the trigger was pulled. Now, I'm not saying that climate change is going to suddenly kill us all, but just giving an analogy to explain the manner in which D&A fundamentally answers the wrong question. (As a secondary point, I believe it is the attempt to pretend that D&A methods can answer the interesting and useful questions that has lead to the uniform prior nonsense, but that isn't really my point here.)
So, until David and the rest can come up with some plausible arguments as to why CO2 actually has no effect, backed up by a sensible climate model which supports this claim, I'll continue to believe that it does in fact have a significant effect which will (with high probability) lead to continued warming. That is, as he more or less admitted at the start, solid science that is more than 100 years old.
Sunday, February 04, 2007
Some IPCC SPM comments
1. First, they predict continued warming for the next 20 years of "about 0.2C per decade", up from a "likely" range of 0.1-0.2C in the TAR. That change is not really surprising - it had been clear for some time that the actual warming rate was closer to the upper than the lower end of the TAR range. I don't know the history of the TAR but I guess that their selection of endpoints for their range owed as much to rounding as a deliberate selection of 0.15C as a central estimate. Even back then the recent trend was above 0.15C and forecast to increase over time, especially under the implicit assumption of no volcanic eruptions (a big one could knock as much as 0.1C off a decadal average temperature). This new estimate is still a long way short of the probabilistic predictions that have been published though (at least, 2 such papers that I recently re-read).
Apparently there is a new Science paper which talks of a recent trend of >0.2C per decade. Every time I've looked at GISTEMP (eg here) it shows just a bit under 0.2C to me, so I'll have to check exactly what this new paper did. Anyway, "about 0.2C per decade" is fine by me.
2. On climate sensitivity, there is the much-leaked change from 1.5-4.5C (TAR) to 2-4.5C (AR4) at the same "likely" level. I think the change at the lower end may be as much due to increasing recognition that 1.5 is a firm limit, as stronger confidence that the value of S is actually greater than 2C. Forster and Gregory's recent estimate was 1.7C and there are several others with a strong likelihood close to the lower end of the range. Anyway, depending on how "likely" is interpreted, this phrase still acknowledges perhaps as much as 15% probability of S sneaking below the 2C threshold, but it cannot do so by much.
What they have said about the upper end of the range is more...interesting. They have added the phrase:
"Values substantially higher than 4.5°C cannot be excluded, ...".A literal interpretation of this is completely vacuous (we can never assign a probability of precisely zero), so I'm not at all sure what they mean by including it. Note how it carefully avoids using the calibrated probabilistic language that has been adopted (likely, very likely etc). I can't help but be amused by RC's comment:
"the governments (for whom the report is being written) are perfectly entitled to insist that the language be modified so that the conclusions are correctly understood by them and the scientists. [...] The advantage of this process is that everyone involved is absolutely clear what is meant by each sentence. Recall after the National Academies report on surface temperature reconstructions there was much discussion about the definition of 'plausible'. That kind of thing shouldn't happen with AR4."I predict discussion about the definition of "cannot be excluded". I will be discussing it, at least! I complained about this ambiguous phrasing (which appeared in similar form in a couple of chapters) in my review of the last draft, and explicitly asked the authors to explain more clearly what they meant by it. I've also asked a couple of authors who used similar phrasing in their papers but have not got a reply out of them. I find it hard to avoid the conclusion that this "cannot be excluded" phrase was deliberately chosen specifically for its meaninglessness, in order to to be able to present a "consensus" rather than a strong disagreement about the credibility of such high values. I'm sure that those who assign a probability of 5% or even more to S greater than 6C will consider that this phrase supports them, even though Stoat parses it as "they do go on to diss > 4.5 oC a bit" due presumably to the "but agreement of models with observations is not as good for those values" which completes the sentence I partially quoted above. "Not as good" also has no probabilistic interpretation of course.
[Update
Based on this first-hand report, the phrase was indeed chosen specifically for its ambiguity.]
3. There is more probabilistic confusion in the discussion of attribution of past climate changes (I've written about this before). This is perhaps most clearly demonstrated in
"It is very unlikely that climate changes of at least the seven centuries prior to 1950 were due to variability generated within the climate system alone."One thing they might have said (and perhaps thought they were) is that an unforced system is very unlikely to exhibit the observed level of climate changes. That is an essentially frequentist statement about ensembles of model runs. But what they have actually said appears to be the Bayesian statement that they believe that there was external forcing in the real world. No shit Sherlock! The cavalier way in which the detection and attribution community freely switches between frequentist and Bayesian approaches to probability, without any clear explanation, gives every impression that they do not understand the difference (or even perhaps realise that there might be a difference). Their writing about the recent warming is similarly clumsy:
"it is extremely unlikely that global climate change of the past fifty years can be explained without external forcing, and very likely that it is not due to known natural causes alone."The first statement again is essentially frequentist, the second Bayesian. The existence of anthropogenic forcing as a contribution to the recent climate changes is not merely "very likely", it is at least "virtually certain"! (I don't believe there is a single working climate scientist who would argue that the anthropogenic forcing has been precisely zero. There is of course debate over its magnitude - some legitimate, some specious.)
I should point out that this criticism doesn't invalidate (or even weaken) the broad thrust of the report. I'm grumbling about the D&A stuff primarily because it forms the basis of the confusion in the climate sensitivity debate, rather than actually mattering in itself. They could have written things in a clear and correct manner without substantively affecting the overall message.
Wednesday, July 19, 2006
More on detection, attribution and estimation 4: The Literature
I'll start with a mild disclaimer: the purpose of this comment is not to have a go at (or embarass) people who have confused confidence intervals with credible intervals. Indeed I've only recently started to think more clearly about what is going on here myself - and I certainly wouldn't promise that my current understanding is complete and correct. Moreover, given that almost everyone has been getting it wrong for years, it would not be reasonable to single out a handful of individuals for blame. Really, the purpose of this comment is just to point out how completely ubiquitous this confusion is in the literature.
First the results of a bit of web-surfing. It's actually really hard to find a good definition of a confidence interval on the web. Here's a typical faulty version which I found linked from Wikipedia:
The confidence interval defines a band around the sample mean within which the true population [mean?] will lie, to some degree of confidence:As I've already demonstrated, this is false in general.
For example, there is a 95% probability that the true population mean will lie within the 95% confidence interval of the sample mean.
The contents of the Wikipedia page itself was misleading until recently - I have had a go at fixing it, rather clumsily. More editing is welcome...
The very first google hit for confidence interval is an interesting case. It is a Lancaster Uni mirror of some widely distributed educational material:, which contains the commendably careful definition:
If independent samples are taken repeatedly from the same population, and a confidence interval calculated for each sample, then a certain percentage (confidence level) of the intervals will include the unknown population parameter.which makes it clear that the probability is based on the frequentist concept of repeated sampling to generate a population of confidence intervals. So far so good. However, their definition actually started off with the at best ambiguous:
A confidence interval gives an estimated range of values which is likely to include an unknown population parameter, the estimated range being calculated from a given set of sample data.and concludes with the rather unfortunate:
Confidence intervals ... provide a range of plausible values for the unknown parameter.(That's not to say that confidence intervals never provide such a range - but it is not necessarily what they are designed to do, and they well may fail to achieve this.) Despite the careful description of repeated sampling in the middle of their text, it seems quite possible that some readers will end up with a rather misleading impression of what confidence intervals are.
The first edition of the highly-regarded book by Wilks ("Statistical methods in the atmospheric sciences") also equated confidence and credible intervals (here), and said things like "H0 [the null hypothesis] is rejected as too unlikely to have been true" - despite the frequentist paradigm explicitly forbidding the attachment of probabilities to hypotheses. I was pleased to find that Prof Wilks quickly agreed with me that this was misleading, and stated that in fact the recently-published second edition (which I have not seen) does not contain the comment about credible intervals.
Perhaps most surprisingly, the mistake is committed by people even when they are railing against the limitations of standard frequentist hypothesis testing! Eg, in this comment, Nicholls recommends reporting confidence intervals:
The reporting of confidence intervals would allow readers to address the question 'Given these data and the correlation calculated with them, what is the probability that H0 is true?'Oops.
So, how well does climate science come out of this? Well, the confusion seems to pop up just about everywhere that you see these words "likely" and "very likely" (eg throughout the IPCC TAR). The one exception would seem to be the climate sensitivity work, the bulk of which is explicitly Bayesian right from the start, with clearly stated priors. One could probably argue that many of the TAR judgements were based on experts carefully weighing up the evidence, but in many cases (especially the D&A chapter), it seems entirely routine to directly interpret a confidence interval (say, from a regression analysis) as a credible interval. I've seen an absolutely explicit occurrence of this in a recent D&A paper, co-authored by 2 prominent figures in the field.
I promised TCO I would say something about the Hockey Stick. Of course, it's the same story here. On the basis of regression coefficients, MBH (specifically in their 1999 GRL paper, which is repeated in the TAR) make a statement about how "likely" it is that the 1990 were warmer than the previous millenium. I can't help but be amused to note that with all the "auditing" and peer-review, including by professional statisticians, none of them seems to have noticed this little detail.
Of course an important question to consider is, how much does this all matter? And the answer is...it depends. In many cases, the answers you get probably won't turn out too different. There was some explicitly Bayesian estimation briefly mentioned in the D&A chapter of the TAR, and broadly speaking it seemed to give results that were similar to (in fact perhaps even stronger than) the more mainstream stuff. Moreover, there is a lot of slack in terms like "likely", so changing the probability might not invalidate such statements anyway. So I am not by any means suggesting that the TAR needs to be thrown out, and therefore maybe some people will claim this is all a fuss about nothing. However, I would argue that it is still surely a good thing to at least understand what is going on, so that people can consider how important an issue it is in each particular case. For instance, equating confidence intervals and credible intervals seems to assume (inter alia) the choice of a uniform prior: this decision can by no means be an automatic one, and equating it with "no prior" or "initial ignorance" is definitely wrong. Eg, no-one, not even the most rabid septic, has actually ever believed that CO2 is as likely to cool as warm the planet (at least not since Arrhenius), so assigning an equal prior probability to positive and negative effects would surely be hard to defend. Like the example of the "negative mass" apple, a confidence interval that covers a wide range does not mean that we think the parameter has a significant probability of taking an extreme value! At an absolute minimum, it certainly needs to be stated clearly that this choice of prior was made.
There are a number of other technical issues such as model error which are intimately related, (eg, what does it mean to determine a parameter or regression coefficient to n decimal places, in a model that is an incomplete representation of the real world?). We can try a bit of hand-waving and claim it doesn't matter too much, but ultimately I think if we are going to try to make useful and credible estimates then there is little alternative but to try to deal with these things more coherently and consistently, even if it does mean more work for the statisticians!
Sunday, July 16, 2006
More on detection, attribution and estimation 3: The Prosecutor's Fallacy
Formally, the exact calculation of P(H|D) when we are given P(D|H) requires Bayes' Theorem:
P(H|D)=P(D|H)P(H)/P(D)
Although the prosecutor's fallacy is generally demonstrated through discrete probability, Bayes' Theoreom applies equally to continuous probability distribution functions, with f(h|d) being related to f(d|h) via f(h|d)=f(d|h)f(h)/f(d). This explains the distinction between confidence intervals and credible intervals demonstrated in the last post, since an experimental observation gives us f(d|h) (a likelihood function), and in order to turn it into a posterior pdf f(h|d) we need to use a prior f(h).
For example, given the previous apple-weighing example, we might have a prior belief that the apple will weight about 100g, plus or minus 20g at 1 standard deviation (and strictly speaking, the prior should be truncated at 0). The likelihod function arising from the measurement is itself a Gaussian shape centred on the observed 40g, with a width of 50g - this function does extend to negative values, as a hypothetical negative-mass apple would have a nonzero probability of a returning a 40g measurement. Applying Bayes' Theorem formally gives us the well-known result of optimal interpolation between two gaussians, which in this case works out to 91.7+-18.6g. In this case, the observation is so poor that it hardly affects our prior belief, but if our scales had an error of only 5g we'd obviously depend far more on their output. In no case would we end up believing that the apple's mass was negative!
Next, and perhaps last (for now at least): what the literature says.
Saturday, July 15, 2006
More on detection, attribution and estimation 2: Incredible confidence intervals
For the first, rather natural example, let's assume we are trying to measure some simple non-negative quantity such as the mass of an apple. We have a set of scales which have a random (but well-characterised) error of +-50g (Gaussian at 1 standard deviation). That is, if we take a calibrated mass of value X, repeatedly use the scales and plot a histogram of the results of each measurements, the outputs will form a nice gaussian shape with mean X and standard deviation 50g. [OK, I know I'm doing this at a very boring pace, but I need to make sure it is all clearly set out.]
One obvious and very natural way to create a confidence interval for the apple's mass is to take a single measurement (call the observed mass m) and then write down (m-50,m+50), which is a symmetric 68% confidence interval for M, the true mass. That is to say, if we were to hypothetically repeat this experiment numerous times, and construct the set of confidence intervals (mi-50,mi+50) where i indexes the measurements generated in the experiments, then 68% of these intervals would include M, 16% would be wholly greater than M and 16% wholly smaller (this is guaranteed by the specified observational uncertainty). We don't, of course, actually do this infinite experiment - but this is precisely what is meant by "(m-50,m+50) is a symmetrical 68% confidence interval for M". [Are you asleep yet? The punchline is coming up...]
It is incorrect to interpret the specific confidence interval (m-50,m+50) as implying that M lies in that range with probability 68% (and above and below with probability 16%).
To see why this is the case, consider the following: what if the reading happens to be m=40g? (Which it might well be, if the true mass is say 80g.) Is the confidence interval (-10,90) really a symmetric credible interval at the 68% level? That is to say, would anyone believe that the apple's mass is <-10g with probability 16% (would they bet on it - if so, please point them this way...)? Of course not. Any symmetric 68% credible interval (ie an interval so that one believes M lies in it with probability 68%, and below and above with probability 16% each) must necessarily be truncated somewhere above zero. Yet it is trivial to show that the confidence interval as constructed above is entirely valid. Under repeated observations, 16% of the measurements will be lower than M-50, 16% greater than M+50, and the remainder in between, so the population of confidence intervals has exactly the statistical properties required of it.
One can, perhaps, state that "negative mass is not statistically inconsistent with the measurement" or maybe even say "negative mass cannot be ruled out by the measurement", but these statements cannot be interpreted as implying that anyone thinks the apple actually has negative mass!
There are other examples of non-credible confidence intervals that are quite striking. Here's one I found on the web (description lightly modified):
Let's say we want to estimate a parameter x. Let's ignore all the available measurements entirely! In their place, start by using a random number generator to generate y uniformly in [0,1]. If y > 0.68, then define the confidence interval to be the empty interval. If y < 0.68, then define the confidence interval to be the whole number line. That's it! Again, this routine trivially generates a 68% CI - that is, exactly 68% of the time, the CI contains x whatever value this takes. But neither of the two possible intervals that the algorithm generates is credible at the 68% level - it should be clear that one of the possible intervals contains x (and the other does not) with certainty, even without knowing what x is.
In the next part, I'll try to reconcile these results with the underlying theory.
More on detection, attribution and estimation 1: Confidence intervals and credible intervals
To recap, D&A is an essentially frequentist procedure, which seeks to determine whether the observational record is "statistically inconsistent" with what we could have expected in the absence of anthropogenic influence, and "not statistically inconsistent" with what we think the anthropogenic influence should have been.
It is fundamentally a frequentist approach (comparing the real data to the population of model outputs which differ due to natural variability). Therefore, it has no direct interpretation as a probabilistic estimate of the magnitude or existence of the anthropogenic influence. Such estimates are absolutely incompatible with the frequentist paradigm. This is the fundamental message of these posts, and I can't stress it too fully.
Unfortunately, this distinction has been badly blurred in some of the literature, including even the IPCC TAR itself. In fact, it seems like there is rather widespread confusion between a (frequentist) Confidence Interval, and a (Bayesian) Credible Interval, so I will expand more on this now. A confidence interval (warning: page may be a bit dodgy, precisely due to the confusion I'm discussing) for a parameter x is an interval constructed according to a specific method, such that if we were to repeat the experiment numerous times, with a new set of observational data (with different random errors) for each experiment, then p% of the confidence intervals we construct using this method would contain the true (fixed) value of x, whatever that is.
This is not the same thing as an interval such that we believe x lies in the interval with probability p! For a frequentist, to make a probabilistic statement about x is to commit a category error: x is a fixed but unknown parameter, it is either in the specific interval or not.
A credible interval (which can also be abbreviated to CI: how confusing) is an inherently Bayesian concept: it is an interval such that the parameter is believed to lie in the interval with probability p. Fundamentally, the belief (probability) attaches to the person who makes the statement, rather than the parameter itself - in other words, it is subjective. In cases where many people agree on a particular credible interval (ie because they share similar judgements about priors, methods and data), it is sometimes called intersubjective. Calling a credible interval objective is potentially rather misleading IMO - it may be taken as implying that there is a true probability that we may be able to discover through careful analysis (analogous to the probability of 4 heads in 5 coin tosses, say). However, the only "true" probability that could apply in this sense is 0 or 1 - the event is either going to happen, or not (NB "going to" can apply to past events which are currently unknown to me, such as the probability that it rained in Birmingham on this day last year). The probability here applies to our belief in the truth of an unknown proposition, not any frequentist limit as the number of replications increases. Different people may quite reasonably differ in their opinion, without any one of them being wrong!
It seems quite clear from a bit of web-surfing (and a literature trawl) that the vast majority of people automatically and intuitively interpret confidence intervals as credible intervals: that is not so surprising, as the precise definition of the confidence interval is rather complex, counterintuitive, and not very useful in real life (in contrast, what people actually want to know is well-encapsulated by the notion of a credible interval). Moreover, even authoritative literature specifically on the subject of statistics frequently gets it wrong (as I'll show later). However, the two concepts are not the same, and it is trivial to construct confidence intervals that are in no way credible. I'll give some examples of this in part 2.


