Showing posts sorted by relevance for query "detection, attribution". Sort by date Show all posts
Showing posts sorted by relevance for query "detection, attribution". Sort by date Show all posts

Sunday, February 12, 2006

Detection, Attribution and Estimation

This is something I've been meaning to blog about for some time. It comes up a lot in the context of the hurricane wars, over at RPJnr's blog. A recent comment of his provides a nice opening:
As you well know much of science works through hypothesis - falsification. Across climate science the null hypothesis used to guide research has been that a human signal is NOT present, and research is then done to falsify this hypothesis. In some cases such research has been done allowing attribution of climate effects to an anthropogenic forcing. This null hypothesis is chosen because of Occam's razor, it is a simpler explanation for what is observed.
He is essentially posing the question as initially one of detection - can we show that the AGW has had an effect, and that the observations are not just the result of climate variability? - before moving on to attribution - how much of a change can we describe as being due to this particular cause?

There is, however, an entirely different but equally valid approach that could also be used from the outset, which is: what is our estimate of the magnitude of the effect? The critical distinction is that the null hypothesis has no particularly priviledged position in this approach.

This distinction between detection and estimation is related to that between a frequentist and Bayesian approach to probability. Ruling out a null hypothesis (or not) at some level of significance is essentially a frequentist approach: estimating the probabilities of various competing hypotheses, about which we have prior beliefs (but not necessarily a strong bias towards a null hypothesis) is Bayesian. The answers that these two approaches provide may be very different in any given situation, and neither is necessarily right or wrong a priori, but it is surely self-evident that the Bayesian approach is more relevant to decision-making. If we have any reasonable expectation that certain policies would have particular bad effects, it would be ridiculous to wait until such effects could be shown to have occurred at some arbitrary level of statistical significance (that's not a point specific to climate change, of course).

IMO, the IPCC slightly muddies the waters with its discussion of D&A. It defines attribution, as Roger implicitly does above, as a strictly stronger statement than detection - one can only attribute once the effect has already been detected. I suppose they can define the term "attribution" in that way if they choose, but it is certainly not then valid to equate this with the general estimation problem as they do further down. It is trivial to create situations in which a currently undetectable effect can be reasonably estimated to be large, and the converse is equally possible - an easily detectable (statistically significant) influence may be wholly irrelevant in practical terms. I suspect that this forms a large part of the difference in presentation between various parties in the hurricane debate - the evidence may not yet rule out the null hypothesis of no effect, but some people estimate that AGW is likely to have a substantial effect (even if the ill-defined error bars on their estimate do not exclude zero). In principle, exactly the same evidence could support both of these conclusions, although I don't personally know enough about hurricanes to make a definitive statement in that particular case.

It is amusing to see Roger, very much at the sharp end of policy-relevant work, promoting the scientifically "pure" but practically less useful detection/frequentist approach rather than the more appropriate estimation/Bayesian angle. It's not surprising, although perhaps a little disappointing, that the IPCC explicitly endorses that view. But by placing the null hypothesis in a priviledged position from which it can only be dislodged by a mountain of observational evidence, this approach provides a strong inbuilt bias for the status quo which cannot be justified on any rational decision-theoretic grounds.

Thursday, September 07, 2017

More on Bayesian approaches to detection and attribution

Timely given events all over the place, this new paper by Mann et al has just appeared.  It's a well-aimed jab at the detection and attribution industry which could perhaps be held substantially responsible for the sterile “debate” over the extent to which AGW has influenced extreme events (and/or will do so in the future). I've argued against D&A several times in the past (such as here, here, here and here) and don't intend to rehash the same arguments over and over again. Suffice to say that it doesn't usefully address the questions that matter, and cannot do so by design.

Mann et al argue that the standard frequentist approach to D&A is inappropriate both from a simple example which shows it to generate poor results, and from the ethical argument that “do no harm” is a better starting point than “assume harmless”. The precautionary versus proactionary principles can be argued indefinitely, and neither really works when reduced ad absurdum, so I'm not really convinced that the latter is a strong argument. A clearer demonstration could perhaps have been provided by a rational cost-benefit analysis in which costs of action versus inaction (and the payoffs) could have been explicitly calculated. This would have still supported their argument of course, as the frequentist approach is not a rational basis for decisions. I suppose that's where I tend to part company with the philosophers (check the co-author list) in preferring a more quantitative approach. I'm not saying they are wrong, it's perhaps a matter of taste.

[I find to my surprise I have not written about the precautionary vs proactionary principle before]

Other points that could have been made (and had I been a reviewer, I'd probably have encouraged the authors to include them) are that when data are limited and the statistical power of the analysis is weak, it is not only inevitable that any frequentist-based estimate that achieves statistical significance will be a large overestimate of the true magnitude of the effect, but there's even a substantial chance it will have the wrong sign! A Bayesian prior solves (or at least greatly ameliorates) these problems. Another benefit of the Bayesian approach is the ability to integrate different sources of information. My favourite example of the weakness of traditional D&A here is the way that we can (at least this was the case a few years ago) barely “attribute” any warming of the world's oceans under this methodology. The reason for this is that the internal variability of the oceans is large (and uncertain) enough that we cannot be entirely confident that an unforced ocean would not have warmed up by itself. On the other hand, it is absurd to believe the null hypothesis that we haven't warmed it, as it has been in contact with the atmosphere that we have certainly warmed, and the energy imbalance due to GHGs is significant, and we've even observed a warming very closely in line with what our models predict should have happened. But D&A can't assimilate this information. In the context of Mann et al, we might consider information about warming sea surface temperatures as relevant to diagnosing and predicting hurricanes, for example, rather than relying entirely on storm counts.

Friday, August 06, 2010

Wiley Interdisciplinary Reviews: Climate Change

A new journal has sprung up recently, I'm not entirely sure why or how, but it seems to be open access for now (not indefinitely) and has some interesting papers so maybe some of you would like to take a look. Called "Wiley Interdisciplinary Reviews: Climate Change" it seems to be a cross between an interdisciplinary journal and collection of encyclopaedic articles on climate change. There are a number of other WIREs journals on unrelated topics such as computational statistics, and nanomedicine and nanobiotechnology.

The Editors seem a bunch of slightly unconventional people, a little removed from the mainstream IPCC stalwarts though eminent enough and with some IPCC links: Hulme, Pielke, von Storch, Nicholls, Yohe are names that many will be familiar with. The others are probably all famous too, but I'm too ignorant to recognise them. I'm sure the journal is not intended as a direct rival to the IPCC, but it may turn out to provide an interesting and slightly alternative perspective.

The articles to date include a mix of authoritative reviews from leading experts - such as Parker on the urban heat island, Stott on detection and attribution, interspersed with perhaps more personal and less authoritative articles. I can safely say that without risk of criticism because one of them is mine - a review on Bayesian approaches to detection and attribution. This article had a rather difficult genesis. I was initially dubious about my suitability for the task and indeed the value of the article, but after declining once (and proposing another author, who also declined) I changed my mind and had a go. My basic difficulty with addressing the concept is that D&A has always seemed to me to be a rather limited and awkward approach to the question of estimating the effects of anthropogenic and natural forcing, which is tortured into a frequentist framework where it doesn't really fit. Eg, no sane person believes these forcings have zero effect, so what exactly is the purpose of a null hypothesis significance test in the first place? However, conventional D&A has such a stranglehold on the scientific conscious that most Bayesian approaches have actually mimicked this frequentist alternative of the Bayesian estimate that you really wanted in the first place. It all seems a bit tortured and long-winded to me.

Anyway, I eventually found some things to say, which hopefully aren't entirely stupid and help to show how a Bayesian approach might actually be useful in answering the questions that (sensible) people might want to know the answer to, rather than the relatively useless questions that frequentist methods can answer, which are then inevitably misinterpreted as answers to the questions that people wanted to answer in the first place (as I argue and document in the article).

Another of the personal and argumentative articles was contributed by Jules, who was invited to say something about skill and uncertainty in climate models. This was actually the article that sparked off our "Reliability" paper, as our discussions kept coming back to the odd inconsistency between the flat rank histogram evaluation that I know is standard in most ensemble prediction, versus the Pascal's triangle distribution that a truth-centred ensemble would generate (ie, if each model is independently and equiprobably greater than or less than the observations, then the obs should generally be very close to the ensemble median). Of course this problem didn't take long to solve once we had set out the issue clearly enough to recognise that there really were two incompatible paradigms in play, and Jules even ended up citing the GRL paper which overtook her WIREs one in the review process.

Perhaps of more widespread interest to other readers, is a simple analysis of the skill of Hansen's forecast which he made back in 1988 to the US Congress. We'd actually had lengthy discussions with several people (listed in the acknowledgements) a year or two ago, trying to resurrect the old model code that was used for this prediction in order to re-run it and analyse its outputs in more detail. But this proved to be impossible. (The code exists but has been updated and gives substantially different results. If only the code had been published in GMD!) Therefore we were left with nothing more than the single printed plot of global mean temperature to look at. This didn't seem much to base a proper peer-reviewed paper on, so the idea died a death. When this WIREs invitation came long it seemed like a good opportunity to publish the one usable result we had obtained, as an example of what skill means. The headline result is that under any reasonable definition of skill, the Hansen prediction was skillful. While no great surprise, I don't think it has been presented in quite those terms before. It's a shame that we weren't able to generate a more comprehensive set of outputs though which might have given a more robust result than this single statistic.


The null hypothesis of persistence (no change in temperature) was found to give best performance over the historical interval, compared to extrapolating a trend. So this is the appropriate zero-skill baseline for evaluating the forecast. Nowadays with the AGW trend well established, probably most would argue that a continuation of that trend is a good bet, though that still leaves open the question of how long a historical interval to fit the trend over. Anyway, the model forecast is clearly drifting on the high side by now - most likely due to some combination of high sensitivity, low thermal inertia and lack of tropospheric aerosols - but is still far closer to the observations than we would have achieved by assuming no change. Furthermore, the observed warming is also very close to the top end of the distribution of historical 20 year trends, meaning that the observed outcome would be very unlikely if the the climate was merely following some sort of random walk. This evidence for the power of climate models is obviously limited by lack of detailed outputs for validation, but what there is is clearly very strongly supportive.

Sunday, July 16, 2006

More on detection, attribution and estimation 3: The Prosecutor's Fallacy

The error of equating P(Data|Hypothesis) and P(Hypothesis|Data) is known as the Prosecutor's Fallacy, due to its frequent appearance in criminal trials (misinterpretation of DNA evidence etc) (the dispute on that wikipedia page seems to refer to the details of its applicability to a particular legal case, not the underlying theory). Typically, it is illustrated via a simple discrete yes/no question along the following lines: if the probability of a random person matching a DNA sample from a crime scene is 1 in 1,000,000, then what is the probability that a suspect is guilty, given only that their DNA matches? The fallacial answer is 999,999 in 1,000,000. An easy way to see the flaw in this is to note that in the UK, there are 60,000,000 people so there will be 60 people whose DNA matches, only 1 of whom will be the guilty one (note various other assumptions I've made, including the fact that a crime actually took place at all, it was committed by one person, and that there is no other evidence as to the suspect).

Formally, the exact calculation of P(H|D) when we are given P(D|H) requires Bayes' Theorem:
P(H|D)=P(D|H)P(H)/P(D)
which requires the specification of a "prior" P(H) (P(D) is a normalisation constant which provides no real theoretical difficulties, although it might be hard to calculate in practice). It must be understood that this equation does not depend on "being a Bayesian" or "being a frequentist". It is simply a law of probability, which follows directly from the axioms (in particular, P(D,H)=P(D|H)P(H)=P(H|D)P(D)). So it's not something we can choose to obey or not - at least, without abandoning any pretence that we are talking about probability as it is usually understood.

Although the prosecutor's fallacy is generally demonstrated through discrete probability, Bayes' Theoreom applies equally to continuous probability distribution functions, with f(h|d) being related to f(d|h) via f(h|d)=f(d|h)f(h)/f(d). This explains the distinction between confidence intervals and credible intervals demonstrated in the last post, since an experimental observation gives us f(d|h) (a likelihood function), and in order to turn it into a posterior pdf f(h|d) we need to use a prior f(h).

For example, given the previous apple-weighing example, we might have a prior belief that the apple will weight about 100g, plus or minus 20g at 1 standard deviation (and strictly speaking, the prior should be truncated at 0). The likelihod function arising from the measurement is itself a Gaussian shape centred on the observed 40g, with a width of 50g - this function does extend to negative values, as a hypothetical negative-mass apple would have a nonzero probability of a returning a 40g measurement. Applying Bayes' Theorem formally gives us the well-known result of optimal interpolation between two gaussians, which in this case works out to 91.7+-18.6g. In this case, the observation is so poor that it hardly affects our prior belief, but if our scales had an error of only 5g we'd obviously depend far more on their output. In no case would we end up believing that the apple's mass was negative!

Next, and perhaps last (for now at least): what the literature says.

Saturday, July 15, 2006

More on detection, attribution and estimation 2: Incredible confidence intervals

Following on from this post, here are a couple of simple examples where perfectly valid confidence intervals are clearly not credible intervals at the same level of probability.

For the first, rather natural example, let's assume we are trying to measure some simple non-negative quantity such as the mass of an apple. We have a set of scales which have a random (but well-characterised) error of +-50g (Gaussian at 1 standard deviation). That is, if we take a calibrated mass of value X, repeatedly use the scales and plot a histogram of the results of each measurements, the outputs will form a nice gaussian shape with mean X and standard deviation 50g. [OK, I know I'm doing this at a very boring pace, but I need to make sure it is all clearly set out.]

One obvious and very natural way to create a confidence interval for the apple's mass is to take a single measurement (call the observed mass m) and then write down (m-50,m+50), which is a symmetric 68% confidence interval for M, the true mass. That is to say, if we were to hypothetically repeat this experiment numerous times, and construct the set of confidence intervals (mi-50,mi+50) where i indexes the measurements generated in the experiments, then 68% of these intervals would include M, 16% would be wholly greater than M and 16% wholly smaller (this is guaranteed by the specified observational uncertainty). We don't, of course, actually do this infinite experiment - but this is precisely what is meant by "(m-50,m+50) is a symmetrical 68% confidence interval for M". [Are you asleep yet? The punchline is coming up...]

It is incorrect to interpret the specific confidence interval (m-50,m+50) as implying that M lies in that range with probability 68% (and above and below with probability 16%).

To see why this is the case, consider the following: what if the reading happens to be m=40g? (Which it might well be, if the true mass is say 80g.) Is the confidence interval (-10,90) really a symmetric credible interval at the 68% level? That is to say, would anyone believe that the apple's mass is <-10g with probability 16% (would they bet on it - if so, please point them this way...)? Of course not. Any symmetric 68% credible interval (ie an interval so that one believes M lies in it with probability 68%, and below and above with probability 16% each) must necessarily be truncated somewhere above zero. Yet it is trivial to show that the confidence interval as constructed above is entirely valid. Under repeated observations, 16% of the measurements will be lower than M-50, 16% greater than M+50, and the remainder in between, so the population of confidence intervals has exactly the statistical properties required of it.

One can, perhaps, state that "negative mass is not statistically inconsistent with the measurement" or maybe even say "negative mass cannot be ruled out by the measurement", but these statements cannot be interpreted as implying that anyone thinks the apple actually has negative mass!

There are other examples of non-credible confidence intervals that are quite striking. Here's one I found on the web (description lightly modified):

Let's say we want to estimate a parameter x. Let's ignore all the available measurements entirely! In their place, start by using a random number generator to generate y uniformly in [0,1]. If y > 0.68, then define the confidence interval to be the empty interval. If y < 0.68, then define the confidence interval to be the whole number line. That's it! Again, this routine trivially generates a 68% CI - that is, exactly 68% of the time, the CI contains x whatever value this takes. But neither of the two possible intervals that the algorithm generates is credible at the 68% level - it should be clear that one of the possible intervals contains x (and the other does not) with certainty, even without knowing what x is.

In the next part, I'll try to reconcile these results with the underlying theory.

More on detection, attribution and estimation 1: Confidence intervals and credible intervals

I wrote this some time ago, and have been occasionally looking into the D&A stuff since then. It all seems rather a lot murkier than I had expected...this will take a lot of writing to explain so I'm splitting it into parts. This part is an introduction to the distinct concepts of confidence intervals and credible intervals.

To recap, D&A is an essentially frequentist procedure, which seeks to determine whether the observational record is "statistically inconsistent" with what we could have expected in the absence of anthropogenic influence, and "not statistically inconsistent" with what we think the anthropogenic influence should have been.

It is fundamentally a frequentist approach (comparing the real data to the population of model outputs which differ due to natural variability). Therefore, it has no direct interpretation as a probabilistic estimate of the magnitude or existence of the anthropogenic influence. Such estimates are absolutely incompatible with the frequentist paradigm. This is the fundamental message of these posts, and I can't stress it too fully.

Unfortunately, this distinction has been badly blurred in some of the literature, including even the IPCC TAR itself. In fact, it seems like there is rather widespread confusion between a (frequentist) Confidence Interval, and a (Bayesian) Credible Interval, so I will expand more on this now. A confidence interval (warning: page may be a bit dodgy, precisely due to the confusion I'm discussing) for a parameter x is an interval constructed according to a specific method, such that if we were to repeat the experiment numerous times, with a new set of observational data (with different random errors) for each experiment, then p% of the confidence intervals we construct using this method would contain the true (fixed) value of x, whatever that is.

This is not the same thing as an interval such that we believe x lies in the interval with probability p! For a frequentist, to make a probabilistic statement about x is to commit a category error: x is a fixed but unknown parameter, it is either in the specific interval or not.

A credible interval (which can also be abbreviated to CI: how confusing) is an inherently Bayesian concept: it is an interval such that the parameter is believed to lie in the interval with probability p. Fundamentally, the belief (probability) attaches to the person who makes the statement, rather than the parameter itself - in other words, it is subjective. In cases where many people agree on a particular credible interval (ie because they share similar judgements about priors, methods and data), it is sometimes called intersubjective. Calling a credible interval objective is potentially rather misleading IMO - it may be taken as implying that there is a true probability that we may be able to discover through careful analysis (analogous to the probability of 4 heads in 5 coin tosses, say). However, the only "true" probability that could apply in this sense is 0 or 1 - the event is either going to happen, or not (NB "going to" can apply to past events which are currently unknown to me, such as the probability that it rained in Birmingham on this day last year). The probability here applies to our belief in the truth of an unknown proposition, not any frequentist limit as the number of replications increases. Different people may quite reasonably differ in their opinion, without any one of them being wrong!

It seems quite clear from a bit of web-surfing (and a literature trawl) that the vast majority of people automatically and intuitively interpret confidence intervals as credible intervals: that is not so surprising, as the precise definition of the confidence interval is rather complex, counterintuitive, and not very useful in real life (in contrast, what people actually want to know is well-encapsulated by the notion of a credible interval). Moreover, even authoritative literature specifically on the subject of statistics frequently gets it wrong (as I'll show later). However, the two concepts are not the same, and it is trivial to construct confidence intervals that are in no way credible. I'll give some examples of this in part 2.

Sunday, May 29, 2005

Climate research workshop

This is why I was wedged on a train on Friday morning - to visit our close partners CCSR for a project workshop on climate research (our usual commute is a cycle ride to our own lab, which is a much more pleasant journey). The main reason for attending is that due to the number of foreign visitors, it is all in English today :-) We arrived in good time for coffee and a tour of the impressive building.

There are various presentations about model developments and how they improve (or not) the model behaviour, also detection/attribution/forecasting of climate change. All fairly typical fare in the climate modelling world, and several bits were very interesting - particularly looking at the global cooling after the Pinatubo eruption and how this can be used to estimate sensitivity to GHG forcing, which I'm hoping to delve into in more detail shortly. The Hadley Centre/CGAM visitors also reported on their progress and plans - mostly computing details so far, with science to follow.

And then the ice-breaker party and an upgrade to first class (Green Car) for the trip home so we don't have to stand all the way.

Wednesday, July 19, 2006

More on detection, attribution and estimation 4: The Literature

Time for a look at the literature.

I'll start with a mild disclaimer: the purpose of this comment is not to have a go at (or embarass) people who have confused confidence intervals with credible intervals. Indeed I've only recently started to think more clearly about what is going on here myself - and I certainly wouldn't promise that my current understanding is complete and correct. Moreover, given that almost everyone has been getting it wrong for years, it would not be reasonable to single out a handful of individuals for blame. Really, the purpose of this comment is just to point out how completely ubiquitous this confusion is in the literature.

First the results of a bit of web-surfing. It's actually really hard to find a good definition of a confidence interval on the web. Here's a typical faulty version which I found linked from Wikipedia:
The confidence interval defines a band around the sample mean within which the true population [mean?] will lie, to some degree of confidence:

For example, there is a 95% probability that the true population mean will lie within the 95% confidence interval of the sample mean.
As I've already demonstrated, this is false in general.

The contents of the Wikipedia page itself was misleading until recently - I have had a go at fixing it, rather clumsily. More editing is welcome...

The very first google hit for confidence interval is an interesting case. It is a Lancaster Uni mirror of some widely distributed educational material:, which contains the commendably careful definition:
If independent samples are taken repeatedly from the same population, and a confidence interval calculated for each sample, then a certain percentage (confidence level) of the intervals will include the unknown population parameter.
which makes it clear that the probability is based on the frequentist concept of repeated sampling to generate a population of confidence intervals. So far so good. However, their definition actually started off with the at best ambiguous:
A confidence interval gives an estimated range of values which is likely to include an unknown population parameter, the estimated range being calculated from a given set of sample data.
and concludes with the rather unfortunate:
Confidence intervals ... provide a range of plausible values for the unknown parameter.
(That's not to say that confidence intervals never provide such a range - but it is not necessarily what they are designed to do, and they well may fail to achieve this.) Despite the careful description of repeated sampling in the middle of their text, it seems quite possible that some readers will end up with a rather misleading impression of what confidence intervals are.

The first edition of the highly-regarded book by Wilks ("Statistical methods in the atmospheric sciences") also equated confidence and credible intervals (here), and said things like "H0 [the null hypothesis] is rejected as too unlikely to have been true" - despite the frequentist paradigm explicitly forbidding the attachment of probabilities to hypotheses. I was pleased to find that Prof Wilks quickly agreed with me that this was misleading, and stated that in fact the recently-published second edition (which I have not seen) does not contain the comment about credible intervals.

Perhaps most surprisingly, the mistake is committed by people even when they are railing against the limitations of standard frequentist hypothesis testing! Eg, in this comment, Nicholls recommends reporting confidence intervals:
The reporting of confidence intervals would allow readers to address the question 'Given these data and the correlation calculated with them, what is the probability that H0 is true?'
Oops.

So, how well does climate science come out of this? Well, the confusion seems to pop up just about everywhere that you see these words "likely" and "very likely" (eg throughout the IPCC TAR). The one exception would seem to be the climate sensitivity work, the bulk of which is explicitly Bayesian right from the start, with clearly stated priors. One could probably argue that many of the TAR judgements were based on experts carefully weighing up the evidence, but in many cases (especially the D&A chapter), it seems entirely routine to directly interpret a confidence interval (say, from a regression analysis) as a credible interval. I've seen an absolutely explicit occurrence of this in a recent D&A paper, co-authored by 2 prominent figures in the field.

I promised TCO I would say something about the Hockey Stick. Of course, it's the same story here. On the basis of regression coefficients, MBH (specifically in their 1999 GRL paper, which is repeated in the TAR) make a statement about how "likely" it is that the 1990 were warmer than the previous millenium. I can't help but be amused to note that with all the "auditing" and peer-review, including by professional statisticians, none of them seems to have noticed this little detail.

Of course an important question to consider is, how much does this all matter? And the answer is...it depends. In many cases, the answers you get probably won't turn out too different. There was some explicitly Bayesian estimation briefly mentioned in the D&A chapter of the TAR, and broadly speaking it seemed to give results that were similar to (in fact perhaps even stronger than) the more mainstream stuff. Moreover, there is a lot of slack in terms like "likely", so changing the probability might not invalidate such statements anyway. So I am not by any means suggesting that the TAR needs to be thrown out, and therefore maybe some people will claim this is all a fuss about nothing. However, I would argue that it is still surely a good thing to at least understand what is going on, so that people can consider how important an issue it is in each particular case. For instance, equating confidence intervals and credible intervals seems to assume (inter alia) the choice of a uniform prior: this decision can by no means be an automatic one, and equating it with "no prior" or "initial ignorance" is definitely wrong. Eg, no-one, not even the most rabid septic, has actually ever believed that CO2 is as likely to cool as warm the planet (at least not since Arrhenius), so assigning an equal prior probability to positive and negative effects would surely be hard to defend. Like the example of the "negative mass" apple, a confidence interval that covers a wide range does not mean that we think the parameter has a significant probability of taking an extreme value! At an absolute minimum, it certainly needs to be stated clearly that this choice of prior was made.

There are a number of other technical issues such as model error which are intimately related, (eg, what does it mean to determine a parameter or regression coefficient to n decimal places, in a model that is an incomplete representation of the real world?). We can try a bit of hand-waving and claim it doesn't matter too much, but ultimately I think if we are going to try to make useful and credible estimates then there is little alternative but to try to deal with these things more coherently and consistently, even if it does mean more work for the statisticians!

Sunday, February 04, 2007

Some IPCC SPM comments

I'm not going to give a broad overview, as there is is already a plethora of rather boring bloggorhea (sorry guys, but it is :-) ) on the subject. Go and read RC if you want the "consensus" view. Or just read the document itself, it's simple enough. I'm just going to pick out a few bits that are particularly interesting to me.

1. First, they predict continued warming for the next 20 years of "about 0.2C per decade", up from a "likely" range of 0.1-0.2C in the TAR. That change is not really surprising - it had been clear for some time that the actual warming rate was closer to the upper than the lower end of the TAR range. I don't know the history of the TAR but I guess that their selection of endpoints for their range owed as much to rounding as a deliberate selection of 0.15C as a central estimate. Even back then the recent trend was above 0.15C and forecast to increase over time, especially under the implicit assumption of no volcanic eruptions (a big one could knock as much as 0.1C off a decadal average temperature). This new estimate is still a long way short of the probabilistic predictions that have been published though (at least, 2 such papers that I recently re-read).

Apparently there is a new Science paper which talks of a recent trend of >0.2C per decade. Every time I've looked at GISTEMP (eg here) it shows just a bit under 0.2C to me, so I'll have to check exactly what this new paper did. Anyway, "about 0.2C per decade" is fine by me.

2. On climate sensitivity, there is the much-leaked change from 1.5-4.5C (TAR) to 2-4.5C (AR4) at the same "likely" level. I think the change at the lower end may be as much due to increasing recognition that 1.5 is a firm limit, as stronger confidence that the value of S is actually greater than 2C. Forster and Gregory's recent estimate was 1.7C and there are several others with a strong likelihood close to the lower end of the range. Anyway, depending on how "likely" is interpreted, this phrase still acknowledges perhaps as much as 15% probability of S sneaking below the 2C threshold, but it cannot do so by much.

What they have said about the upper end of the range is more...interesting. They have added the phrase:
"Values substantially higher than 4.5°C cannot be excluded, ...".
A literal interpretation of this is completely vacuous (we can never assign a probability of precisely zero), so I'm not at all sure what they mean by including it. Note how it carefully avoids using the calibrated probabilistic language that has been adopted (likely, very likely etc). I can't help but be amused by RC's comment:
"the governments (for whom the report is being written) are perfectly entitled to insist that the language be modified so that the conclusions are correctly understood by them and the scientists. [...] The advantage of this process is that everyone involved is absolutely clear what is meant by each sentence. Recall after the National Academies report on surface temperature reconstructions there was much discussion about the definition of 'plausible'. That kind of thing shouldn't happen with AR4."
I predict discussion about the definition of "cannot be excluded". I will be discussing it, at least! I complained about this ambiguous phrasing (which appeared in similar form in a couple of chapters) in my review of the last draft, and explicitly asked the authors to explain more clearly what they meant by it. I've also asked a couple of authors who used similar phrasing in their papers but have not got a reply out of them. I find it hard to avoid the conclusion that this "cannot be excluded" phrase was deliberately chosen specifically for its meaninglessness, in order to to be able to present a "consensus" rather than a strong disagreement about the credibility of such high values. I'm sure that those who assign a probability of 5% or even more to S greater than 6C will consider that this phrase supports them, even though Stoat parses it as "they do go on to diss > 4.5 oC a bit" due presumably to the "but agreement of models with observations is not as good for those values" which completes the sentence I partially quoted above. "Not as good" also has no probabilistic interpretation of course.

[Update

Based on this first-hand report, the phrase was indeed chosen specifically for its ambiguity.]

3. There is more probabilistic confusion in the discussion of attribution of past climate changes (I've written about this before). This is perhaps most clearly demonstrated in
"It is very unlikely that climate changes of at least the seven centuries prior to 1950 were due to variability generated within the climate system alone."
One thing they might have said (and perhaps thought they were) is that an unforced system is very unlikely to exhibit the observed level of climate changes. That is an essentially frequentist statement about ensembles of model runs. But what they have actually said appears to be the Bayesian statement that they believe that there was external forcing in the real world. No shit Sherlock! The cavalier way in which the detection and attribution community freely switches between frequentist and Bayesian approaches to probability, without any clear explanation, gives every impression that they do not understand the difference (or even perhaps realise that there might be a difference). Their writing about the recent warming is similarly clumsy:
"it is extremely unlikely that global climate change of the past fifty years can be explained without external forcing, and very likely that it is not due to known natural causes alone."
The first statement again is essentially frequentist, the second Bayesian. The existence of anthropogenic forcing as a contribution to the recent climate changes is not merely "very likely", it is at least "virtually certain"! (I don't believe there is a single working climate scientist who would argue that the anthropogenic forcing has been precisely zero. There is of course debate over its magnitude - some legitimate, some specious.)

I should point out that this criticism doesn't invalidate (or even weaken) the broad thrust of the report. I'm grumbling about the D&A stuff primarily because it forms the basis of the confusion in the climate sensitivity debate, rather than actually mattering in itself. They could have written things in a clear and correct manner without substantively affecting the overall message.

Saturday, September 08, 2007

Pile on!

It seems like everyone is having a go at Steven Schwartz these days: based on the links to my earlier post - and the "lurkers who support me via email" :-) - my commentary on his recent climate sensitivity estimate met with general approval (more on that later), but this post is about his own commentary on the IPCC report that I noticed previously in Nature (available here).

I didn't want to distract myself before by fisking it in detail, but it seems largely misguided. His opening sentence "The latest report from the Intergovernmental Panel on Climate Change assesses the skill of climate models by their ability to reproduce warming over the twentieth century..." is obviously wrong and the article does not improve much from there. There is a whole chapter (number 8) explicitly dedicated to assessing model performance, there is another chapter (6) on paleoclimate, and large parts of chapters 9 and 10 are also concerned with how we gain confidence in projections/probabilistic predictions. While 20th century climate change detection and attribution does get (IMO) a slightly unhealthy prominence in the report as a whole, it is hardly the whole story - indeed it has long been known that merely matching the general temperature trend is a rather weak test of model performance at least for the longer term (one detail that is worth noting is that all plausible models indicate a roughly linear trend over say 1980-2030, which does contribute to our confidence in a continuation of the recent warming trend).

A response from several IPCC authors has just appeared here, together with a brief rebuttal from Schwartz. The IPCC authors point out a number of reasons why Schwartz is wrong, and in his final rebuttal all he does is repeat his original erroneous claim that "in assessing the skill of climate models by their ability to reproduce warming over the twentieth century, the latest report from the IPCC may give a false sense of their predictive capability". It's just not true.

Coming soon: more on Schwartz and that sensitivity estimate.

Sunday, December 19, 2010

AGU part 2

On with the show!

Perhaps it's best to blog a little after the event, as it gives time for the boring and disappointing bits to fade away, leaving a more positive overall impression. And with an 11h flight I have ample time to wax lyrical about the good bits too, rather than struggling to scribble something down in a spare moment. Indeed I now see on checking through my notes that there were a couple of interesting talks on Tuesday, that I didn't mention before, in the section on uncertainty quantification. First Carol Snyder presented an unusually high estimate of climate sensitivity based on paleo data, which appears to hinge on a strong estimated LGM cooling. Not that this makes her wrong of course, but the previous week, Andreas Schmittner had presented a somewhat contrary result at the PMIP meeting, so I will email them to try to identify the discrepancy. Derek Lemoine also presented some evidence, probably compatible with Frank et al (who spoke in the morning), supporting a lowish (but positive) carbon cycle feedback at the bottom end of the model range. Jules went to a session on geoengineering where everyone seems to be trying to prove that even if we could restrain the rise in global mean temperature, it was only by screwing up all regional patterns of rainfall.

Wednesday

On Wednesday morning I started off at writers corner, where several authors of popular science books discussed their work and/or their lives. And then at the end of this session we had Greg Craven. It tended rather to the hysterical, and I don't mean that in a good way. My complete unabridged notes on his presentation read "I am insane. Apocalpyse soon." although in the interests of fairness I should point out that only the first sentence was a direct quote, the second was merely my personal summary of events. Perhaps the most I should say is that since Stephen Mosher wrote in detail about how much he didn't like the panel discussion later that day, I can only assume that he did not attend the prior presentations. I actually walked out of the panel discussion at the point that Greg started telling an uppity woman (president of some equal opportunities organisation, no less) that she had said her piece and could she please shut up and sit down. Pot, meet kettle. Oppenheimer, on the other hand, gave a sensible talk that I found interesting, if not too earth-shattering. He argued that scientists have a general duty to engage with the public, and even if we didn't want to do it individually, we can't necessarily avoid it in a world where merely being a climate scientist puts us in the firing line. However, under the rubric of "engaging" he discussed a wide range of options, and I was amused to see him list blogging as ranking higher than merely participating in assessments such as the IPCC and NAS :-) He also emphasised the importance of only speaking in areas where you had earnt credibility based on your published record, which formed an interesting backdrop to Judith Curry's talk later that day. She devoted her time to accusing the IPCC of ignoring the tails of the pdfs of climate sensitivity that were clearly presented in the very figure that she repeatedly referred to and explicitly emphasised in the summary ("values substantially higher than 4.5C cannot be excluded"), then read out a few cartoons and finally, literally out of nowhere, concluded that therefore they had underestimated the magnitude of decadal variability and that their detection and attribution results were unsound! Really, I'm not making this up, it was actually how it happened. These latter topics were first introduced on her concluding slide and there was no hint of supporting argument. She also talked about the "modal falsification" of Betz 2009, (which I haven't read but just googled now, is there a free version somewhere?) so I asked if and how this "falsification" (and she used the scare quotes herself) was distinct from assigning a low posterior probability in a Bayesian sense. She replied that it could be considered the same, at which point some of the audience were shaking their heads and others were nodding in agreement. From which I conclude that nobody, including Judith, knows what Judith means. Unfortunately, she didn't seem to be anywhere to be found at the end of the session and I didn't see her at any of the other relevant sessions where people actually dealing with these sorts of issues were actually presenting concrete results.

Thursday

There was more communication stuff on Thursday morning, much along the lines of wailing about how nasty everyone (well, Republicans and/or denialists, at least) was being to climate scientists, and how we need to educate everyone about the Truth of climate change. Which of course is true to some extent, but I'm not really convinced it is worth the time and energy that was devoted to it at the AGU. Tim Palmer gave a really good Bjerknes Lecture, which I nearly didn't go to as I've seen most of the content before (at the INI) but I'm glad I did as it seemed much better this time round. Of course, as he mentioned, it was a sort of anti-Bjerknes lecture in some ways, because the eponymous scientist was firmly rooted in the deterministic world and Tim is very much in the probabilistic/ensemble forecasting mould (as everyone in NWP has to be, of course, and I'm sure Bjerknes would agree were he around today). In fact Tim is a strong and convincing advocate of the use of stochastic parameterisations in climate models, which formed the main content of his talk. I agree they are a good idea and must remember to mention some time that Jim Hansen used a random number generator in the cloud scheme of his 1984 model :-) One of the numerous modelling groups here is using a more modern equivalent, so I don't think it's something that people are particularly hostile to. Where I part company with Tim is in his advocacy of a single coordinated model-building effort, which to be fair he hardly mentioned this time. I think it is clear that even with stochastic physics, we would still need a range of different models to investigate our uncertainty in long-term climate change meaningfully, and Suki Manabe made exactly this point in his inimitable style in the questions after the talk.

After lunch there was even more on communication. I'm not sure if it was deliberate or not, but the session discussing how to cope with this blogospheric "other" that scientists don't really understand was running in parallel with a panel of science bloggers offering advice to those who were prepared to join in. I only stayed for a little of the latter, it was pretty anodyne stuff. As jules noted, they all seemed to be "normal" geoscientists, none of them were involved in the post-normal world of climate science so their take on questions of debate, argument and abuse seemed somewhat rose-tinted to me. Not that that should put anyone off who is considering getting involved. I also thought their attitude towards discussing on-going and unpublished work was rather 20th century.

Then it was time for us to head off to defend our posters, which we had planned for beer o'clock to ease the pain. Unlike the EGU where the poster session was arranged every evening, here the poster session was scheduled for a full half-day (4h20m) out of which you were supposed to pick at least an hour that you would be there. Of course this is horribly inefficient and made it very hard to meet specific people, though perhaps there were enough random encounters that it didn't matter and a good poster should be pretty self-explanatory anyway. I didn't have that many visitors but made up for that by having long chats with several of them.

Friday

Friday was probably the most fun, with lots of sessions on using data together with models. First the reanalysis session, where various improvements and new analyses were presented. There are still lots of problems with inhomogeneities in observations which make the trends in precipitation extremely dubious - even at the global average scale, different analyses gave substantially different results. Jules heard about various attempts at carbon sequestration, but it wasn't clear if they would be practical and effective. Then in the afternoon there was a big session on using observations to test and validate models, which is of course right up our street. There is a big push for more of this in the context of the CMIP5 model runs, and it seems there will be some basket of comparisons proposed although it must be remembered that no-one actually knows what particular features are important in improving model predictions. As well as a number of fairly detailed and specific studies Wendy Parker spoke about the broader context of how we could think about model adequacy and performance. There had already been a whole session on more philosophical aspects of model interpretation which I did not attend but jules did, I think many people are generally aware of the issues, but struggling to find concrete solution. Anyway, I didn't think that anyone presented anything particularly earth-shattering in this session but unlike some of the other more disappointing parts earlier in the week, they were at least presenting some new ideas and results, rather than either merely re-reading an old long-published paper I already knew about (which are few had done) or else talking vaguely about what they were hoping to do in the next few years. I also stuck my oar in with a few questions and comments, which at least made it interesting to me :-) I also have a few more things to chase up with some of the speakers via emails as I didn't quite understand what they had done, or perhaps why they had done it.

So in the end I suppose the whole thing was quite fun and I'm certainly happier on the plane on the way home than I was a few days ago when I wrote the previous post. Given that we were tired before we even started (following on directly from the PMIP meeting in Kyoto last week) and that jules came down with a cold that hung around all week, we are glad to be heading home. In fact our devotion to duty is such that we turned down the offer of $800 each to get bumped off the plane which was over-booked. Maybe we should have taken the money and gone straight back to the Mac store but we are sufficiently worn out that another day of sightseeing in SF (especially given the rain) didn't really appeal. Jamstec would also have had a fit of course, which almost tipped the balance in favour of staying, but not quite. And anyway, they've already bought us everything that the Mac store sells :-)

Wednesday, December 29, 2010

Snowfalls are now just a thing of the past

Cheap shot I know, but hard to resist. This is the Indescribablyboring from March 2000:

According to Dr David Viner, a senior research scientist at the climatic research unit (CRU) of the University of East Anglia,within a few years winter snowfall will become "a very rare and exciting event".

"Children just aren't going to know what snow is," he said.

[...]

Professor Jarich Oosten, an anthropologist at the University of Leiden in the Netherlands, says that even if we no longer see snow, it will remain culturally important.

"We don't really have wolves in Europe any more, but they are still an important part of our culture and everyone knows what they look like," he said.

David Parker, at the Hadley Centre for Climate Prediction and Research in Berkshire, says ultimately, British children could have only virtual experience of snow. Via the internet, they might wonder at polar scenes - or eventually "feel" virtual cold.

Heavy snow will return occasionally, says Dr Viner, but when it does we will be unprepared. "We're really going to get caught out. Snow will probably cause chaos in 20 years time," he said.

Of course snow falling always has and always will cause chaos in Britain. Until it stops falling completely, which may now be a few years further away than was previously thought :-)

(for those who've been living in a cave for the last few years, or at least outside the UK, it appears that in fact reports of the demise of snowfall are greatly exaggerated, at least according to the winters of 2008, 2009 and 2010).

Actually, it's better than that, because the latest research is that all this snowfall is actually more proof of global warming after all! I await with amusement the reaction of the detection and attribution community to this proof that all their results are bogus, since they have already proved that the "warmer winters" are being caused by anthropogenic global warming (eg here, a paper I just saw today). Personally, while I accept it's theoretically possible for AGW to cause some localised cooling (at least on a temporary basis), my money is on the D&A results for the time being. Invest in Scottish skiing resorts (at least for the long term) at your peril...

(and as a footnote to the pedants, I know there doesn't necessarily have to be a contradiction between snowfall and warmth, but in fact this December has been not only snowy but also perhaps the coldest since records began in the UK).

Wednesday, March 21, 2012

HadCRUT4: 1998 and all that

So the long-awaited HadCRUT4 paper is now published, and the UMKO web site has a press release, though the full data set does not seem to be available yet. I'm sure that won't be long.

As was exclusively revealed on this blog some time ago, according to this updated analysis, 1998 is no longer regarded as the warmest year on record, having been overtaken first by 2005 and then again in 2010. Of course, this was not really controversial in scientific circles, as the small cold bias of HadCRUT3 (due to data voids in the most rapidly warming areas) had been well documented. Competing analyses (NCDC and GISTEMP) that smooth over the gaps already showed these results.

Some of the major differences between the data sets are illustrated by this figure from the paper:


I've added two green ovals to each map to highlight some differences - firstly, filling in the big gap in Siberia, where it was clear in HadCRUT3 that there had been lots of warming even though there were gaps in the coverage. The addition of more observations has merely confirmed what everyone knew, and the previous data has not been changed in any significant way. Moreover, the additional data in the relatively cool area of the south Indian ocean shows that they didn't just try to collect data in the hottest regions. So whatever desperate attempts the sceptics make to discredit this update, they simply don't have a leg to stand on.

The problem really is in the presentation of the data average as representing "global mean temperature" in the first place, when it doesn't, as the data are not missing at random. One good way to deal with missing data in this sort of situation is to fill the gaps with some sort of smoothing (there are a wide range of options of varying sophistication) to generate a global field before averaging, as NCDC and GISTEMP have always done. However, this doesn't matter in the context of the sort of detailed model-data comparisons that underpin detection and attribution, since they generally work at the level of the gridded data and voids are ignored. The mid-century changes to ocean data may however have a modest impact here.

David Whitehouse (or someone impersonating him) has been quick off the mark to bluster about how he would have really won the bet anyway. Unfortunately his comments are highly misleading.

Firstly, we never clearly specified HadCRUT3. I did try to get David to confirm exact details via email, but he refused, and I believe the precise phrase used by Tim Harford was "the Hadley Centre analysis". Maybe his behaviour should have been a red flag that he would resort to rewriting history whenever possible, but I thought it was unlikely to be ambiguous. Of course I didn't know the HC were planning to change their analysis, even though it is inevitable that these sort of changes do take place over time (hence the "3" in HadCRUT3). For how a bet of this type is more clearly specified, see here for example: "in the data set HadCRUTv. or successor data set. Successor data set is the data set used by the Hadley Centre to compare hottest years to the media at that time." Incidentally, Gabi has not won that bet yet, as it specifically refers to the setting of a new record, and 2010 was not.

Secondly, he claims that because 2010 is no hotter than 2005 in the new analysis, he wins anyway. This is more than a bit ridiculous, as the entire bet was predicated on the usually long interval after 1998 in which the record had (apparently) not been broken.

(Of course it will only be a matter of weeks before we start seeing sceptical arguments along the lines of "no global warming since 2005").

Friday, January 29, 2010

Adapt or die

There's an interesting article in Climatic Change which looks at temperature-related mortality in England and Wales in recent decades. They use detection-and-attribution methods to look at the influence of climate change and adaptation.

Given the lengthy introduction citing numerous hyperbolic warnings about the primarily negative and potentially disastrous health effects anticipated due to climate change, it is mildly amusing (though hardly surprising to anyone familiar with the literature) that they observe that temperature-related mortality has decreased sharply over the interval studied. More interestingly, they find that a major cause of this decrease is not even the warmer winters the UK has seen, but rather better adaptation to cold weather. This of course directly refutes the "optimally adapted" meme, since if that was the case, we would expect to see a decrease rather than the observed increase in tolerance to cold extremes as the winters got warmer. But in reality, wealth and technology (eg insulation, heating and clothing) act to increase our tolerance whether or not the climate changes, although the latter may in theory act as an additional stimulus if it were sufficiently important (which it clearly is not, in the UK).

Of course they have to finish off by talking about increases in heatwaves, wondering "whether adaptation will manage to keep pace with such changes". I think it is patently obvious that it will, and would happily bet against any predictions of increasing temperature-related deaths in coming decades.

Monday, November 07, 2011

The null hypothesis in climate science

Three papers have just appeared in WIREs Climate Change (here, here and here) discussing the role of the null hypothesis in climate science, especially detection and attribution.

Trenberth argues that, since the null (that we have not changed the climate) is not true, we should try to test some other null hypothesis. He sounds like someone who has just discovered that the frequentist approach is actually pretty useless in principle (as I've said many times before, it is fundamentally incapable of even addressing the questions that people want answers to), but although he seems to be grasping towards a Bayesian approach, he hasn't really got there, at least not in a coherent and clear manner. Curry is just nonsense as usual, and beside noting that she has (1) grossly misrepresented the IAC report and (2) abjectly failed to back up the claims that Curry and Webster made in a previous paper, there isn't really anything meaningful to discuss in what she said.

Myles Allen's commentary is by some distance the best of the bunch, in fact I broadly agree (shock horror) with what he has said. If one is going to take a frequentist approach, the null hypothesis of no effect is often an entirely reasonable starting point. It is important to understand that rejecting the null does not simply mean learning that there has been some effect, but it also indicates that we know (at least at some level of confidence) the direction of the effect! That is, it is not only an effect of zero which is rejected, but all possible negative (say) effects of any magnitude too - this generalisation may not be strictly correct in all possible applications of this sort of methodology, but I'm pretty sure it is true in practice for the D&A field. Especially when we are talking about the local incidence of extreme weather, there really are many cases when we have little reason for a prior belief in an anthropogenically-forced increase versus a decrease in these events, so a reasonable Bayesian approach would also start from a prior which was basically symmetric around zero. The correct interpretation of a non-rejection of the null here is not "there has been no effect" but rather "we don't know if AGW is making these events more or less likely/large". Much of Trenberth's complaint could be more productively aimed at the routine misinterpretation of D&A results, rather than the method of their generation. Trenberth also sometimes sounds like he is arguing that we should always assume that every bad thing was caused by (or at least exacerbated by) AGW, but this simply isn't tenable. Even if storminess increases in general, changes in storm tracks might lead to reduction in events in some areas, with Zahn and von Storch's work on polar lows an obvious example of this. On the other hand, there are also some types of event where we may have decent prior belief in the nature of the anthropogenically-forced change (such as temperature extremes) and in these cases it would be reasonable for a Bayesian to use a prior that reflects this belief.

I can find one thing to object to in Myles' commentary though, and that's the manner in which he tries to pre-judge the "consensus" response to Trenberth's argument. Noting that he (Allen) is in fact a major figure in forming the "consensus" in these private meetings where the handful of IPCC authors decide what to say, it sounds to me rather like a pre-emptive strike against anyone who might be tempted to take the opposing view. I would prefer it if he restricted himself to arguing on the basis of the issues rather than that he holds/forms the majority view. His behaviour here is reminiscent of the way he (and others) tried to reject our arguments about uniform priors, on the basis that everyone had already agreed that his approach was the correct solution. All that achieved was to slow the progress of knowledge by a few years.

Wednesday, May 02, 2007

Why David Evans is wrong (along with all the other sceptics)

David Evans has a post up on Backseat driving, explaining why he has taken Brian on with a series of bets about global warming.

Usually, I don't bother addressing the sceptic stuff that can be found on a thousand blogs: people who want the facts can come looking for them, and I've got limited patience for wrestling with pigs (you both get mucky, but the pig enjoys it). However, David is a bit of a special case as he's actually been prepared to put a significant amount of his money where his mouth is. It turns out that his post is a fairly standard laundry list of excuses as to why he thinks climate science is a bit ropey. I could have some sympathy with some of his points, but there is one gaping hole in his analysis.

He starts off by acknowledging:
1. Carbon dioxide is a greenhouse gas. Proved in a laboratory a century ago.
and then continues with a string of tenuous arguments as to how it is possible that other factors could be behind most of the recent warming of the Earth.

At no point does he actually produce any evidence for his implied belief that CO2 has a negligible effect.

This, to me, is the crux of the matter. With no feedbacks at all, the sensitivity to doubled CO2 is about 1C, based on the well-understood radiative physics. It is also obvious that a warmer atmosphere has the potential to hold more water vapour (itself a GHG of course), and although the magnitude of this effect isn't known with certainty, the most plausible first-order estimate (supported by models and data) would be that relative humidity will stay roughly constant. This gives another 1C, making 2C in total. [The numbers here are intended as ballpark estimates, please let's have no quibbles about the precision.] Clouds may have a significant effect to enhance or offset warming. We know the climate has varied plenty in the past (indeed this is generally one of the septics' favourite talking points), so it seems implausible that they are a very strong stabilising force. Almost all models, using a wide range of physical parameterisations, suggest a significant positive amplification, giving the typical range of 2-4.5C for sensitivity. All analyses of observational evidence also point towards a value of close to 3C (exactly how close is still subject to some debate).

So we are left with the question: why on Earth would anyone believe that CO2 has almost no effect?

The claim that a world without anthropogenic forcing could possibly have warmed this much, whether true or not, is almost entirely irrelevant to the question of what we expect the anthropogenic effect to be. It's true that in the event of stronger natural variability, that might suggest a slightly larger possibility that a future downturn in the natural component could exceed the anthropogenically-forced response in the short term. But as I mentioned above, more variability also implies smaller stabilising feedbacks, so we'd also expect to see a larger sensitivity to CO2 in this case. Are we really supposed to believe that the planet is highly sensitive to some speculative and unquantified mechanism such as cosmic rays, and simultaneously insensitive to an effect that's been reasonably well understood for over 100 years? Why?

Detection and attribution has a lot to answer for in respect of this confusion. D&A essentially addresses the question "could an unforced planet have warmed as much as the observations"? However, this (frequentist) question has only tangential relevance to a (Bayesian) estimate of future warming. IMO, the arguments over whether or not we have "detected" AGW, and at what level of confidence, entirely misses the point. Imagine that someone points a gun in roughly your direction, and pulls the trigger. According to D&A, nothing interesting is going on until the bullet hits you, but at that point it's too late. An intelligent Bayesian would believe that there was a significant probability of serious harm before the bullet arrived - hopefully even before the trigger was pulled. Now, I'm not saying that climate change is going to suddenly kill us all, but just giving an analogy to explain the manner in which D&A fundamentally answers the wrong question. (As a secondary point, I believe it is the attempt to pretend that D&A methods can answer the interesting and useful questions that has lead to the uniform prior nonsense, but that isn't really my point here.)

So, until David and the rest can come up with some plausible arguments as to why CO2 actually has no effect, backed up by a sensible climate model which supports this claim, I'll continue to believe that it does in fact have a significant effect which will (with high probability) lead to continued warming. That is, as he more or less admitted at the start, solid science that is more than 100 years old.

Friday, August 12, 2011

How many of Roger's findings about probability manage to be wrong? Answer: he's more inventive than you might expect.

Roger Pielke has a new post up asserting that 28% of the IPCC's findings are incorrect. Although it's obviously a rather implausible figure, I was expecting this claim to be backed up with some sort of evidence of errors, or at least sloppiness, or something, so I had a look at his paper that he cites to justify the claim.

It turns out that 100%-28% = 72% is merely the average (lower bound) probability level associated with the statements they made. Such as "It is very likely that hot extremes, heat waves and heavy precipitation events will continue to become more frequent." Here "very likely" means greater than 90%. So, given 10 such statements, the IPCC is saying that they would expect the "very likely" outcome to occur about 9 times, and not occur about once. And similarly for "likely" (66%). Averaging over all the probabilistic statements, it should be expected that in about 28% of cases, the (probabilistically) preferred outcome will not actually happen.

And in Roger-world, this means that 28% of the statements are "incorrect". Note, however, that he does not make this silly claim in the paper itself, but only in his blog post.

To see why this interpretation is nonsensical, consider a single roll of a fair die. I state (accurately) that it is "likely" to lie in the range 1-5. If I roll a 6, then in Roger-world, my statement was incorrect. However, it was not incorrect, and Roger is simply wrong to claim so.

As you can see from the comments, I challenged Roger on this, and his response (entirely in character) is to duck and weave. In his comment #5, for example, he shamelessly misrepresents what I said, and brings up the red herring of a definitive prediction (when in fact I had clearly made a probabilistic one, and the distinction is of course absolutely fundamental to the point). The obvious elephant in the room that Roger cannot bring himself to acknowledge is that the statement is correct irrespective of the outcome of the roll. "Correctness" of a single statement simply isn't something that can be directly validated (or invalidated) by the outcome, and the accurate calibration of a probabilistic prediction system actually relies on having the appropriate number of "failures" for each level of probability.

I realise of course that having done some rather boring textual analysis that in his own words amounts to "Nothing too interesting, really", Roger is just rabble-rousing on his blog. I'm confident that any competent scientist will see straight though it, but that's hardly his target audience.

As for what the 72%/28% average actually does mean, it doesn't actually tell us anything except that the IPCC makes a lot of statements about things that it is only (by its own admission) moderately confident about. It might in principle be interesting to see how the confidence level changes over time, but only if the set of statements were to be held fixed from one assessment to the next. People have looked at climate sensitivity estimates (hardly changed) and detection and attribution (increased markedly in confidence) but not a lot else AIUI. I suppose we can anticipate Roger claiming that the next report is either more correct, or less, depending on what mix of statements they happen to include :-)

Incidentally, and although it's a minor point it is perhaps telling in terms of his overall level of competence, Roger is also wrong where he claims that if the statements are not independent, then the proportion of "incorrect" will be higher than 28%. Actually, if the statements are not independent (while still being correctly calibrated), then the proportion that do not come to pass would still be 28% in expectation, just with higher variance, meaning that either a larger or smaller proportion would not be surprising. Unlike the simple misinterpretation in his blog post title, this elementary error is actually made in the paper itself.

Sunday, March 31, 2013

Decadal prediction stuff part 2

Ok, having got various things out of the way, on with the show.

I liked this letter which appeared in Nature recently. Not just because I'd done something similar myself with the earlier Hansen forecast :-) In general, I think it's important to revisit historical statements to see how well they held up. Allen et al have gone back to a forecast they made about 10 years ago, and checked how well it matches up to reality. The answer is...



really well. On the left, is the original forecast with new data added, and the right is the same result re-expressed relative to a 1986-96 baseline. The forecast was originally expressed in terms of decadal means, so I don't think there is anything untoward in the smoothing. The solid line black in the left plot is the original HadCM2 output, with the dashed line and grey region representing the adjusted result after fitting to recent (at that time) obs using standard detection and attribution tchniques.

They also compared their forecast to a couple of alternative approaches:


This plot shows the HadCM2 forecast (black), CMIP5 models (blue) and another possible forecast of no forced trend, just a random walk (green). The red line is the observed temperature. They point out that their forecast performed better than the alternatives, in the sense that it assigned higher probability (density) to the observations.

So far, so good. However, I disagree with their statement that "the CMIP5 forecast also clearly outperforms the random walk, primarily because it has better sharpness" (my emphasis). Actually, the CMIP5 forecast outperforms the random walk simply because it is clearly much closer to the data. The CMIP5 mean is about 0.41 in these units (all these numbers are just read off the graph, and may not be precise), the random walk is of course 0, and the observed anomaly is 0.27. The only ways a forecast based on the CMIP5 mean could have undererformed the random walk would have been if it was either so sharp that it excluded the obs (which in practice would mean a standard deviation of 0.06 or less, resulting in a 90% range of 0.31-0.51), or so diffuse that it assiged low probability across a huge range of values (ie, a standard deviation of 0.75 or greater, with associated 90% range of -0.8 to 1.6). The actual CMIP5 width here seems to be close to 0.1, well within those sharpness bounds.

I do think I know what the authors are trying to say, which is that if you are going to be at the 5th percentile of a distribution, it's better to be at the 5th percentile of a sharp forecast than a broad one. But changing the sharpness of the forecast based on CMIP5 would obviously mean the obs were no longer at the 5th percentile! In fact, despite not quite hitting the obs, the CMIP5 forecast is not that much worse than the tuned forecast (black curve), thanks to being quite sharp. And according to the authors' own estimation of how much uncertainty they had in their original forecast, they obviously got extraordinarily lucky to hit the data so precisely. With their forecast width, it would have been almost impossible to miss the 90% interval - this would have required a very large decadal jump in temperature. I don't think it is reasonable to say that one method is intrinsically better than the other, on the basis of a single verification point that both methods actually forecast correctly. If the obs had come in at say 0.4 - which they forecast with high probability - I hardly think they would have been saying that the CMIP5 ensemble of opportunity was a superior approach.

(For what it's worth, I think the method used in this forecast intrinsically has exaggerated uncertainty, but that's another story entirely.)

Saturday, April 25, 2015

BlueSkiesResearch.org.uk: The EGU review 2015


If it’s April, it must be Austria. So far, our plan to travel less now we are back in the UK seems to be being mostly honoured in the breach. To be fair, we are travelling fewer miles and suffering less jet lag, and both these factors probably help to explain why we are still spending quite a lot time away from home. As you may have seen, we entered the Vienna Marathon (and half), and the subsequent need for recovery gave us a good excuse for week of lazily sitting around, eating and drinking. It didn’t turn out quite as lazy as we’d expected…

Somehow we managed to get in for 8:30 on Monday morning, where I started off in a hydrology session about data, models, model building and prediction. I had put my poster into this, not because it was really that great a fit but because hydrologists have been thinking about these issues for a lot longer than most climate scientists and have a range of interesting ideas. Keith Beven is one of the famous names in the field and he gave an good general talk about determining when and whether models might be fit for purpose. The session was a good choice for me, I picked up a couple of new (to me at least) ideas that might well come in handy. There was an unfortunate clash with an NP session on building models from data, so jules attended that one. In the afternoon there was nonlinear time series analysis, wavelets and the like, including Bo Christiansen explaining his 2014 paper on regression, which I found much more digestible as a presentation than I had when reading the paper.

Tuesday started with the data assimilation session, including some steady progress towards the practical use of particle filters. To be honest, I’m surprised how long this has been in the pipeline, as some bold claims were being made for it several years ago. Another interesting talk was Alexis Hannart on the use of data assimilation in detection and attribution. Two relevant papers are here and here. I’m not sure exactly what it all means yet but it’s a promising idea at least. jules went to the seasonal/decadal prediction session and reported that it was the same as usual…some spots of skill but at longer time scales this is essentially due to the forced response rather than the precise initialisation.

Wednesday was the big paleoclimate day, with jules’ co-organised session including the Milankovitch Medal lecture, all the latest paleo research and some work specifically linking past and future climate change. However, before that, we both managed to see Bjorn Stevens’ talk on why aerosol forcing is lower than most people had previously thought. This was a longer and more detailed version of what he had said at Ringberg, and seemed fairly convincing to me. Then on to the paleo, which started with Paul Valdes giving a great medal lecture covering several new pieces of work from his group covering multiple time scales over the last 500 million years. We were amused by one of the questioners afterwards managing to drop in the line “As I said in my medal lecture a few years ago…”. In the rest of the session, there was interesting stuff on re-interpreting some ocean proxies which brought them more into line with other evidence from data and models, and some analysis of model simulations.

By Thursday i was flagging a bit, and the program was a bit less busy. I went to the sesssion on the Last Millennium, where someone had discovered a new volcanic eruption, someone else was debating about the significance and cause of various wiggles, and several others were considering the question of what we can really expect to learn from the proxy data anyway. Rob Wilson had a potentially provocative title along the lines of “Are the tree rings fit for purpose?” and I was hoping for a bit of a bunfight but he pulled his punches a bit. After a long, large and loud dinner on Thursday evening with the Bristol crowd, Friday was fortunately very quiet. The highlight was the EGU Great Debate on open access publishing, which can be streamed from the EGU web site. A couple of commercial publishers on the panel tried their best to justify the continuation of closed journals with the rather bizarre justification that scientists wanted them (only because they can’t/won’t pay the extortionate open access fee!), but even they acknowledged the inevitable growth of the open access movement. Uli Poschl was fairly direct without being too aggressive, saying that we (scientists) were going to do it anyway, so commercial publishers could either adapt or fail. There were a few red herrings raised in the discussion, but overall it was an interesting event. After 4 consecutive evenings of dining out with various people, we gave the convenors’ party a miss and had an early night instead.

Throughout the week, we gained the impression that a lot of younger scientists had not attended and their work was instead being presented by their supervisors. This rather undermines the EGU goal of trying to allow younger scientists to gain some visibility. No good picking a couple of early career scientists for your oral session if they don’t have the funding to turn up! But of course there have always been absentees and I don’t know if this is really a trend or just chance in what we happened to see. The conference attendance was marginally down on previous years, but there’s no sign of a trend there either.

Throughout the week the poster sessions were enjoyable and well-attended, though I wish all divisions would adopt the CL convention of using the evening session for posters only, and not scheduling talks. It would have been more convenient to be able to talk to all the poster presenters at the same time.

postas-1

Viennese food was fun as always, perhaps a bit less exciting now we are not in Japan and can get big hunks of meat and good beer any time we like. What with no longer being bound by Japanese rules, we decided to stay an extra day in Vienna at the end. As it happened, it was a cold and windy day, we didn’t really have the energy for much sightseeing but did walk most of the way round the town. In the evening, we found a concert in Karlskirche – Mozart’s Requiem with an orchestra on traditional instruments. The singers were very good but the pews were very hard and the accoustics were rather odd. It was an interesting experience that I wouldn’t rush to repeat. Next year the Vienna marathon is a week earlier than the EGU, so we certainly won’t be doing both of them, and quite possibly neither.