Showing posts with label probability. Show all posts
Showing posts with label probability. Show all posts

Monday, December 23, 2019

Outcomes

At the start of the year I made some predictions. It's now time to see how I did.

In reverse order....


6. The level of CO2 in the atmosphere will increase (p=0.999).

Yup. I don't really need to wait for 1 Jan for that.

5. 2019 will be warmer than most years this century so far (p=0.75 - not the result of any real analysis).

As above, I know we've a few days to go but no need to wait for this one, which has been very clear for a good while now.

4. We will also submit a highly impactful paper in collaboration with many others (p=0.85).

Done, reviews back which look broadly ok, revision planned for early next year when we all have a bit of time (31 Dec is a stupid IPCC submission deadline for lots of other stuff).

3. Jules and I will finish off the rather delayed work with Thorsten and Bjorn (p=0.95).

Yes, the project is done, through the write-up continues. Actually we hope to submit a paper by 31 Dec but there will be more to do next year too.

2. I will run a time (just!) under 2:45 at Manchester marathon (p=0.6).

Nope, 2:47:15 this time. The prediction was made just a few days after I'd run a big PB in a 10k but even then I thought it was barely more likely than not, and it got less likely as the date approached. 

1. Brexit won't happen (p=0.95).

On re-reading the old post, I have to admit I cannot remember the precise intention I had when I wrote this. Given the annual time frame of the remainder of the bets, and the narrative of that time being that we were certainly going to leave on the 29th March (as repeated over 100 times by May - remember her? - and the rest of them) I do believe I must have been referring to leaving during 2019. After all, I could never hope to validate a bet of infinite duration. So yes, I'm going to give myself this one.

On the other hand, I did actually think that we would probably not be stupid enough to leave at all, and clearly I misunderestimated the electorate and also the dishonesty of the Conservative Party, or perhaps as it should be known, the English National Party.

I have learnt from that misjudgment and will not be offering any predictions as to where we end up at the end of next year. Which is sort of inconvenient, as we are trying to arrange a new contract with our European friends for work which could extend into 2021. Our options would seem to include: limiting the scope of the contract to what we can confidently complete strictly within 2020, which is far from ideal, or shifting everything to Estonia (incurring additional costs and inconvenience for us, though it may be the best option in the long term). Or just take a punt and cross our fingers that it all turns out ok, despite there being as yet no hint of a sketch of a plan as to how the sales of services into the EU will be regulated or taxed past 2020. It is quite possible that we'll just shut down the (very modest) operation and put our feet up. 

I'm still waiting for the brexiters to tell me how any of this is in the country's interests. But that's a rant for another day. Perhaps it's something to do with having enough of experts.

As for scoring my predictions: the idea of a “proper scoring rule” is to provide a useful measure of performance for probabilistic prediction. A natural choice is the logarithmic scoring rule L = log(p) where p is the probability assigned to the outcome, and with all of my predictions having a binary yes/no basis I'll use base 2 for the calculation. The aim is to maximise the score (ie minimise its negativity, as the log of numbers in the range 0 to 1 is negative). A certain prediction where we assign a probability of p=1 to something that comes out right scores a maximum 0, a coin toss is -1 whether right or wrong but if you predict something to only have a p=0.1 chance and it happens, then the score is log(0.1) which in base 2 is a whopping -3.3. Assigning a probability of 0 to the event that happens is a bad idea, the score is infinitely negative...oops.

My score is therefore:
0 - 0.42 - .23 - .07 -1.32 - .07 = 2.11

or about 0.35 per bet, which is equivalent to assigning p=0.78 to the correct outcome each time (which is just the geometric mean of the probabilities I did assign). Of course some were very easy, but that's why I gave them high p estimates which means high score (but a big risk if I'd got them wrong). I could have given a higher probability to the temperature prediction if I'd bothered thinking about it a bit more carefully. The running one was the only truly difficult prediction, because I was specifically calibrating the threshold to be close to the border of what I might achieve. It might have been better presented as a distribution for my finish time, where I would have had to judge the sharpness of the pdf as well as its location (ie mean).

Monday, December 10, 2018

The failure of brexit

This post is long overdue, I thought I should at least write it down before the vote on Tuesday. And now this post has been overtaken by events during its gestation, and it looks like we won't even have that vote. No matter. This isn't really going to be about the failure of the brexit process, that would be too easy. OK, just a few words about that first. It failed because the 52% who voted to leave were all promised such ridiculous and contradictory things, ranging from the Bangladeshi Caterers Associating being conned into believing they might find it easier to recruit curry house chefs by none other than Priti Patel (see how well that turned out), to Scottish fishermen believing we'd get all "our" fish back (hint: a large proportion of them have to be sold into the EU anyway as they are not eaten in the UK), the lies about more money for the NHS, the forthcoming trade deal being the "easiest in human history", to the general fabulists promising three-quarters of the Single Market (but none of those pesky foreigners) and unlimited free trade with no Customs Union, and by the way let's pretend the Irish problem doesn't exist. It was obvious from the outset that there never was an actual real brexit that would be supported by a majority, either in the country or in parliament. Moreover, there was no way we were going to build the necessary infrastructure in the time available for things that would be required by a "real" brexit like customs checkpoints, let alone replicating all the other functions that the EU currently performs for us (EURATOM being one notable example), this would be a humungous planning exercise and expense that would probably take a decade to achieve even if the govt pulled its finger out and went full steam ahead on it.

Therefore, by the time I'd had my breakfast on the morning of the 24th June 2016 I had worked out that brexit probably wasn't going to happen, and I was feeling a bit stupid that I hadn't actually realised this before the vote. I think the first time I actually wrote this on the blog was June 2017 but I'd already bored Stoat in the pub on the topic rather earlier (must have been August 2016?), and a few others besides. I shouldn't big myself up too much: I have never been 100% certain that brexit was not going to happen, and indeed there are still some mechanisms by which it could happen, but I was always quite confident that it was unlikely. Once the fantasies unraveled, the reality was never going to be attractive, and as long as there's still a way out at that point, we will probably take it.

What I'm really interested in, for the purposes of this blog post, is how and why the rest of the country hasn't allowed itself to work this out: why have we failed to analyse and understand the brexit process adequately?

Since that fateful day in 2016, there has been a rapidly growing cottage industry of experts pontificating on "which brexit" and "consequences of brexit" and "types of brexit" and "routes to brexit". Journalists have breathlessly interviewed any number of talking heads who have come out with their vacuous slogans of "brexit means brexit" and "red, white and blue brexit" and "jobs-first brexit" and...it's all just hot air. It really does seem like they have all been so firmly embedded in their own little self-referential bubble of groupthink that none of them ever stopped to consider...is this really going to happen? There has been an utter failure of the journalistic principle of holding power to account, and also an utter failure of academic research to explore possibilities, to test the boundaries of our knowledge. Instead there has been little more than non-stop regurgitation of the drivel that "brexit means brexit" and that the govt is going to "deliver a successful brexit". I wonder if, to a journalist or political scientist, the new landscape of a post-brexit world is so enticing and exciting that they have wished themselves there already?

The BBC had an official policy of "don't talk about no brexit", right up to and beyond the million-person march in London at the end of October. Humphries enjoyed sneering about the "ludicrous Peoples' Vote" when forced to mention it, though of course he sneers about most things these days. Note that the last time so many people marched together London, it was in opposition to the Iraq war. There could be a lesson there, if anyone was prepared to think about it...

Theresa May of all people took everyone by surprise when she was the first person in any position of authority to utter the words "no brexit at all" earlier this year. At which point at least half the country issued a huge sigh of relief, even though it was only intended as a scare tactic to bring the brexiters into line. Amusingly, even many of them agreed that her threatened no brexit at all was actually better than the dog's dinner she was in the process of negotiating.

And how about the academics and think-tanks? Of course some of these are nakedly political and cannot be taken seriously, but some are suppose to be independent and authoritative. Such as "The UK in a Changing Europe". (Disclaimer: I was at university with Anand Menon, who was a clever and interesting person back then too, so I'm sure he won't mind a bit of gentle criticism.) 




Watch this short video, which was published to widespread acclaim just 6 weeks ago at the end of September. In it, Anand promises that he will tell you "everything you need to know" about brexit. He even emphasises the "everything". And then proceeds to talk about different types of brexit and how they might arise. What is telling here is that he didn't discuss the possibility of Article 50 being revoked and the UK staying in the EU - this outcome simply was not in the scope of possible outcomes for him as recently as September! An impressive failure of foresight. (Even those who don't yet think it is the front-runner must surely agree it is now reasonably plausible.) If academics are not able to think the unthinkable and explore the range of possible outcomes, then I have to wonder what they are actually for. It is the political equivalent of a world in which climate scientists had determinedly ignored the possibility that CO2 might be an influence on climate, and had instead devoted themselves to arguing fruitlessly over whether the observed warming was due to the sun, or aerosols instead.

So here's my real point, and the reason for my rant. Journalists and academics, by studiously avoiding speaking truth to power and colluding with this false brexit certainty, have done a great disservice to the British public. Their unwillingness to challenge politicians on both sides has permitted an entirely fake debate about a blue Tory unicorn brexit on one side, and a red Labour unicorn brexit on the other. As a result, the miserable deal that May has produced - pretty much the only one possible, if you insist that keeping out foreigners is the top priority - seems shockingly poor to everyone. We were promised sunlit uplands, and jokers like Johnson and Rees-Mogg are still promising sunlit uplands to all and sundry with no fear of an intelligent challenge from a journalist. (Note how affronted Peter Lilley was recently when the BBC actually did produce a "fact-checker" to opine on his interview.) Meanwhile Labour promises to do the same only better, just because.  The entire Govt policy over the last two years has been nothing more than "let's get through the week and see what turns up". And when the plan falls down and we end up staying, roughly half the country will be shocked and feel betrayed, because they were told their vote would be acted on, and have been told for the past two years that their votes were being acted on, and everyone pretended that things were going full steam ahead when in fact there never was a plan, or a plausible way forward.

Well, the public were told lots of things, many of them were lies, and this was enabled by the journalists and academics not doing their jobs. Whether it is collusion, group-think, cowardice or stupidity, it has greatly damaged the country, and we will all have to suffer the consequences. I like to think that lessons may be learnt, but in all probability they will all pat each other on the back and utter meaningless aphorisms: "prediction is difficult: especially about the future". Maybe so, but I predicted it, and I was not alone in doing so.

Better post this before it's overtaken by events again :-)

Thursday, February 12, 2015

The importance of priors!

I was amused to see the different publicity, and reaction, to two pieces of research recently. The first was that silly article about running, which claimed that too much running is bad for you. Interestingly, that BBC page seems to have changed from the tautologous "too much" to the term "running hard", which if anything makes it worse, becaue too much being bad is just a definition of what too much means, whereas running hard...well that's where the research falls down. It only took a couple of mins to find the relevant paper, which shows...huge error bars on estimated risk for hard runners, such that the confidence interval on the hazard ratio actually goes below 1 (at which point running hard is good for you!). The underlying problem is that study was small, there were only 36(?) runners in this group, and this simply isn't enough to show conclusively what the health effects would be. I believe that More or Less has dealt with this, though I haven't listened to it yet.

Then just yesterday, David Spiegelhalter drew my attention to a study on the effects of alcohol, which claimed that modest drinking had no benefits (in opposition to the widely held view that it did). He explains that the study was again underpowered, such that any modest effect would by construction not be "statistically significant". The underlying problem is that, as Andrew Gelman often mentions, where an effect is probably small (but non-zero) and only weak studies with small samples are used, any "significant" result will necessarily be a huge overestimate of the effect (ie, if the true value is x but the error bar is ±10x then only estimates that come out to as much larger than x, and perhaps even with the wrong sign, can be reported as significant), and any realistic estimate close to the true value of x will be found "insignificant" and therefore be liable to being discarded or denied by silly scientists.

One obvious solution is to use a Bayesian approach with a reasonable prior, which in both cases would have found that the new data were insufficient to overturn what was previously believed to be the case...but that won't publish high profile papers or sell newspapers.

Friday, December 19, 2014

BlueSkiesResearch.org.uk: Project ICAD and UKCP09




One of the more interesting talks for us in the Paris conference mentioned previously here was James Porter talking about UKCP09. It turns out there has been a social sciences project ("Project ICAD") part of which involved looking at the UKCP09 project (and they are based in Leeds University, perhaps an additional reason for a visit there some time?). We were in Japan over the entirety of the interval in which UKCP09 took place, and only had limited contact with the relevant parties, but perhaps know enough about the issues for our perspective to have some relevance. The speaker had spent some time embedded with the Hadley Centre and had talked to a lot of people involved in the production and review of the UKCP09 project.

A significant part of Porter’s talk looked at the question of how the probabilistic predictions were made, and in particular the UKCP choice of basing their probabilities primarily on their ensemble of HadCM3 simulations with different parameter values (perturbed parameter ensemble or PPE), rather than basing their results on the CMIP3 ensemble of different models (multi-model ensemble, MME). I was surprised to see this presented as such a major decision, as my recollection is that most of the critics at the time were really complaining about the willingness of UKMO to generate probabilistic predictions at all since (the critics argued) there was not really a sound basis for assigning numerical values by any method.

The main UKCP09 proposal was (according to their web page) funded in 2004 and at that time, it seemed quite widely accepted that PPEs were a better foundation for probabilistic prediction than the MME. In fact this era was very much the  heyday of PPEs, with climateprediction.net, the Hadley Centre’s QUMP group and our own rather smaller ensemble research activity all making rapid progress. The UKCP09 approach was externally reviewed back in 2008/9. The full review doesn’t seem to be available (anyone know where it is?) but I don’t see any evidence in either the summary or response that the question of MME vs PPE was seriously raised by anyone even at that later time.

I believe (though I could be wrong and would welcome references) that we were actually the first to argue the contrary. The roots of our argument can be found in this Yokohata et al paper (which although published in 2010 was submitted back in 2008), which pointed to substantial inconsistencies between two PPEs based on our two different GCMs (MIROC and HadCM3). However it was actually our series of papers on ensemble analysis starting in 2010 (eg here, here, here and here) that most clearly argued not only that PPEs had serious problems, but also that the MME was much better than previously believed. So while I’m encouraged to see that this question is now high on the agenda, I really don’t think it was on the table at the outset of UKCP09 and it doesn’t really seem fair to use it as evidence of insularity or reflexive dismissal of outside ideas, which seemed to be the speaker’s point. Given the work they had already done by 2010 or so, the UKCP09 researchers actually made quite substantial and constructive efforts to account for the (by then) emerging failings of the PPE approach by effectively adding on the MME’s uncertainty to their results. While this may satisfy neither the resolutely anti-Bayesian nor the most purist pro-Bayesian, in my view it certainly improved the credibility of their results.

Some of the interviewees gave excuses for their apparent reticence to air their doubts openly at the time. According to Porter, some of them said they were scared of being labelled sceptics! What a feeble excuse. Perhaps more plausible, is the additional argument that the incestuous and cliquey nature of climate science in the UK made it a bit of a career risk in terms of future funding. But in any case, I certainly recall some people making their criticisms very plain. In particular, Lenny Smith argued eloquently about the risks of assigning probabilities where there was not really a sound basis for them. If the next model generates different results (which is entirely plausible) then someone is going to end up looking rather silly.

So I’m not going to stick the knife into the Hadley Centre for proposing in 2004 to base their probabilistic predictions on a PPE methodology that they had already started to work on. You would have had to be unusually prescient to anticipate our research by several years, although I’m encouraged to see it is now obviously high on peoples’ minds. On the other hand, the Hadley Centre’s apparent continuing preference for PPEs is hard to defend, now that they have a chance to regroup. To that extent, perhaps this Project ICAD analysis contains a truth that is deeper than the actual story they purported to tell :-)

Friday, October 03, 2014

Much ado about sensitivity

Well, perhaps not really very much ado. There's a new paper in Climate Dynamics, by Lewis and Curry, with a central sensitivity estimate of 1.6C with a 90% range of 1-4C, based on energy budget analyses over the instrumental period, updated to the present day, also taking account of the newer AR5 forcing estimates. I don't find it particularly exciting, the authors cite several recent papers with similar results including Aldrin and Otto et al. I wrote about those papers some time ago, and I think these posts (1, 2, 3) still stand. I've commented before on my objections to Lewis' method, and especially the sleight-of-words with which it is described, but (as I've also emphasised) I don't think this substantially affects the results in this application.

Clearly, the longer the relatively slow warming continues, the lower the estimates will go. And despite what some people might like to think, the slow warming has certainly been a surprise, as anyone who was paying attention at the time of the AR4 writing can attest. I remain deeply unimpressed by the way in which this embarrassment has been handled by the climate science insiders, and IPCC authors in particular. Their seemingly desperate attempts to denigrate anything that undermines their storyline (even though a few years ago the same people were using markedly inferior analyses of this very type to bolster it!) do them no credit.
One weakness of these energy-budget type of analyses, that I believe Lewis and others could easily address, is to demonstrate how well it works in application to GCM output. That is, can the method accurately diagnose the sensitivity of a model given equivalent information to that which we have for the real world? Aldrin et al addressed this rather briefly and in a very limited way, using a far stronger forcing scenario (1%pa CO2 enrichment) than what has occurred in reality. It would be easy to investigate the precision of the method, and whether it gives rise to any systematic biases, by using output from the more realistic 20th century simulations. It is also noteworthy that the Aldrin method struggles to cope with hemispheric differences, which may point to some limitations of the energy balance concept. While the climate system certainly does obey the fundamental conservation laws, supposedly “fixed” parameters (in simple models) are not actually constants in reality. And no matter now precisely we can determine the historical transient response to the current radiative imbalance, there will always be a bit of additional uncertainty in extrapolating that to an equilibrium 2xCO2 state.
Finally, it is also amusing to see Judith “we don't know anything” Curry to put her name to this new paper: it is unclear what she might have added, as Nic has been presenting analyses of this nature for some time now. But that's a minor matter.

Tuesday, April 22, 2014

Objective probability or automatic nonsense?

A follow-up to the previous probability post.

Perhaps this will provide a clearer demonstration of the limitations of Nic's method. In his post, he conveniently provided a simple worked example, which in his view demonstrates how well his method works. A big advantage of using his example is that hopefully no-one can argue I've misapplied his method :-) This is Figure 2 from his post:



This example is based on carbon-14 dating, about which I know very little, but hopefully enough to explain what is going on. The x-axis in the above is real age with 0 corresponding to the "present day", which I think is generally defined as 1950 (so papers don't need to be continually reparsed as time passes). The y-axis is "carbon age" which is basically a measure of the C14 content of something under investigation, typically something organic (plant or animal). The basic idea is that the plant or aminal took up C14 as it grew, but this C14 slowly decays so the proportion in the sample declines after death according to the C14 half-life. So in principle you would think that the age (at death) can be determined directly from measurement of the proportion of carbon that is C14. However, the proportion of C14 in the original organism depends on the ambient concentration of C14 which has varied significantly in the past (it's created by cosmic rays and the like), so there's quite a complicated calibration curve. The black line in the above is a simplified and stylised version of what a curve could look like (Nic's post also has a real calibration curve, but this example is clearer to work with).

So in the example above, the red gaussian represents a measurement of radiocarbon which represents a "carbon age" of about 1000y, with some uncertainty. This is mapped via the calibration curve into a real age distribution on the x-axis, and Nic has provided two worked examples using a uniform prior and his favoured Jeffreys prior.

As some of you may recall, I lived in Japan until recently. Quite by chance, my home town of Kamakura was the capital of Japan for a brief period roughly 7-800y ago. Lots of temples date from that time, and there are numerous wooden artefacts which are well-dated to the Kamakura Era (let's assume, carved out of conteporaneous wood, though of course wood is generally a bit older than the date of the tree felling). Let's see what happens when we try to carbon-date some of these artefacts using Nic's method.

Well, one thing that Nic's method will say with certainty is "this is not a Kamakura-era artefact"! The example above is a plausible outcome, with the carbon age of 1000y covering the entire Kamakura era. Nic's posterior (green solid curve) is flatlining along the axis over the range 650-900y, meaning zero probability for this whole range. The obvious reason for this is that his prior (dashed line) is also flatlining here, making it essentially impossible for any evidence, no matter how strong, to overturn the prior presumption that the age is not in this range.

It is important to recognise that the problem here is not with the actual measurement itself. In fact the measurement shown in the figure indicates very high likelihood (in the Bayesian sense) of the Kamakura era. The problem is entirely in Nic's prior, which ruled out this time interval even before the measurement was made - just because he knew that a measurement of carbon age was going to be made!

Nic uses the emotionally appealing terminology of "objective probability" for this method. I don't blame him for this (he didn't invent it) but I do wonder whether many people have been seduced by the language without understanding what it actually does. You can see Richard Tol insisting that the Jeffreys prior is "truly uninformative" in a comment on my previous post, for example. Well, that might be true, but only if you define "uninformative" in a technical sense not equivalent to common english usage. If you then use it in public, including among scientists who are not well versed in this stuff, then people are going to get badly misled. Frame and Allen went down this rabbit hole a few years ago, I'm not sure if they ever came out. It seems to work for many as an anchoring point, when you discuss in detail, they acknowledge that yes, it's not really "uninformative" or "ignorant" really, but then they quickly revert back to this usage, and the caveats somehow get lost.

I propose that it would be better to use the term "automatic" rather than "objective". What Nic is presenting is an automatic way of generating probabilities, though it remains questionable (to put it mildly) whether they are of any value. Nic's method insists that no trace remains of the Kamakura era, and I don't see any point in a probabilistic method that generates such obvious nonsense.

Friday, April 18, 2014

Coverage

Or, why Nic Lewis is wrong.

Long time no post, but I've been thinking recently about climate sensitivity (about which more soon) and was provoked into writing something by this post, in which Nic Lewis sings the praises of so-called "objective Bayesian" methods.

Firstly, I'd like to acknowledge that Nic has made a significant contribution to research on climate sensitivity, both through identifying a number of errors in the work of others (eg here, here and most recently here) and through his own contributions in the literature and elsewhere. Nevertheless, I think that what he writes about so-called "objective" priors and Bayesian methods is deeply misleading. No prior can encapsulate no knowledge, and underneath the use of these bold claims there is always a much more mealy-mouthed explanation in terms of a prior having "minimal" influence, and then you need to have a look at what "minimal" really means, and so on. Well, such a prior may or may not be a good thing, but it is certainly not what I understand "no information" to mean. I suggest that "automatic" is a less emotive term than "objective" and would be less likely to mislead people as to what is really going on. Nic is suggesting ways of automatically choosing a prior, which may or may not have useful properties.
[As a somewhat unrelated aside, it seems strange to me that the authors of the corrigendum here concerning a detail of the method, do not also correct their erroneous claims concerning "ignorant" priors. It's one thing to let errors lie in earlier work - no-one goes back and corrects minor details routinely - but it is unfortunate that when actually writing a correction about something they state does not substantially affect their results, they didn't take the opportunity to also correct a horrible error that has seriously mislead much of the climate science community and which continues to undermine much work in this area. I'm left with the uncomfortable conclusion that they still don't accept that this aspect of the work was actually in error, despite my paper which they are apparently trying to ignore rather than respond to. But I'm digressing.]

All this stuff about "objective priors" is just rhetoric - the term simply does not mean what a lay-person might expect (including a climate scientist not well-versed in statistical methodology). The posterior P(S|O) is equal to to the (normalised) product of prior and likelihood - it makes no more sense to speak of a prior not influencing the posterior, as it does to talk of the width of a rectangle not influencing its area (= width x height). Attempts to get round this by then footnoting a vaguer "minimal effect, relative to the data" are just shifting the pea around under the thimble.

In his blog post, Nic also extolls the virtue of probabilistic coverage as a way of evaluating methods. This initially sounds very attractive - the idea being that your 95% intervals should include reality, 95% of the time (and similarly for other intervals). There is however a devil in the detail here, because such a probabilistic evaluation implies some sort of (infinitely) repeated sampling, and it's critical to consider what is being sampled, and how. If you consider only a perfect repetition in which both the unknown parameter(s) and the uncertain observational error(s) take precisely the same values, then any deterministic algorithm will return the same answer, so the coverage in this case will be either 100% or 0%! Instead of this, Nic considers repetition in which the parameter is fixed and the uncertain observations are repeated. Perfect coverage in this case sounds attractive, but it's trivial to think of examples where it is simply wrong, as I'll now present.

Let's assume Alice picks a parameter S (we'll consider her sampling distribution in a minute) and conceals it from Bob. Alice also samples an "error" e from the simple Gaussian N(0,1). Alice provides the sum O=S+e to Bob, who knows the sampling distribution for e. What should Bob infer about S? Frequentists have a simple answer that does not depend on any prior belief about S - their 95% confidence interval will be (S-2e,S+2e) (yes I'm approximating negligibly throughout the post). This has probabilistically perfect coverage if S is held fixed and e is repeatedly sampled. Note that even this approach, which basically every scientist and statistician in the world will agree is the correct answer to the situation as stated, does not have perfect coverage if instead e is held fixed and S is repeatedly sampled! In this case, coverage will be 100% or 0%, regardless of the sampling distribution of S. But never mind about that.

As for Bayesians, well they need a prior on S. One obvious choice is a uniform prior and this will basically give the same answer as the frequentist approach. But now let's consider the case that Alice picks S from the standard Normal N(0,1), and tells Bob that she is doing so. The frequentist interval still works here (i.e., ignoring this prior information about S), but Bayesian Bob can do "better", in the sense of generating a shorter interval. Using the prior N(0,1) - which I assert is the only prior anyone could reasonably use - his Bayesian posterior estimate for S is the Normal N(O/2,0.7), giving a 95% probability interval of (O/2-1.4,O/2+1.4). It is easy to see that for a fixed S, and repeated observational errors e, Bob will systematically shrink his central estimates towards the prior mean 0, relative to the true value of S. Let's say S=2, then (over a set of repeated observations) Bob's posterior estimates will be centred on 1 (since the mean of all the samples of e is 0) and far more than 5% of his 95% intervals (including the full 27% of cases where e is more negative than -0.6) will fail to include the true value of S. Conversely, if S=0, then far too many of Bob's 95% intervals will include S. In particular, all cases where e lies in (-2.8,2.8) - which is about 99.5% of them - will generate posteriors that include 0. So coverage - or probability matching, as Nic calls it - varies from far too generous, when S is close to 0, to far too rare, for extreme values of S.

I don't think that any rational Bayesian could possibly disagree with Bob's analysis here. I challenge Nic to present any other approach, based on "objective" priors or anything else, and defend it as a plausible alternative to the above. Or else, I hope he will accept that probability matching is simply not (always) a valid measure of performance. These Bayesian intervals are unambiguously and indisputably the correct answer in the situation as described, and yet they do not provide the correct coverage conditional on a fixed value for S

Just to be absolutely clear in summarising this - I believe Bayesian Bob is providing the only acceptable answer given the information as provided in this situation. No rational person could support a different belief about S, and therefore any alternative algorithm or answer is simply wrong. Bob's method does not provide matching probabilities, for a fixed S and repeated observations. Nothing in this paragraph is open to debate.

Therefore, I conclude that matching probabilities (in this sense, i.e. repeated sampling of obs for a fixed parameter) is not an appropriate test or desirable condition in general. There may be cases where it's a good thing, but this would have to be argued for explicitly.

Tuesday, October 23, 2012

New Italian Earthquake Prediction

Another Empty Blog exclusive: Italian seismologists have just issued a new assessment of earthquake risks. You can't say you weren't warned now!

Among all the outrage (surrounding this, for anyone who slept through it), there are more nuanced views (expressed prior to the verdict) about whether the scientists' statements were negligently falsely confident rather than just being unfortunate. Irrespective of whether they could have been expected to predict the quake, "absolutely no risk" is an unfortunate choice of words.

One predictable outcome is that Italian seismologists (and presumably scientists in other fields) will be rather less willing to proffer risk-relevant advice in any sort of official capacity, at least in Italy. Hard to see this sort of trial catching on in Japan, where such risk management failures are seen as a cultural imperative.

Monday, June 25, 2012

BayesComp2012

David Benson asked what we were doing in the company of mathematicians at Tachikawa...

There's a huge Bayesian stats meeting in Kyoto this week, which would probably have been quite interesting but which I thought was a little too far away both geographically, and topically, to be worth attending, especially right now as I'm pretty busy.

Fortunately, a satellite meeting was arranged last Friday/Saturday, hosted by the Institute of Statistical Mathematics in Tachikawa (which I once visited at its previous site in Hiroo in central Tokyo, before they moved). Some people there are also working on climate change related projects, possibly the same projects we are working on though there seems to be some overlap/duplication between the various ministries and institutes as to who is doing what! Several of the eminent attendees of the Kyoto meeting were somehow persuaded to come a few days early and visit the fleshpots of Tachikawa - I hope they thought it was worthwhile - but for us it was a great opportunity to hear what is going on in the latest research into computational methods for the Bayesian paradigm (mostly Markov Chain and Sequential Monte Carlo methods). And it was also a good excuse to visit the new Tachikawa site of ISM, and realise that it's one more place where we really wouldn't much want to work when our time here in Yokohama is up :-)

The meeting was, as expected, a little obscure and distant from our work - which confirmed my decision to not go to Kyoto - but was well worth spending a couple of days on. One or two bits were particularly interesting - especially the methods for estimating very small probabilities (down to 10-120 or even 10-200), which may be relevant to our future plans, now we are in a post-Fukushima world and being urged to plan for the unimaginable...

Monday, November 07, 2011

The null hypothesis in climate science

Three papers have just appeared in WIREs Climate Change (here, here and here) discussing the role of the null hypothesis in climate science, especially detection and attribution.

Trenberth argues that, since the null (that we have not changed the climate) is not true, we should try to test some other null hypothesis. He sounds like someone who has just discovered that the frequentist approach is actually pretty useless in principle (as I've said many times before, it is fundamentally incapable of even addressing the questions that people want answers to), but although he seems to be grasping towards a Bayesian approach, he hasn't really got there, at least not in a coherent and clear manner. Curry is just nonsense as usual, and beside noting that she has (1) grossly misrepresented the IAC report and (2) abjectly failed to back up the claims that Curry and Webster made in a previous paper, there isn't really anything meaningful to discuss in what she said.

Myles Allen's commentary is by some distance the best of the bunch, in fact I broadly agree (shock horror) with what he has said. If one is going to take a frequentist approach, the null hypothesis of no effect is often an entirely reasonable starting point. It is important to understand that rejecting the null does not simply mean learning that there has been some effect, but it also indicates that we know (at least at some level of confidence) the direction of the effect! That is, it is not only an effect of zero which is rejected, but all possible negative (say) effects of any magnitude too - this generalisation may not be strictly correct in all possible applications of this sort of methodology, but I'm pretty sure it is true in practice for the D&A field. Especially when we are talking about the local incidence of extreme weather, there really are many cases when we have little reason for a prior belief in an anthropogenically-forced increase versus a decrease in these events, so a reasonable Bayesian approach would also start from a prior which was basically symmetric around zero. The correct interpretation of a non-rejection of the null here is not "there has been no effect" but rather "we don't know if AGW is making these events more or less likely/large". Much of Trenberth's complaint could be more productively aimed at the routine misinterpretation of D&A results, rather than the method of their generation. Trenberth also sometimes sounds like he is arguing that we should always assume that every bad thing was caused by (or at least exacerbated by) AGW, but this simply isn't tenable. Even if storminess increases in general, changes in storm tracks might lead to reduction in events in some areas, with Zahn and von Storch's work on polar lows an obvious example of this. On the other hand, there are also some types of event where we may have decent prior belief in the nature of the anthropogenically-forced change (such as temperature extremes) and in these cases it would be reasonable for a Bayesian to use a prior that reflects this belief.

I can find one thing to object to in Myles' commentary though, and that's the manner in which he tries to pre-judge the "consensus" response to Trenberth's argument. Noting that he (Allen) is in fact a major figure in forming the "consensus" in these private meetings where the handful of IPCC authors decide what to say, it sounds to me rather like a pre-emptive strike against anyone who might be tempted to take the opposing view. I would prefer it if he restricted himself to arguing on the basis of the issues rather than that he holds/forms the majority view. His behaviour here is reminiscent of the way he (and others) tried to reject our arguments about uniform priors, on the basis that everyone had already agreed that his approach was the correct solution. All that achieved was to slow the progress of knowledge by a few years.

Friday, November 04, 2011

Curry on fuzzy logic

Before I get on to the meat of some more new papers...

I noticed not so long ago Curry and Webster flying a kite about fuzzy logic being a better alternative to Bayesian probability, in the context of D&A:

The logic of the IPCC AR4 attribution statement is discussed by Curry (2011b). Curry argues that the attribution argument cannot be well formulated in the context of Boolean logic or Bayesian probability. Attribution (natural versus anthropogenic) is a shades-of-gray issue and not a black or white, 0 or 1 issue, or even an issue of probability. Towards taming the attribution uncertainty monster, Curry argues that fuzzy logic provides a better framework for considering attribution, whereby the relative degrees of truth for each attribution mechanism can range in degree between 0 and 1, thereby bypassing the problem of the excluded middle.


As you will recall, I've been waiting for a year now for Curry to explain her muddled and confused approach to probability, in particular her nonsensical "Italian Flag" analysis which she seems to be recasting as "fuzzy logic" (as an aside, I do agree that her logic is fuzzy, but perhaps not in the way she intended).

So I was eagerly awaiting "Curry (2011b)", which has just appeared. And what does it say about fuzzy logic?

[fx: tumbleweed]

Not one single mention, that's what. No mention of Bayesian probability, either. Or Boolean logic. These terms are completely absent from the paper, so this whole line of specious assertions has simply been abandoned without any support whatsoever.

Solution to the paradox of climate sensitivity

A lot of bloggable papers have suddenly appeared, so I will work through them over the next few days.

First, a quick comment about this interesting paper: "Solution to the paradox of climate sensitivity" by Salvador Pueyo. In it, he argues that we should use a log-uniform prior for estimating climate sensitivity. This is fundamentally an "Objective Bayes" approach, that "non-informative" can be interpreted in a unique way. I don't much like this point of view, but if one is going to take it, then it should at least be done properly, and he seems to have provided decent arguments in that direction. Readers may recall that IPCC authors have in the past claimed that a uniform distribution was the unique correct representation of ignorance, which formed one of the planks of their assessment of the literature in the AR4.

As we showed here, all this talk of a long tail basically vanishes when anything other than a uniform prior is used, so in that sense this new paper is broadly compatible with our existing results which were based on a subjective paradigm. However, I'm not sure how it would work with a more complex multivariate approach, as has been common in this sort of work (eg simultaneously considering the three major uncertainties of ocean heat uptake, aerosol forcing and sensitivity).

What the new IPCC authors will make of it all is anyone's guess. Perhaps we will find out in December some time, when the first draft is scheduled to be opened for comments.

Sunday, August 21, 2011

Roger Rabbit

He just can't stop himself, burrowing away (don't much like the idea of a Curry hole, euuuurgh.)

It's quite amusing to watch the contortions he'll go to in order to avoid admitting a mistake. Recall that this started with his novel idea that one could determine the "correctness" of a probabilistic prediction of an event, by whether the event in question actually happens. Eg the prediction "likely to rain tomorrow" is correct if and only if the rain actually falls tomorrow.

While this might sound intuitively appealing, it quickly falls apart under any careful examination (as Doswell and Brooks warn). That is, it leads to conclusions that are obviously nonsensical and/or inconsistent. For example, if we say that a roll of a fair die is likely to come up 1-5, then this statement is correct in the sense of, well, being correct, but Roger's analysis would determine it to have been false if the roll actually turned out to be 6.

Oh, but at this point, rather than admitting that his usage of "correct" made no sense, Roger decided that for some reason his method only applies in truly epistemic cases where probability is a state of belief rather than a long-run property. It's funny that while (dishonestly) accusing me of making the IPCC out to be infallible, he then tries his best to ensure that his personal "correctness" theory is unfalsifiable. But I'm sure he is blind to that irony. Of course, no explanation is forthcoming as to why his theory, if it is useful and valid, should fall flat so quickly when confronted with a simple example. I tried again with a handmade imperfect die which is initially not known to be fair, but for which I still make the same prediction and again throw a 6. In Roger-world the probabilistic prediction is incorrect. However, in this case the long-run frequency of a 6 can subsequently found by experiment, and let's assume it turns out to be 20±0.1%. Was the original probabilistic statement still Roger-incorrect? Answer came there none...

Best of all, entirely unprompted, he came up with an example based on an asteroid falling on Boulder. While he had several times insisted that a prediction at the 90% level should be considered "incorrect" if the event did not occur, he then stated that if I predict that it is 10% probable that an asteroid hits Boulder tomorrow (ie 90% probable that it does not), then my prediction is correct if the asteroid DOES hit! This, he explains, is due to the "baseline expectation" which apparently allows Roger to invert his original definition of "correctness" whenever he feels like it. It's a bit odd that he came up with this new twist completely unprompted, as it blows apart all his previous analysis, but it's not as if his theory made any sense anyway.

Naturally, the actual paper that he co-authored contains no mention of this "baseline expectation".

With his latest post on aleatory and epistemic uncertainty, one might hope that he could have at last been starting to realise that the concept of "correctness" of a probabilistic prediction cannot in general be determined from the occurrence - or otherwise - of the predicted event (the occurrence of an event assigned a probability of zero is of course an exception). But based on the comments, it seems that this insight still eludes him.

It does seem that one infallible guide to "Pielkeian correctness" has emerged, though. If Roger says it, then it is correct, no matter how many impossible or ridiculous contortions and evasions are required to avoid admitting error.

Thursday, August 18, 2011

Probabilistic Forecasting - A Primer

Roger continues to flounder away, trying to salvage something from his latest statistical train-wreck. It's all remarkably trivial for someone who claims to "fully understand verification of probabilistic forecasts". Commenter Steve Scolnik skewers him neatly with a quotation from Doswell and Brooks:

"An important property of probability forecasts is that single forecasts using probability have no clear sense of "right" and "wrong." That is, if it rains on a 10 percent PoP forecast, is that forecast right or wrong? Intuitively, one suspects that having it rain on a 90 percent PoP is in some sense "more right" than having it rain on a 10 percent forecast. However, this aspect of probability forecasting is only one aspect of the assessment of the performance of the forecasts. In fact, the use of probabilities precludes such a simple assessment of performance as the notion of "right vs. wrong" implies. This is a price we pay for the added flexibility and information content of using probability forecasts. Thus, the fact that on any given forecast day, two forecasters arrive at different subjective probabilities from the same data doesn't mean that one is right and the other wrong! It simply means that one is more certain of the event than the other. All this does is quantify the differences between the forecasters."
Of course, there isn't a cigarette-paper of difference between what I was saying, and what Doswell and Brookes are saying, because this is all well-established basic stuff.

Nevertheless, RP chooses to make up nonsense and misrepresent what I said, without even having the decency to link to my post. I nowhere say, or imply that the IPCC statements "could not be judged to be wrong because of their probabilistic nature", indeed as he well knows I have explicitly contradicted this nonsense claim of his multiple times in the past. A single probabilistic statement at the "likely" level cannot generally meaningfully be validated because no outcome is sufficiently improbable to falsify it (under the standard significance testing paradigm). Once you have a large enough ensemble of statements, such as those the IPCC make, their judgement as a whole can easily be validated because it is highly improbable that either a small or large number of the particular events should occur, if the probability was accurate.

(Even this approach suffers from the usual problems of frequentist statistics, in that it does not actually address the issue of "how likely is it that the probabilistic system is well calibrated, given these results" but rather answers "how likely are these results, if the probabilistic system is well calibrated". However, if the probability level is small enough, we can safely reject the system anyway. This digression is probably best ignored by all readers, I just put it in to head off another avenue for nit-picking.)

Even his own analogy, he fails to be consistent with himself. Having stated unequivocally that a large proportion of the findings of the IPCC are "incorrect" he admits wrt some hypothetical bet on a football game:

"It is important to understand that the judgment [A] may have been perfectly sound and defensible at the time that it was made ... Perhaps then the outcome was just bad luck, meaning that the 10% is realized 10% of the time. Actually, we can never know the answer to whether the expectation was actually sound or not"

So he's prepared to consider probabilistic judgements "perfectly sound" when it suits him, but "incorrect" whenever the IPCC make them. Uh-huh.

Friday, August 12, 2011

How many of Roger's findings about probability manage to be wrong? Answer: he's more inventive than you might expect.

Roger Pielke has a new post up asserting that 28% of the IPCC's findings are incorrect. Although it's obviously a rather implausible figure, I was expecting this claim to be backed up with some sort of evidence of errors, or at least sloppiness, or something, so I had a look at his paper that he cites to justify the claim.

It turns out that 100%-28% = 72% is merely the average (lower bound) probability level associated with the statements they made. Such as "It is very likely that hot extremes, heat waves and heavy precipitation events will continue to become more frequent." Here "very likely" means greater than 90%. So, given 10 such statements, the IPCC is saying that they would expect the "very likely" outcome to occur about 9 times, and not occur about once. And similarly for "likely" (66%). Averaging over all the probabilistic statements, it should be expected that in about 28% of cases, the (probabilistically) preferred outcome will not actually happen.

And in Roger-world, this means that 28% of the statements are "incorrect". Note, however, that he does not make this silly claim in the paper itself, but only in his blog post.

To see why this interpretation is nonsensical, consider a single roll of a fair die. I state (accurately) that it is "likely" to lie in the range 1-5. If I roll a 6, then in Roger-world, my statement was incorrect. However, it was not incorrect, and Roger is simply wrong to claim so.

As you can see from the comments, I challenged Roger on this, and his response (entirely in character) is to duck and weave. In his comment #5, for example, he shamelessly misrepresents what I said, and brings up the red herring of a definitive prediction (when in fact I had clearly made a probabilistic one, and the distinction is of course absolutely fundamental to the point). The obvious elephant in the room that Roger cannot bring himself to acknowledge is that the statement is correct irrespective of the outcome of the roll. "Correctness" of a single statement simply isn't something that can be directly validated (or invalidated) by the outcome, and the accurate calibration of a probabilistic prediction system actually relies on having the appropriate number of "failures" for each level of probability.

I realise of course that having done some rather boring textual analysis that in his own words amounts to "Nothing too interesting, really", Roger is just rabble-rousing on his blog. I'm confident that any competent scientist will see straight though it, but that's hardly his target audience.

As for what the 72%/28% average actually does mean, it doesn't actually tell us anything except that the IPCC makes a lot of statements about things that it is only (by its own admission) moderately confident about. It might in principle be interesting to see how the confidence level changes over time, but only if the set of statements were to be held fixed from one assessment to the next. People have looked at climate sensitivity estimates (hardly changed) and detection and attribution (increased markedly in confidence) but not a lot else AIUI. I suppose we can anticipate Roger claiming that the next report is either more correct, or less, depending on what mix of statements they happen to include :-)

Incidentally, and although it's a minor point it is perhaps telling in terms of his overall level of competence, Roger is also wrong where he claims that if the statements are not independent, then the proportion of "incorrect" will be higher than 28%. Actually, if the statements are not independent (while still being correctly calibrated), then the proportion that do not come to pass would still be 28% in expectation, just with higher variance, meaning that either a larger or smaller proportion would not be surprising. Unlike the simple misinterpretation in his blog post title, this elementary error is actually made in the paper itself.

Wednesday, July 06, 2011

Priors and climate sensitivity again

Several people have email about this article. I don't have anything particularly novel or interesting to say, so I'll just repeat an email that I sent regarding it...




From my point of view, the problem is not particularly in the treatment of the Forster and Gregory result - the authors had already in that paper pointed to the choice of prior as an important factor in the specific results they generated. More, the error was in the IPCC's endorsement and rigid adherence to the use of uniform prior, despite the existence of very straightforward arguments that this approach is simply not tenable:

http://www.springerlink.com/content/7np5t35mq27p3q24/
(also here:
http://www.jamstec.go.jp/frsgc/research/d5/jdannan/probrevised.pdf )

These arguments (which as you saw I made during the IPCC review process [here here here]) were basically brushed aside. The IPCC authors exclusively relied on and highlighted the results that had been generated using uniform priors, and downplayed alternative results, which already existed in the literature, that had been generated with different priors.

However, with the passage of time I believe my arguments have now become more widely (if grudgingly) accepted, so I look forward with some interest to see how the IPCC authors deal with the subject this time.





I should also add that I'm not at all convinced by the author's claims that a prior which is uniform in feedback (1/sensitivity) is "correct", rather, it is something that people have to think about, and may reasonably disagree. Such is life. It is theoretically possible that someone could even present a plausible argument for a uniform prior in sensitivity, but I've not yet seen one...

Tuesday, June 21, 2011

Statistically significant

Apparently global warming is statistically significant again.

But we all know that the difference between "significant" and not significant is not itself statistically significant, don't we?

Richard Black is usually pretty good, so it's a shame to see the old canard "If a trend meets the 95% threshold, it basically means that the odds of it being down to chance are less than one in 20." Of course, you all know why that's not true (at least, if you don't, you will after reading this).

Friday, April 29, 2011

The Cult of Statistical Significance?


I was listening to a recent More or Less which had a piece about statistical significance. The guest was Stephen Ziliak who has a book on the topic. I actually thought he gave a slightly confusing account of the limitations of significance testing ("likelihood of the magnitude"?). His book also has a lot of hostile reviews on Amazon suggesting it reads a bit like a blog rant. Perhaps this Gerd Gigerenzer article is better written.

The reason for the More or Less article was a recent US Supreme Court decision that medical trial results could not be brushed under the carpet simply due to their being "statistically insignificant". In the case in question, it seems that there might have been prior reasons to suspect side effects of the type observed, so the fact that they had not (at that time) reached an arbitrary threshold is not adequate justification for concealing them.

I've mentioned before, IMO most of the confusion over significance testing is that the p-value actually doesn't answer the question people are interested in (probability of a hypothesis being true), but is routinely misinterpreted in that way. The same confusion extends to confidence intervals, of course, and these errors are routinely found even in articles that claim to be authoritative (eg and of course also here). But I wouldn't call it a cult, it's more likely to be down to confused thinking and laziness on the whole.

And also, as several people spotted, on xkcd:



Friday, March 04, 2011

Pigeons vs sciencebloggers, round two

And now for the follow-up.

The question was, if someone has two children, and tells you that (at least) one of them is a boy born on a Tuesday, what is the probability that the other child is also a boy? And the intended answer, of course, is 13/27. Why? Well, including the birth day (of week) as a variable, at the outset for a two-child family there are 142 = 196 possible and equiprobable pairs such as (B-Mon, G-Wed) where each child has a gender and birthday (of week) and the ordering denotes birth order. The vast majority of these pairs don't have a Tuesday boy and can be ignored. Of those that do, 7 of them look like (B-Tue, G-any) and another 7 are (G-any, B-Tue). Similarly, there are 7 pairs (B-Tue, B-any) and 7 (B-any, B-Tue). But wait! The case (B-Tue, B-Tue) has been counted twice! In fact there are only 13 cases where one child is a Tuesday boy and the other is also a boy. So that makes the probability 13/(13+14) = 13/27 where one child is a Tuesday boy and the other is also a boy.

However, there's a sting in the tail. Let's say you take a bunch of people with two children, and ask them if one of their children is a boy. Pick such a parent at random. We have already seen that in 1/3 of cases where the answer was yes, the other child will also be a boy. Now ask this parent for the day of the week that their son was born on (if they have a choice, they can flip a coin and pick one son at random, preferably out of sight so as not to give the game away). You get the reply..."Tuesday". Or some other day. Maybe they will say Saturday. In this case, hearing the day of the week on which a son was born does *not* change the probability that the other child is also a boy. So anyone who interpreted the original situation as meaning "a parent of two children, at least one of whom is a boy, tells you the day of the week on which their son (or a random son if applicable) was born, and their answer was `Tuesday'" would be completely justified in concluding that the probability of the other child being a boy was still 1/3. This solution was mentioned in More or Less on 11 June, and it was (re)listening to this old podcast (probably no longer available, but here's a related web page) that prompted me to blog it at last. What Tim Harford could have gone on to point out, but didn't (IIRC), is that even the original "two children" problem is typically under-determined in the way that it is presented. If a parent of two children is asked for the gender of one randomly-selected child, then irrespective of their reply, the probability of their other child being a boy (or alternatively, being the same gender as the one they gave) is...50%. So the solution to this problem also depends on how this "one child is a boy" parent is found.

Similar weaknesses can be found in most statements of the Monty Hall problem too. If all that is reported is the observed actions of the game show host on one occasion, then we don't really have enough information to generalise to a rigorous frequency-based calculation. Maybe Monty only opens another door on the occasions that the player originally picked the car, in which case swapping will lose. Maybe Monte opens a random door (and might expose the car)...in which case the player should still swap, and in fact the probabilities are unchanged from the original problem, in fact. Wikipedia discusses several alternative interpretations in some detail.

Contrary to Gary Foshee's statement on this web page, I don't think this sort of ambiguity is particularly controversial, it's just the result of trying to dress up a mathematical problem in natural-sounding (but slightly imprecise) English. When the problem is stated unambiguously, it's not particularly difficult. At least for pigeons.

And to any pigeons still reading, all I can say is: coo.

Wednesday, March 02, 2011

Are pigeons smarter than sciencebloggers?

I've been meaning to write about Monty Hall for ages. This is the well-worn probability problem based on a game show. There are three closed doors, with a car hidden behind one of them, and goats behind the other two. The contestant chooses a door, and the game show host then opens a different door, which has a goat behind it. The contestant is then offered the choice to swap to the other door, or stick with the original one. If the contestant settles on the door with the car behind it, they win it.

The obvious question is, should they swap, and what is the probability of winning if they do/do not swap? This problem has been published in magazines, resulting in decades of correspondence and controversy (ok, not decades, probably).

The "wrong" answer (and the scare quotes are appropriate) is that once one door has been opened, the prize is equiprobably behind the other two and it doesn't make any difference if the contestant swaps or not. The "correct" answer is that they should swap, as the probabilities are easily calculated to be 2/3 and 1/3 respectively.

The real answer is that it depends. In particular, it depends on the problem being very precisely stated, or alternatively, it depends on the assumptions that people make in the absence of a clear statement. As I wrote it above, it's actually a little unclear, though I wouldn't criticise anyone who made the obvious assumptions and who came up with the "right" answer. Wikipedia has a decent discussion of the ambiguities and history of the problem.

Notably, even when the problem is precisely specified, a large majority of people still get it wrong. Which only goes to show that people, even supposedly intelligent and educated ones, can still have a complete mental block where probability is concerned.

Another popular problem is the following: a person has two children, at least one of which is a boy. What is the probability that the other child is also a boy? Again, there is a wrong (but common) answer of 1/2, and a "right" answer of 1/3. Again, the problem is typically ambiguous in its statement. A more interesting version of it appeared on blogs and the "More or less" BBC Radio4 program recently: a person has two children, at least one of which is a boy born on a Tuesday. What is the probability that the other child is also a boy?

To give you the fun of thinking about it yourselves, I'll not not include any links and continue this in a later post...

And the title? I couldn't resist the juxtaposition of the fact that someone thought it was worth writing a whole book about this simple maths problem that pigeons can solve!