Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Friday, 15 June 2012

Information and statistics - thoughts following Kevin McConway's Inaugural Lecture

Kevin McConway, Open University Professor of Applied Statistics, in his Inaugural Lecture: Statistical thinking: the good, the bad and the ugly, "explored how, for many people statistics is about firm, objective answers that are either right or wrong while one of its main goals is coping with uncertainty, not certainty".

The lecture should be available here http://stadium.open.ac.uk/berrill/ in due course, and is well worth watching: informative, authoritative and entertaining.

I think he only used the word information once and that in reference to the content of a website, but, for me, his talk was all about information. But then I see information ideas everywhere these days.

For one thing, he talked of data and knowledge (and 'facts', 'theory', 'opinion', 'evidence', 'truth', 'lies') which immediately put me in mind of the DIKW* hierarchy with information a notable absence (my issues with DIKW notwithstanding). Data in this context is largely taken to be numbers, and it is the job of statistics and statisticians to turn this into something more useful: knowledge, meaning, or, maybe, information. So perhaps we have a 'special case' of the trapezium, with the input at the bottom as numerical data and the trapezium as the statistics/statistician. I'm increasingly thinking, though, of information as somehow this process of getting meaning out of something. In parallel with a semiotic triangle, where the sign is emphatically not just the signifier but is the combination of signifier, signified and signification (or whatever you have chosen to call the vertices), so, too, information is defined by the input, the output and the trapezium itself. "The difference that makes a difference": "The difference", yes, but only together with "a difference" and "that makes a".

* DIKW: Data-Information-Knowledge-Wisdom. I thought I'd written about that previously, but I can't find anything. If I find it I'll add a link the post. If not, I'll blog about it at a later date! My key point is that none of these are absolute levels. Everything is relative.

A key message of Kevin's talk was about the misunderstanding/misuse/misinterpretation of the p-value in the null hypothesis, or rather the use of the null hypothesis generally. He referred to papers with titles like "Why most published research findings are false".  I'd thought, broadly speaking, that I understood how hypothesis testing works, but Kevin said something along the lines of 'if you think you understand it you probably don't', and a recurring theme of my life as an academic (and not just as an academic) has been to discover that I don't understand things I'd previously thought I did understand, so I'm sure Kevin is right.  The one thing I want to pick up on, though, is that Kevin said - I think - that testing against the null hypothesis is about looking for differences, and of course that word 'difference' brings us back to information!  The test is to see if there is a difference that can make a difference in the data. That is, to see if the data combined with the statistics and the conclusion constitute information.

Thursday, 14 April 2011

Statistical significance and information

Statistical significance is one thing, but the information you get from the statistics is something else - and, crucially, depends on context.
Amplify’d from www.open.ac.uk
Statistically significant: The US Supreme Court takes a view
By: Professor Kevin McConway1 (The Open University)
A lawsuit over a cold remedy has forced the US Supreme Court into deciding what might be statistically significant.
The rulings of the US Supreme Court aren't the place most people would start looking for a discussion of what the phrase "statistically significant" might mean. But in March 2011, America's highest court of justice made a decision that involved exactly that. [...]
If a feature in data is not statistically significant, it's reasonable to take that to mean that the feature might plausibly be due solely to the workings of chance.
But just because it might be due to chance, that doesn't mean it is due to chance. That's only one possible explanation.
Read more at www.open.ac.uk

Sunday, 7 March 2010

Info content of the FA cup draw

For the semi-final draw they dramatically select each of the four balls for the four clubs, but in fact there's only log2(3) = 1.58 bits of information, and it could be done with a single draw. Since both matches play at Wembley, there's no 'home' or 'away', so there's only three possibilities:

A - B and C - D
A - C and B - D
A - D and B - C

During the draw:

- first team out of the box: gives no information
- second team out, select one from the three remaining, and gives all of the information
- third and fourth teams out, tell us nothing, because we already knew they were to play each other

Unless I suppose it makes a difference which match is played first, in which case there are six possibilities, adding one bit of information, 2.58 bits. (log2(6) = 2.58)

A - B then C - D
A - C then B - D
A - D then B - C
C - D then A - B
B - D then A - C
B - C then A - D

- first team out. Tells us that team plays in the first match. That is a selection from two equal possibilities (that team could have played in the first or second match), so contributes 1 bit of information
- second team out. This is a selection of one from three (the three teams still in the hat) so is 1.58 bits of information. This is now all the information, because we know the other two teams play each other in the second match. Information adds, so total information is 2.58 bits.
- third and fourth teams out contribute no information.

Tuesday, 29 September 2009

Estimation and compression

In the current (September 09) IEEE Information Theory Society Newsletter, (Available online, but the September 09 one isn't there yet), J Rissanen writing on 'Optimal Estimation':
Soon after I had studied Shannon's formal definition of information in random variables and his other remarkable performance bounds for communication, I wanted to apply them to other fields - in particular to estimation and statistics in general. After all, the central problem in statistics is to extract information from data. After having worked on data compression and introduced arithmetic coding it seemed evident that both estimation and compression have a common goal: in data compression the shortest code lenth cannot be achieved without taking advantage of the regular features in data, while in estimation it is these regular features, the underlying mechanism, that we want to learn. This led me to introduce the MDL or Minimum Description Length principle, and I thought that the job was done. [However...]
My emphasis.

Thursday, 16 April 2009

Information Engineering, BCS/Turing lecture by Sir Michael Brady

I've just watched this on IET TV.



Information Engineering and its future

Sir Michael Brady

Presentation from BCS/IET Turing Lecture 2009

2009-01-28 12:00:00.0 IT Channel

>> go to webcast>> recommend to friend



There's lots of fascinating ideas in there, but what initially caught my attention was the story of reading the stylus tablets from Vindolanda. More information in a paper by Melissa Terras (Interpreting the image: using advanced computational techniques to read the Vindolanda texts, ASLIB PROCEEDINGS v58 n1-2 pp102-117 2006):
This paper describes the developmental stages undertaken to construct a system that can read in images of an ancient document and produce plausible interpretations of the document, to aid the historians in the lengthy process of reading an ancient text. In carrying out the development, an explicit representation of how experts approach and reason about damaged and deteriorated texts was formulated, and a large corpus of letter forms and linguistic data were captured. Preliminary results from the resulting computer system are presented which demonstrate the usefulness of the technique, although more work is needed to develop this into a stand-alone computer system.
The multidisciplinary nature of the task stood out.

See also this on Michael Brady's Oxford website.

Sunday, 7 September 2008

Probability of the end of the world

From The Independent on Saturday (in the print edition, I can't find it online):
The Large Hadron Collider...the possibility that it could create an apocalyptic black hole... . One estimate has put the chance at 1 in 1,019. To put that in context, you have a one in 1,011 chance of spontaneously evaporating, while you shave.
I trust that is wrong about the probability of the end of the world, and, as to the spontaneous evaporation, with millions of people shaving each morning thousands of them must be disappearing each day. I've never seen that reported in The Independent!
I presume the numbers should have been 1019 and 1011, and the superscipt formatting has got lost - a common enough event. But, clearly someone has at a later stage put it the commas for the thousands, unaware of the meaning or significance of the numbers.

Friday, 29 August 2008

Telling stories and lies with statistics

An article in the 30/8/08 New Scientist "Keep your head" discusses how our emotions override rational decision-making.

It is fine as far it goes, and the example of the vast increase in deaths as a consequence of people driving long distances instead of flying, due to fear caused by the terrorist attacks, is fine. But, the chart that it uses to support the argument - and the general story-line of "when it comes to risks, feel the numbers" - raises other problems.

Here is a reproduction of the chart:

OK, so flying is much safer than driving. But, look, driving is much safer than walking or cycling? Is that right? Can I really say 'driving is much safer than walking or cycling' based on this graph? I don't think so, and let me give three problems with drawing simplistic conclusions like that:

1) They are not simple alternatives. You can't choose between flying or walking to the local shop. If you need to do some food shopping you might choose between walking ten minutes to the corner shop, cycling ten minutes to the 'mini-market' or driving ten minutes to the supermarket. The comparison then is between deaths per hour, not deaths per kilometer. You could make a similar choice for you holiday destination: drive, train or bus to Blackpool, or fly to Spain. Again, the meaningful comparison is per hour, not per km.

2) These only include deaths to the passenger during the travel. What about killing other people? Cars kill other people - including cyclists and pedestrians, bicycles rarely do and pedestrians never do (not as a result of walking into people anyway - I presume!). And of course cycling and walking keeps you fit. It is said that the health benefits of cycling far outweigh the increased risk from accidents, compared to driving. I don't have the figures to confirm that, but it is a consideration.

3) What about wider systems issues - damage to the environment (global warming), impact on the economy. These are really difficult to evaluate, but it doesn't mean that don't exist.

To look at (1), I did a rough and ready conversion of the numbers from the article to 'deaths per million hours' based on a guess of the average speed for each mode of transport. (Motorcycle, car and bus 70 km/hour, walking 4 km/hour, bike 15 km/hour, van 60 km/hour, water 50 km/hour, rail 100 km/hour and air 800 km/hour. I also needed a figure for deaths per billion km for air travel, because the New Scientist article just said ‘less than 0.1'. A search on the web turned up a figure of 1 death per 15 billion passenger km, so I used that.) This gave the chart below (note that the column for motorcycle is shortened).

I admit I was disappointed to find that cycling still came out worse than driving, but the difference is now much smaller, and surely more than compensated by the health benefits of cycling. The thing that surprised me in this, was that flying came out worse than water, bus or rail. So a train trip to Blackpool looks like your best bet for the hols! (And why would anyone ever go anywhere on a motorbike!)