20 July 2012
Redefining productivity
28 September 2009
The Greatest Show on Earth
To whom, then, is it addressed? Presumably to the fair minded but not so well informed person who has heard a lot about the evolution wars and wants to know just why scientists are so confident that they are right. If this is the case, I would have thought it best if Dawkins had left out the snarky and condescending comments about creationists. Admittedly, they are easy targets, and perhaps Dawkins by temperament cannot resist taking pot shots (I do it myself), but the book would work best, I believe, as a persuader of the unpersuaded if he had let the evidence stand on its own, untainted by polemic.
14 July 2009
Meeting Royalty

His Royal Highness, the Prince Andrew, Duke of York, came to visit the Sanger campus last week. He had previously developed an interest in the Institute when discussing it (among other things) at a meeting with our principal funder, the Wellcome Trust.
29 May 2009
30 June 2008
09 July 2007
<i>The Daily Show</i>
Scotticus made an observation recently on his webpage, and since he doesn't allow comments I'm commenting here:
A recent Pew Research Center study shows that people who watch The Colbert Report and The Daily Show correctly answered 54% of questions about current affairs, whereas viewers of regular TV news correctly answered only 35%. I'm not sure what to make of that finding.
This is a classical example of statistical confounding. While I value the real-news perspective of these programs, watching them probably has no effect on current events knowledge. The Daily Show is a comedy program aimed at well-educated, mainly liberal young people with an interest in current affairs. TV news programs are aimed at the typical local viewer. It's very likely the existing differences in the audiences accounts for the difference in news-awareness.
07 June 2007
18 April 2007
Get in!
Introduction to my transfer report done. Now just to cut and paste my paper and write up some future work BS...
12 April 2007
Obesity and awesomeness
The first (of many) papers stemming from the Wellcome Trust Case Control Consortium (over which I have laboured for two years) was published today, elucidating the first convincing genetic link to obesity. We all went to the pub at 7PM to watch my semi-boss Mark McCarthy appear on the BBC World News to discuss the finding. And in a cardinal example of how sometimes England is awesome, the landlord of the Butcher's Arms (who knows we often stumble in on a Friday after a long week at work) surprised us with a bottle of champagne on the house. Cheers, and watch out for rs9939609 (if you have two bad copies you have a 70% increased chance of being obese!).
25 February 2007
Ulf Milton
I discovered some file errors today, which means that I need to redo a lot of work before sleeping tonight. So I've brewed up the strong coffee and put on the music. Only a few more days of this before my escape to the USA.
29 January 2007
CNN sucks, part N+1
CNN.com is featuring a "How it works" animation about genetic testing which gets incorrect nearly every basic fact about biology. It includes the following caption above a picture of a double helix:
The double helix is two exact copies of chromosomes wound together — a strand from each parent.
After these "Basics" (the title of that slide) we move on to the discussion of disease causing variants, where we learn that DNA mutations result in "abnormal pairings" of A and T to C or G and vice versa.
How is it possible that they couldn't find someone with a passing understanding of high school biology to work on this feature? I'm not even asking for an actual biologist to do their research, but believe me when I say that this is easily as egregious as:
05 January 2007
Emerald City
I'm really enjoying working here at the Fred Hutchinson Cancer Research Center in Seattle, where I'm treated to the grey but pleasant view seen here. When the weather is warmer, sailboats cross Lake Union, and seaplanes land and take off regularly. In addition I get a fairly spacious cube all to myself, along with a big wall size whiteboard to use. Plus the support services here are just amazing. Lon's PA has set us up with normal stuff like pens and staplers and paper clips, but he really goes above and beyond the call of duty: the office bookshelf has Lonely Planet guides to Seattle and Vancouver, plus some nice laminated maps of Seattle.
28 December 2006
12 October 2006
655,000
A paper published in the Lancet, "Mortality after the 2003 invasion of Iraq: a cross-sectional cluster sample survey", has been much in the news lately. Researchers from Johns Hopkins, funded by MIT, have conducted a statistical survey of households in Iraq to estimate the number of deaths in the period immediately preceding, and following the 2003 invasion. Their estimate for the total number of excess deaths in the years since the invasion that might otherwise not have happened is more than an order of magnituded higher than any previous estimate: roughly 655,000 (95% confidence interval of 393,000 - 943,000).
Having read the article I can't find anything wrong with their methods -- indeed the paper is extremely well-written and carefully considered (as one would expect the editors to enforce on such a controversial topic). Furthermore, every single piece of media coverage I've seen that actually quotes an expert in statistics, polling or epidemiology compliments the study as being the best designed and most comprehensive to date on the topic.
Of course, other people take a different view:
PRESIDENT BUSH: "I don't consider it a credible report...the methodology was pretty well discredited."
GEN'L GEORGE CASEY (commander of US ground forces in Iraq): "[the death toll] seems way, way beyond any number that I have seen. I've not seen a number
higher than 50,000. And so I don't give that much credibility at all."ALI AL DABBAGH (Iraqi gov't spokesman): "The report is unbelievable. These numbers are exaggerated."
It's a little worrying that the people in charge treat scientific research as lacking credibility just because it doesn't jive with what they expected the answer to be. Isn't this the point of science? To ask questions we don't know the answers to and then try to build upon the knowledge we gain? Shouldn't this report lead to an immediate follow-up to narrow that confidence interval and try to replicate the findings? Especially aggravating is the President claiming that the methodology is discredited, which is not only patently false, but sounds idiotic coming from a man who is clearly not an expert in statistical survey techniques.
In a broader sense, this is what drives me insane about politics and government: everything is driven by policy and electoral math, rather than by facts. Some researchers go out and do this incredibly dangerous study (it's not often that you read an academic paper with sentences such as "No interviewers died or were injured during the survey." in the Results section) and when they publish their shocking results, they're dismissed out of hand because they're inconvenient. Bah, I'm staying in science.
10 October 2006
Expertise
NERD ALERT — this post will only be interesting to grad students.
It's pretty satisfying to be able to read a paper, and when you come to a numbered reference in the text to know what paper they're referencing without even checking the list at the end. For me, at least, it serves as a signpost that I've understood the preceding sentence well enough to not only grasp the concept, but also to be familiar with the most relevant literature.
09 October 2006
Education
I received an email sent out to the DPhil mailing list today advertising an "Online Plagiarism Course". I really hope it teaches how to efficiently plagiarise and not get caught...
22 August 2006
Segue
The funny thing about printing out a paper from a major journal like Science or Nature is that you usually get a snippet of the next article on the last page of whatever it is you're looking for. Not many publications can get away with transitioning from "The Fine-Scale Structure of Recombination Rate Variation in the Human Genome" (a perfectly legitimate arena of research) to "The Rise of Rhizosolenoid Diatoms" (which sounds like a bad sci-fi movie).
11 July 2006
Philly
I've got another couple of hours in Philly (I'm currently on a brief break between meetings) before flying back to the UK. So far the trip has been very cool. My friend Brendan (who recently finished his DPhil at Oxford and is now post-doc'ing here) has graciously hosted me in the mansion-like home he's house-sitting. The place is a four floor townhouse in one of Philly's trendiest areas (Rittenhouse Square) and features (among many other things) a huge entry hall with a sweeping grand staircase, a 19th century period banquet hall, an immaculately appointed kitchen big enough to run a medium sized restaurant out of and a two-room master bedroom with walk in closets bigger than my bedroom in Oxford. The owners moved out about 6 months ago and have been looking for a buyer ever since. In the meantime Brendan is house-sitting (there's a fairly autonomous apartment set up on the third floor) for them. All the furniture is gone, so the place is kind of spooky and echo-y, but still it's pretty sweet. I would've taken photos, but alas my camera was destroyed in the carboat adventure.
As far as the actual purpose of my trip, it has gone really well. I've managed to meet some collaborators and make progress on where my portion of the project is headed. My talk was well-received (it's pretty straightforward when you've given it a dozen times) and I've made some additional contacts who have pretty exciting projects starting at Penn and CHoP. A good vacation overall and I'm ready to head back to Oxford.
14 June 2006
More Hardball Stats
In response to a RZA comment on the previous entry I have to strenuously disagree that this is a large sample size to make small distinctions. If we split Papi's numbers, for instance by odd innings vs even innings, we get almost as big a difference as the early vs. late split: 0.299 odd vs 0.281 even. I'm not saying Papi isn't clutch, I'm saying that the statistical evidence isn't overwhelming.
13 June 2006
Stats and Baseball
After reading this excellent post at Yanksfan vs Sox fan (via 2GD) on the statistical analysis of whether Big Papi is a "clutch" hitter, I can't help but chime in as both a statistician and baseball fan. More than any other sport, baseball fans and pundits alike love to troll through huge quanitites of numerical data to try to find interesting trends and observations about their favourite players.
Unfortunately nearly all such analyses are statistically bogus. In this case, the question is whether David Ortiz is a great "clutch" hitter, that is, he steps up his game in situations that are most important. The author then proceeds to parade a lot of numbers to argue his point. He makes the common mistake of presenting a trend (e.g. someone performs better on Wednesdays vs. Thursdays) without asking whether the data at hand are enough to prove that that trend is significant. In statistics it's all about sample size — whether you have sufficient observations to draw confident conclusions from your data.
If, for instance, I told you that it rained today, a Tuesday and was sunny yesterday, a Monday. Nobody would believe me if I then turned around and said, "It rains way more often on Tuesday than Monday!" In small samples, of course, random chance creates perceived patterns (such as rain correlating with Tuesday) where none actually exist.
In baseball we fool ourselves into thinking that we have enough observations to make all kinds of statements in which we have no confidence. In this specific case of clutch hitting, the author makes a whole series of claims with fairly small sample sizes, but lets look at his best case: batting average with runners on base vs. batting average with the bases empty. Ortiz has had 1861 total Red Sox at bats, distributed pretty evenly between these two scenarios (947 bases empty, 914 with one or more runners). He has had 265 hits in the first case (for a 0.280 average) and 280 hits in the second case (for a 0.306). So he's got more hits in fewer tries, thus the higher average with runners on. Regardless of whether this is a good measure of clutch performance (which is an entirely separate argument) we can ask whether these numbers actually mean something or whether they could've arisen by chance. Does Ortiz hit better with runners on?
In short, these numbers can't answer the question. When performing a simple test of statistical signficance, these values could easily have arisen by chance. We could easily have seen this discrepancy by dividing his at bats into those on where an odd number of fans were in the stadium vs. those where an even number of fans were watching. And this really makes sense when you think about it carefully. Even with nearly 2000 observations we're trying to gain insight into a very tiny difference: 0.280 vs 0.306! In baseball the difference between a guy with a career 280 average and a guy with a 306 is pretty big, but in almost any other circumstance we'd round both of these to an even 30% and call it a day. You would need tens of thousands of observations to demonstrate that Ortiz hits better with men on base with even modest confidence.
Keep in mind that this is actually a pretty big sample size for baseball. Many times people quote some 1-for-10 and say that Pitcher X "owns" Hitter Y. This is an even bigger joke, since a guy hitting .300 is likely to have only one hit in any given ten at-bats! I certainly hope that the guys actually working for ball clubs have a better handle on this than the average pundit.