Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Thursday, 15 February 2024

How Not To Do Great Science (The Lost Post)

This post was originally published on Discover Magazine on September 16th 2013, but has since vanished (although most of my other Discover posts are still available). Luckily, I saved a backup. So here's the original "How Not To Do Great Science".

---

This post is a bit special. For the first time ever, I've collaborated with an artist, Erene Stergiopoulos. Her webcomic is here and she's on Twitter here. I think you'll agree that the artistic standard is a little higher than I usually achieve. Anyway, here's what we did:



It would be silly to expect that every architect should finish buildings at a certain rate. That would make it impossible to anyone to build certain things. Some things take longer to build than others, and most great things take a great deal of time. Faced with a sufficiently demanding quota, builders might be reduced to rushing out follies that might look impressive from a distance, but that are no more than hollow shells. Yet, as silly it would be to make uniform demands of architects, this is what is happening to scientists.

Rather than build, scientists are expected to publish - and publish fast - or perish. My worry (and that of many others) is that the pressure to publish often fundamentally changes not just how much scientists write, but what they can write about. It turns researchers into prolific doers of small deeds, but it leaves them little time to think about, let alone complete, great works. Though the mills of God grind slowly...

Yet the problem is not just the speed of science today, but also the direction: go to a scientific conference and you'll see perfectly good data in the process of being oversold, misinterpreted, and p-hacked into a 'publishable' form.

Much has been said about how this leads to false positives - impressive follies that don't stand up to scrutiny. What's less discussed - and the point of this piece - is the opportunity cost. New theories come out of attempts to explain 'negative' data - negative from the perspective of the old theory. Null results are the foundations of future progress, but only if they are allowed to lie there awhile; not if they are torn up and used to prop up tottering old structures.

Thursday, 24 January 2013

Is Medical Science Really 86% True?

The idea that Most Published Research Findings Are False rocked the world of science when it was proposed in 2005. Since then, however, it's become widely accepted - at least with respect to many kinds of studies in biology, genetics, medicine and psychology.

Now, however, a new analysis from Jager and Leek says things are nowhere near as bad after all: only 14% of the medical literature is wrong, not half of it. Phew!

But is this conclusion... falsely positive?

I'm skeptical of this result for two separate reasons. First off, I have problems with the sample of the literature they used: it seems likely to contain only the 'best' results. This is because the authors:
  • only considered the creme-de-la-creme of top-ranked medical journals, which may be more reliable than others.
  • only looked at the Abstracts of the papers, which generally contain the best results in the paper.
  • only included the just over 5000 statistically significant p-values present in the 75,000 Abstracts published. Those papers that put their p-values up front might be more reliable than those that bury them deep in the Results.
In other words, even if it's true that only 14% of the results in these Abstracts were false, the proportion in the medical literature as a whole might be much higher.

Secondly, I have doubts about the statistics. Jager and Leek estimated the proportion of false positive p values, by assuming that true p-values tend to be low: not just below the arbitrary 0.05 cutoff, but well below it.

It turns out that p-values in these Abstracts strongly cluster around 0, and the conclusion is that most of them are real:

But this depends on the crucial assumption that false-positive p values are different from real ones, and equally likely to be anywhere from 0 to 0.05.
"if we consider only the P-­values that are less than 0.05, the P-­values for false positives must be distributed uniformly between 0 and 0.05."

The statement is true in theory - by definition, p values should behave in that way assuming the null hypothesis is true. In theory.

But... we have no way of knowing if it's true in practice. It might well not be.

For example, authors tend to put their best p-values in the Abstract. If they have several significant findings below 0.05, they'll likely put the lowest one up front. This works for both true and false positives: if you get p=0.01 and p=0.05, you'll probably highlight the 0.01. Therefore, false positive p values in Abstracts might cluster low, just like true positives.

Alternatively, false p's could also cluster the other way, just below 0.05. This is because running lots of independent comparisons is not the only way to generate false positives. You can also take almost-significant p's and fudge them downwards, for example by excluding 'outliers', or running slightly different statistical tests. You won't get p=0.06 down to p=0.001 by doing that, but you can get it down to p=0.04.

In this dataset, there's no evidence that p's just below 0.05 were more common. However, in many other sets of scientific papers, clear evidence of such "p hacking" has been found. That reinforces my suspicion that this is an especially 'good' sample.

Anyway, those are just two examples of why false p's might be unevenly distributed; there are plenty of others: 'there are more bad scientific practices in heaven and earth, Horatio, than are dreamt of in your model...'

In summary, although I think the idea of modelling the distribution of true and false findings, and using these models to estimate the proportions of each in a sample, is promising, I think a lot more work is needed before we can be confident in the results of the approach.

Thursday, 3 January 2013

Flawed Statistics Make Almost Everyone's Brain "Abnormal"

A popular method for detecting abnormalities in the shape and size of individual brains is seriously flawed, and is almost guaranteed to find 'differences' even in normal people.

So say Italian neuroscientists Scarpazza and colleagues in an important new report: Very high false positive rates in single case Voxel Based Morphometry.

Voxel Based Morphometry (VBM) is a way of analyzing brain scans to detect structural differences. It's most commonly used to compare groups of brains to find average differences, but some neuroscientists have started using VBM to check for abnormalities in a single brain. Scarpazza et al list 34 pieces of research about that, including 13 since 2010.

So it would suck if there were a problem with individual VBM... but there is. This pic tells the tale:


The authors took 200 normal brains and compared each one of them in turn to a control group of 16 normal brains. Because all of them were healthy, the comparisons ought to show no significant differences.

The technique was set up so that, in theory, only 5% of the brains should have been wrongly labelled as containing an abnormality. But in fact, a full 93.5% of the normal brains gave at least one false positive.

So 5% is more like the rate of not being wrong. Oops.

The image shows that in some brain areas, almost 25% of the normal brains were branded as 'abnormal' just in that region alone - the hotter the colour, the higher the proportion of false 'hits'. The top row is for false reports of brain volume increases, while the bottom row is decreases; false 'increases' were more common.

So what's going wrong? It's not entirely clear and several factors are probably at play, but the authors say that the main issue is that VBM makes the assumption of statistical normality which doesn't in fact hold.

Either way, it's a serious problem, and Scarpazza et al point to one especially worrying implication: some people have proposed using single-subject VBM in a legal context, to reinforce insanity pleas by showing subtle 'brain abnormalities' not obvious to the naked eye. Yet if this paper's right, such evidence could be entirely meaningless, almost guaranteed to give a positive result.

P.S. Last time I posted about this kind of analysis flaw, the internet went crazy because they didn't understand it. So just to be clear, this is not a problem for clinical scans - the kind you'd get to check whether you have a brain tumour.

ResearchBlogging.orgScarpazza, C., Sartori, G., De Simone, M., and Mechelli, A. (2013). When the single matters more than the group: Very high false positive rates in single case Voxel Based Morphometry NeuroImage DOI: 10.1016/j.neuroimage.2012.12.045

Saturday, 20 October 2012

When Replication Goes Bad

How to ensure that results in psychology (and other fields) are replicated has become a popular topic of discussion recently. There's no doubt that many results fail to replicate, and also, that people don't even try to replicate findings as much as they should.


Yet psychologist Gregory Francis warns that replication per se is not always a good thing: Publication bias and the failure of replication in experimental psychology
Among experimental psychologists, successful replication enhances belief in a finding, while a failure to replicate is often interpreted to mean that one of the experiments is flawed. This view is wrong.

Because experimental psychology uses statistics, empirical findings should appear with predictable probabilities. In a misguided effort to demonstrate successful replication of empirical findings and avoid failures to replicate, experimental psychologists sometimes report too many positive results.

Rather than strengthen confidence in an effect, too much successful replication actually indicates publication bias, which invalidates entire sets of experimental findings...

Even populations with strong effects should have some experiments that do not reject the null hypothesis. Such null findings should not be interpreted as failures to replicate, because if the experiments are run properly and reported fully, such nonsignificant
findings are an expected outcome of random sampling... If there are not enough null findings in a set of moderately powered experiments, the experiments were either not run properly or not fully reported. If experiments are not run properly or not reported fully, there is no reason to believe the reported effect is real.
Say you took a pack of playing cards and removed half the red cards. Your pack would now be 2/3rds black, so if you took a random sample of cards, say a poker hand of 5 cards, then you'd expect more blacks than reds (a significant 'effect' of color). But you'd still expect some reds, and some random hands would in fact be entirely red, just by chance. If someone claimed to have drawn 10 random hands and they'd all been mainly black, that would be implausible - "too good".

Francis's approach is a bit like Uri Simonsohn's method for detecting fraudulent data - they both work on the principle that "If it's too good to be true, it's probably false" - but they differ in their specifics, and I believe that we should not conflate fraud with publication bias... so let's not get carried away with the parallels.

Earlier this year, Francis wrote a critical letter about a paper published in PNAS purporting to show that wealthier Americans are less ethical. He argued that the paper's results were "unbelievable" - it reported on the results of seven separate experiments, all of which showed a small, but significant, effect in favour of the hypothesis.

Even if rich people really were meaner, Francis said, the chance of 7/7 experiments being positive is very low: just by chance, you'd expect some of them to show no difference (given that the size of the difference in those seven was low, with a lot of overlap between the groups). Francis suggested that the authors may have run more than seven experiments, and only published the positive ones; the authors denied this in their Letter.

Anyway, in the new paper, Francis expands on this approach in much more detail, drawing from this 2007 paper, and suggests a Bayesian approach that might help mitigate the problem.

ResearchBlogging.orgFrancis G (2012). Publication bias and the failure of replication in experimental psychology. Psychonomic Bulletin and Review PMID: 23055145

Saturday, 25 August 2012

Replication Alone Is Not Enough

Psychology has lately been hit by high-profile fraud scandals, and broader concerns over questionable research practices. Now the Society for Personality and Social Psychology (SPSP) has released a statement on "Responsible Conduct", and a task force has produced a report.

This is a start, and the SPSP is to be commended for facing up these problems (which affect many other fields) relatively early. However, neither of their documents contains much meat in my view.

Point One on the task force report is that "Replication is the key to building our science" and they suggest a "web site for depositing replications and failures to replicate" - but don't mention that various enterprising researchers have already made one. Nor do they tip their hats to the Open Science Initiative addressing just this issue. This makes me worried that they're planning to reinvent the wheel.

More fundamentally I disagree that replication is key to psychology or any field. Our goal should be replicability. Failure to replicate findings is a symptom of problems with those original findings, rather than being a problem in and of itself. Good results replicate; we want better results to be published.

In other words, we should strike at the root cause of invalid research, namely, the perverse incentives towards publishing as many eye-catching positive results with p values below 0.05 as possible by any means necessary. P-value fishing, selective reporting, post-hoc "prior hypotheses" and other questionable practices are a large part of what make unreplicable results.

We should encourage replication, but it's no panacea.

An overemphasis on replication, without addressing the incentives, could actually harm science. It could lead to scientists spending all their time worrying about the political drama of who's replicating who and why, and which questionable practices they can use to replicate their friends' data - rather than actually doing science.

This is why we shouldn't be satisfied with any reform effort that puts replication before replicability. If you can fudge a result, you can fudge the data a replication. How to fight questionable practices is another question but I've proposed reforms that I think would work, namely pre-registration of hypotheses, methods, and statistical analyses. Others have their own ideas.

A lesson from clinical medicine here. Clinical trials of new drugs adopted pre-registration, but only after they tried replication and it didn't work. Pharmaceutical regulators have long required multiple demonstrations of drug efficacy. One trial was not enough. Sounds good - but the problem was that drug companies just did lots of trials and analyses, picked the positive ones, and used them.

So in summary: replication is important, and we don't do enough of it, but replication alone is not enough to fix psychology.

Saturday, 28 July 2012

Catching Fraud: Simonsohn Says

Everyone's been talking about psychologist Uri Simonsohn and his role in the downfall of two scientific fraudsters.


When the story first broke, the methods Simonsohn used that allowed him to spot the dodgy data were mysterious - which only added to the buzz. The paper revealing the approach is now up online and it's a must-read. It's not often a statistics paper offers the train-wrecky schadenfreude of watching two fraudsters' careers come to a well-deserved end.

What's rather disturbing about the article, however, is that it doesn't really contain much that's new, in principle. Simonsohn used statistics to spot data in published papers that was, in effect, 'too good to be true'. He then followed up seemingly dodgy cases with some more stats, using simulations of what real data ought to look like, to verify that it was in fact made up. A simple idea in retrospect but one that's never been tried before. I don't think there's a single "Simonsohn method", rather, the paper uses multiple techniques, each one tailored to the particular data in question.

But it shouldn't have come to this. Someone else ought to have spotted that the data looked dodgy.

Take this table from one of Simonsohn's conquests, a soon-to-be-retracted paper by Lawrence J Sanna et al:

We now know that the data from Studies 2,3 and 4 were all made up. Each study compared 3 conditions, and what makes these data dodgy is that the standard deviations of the 3 sets of results for each study were almost identical. The chances of that happening are very low and it suggests that someone has (clumsily) made the data up.

I'm going to say that these data are obviously suspicious, at least to anyone who has worked with real data. Maybe you'll say that hindsight is 20/20, but Simonsohn didn't need hindsight and the stats he used were nothing remarkable. I'm not saying that to belittle his achievements, he deserves plenty of credit. But other people deserve blame.

Namely, whoever peer reviewed this paper should have spotted that these data looked unusual - and they should not have needed any special statistical tools to do so.

Simonsohn calls for journals to require that the raw data be made available for all published work, on the grounds that. That's a great idea - and not just because it would help catch bad science: it would facilitate proper research and teaching no end. But Simonsohn didn't need the raw data to detect these cases of fraud - he only checked the raw results to confirm the suspicions based on the published data.

Checking that the data are valid is the job of peer reviewers, and they dropped the ball. Instead Simonsohn had to conduct his own private crusade against fraud... a bit like Batman. Batman is awesome, but the point about Batman is that he's only needed because the police can't or won't cope on their own. He's not a superhero, he's just a guy with the will.

Peer reviewers are the police of science, but all too often, they're asleep on the job. Not just in psychology. Retraction Watch provides plenty of examples of published results in biology that were faked, often in comically crude fashion, and should have been obvious to anyone paying attention.

Peer reviewers are usually anonymous. I wonder if a policy of retrospectively naming and shaming the reviewers when a paper turns out to have been fraudulent, might help motivate them...?

Saturday, 21 July 2012

A Case Study in Voodoo Genetics

A new review of published studies looking at the relationship between a gene and brain structure offers a sobering lesson in how science goes wrong.

Dutch neuroscientists Marc Molendijk and colleagues took all of the studies that compared a particular variant, BDNF val66met, and the volume of the human hippocampus. It's a long story, but there are various biological reasons that these two things might be correlated.

It turns out that the first published reports found large genetic effects, but that ever since then, the size of the effects has dropped, with the latest studies finding no effects at all -


A cumulative meta-analysis confirms that as more studies on BDNF val66met have appeared, the overall effect estimate has steadily declined -


Finally, the authors found signs of publication bias: there were three small, imprecise studies that reported very large effects of the gene, but no such studies finding no effect (or a reverse effect). You'd expect that small and noisy studies would have a lot of random variation so they wouldn't all be positive even if there was a true effect; that all of the published ones were positive, suggests that null findings are out there, unreported.

Overall, this suggests that val66met probably isn't associated with hippocampus volume after all, and that the early studies showing that it was, were misleading. There's no reason to think that the early studies were wrong as such - they may have accurately reported an effect in the small sample of people they looked at, but it was only a chance finding.

I know a lot of neuroscientists who are now fairly skeptical of this whole genre of candidate gene studies; there was much excitement 5 or 10 years ago, but in retrospect, most of these studies were too small, and the publication process meant that it was the random chance findings that were most likely to get published.

However, we need to avoid any sense of complacency. Until we fix the scientific process, the same thing will happen again.

ResearchBlogging.orgMolendijk ML, Bus BA, Spinhoven P, Kaimatzoglou A, Voshaar RC, Penninx BW, van Ijzendoorn MH, and Elzinga BM (2012). A systematic review and meta-analysis on the association between BDNF val(66) met and hippocampal volume American journal of medical genetics B Neuropsychiatric genetics PMID: 22815222

Saturday, 30 June 2012

False Positive Neuroscience?

Recently, psychologists Joseph Simmons, Leif Nelson and Uri Simonsohn made waves when they published a provocative article called False-Positive Psychology


The paper's subtitle was "Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant". It explained how there are so many possible ways to gather and analyze the results of a (very simple) psychology experiment that even if there's nothing interesting really happening, it'll be possible to find some "significant" positive results purely by chance. Then you could publish those 'findings' and not mention all the other things you tried.

It's not a new argument, and the problem has been recognized for a long time as the "file drawer problem", "p-value fishing", "outcome reporting bias", and by many other names. But not much has been done to prevent it.

The problem's not just seen in psychology however, and I'm concerned that it's especially dangerous in modern neuroimaging research.

*

Let's assume a very simple fMRI experiment. The task is a facial emotion visual response. Volunteers are shown 30 second blocks of Neutral, Fearful and Happy faces during a standard functional EPI scanning. We also collect a standard structural MRI as required to analyze that data.

This is a minimalist study. Most imaging projects have more than one task, commonly two or three and maybe up to half a dozen, as part of one scan. If one task failed to show positive results, it need never be reported at all, so additional tasks would compound the problems here.

Our study is comparing two groups: people with depression, and healthy controls.

How many different ways could you analyze that data? How much flexibility is there?

*
First some ground rules. We'll stick to validated, optimal approaches. There are plenty of commonly used less favoured approaches, like using uncorrected thresholds (and then, which ones?) or voodoo stats, but let's assume we want to stick to 'best practice'.

As far as I can see, here's all the different things you could try. Please suggest more in the comments if you think I've missed any:

First off, general points:
  • What's the sample size? Unless it's fixed in advance, data peeking - checking whether you've got a significant result after each scan, and stopping the study when you get one - gives you multiple bites at the cherry.
  • Do you use parametric, or nonparametric analysis?
Now, what do you do with the data?
  • Preprocessing
    • How much smoothing?
    • Do you reject subjects for "too much head movement"?  If so, what's too much?
  • Straightforward whole-brain general linear model (GLM) analysis followed by a group comparison.
    • What's the contrast of interest? You could make a case for Fear vs Neutral, Happy vs Neutral, Happy vs Fear, "Emotional" vs Neutral.
    • Fixed effects or random effects group comparison?
    • Do you reject outliers? If so, what's an 'outlier'?
    • Do you consider all of the Fear, Happy and Neutral blocks to be equivalent, or do you model the first of each kind of block seperately? etc.
  • "Region of Interest" (ROI) GLM analysis. Same options as above, plus:
    • Which ROI(s)?
      • How do you define a given ROI?
  • Functional connectivity analysis
    • Whole-brain analysis, or seed region analysis?
      • If seed region, which region(s)?
    •  Functional connectivity in response to which stimuli?
  • Dynamic Causal Modelling?
    • Lots and lots of options here.
  • MVPA?
    • Lots and lots of options here.
But remember we also got structural MRIs, and while they may have been intended to help analyze the functional data, you could also examine structural differences between groups. What method?
    • Manual measurement of volume of certain regions.
      • Which region(s)?
    • VBM.
    • Cortical morphometry.
      • What measure? Thickness? Curvature...?
That's just the imaging data. You've almost certainly got some other data on these people as well, if only age and gender but maybe depression questionnaire scores, genetics, cognitive test performance...
  • You could try and correlate every variable with every imaging measure discussed above. Plus:
    • Do you only look for correlations in areas where there's a significant group difference (which would increase your chances of finding a correlation in those areas, as there'd be fewer multiple comparisons)?
  • You could define subgroups based on these variables.
*

So even a very straightforward experiment could give rise to hundreds or thousands of possible analyses. 1 in 20 of these would give a statistically significant result at p=0.05 by chance alone, and even if you throw out half those for being "in the wrong direction" (and that's subjective in most cases) you've got plenty of false positives.

This problem is growing. As computation power continues to expand, running multiple analyses is cheaper and faster than ever, and new methods continue to be invented (DCM and MVPA were very rarely used even 5 years ago.)

I want to emphasize that I am not saying that all fMRI studies of this kind are in fact junk. My worry is that it's hard to be confident that any given published study is sound, given that papers are written only after all the data has been collected and analyzed.

I'll also point out that some imaging research, especially what might be called "pure" neuroscience investigating brain function per se rather than "clinical" studies looking at differences between groups, has many fewer variables to play with, but still quite a lot.

As to how to solve this problem, the one solution I believe would work in practice is to require pre-approval of study protocols.
ResearchBlogging.orgSimmons JP, Nelson LD, and Simonsohn U (2011). False-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological science, 22 (11), 1359-66 PMID: 22006061

Tuesday, 19 June 2012

Beware Stimulus Effects in Psychology

Recently I've blogged about methodological problems in neuroscience research but just to even things out a bit, here's a paper that highlights a potentially serious issue for psychologists - Treating Stimuli as a Random Factor in Social Psychology: A New and Comprehensive Solution to a Pervasive but Largely Ignored Problem

Suppose you want to find out whether people react differently to stimuli from two different groups. The reactions, stimuli, and groups could be anything: maybe you want to see if people prefer listening to sound clips of cats as opposed to dogs. Or maybe you show people photos of blonde men vs dark-haired men and see whether people judge guys with one colour as less trustworthy.

A lot of psychology studies amount to this.

Going with the blonde vs. dark example, suppose you take 1000 volunteers, show them some pictures of blonde guys and dark guys, and get them to rate them on trustworthiness. You find a significant difference between the two groups of stimuli. You conclude that your volunteers are hair-bigots and submit it as a paper. The reviewers think, 1000 volunteers? That's a big sample size. They publish it.

Now that study I just described might be perfectly valid. But it might be seriously flawed. The problem is that while your sample size may be large in terms of volunteers, it might be very small in another way. Suppose you have just 10 photos per group. Your 'sample size', as regards the sample of stimuli, is only 20. And that sample size is just as important as the other one.

It might be that there's no real hair difference in perceived trustworthiness, but there are individual differences - some men just look dodgy and it's nothing to do with hair - and in your stimuli, you've happened to pick some dodgy looking blonde guys. Or whatever.

Now you can run your statistical analyses taking these possible stimulus variation effects into account. But according to Judd, Westfall and Kenny, authors of this paper, this is rarely done. They show with both real and hypothetical data, that unless you take care of this, you can find "statistically significant" differences from pure random noise. This is not a new argument, but they say it's been ignored for too long.

The worst part is that increasing the number of volunteers actually makes it more likely that you'll fall foul of this, not less. Only increasing the stimulus sample size can prevent it.

The paper goes into lots of detail, and tackles various hot potatoes, including one of Daryl Bem's notorious precognition "retroactive priming" experiments. Bem claimed that college students were able to predict the future - they responded differently to different pictures... before the pictures appeared on the screen. The effect was statistically significant and he published it. But Judd et al say that accounting for stimulus variation removes the effect.

ResearchBlogging.orgJudd CM, Westfall J, and Kenny DA (2012). Treating Stimuli as a Random Factor in Social Psychology: A New and Comprehensive Solution to a Pervasive but Largely Ignored Problem. Journal of Personality and Social Psychology PMID: 22612667

Wednesday, 23 May 2012

Rich People May Not Be So Unethical

There was quite the stir a few weeks back about a psychology paper claiming that rich people aren't very nice: Higher social class predicts increased unethical behavior.

The article, in PNAS, reported that upper class individuals were more likely to lie, cheat, and break traffic laws.

However, these results have been branded "unbelievable" in a Letter to PNAS just published. Psychologist Gregory Francis notes that the paper contains the results of 7 seperate experiments, and they all found statistically significiant socioeconomic effects on unethical behaviour.

Those 7 replications of the effect "might appear to provide strong evidence for the claim" - one study good, 7 studies better, right? - but Francis says that actually, it's too good to be believed.

Each of the studies was fairly small, and the effects they found were modest, and only just significant. So the observed power of the studies - the probability that a study of that size would detect the effect that they did, in fact, find - was only about 50-88% in each case.

Think of it this way: if you took a pack of cards and discarded half of the black ones, then shuffled the remainder, a random card from the deck would most likely be red. But even so, it would be unlikely that you'd pick seven reds in a row.

The chances of all 7 studies finding a positive result - even assuming that the effect claimed in the paper was real - is just 2%, by Francis's calculations.

Ow.

He concludes "The low probability of the experimental findings suggests that the data are contaminated with publication bias. Piff et al. may have (perhaps unwittingly) run, but not reported, additional experiments that failed to reject the null hypothesis (the file drawer problem), or they may have run the experiments in a way that improperly increased the rejection rate of the null hypothesis (4)".

What might have happened? Maybe there were more than 7 studies and only the positive ones were published. Maybe the authors peeked at the early data before settling on the sample size, or took other outcome measures that showed no effect and went unreported. See also the 9 Circles of Scientific Hell.

Or maybe not. Piff et al respond in their own Letter, firmly denying that they ran any other unpublished experiments, and saying that they "scrutinized our data collection procedures, coding protocols, experimental methods, and debriefing responses. In no case have we found anything untoward." They go on to criticize the method Francis used to get his magic 2% figure, which they point out relies on some debatable assumptions.

Even if you buy the 2% figure, it doesn't mean that the true effect is zero; it might be real, but exaggerated. Ultimately it all becomes rather murky and subjective, which is why I think we need preregistration of research, which would prevent any possibility of such data fiddling, and also remove the possibility of false accusations of it... but that's another story.

ResearchBlogging.orgFrancis, G. (2012). Evidence that publication bias contaminated studies relating social class and unethical behavior Proceedings of the National Academy of Sciences DOI: 10.1073/pnas.1203591109

Thursday, 17 May 2012

Another Antidepressant Crashes & Burns


Yet another "promising" novel antidepressant has failed to actually treat depression.

That's not an uncommon occurrence these days, but this time, the paper reporting the findings is almost as rubbish as the drug: Translational evaluation of JNJ-18038683, a 5-HT7 receptor antagonist, on REM sleep and in major depressive disorder

So, Pharma giant Janssen invented JNJ-18038683. It's a selective antagonist at serotonin 5HT-7 receptors, making it pharmacologically rather unusual. They hoped it would work as an antidepressant. It didn't - in a multicentre randomized controlled trial of 230 depressed people, it had absolutely no benefits over placebo. A popular existing drug, citalopram, failed as well:

About the only thing JNJ-18038683 did do in humans was to reduce the amount of dreaming REM sleep per night. This REM suppressing effect is also seen with other antidepressants and this is evidence that the drug does do something - just not what it's meant to. Being charitable you could call this a failed trial.

Ouch! But it gets better. Unhappy that JNJ-18038683 bombed, Janssen reached for their copy of the Cherrypicker's Manifesto. This is a new statistical method, proposed by fellow Pharma company GSK in a 2010 paper, which consists of excluding data from study centres with a very high (or very low) placebo response rate.

Anyway, after applying this "filter" JNJ-18038683 seemed to do a bit better than placebo, but the benefit over placebo still wasn't statistically significant - with a p value of 0.057, the wrong side of the sacred p=0.05 line (on page 33).
Yet Page 33's "trend towards statistical significance" magically becomes "significant" - in the Abstract:
[with] a post hoc analyses (sic) using an enrichment window strategy... there was a clinically meaningful and statistically significant difference between JNJ-18038683 and placebo.
Well, no, there wasn't actually. It was only a trend. Look it up.

That aside, the problem with the whole filter idea is that it could end up biasing your analysis in favour of the drug, leading to misleading results. The original authors warned that "data enrichment is often perceived as a way of improperly introducing a source of bias... In conventional RCTs, to overcome the bias risk, the enrichment strategy should be accounted for and pre-planned in the study protocol." They should know, as they invented it, but Janssen rather oddly say the exact opposite: "This methodology cannot be included in a protocol prospectively as it will introduce operational bias in that scheme."

Hmm.

Anyway, even after the filter technique, citalopram didn't work either... bad news for citalopram, except, was it citalopram at all? This is really unbelievable: Janssen don't seem clear on whether they compared their drug to citalopram, or to escitalopram - a quite different drug.

They say "citalopram" in most cases, but they have "escitalopram" instead, in three places, including, mysteriously, in a "hidden" text box in that graph I showed earlier:

I'm not making this up: I stumbled upon a text box which is invisible, but if you select it with the cursor, you find it contains "escitalopram"! I have no idea what the story behind that is, but at best it is seriously sloppy.

Come on Janssen. Raise your game. In the glory days of dodgy antidepressant research, your rivals were (allegedly) concealing data on suicides and brushing whole studies under the carpet, to make their drugs look better. Despicable, but at least it had a certain grandeur to it.

ResearchBlogging.orgBonaventure, P., Dugovic, C., Kramer, M., De Boer, P., Singh, J., Wilson, S., Bertelsen, K., Di, J., Shelton, J., Aluisio, L., Dvorak, L., Fraser, I., Lord, B., Nepomuceno, D., Ahnaou, A., Drinkenburg, W., Chai, W., Dvorak, C., Carruthers, N., Sands, S., and Lovenberg, T. (2012). Translational evaluation of JNJ-18038683, a 5-HT7 receptor antagonist, on REM sleep and in major depressive disorder Journal of Pharmacology and Experimental Therapeutics DOI: 10.1124/jpet.112.193995

Wednesday, 2 May 2012

Spurious Positive Mapping of the Brain?

Many fMRI studies could be giving false-positive results according to an important new paper from Anders Eklund and colleagues: Does parametric fMRI analysis with SPM yield valid results?—An empirical study of 1484 rest datasets.

The authors examined the SPM8 software package, probably the most popular tool for analyzing neuroimaging data.

Their approach was beautifully simple. They wanted to check how often conventional analysis of fMRI would "find" a signal when there wasn't really anything happening. So they took data from nearly 1,500 people who were scanned when they were just resting, and saw what would happen if you looked for "task related" activations in those scans, even though there was in fact no task. It's a very clever use of the resting state data.

Eklund et al ran the analysis many thousands of times, under various different conditions. This is the key finding:

This shows the proportion of analyses which produced significant "activations" associated with various different "tasks". In theory, the false positive rate should be way down at the bottom at 5% in each case. That's the error rate they told SPM8 to provide. As you can see, it was often much higher. Oh dear.

The error rate depended on two main things. Most important was the task design. Block designs were much worse than event-related designs (see the labels at the bottom: B1,2,3,4 are block, E1,2,3,4 are event.) The longer the blocks, the more errors. B4, the most error-ridden design of all, corresponds to 30 second blocks.

That's bad news because that's a very common design.

Secondly, the repeat time (TR) mattered, especially for block designs. The TR is how long it takes to scan the whole brain once. The longer the TR, the better, the data showed: 1 second TRs are really dodgy. Luckily, they are rarely used. 2 seconds is OK for most event-related designs, but block designs really suffer. 3 seconds is even better.

Because most fMRI studies today use 2-3 second TRs, this is somewhat reassuring, but for block design B4 the error rate was still up to 30% even with TR=3. Oh dear, oh dear.

So what went wrong? It's complicated, and you should read the paper, but in a nutshell the problem is that fMRI data analysis assumes that there are only two sources of data: the real brain activation signal, and white noise. The key assumption is that it's white noise, which essentially means that it is random at any moment in time: knowing about what the noise did in the past tells you nothing about what it will do in the future. "Random" noise that's actually correlated with itself over time is not white noise.

Now noise in the brain is certainly not white, for various reasons, including the effects of breathing and heart rate (which of course are cyclical, not random.) All fMRI analysis packages try to correct for this - but Eklund et al have shown that SPM8's approach doesn't manage to do that, at least for many designs.

What about rival fMRI software like FSL or BrainVoyager? We don't know. They use different approaches to noise modelling, which might mean they do better, but maybe not.

And the really big question: does this mean we can't trust published SPM8 results? Does SPM stand for Spurious Positive Mapping? Well, that's also not clear. All of Eklund et al's analyses were based on single subject data. But most fMRI studies pool the results from more like 20 or 30 subjects. Averaging over many subjects might make the false positives cancel out, but we don't yet know if that would solve the problem or only lessen it.

ResearchBlogging.orgEklund, A., Andersson, M., Josephson, C., Johannesson, M., and Knutsson, H. (2012). Does parametric fMRI analysis with SPM yield valid results?—An empirical study of 1484 rest datasets NeuroImage DOI: 10.1016/j.neuroimage.2012.03.093

Wednesday, 4 April 2012

Co-Vary Or Die

I've just come across a striking example of why correcting for confounding variables in statistics might not sound exciting, but can be a matter of life and death.

Imagine you're a doctor or researcher working with HIV/AIDS. You're taking a sample of blood from a HIV+ patient when you slip and, to your horror, jab yourself with a bloodied needle. What do you do?

In a 1997 study, researchers Cardo et al studied hundreds of cases of this kind of accidental HIV exposure ("needlestick injuries") in medical and scientific workers. They wanted to find differences between the people who contracted the virus, and the ones who didn't.

One factor they considered was post-exposure prophylaxis - taking HIV drugs as soon as possible after a suspected exposure. Now these drugs were still pretty new in 1997, and it wasn't clear how well they prevented infection, as opposed to just delaying symptoms. Many people with needlestick injuries were offered a course of drugs - but did they work?

Cardo et al's raw data found no significant benefit
By univariate analysis, there was no significant difference between case patients and controls in the use of zidovudine [AZT, the first HIV drug] after exposure.
But it turned out that this was due to confounding variables. When they corrected for other factors...
Infected case patients were significantly less likely to have taken zidovudine than uninfected controls (odds ratio 0.19, P=0.003). This is a classic example of confounding, since the adjusted odds ratio differed from the crude odds ratio (0.7) because zidovudine use was more likely among both case patients and controls after exposure characterized by one or more of the four risk factors in the model.
So while people who took zidovudine were just as likely to catch HIV than ones who didn't, they were also more severely exposed to the virus i.e. by being exposed to a greater quantity of blood, or a deeper wound. People were more likely to decide to take it after severe exposures. Zidovudine actually dramatically reduced the risk.

Post-exposure prophylaxis has since become standard procedure and it has undoubtedly saved many lives since. Without statistical correction, it might have taken longer for people to see the benefits.

In summary, I guess what I'm saying is, remember to correct for confounds - or die.

ResearchBlogging.orgCardo DM, Culver DH, Ciesielski CA, Srivastava PU, Marcus R, Abiteboul D, Heptonstall J, Ippolito G, Lot F, McKibben PS, and Bell DM (1997). A case-control study of HIV seroconversion in health care workers after percutaneous exposure. Centers for Disease Control and Prevention Needlestick Surveillance Group. The New England journal of medicine, 337 (21), 1485-90 PMID: 9366579

Friday, 30 March 2012

The Geography of Faces

How much can you tell about where someone comes from, just from their face?

The other day I was in London and came across a group of young people in Muslim attire who were waving (or in some cases wearing) a particular flag. I thought it was the Iranian flag, but, I thought, they didn't look Iranian. They looked more like Somalis, but it certainly wasn't the blue and white Somali flag. I decided that maybe they were some kind of pro-Iranian demonstrators, but I later worked out that it was the flag of the unrecognised state of Somaliland.

This got me thinking about how reliable these "they look they're from..." judgements are.

Clearly on a basic level, we can usually tell which continent someone's ancestors were from, in terms of the familiar "races" of Europeans, Africans, East Asians etc. But what about shorter distances?


Could you tell, just from looking at them (and setting aside dress, hairstyle, jewellery etc.) whether someone was from Spain as opposed to France? Korea or Japan? Russia or Germany?

I can only speak for England, but there's certainly a vague but widespread belief that every part of Europe has a  distinct 'look'. In the past, people were very fond of talking about that kind of thing; today, we're rather embarrassed by the idea but the belief lives on.

I don't know, but I'd be very surprised if there weren't analogous beliefs in other countries.

But how accurate are these folk beliefs, really?

Supposing you were the world expert on human faces - or suppose you were a supercomputer with face-recognition software and access to Facebook's entire dataset. How accurately could you place someone's origins on the map, on average? To within 1000 km? 100? With what degree of accuracy? In an ideal world, could the ultimate face-placer judge someone as French vs German 75% of the time? 90%? Or only slightly better than chance?

I suspect that if you researched this, you'd find that a supercomputer could do very well, in most parts of the world, but that the majority of actual people are less accurate than they think they are.

Wednesday, 21 March 2012

Brain Scanning - Just the Tip of the Iceberg?

Neuroimaging studies may be giving us a misleading picture of the brain, according to two big papers just out.


By big, I don't just mean important. Both studies made use of a much larger set of data than is usual in neuroimaging studies. Thyreau et al scanned 1,326 people. For comparison, a lot of fMRI studies have more like n=13. Gonzalez-Castillo et al, on the other hand, only had 3 people - but each one was scanned while performing the same task 500 times over.

Both studies found that pretty much the whole brain "lit up" when people are doing simple tasks. In one case it was seeing videos of people's faces, in the other it was deciding whether stimuli on the screen were letters or numbers.

With all that data, the authors could detect effects too small to be noticed in most fMRI experiments, and it turned out that pretty much everywhere was activated. The signal was stronger in some areas than others, but it wasn't limited to particular "blobs".

So conventional fMRI experiments may just be showing us the tip of the iceberg of brain activity. In a small study, only the strongest activations pass the statistical threshold to show up as blobs, but that doesn't mean the rest of the brain is inactive. It just means it's less active. The idea that only small parts of the brain are 'involved' in any particular task may be a statistical artefact.

In fact, I wonder if the whole idea of treating statistically significant blobs as different from nearly-significant areas is itself a form of the error of interacting effects?

As if that wasn't enough, Gonzalez-Castillo further show that there are lots of activations in the brain - even to very simple stimuli - that might go undetected in conventional studies, because they don't follow the time-course predicted by the usual models.

Have a look -


This shows the average neural activation from various regions of the brain during a letter-number task. The two areas I've highlighted in red are the primary visual cortex, and they do follow the expected 'boxcar' pattern - the brain is active when the stimuli are on the screen, inactive when they're not. But you can see that all kinds of other brain areas are also responding to the stimuli - just in different ways.

For example, the left primary motor cortex was activated during the task. That area controls the right hand, and that makes sense, as people responded by pressing buttons with the right hand. But interestingly, the same area on the other side of the brain was deactivated at exactly the same time, even though people weren't doing anything with their left hand.

These papers illustrate the fact that conventional fMRI is a blunt instrument that often only tells us about the most straightforward events that happen in the brain. A bit like how we only hear the shouts and screams from through our neighbor's walls, not their normal conversations, which aren't loud enough to reach our ears.

That's the bad news, but every blob has a silver lining. fMRI is clearly more powerful than most neuroscientists have realized, and this holds out hope for cracking some of the trickiest questions. As Gonzalez-Castillo et al put it
This result helps narrow the gap between thousands of fMRI manuscripts showing limited activation in response to tasks and cognition theories that defend that cognition—understood as the process of “configuring the way in which sensory information becomes linked to adaptive responses and meaningful experiences”—can only result from the distributed collaboration of primary sensory, upstream and downstream unimodal, heteromodal, paralimbic, and limbic regions... [we were able to] switch from a regime where activity detection relates primary to sensory processing to a more sensitive regime, where activity detection includes also cognitive processes with subtler BOLD signatures.
Link: See also the interesting discussion here: Surely, God loves the .06 (blob) nearly as much as the .05.


ResearchBlogging.orgThyreau, B., Schwartz, Y., Thirion, B., Frouin, V., Loth, E., Vollstädt-Klein, S., Paus, T., Artiges, E., Conrod, P., Schumann, G., Whelan, R., and Poline, J. (2012). Very large fMRI study using the IMAGEN database: Sensitivity–specificity and population effect modeling in relation to the underlying anatomy NeuroImage DOI: 10.1016/j.neuroimage.2012.02.083

Gonzalez-Castillo, J., Saad, Z., Handwerker, D., Inati, S., Brenowitz, N., and Bandettini, P. (2012). Whole-brain, time-locked activation with simple tasks revealed using massive averaging and model-free analysis Proceedings of the National Academy of Sciences DOI: 10.1073/pnas.1121049109

Tuesday, 31 January 2012

Voodoo Neuroscience Revisited

Two years ago, neuroscientists were shaken by the appearance of a draft paper showing that half of the published work in a particular field had fallen prey to a major statistical error.


Originally called "Voodoo Correlations in Social Neuroscience", it ended up with the less snappy name of Puzzlingly high correlations in fMRI studies of emotion, personality, and social cognition. I prefer the old title.

The error in question is now known variously as the "circular analysis problem", "non-independence problem" or "double-dipping" although I still call it the "voodoo problem". In a nutshell it arises whenever you take a large set of data, search for data points which are statistically significantly different from some baseline (null hypothesis), and then go on to perform further statistics only on those significant data points.

The problem is that when you picked out the statistically significant observations, you selected the data points that were especially "good", so if you then do some more analyses only on those data, you are almost guaranteed to find something "good". To avoid this you need to make sure that your second analysis is truly independent of your first one.

Anyway, Vul and Pashler, the main authors of the original voodoo article, have just written a short piece in NeuroImage offering some reflections on the paper and the aftermath. They don't make any major new arguments but it's a good read. Particularly fun is their explanation of what inspired them to look into the voodoo problem:
In early 2005 a speaker in our department reported that BOLD activity in a small region of the brain can account for the great majority of the variance in speed with which subjects walk out of the experiment several hours later (this finding was never published as far as we know). The implications of this result struck us as puzzling, to say the least: Are walking speeds really so reliable that most of their variability can be predicted? Does a focal cortical region determine walking speeds? Are walking speeds largely predetermined hours in advance? These implications all struck us as far-fetched...
But they reveal that it was one paper in particular that set them off voodoo-hunting
Our interest in probing the matter was further whetted by an episode occurring a short while later: Grill-Spector et al. (2006) reported that individual voxels in face selective regions have a variety of stable stimulus preferences; in a critical commentary, Baker et al. (2007) found that the analysis used to ascertain this fact implicitly built these conclusions into the method, such that the same analysis applied to noise data (voxels from the nasal cavity) revealed a similar variety of stable preferences. It occurred to us that a similar circularity might underlie the puzzlingly high correlations.

To their credit, Grill-Spector et al quickly accepted Baker et al's criticism and admitted that some of their original conclusions had been wrong.

ResearchBlogging.orgVul, E., and Pashler, H. (2012). Voodoo and circularity errors NeuroImage DOI: 10.1016/j.neuroimage.2012.01.027

Sunday, 11 December 2011

Do Antidepressants Make Some People Worse?

Antidepressants may help depression in some people but make it worse for others, according to a new paper.

This is a tough one so bear with me.

Gueorguieva, Mallinckrodt and Krystal re-analysed the data from a number of trials of duloxetine (Cymbalta) vs placebo. Most of the trials also had another antidepressant (an SSRI) as well. And the SSRIs and duloxetine seemed to be indistinguishable so from now on I'll just call it antidepressants vs. placebo as the authors did.

People on placebo got, on average, moderately better over 8 weeks.

People on antidepressants fell into two classes. The largest class got, on average, a lot better. But about 25% did poorly, staying just as depressed as before. This "nonresponder" group did much worse than the placebo group - again on average. Here you can see the mean "trajectories" of depression symptoms (HAMD scores) in the three groups:

This raises the scary possibility that while antidepressants are helping some people, they're harming others. But hang on. It's complicated.

First off, maybe this is all a statistical illusion. When the authors say that the people on drug fell into two classes, what they mean is that when you try to model the data according to a certain mathematical model, assuming either 1, 2, 3 or 4 underlying classes, the 2 class solution was the best fit. While for placebo a 1 class solution was best.
We considered linear, quadratic, and cubic trends over time, with between 1 and 4 trajectory classes. We also considered piecewise models with a change point at 2 weeks, linear change before week 2, and quadratic change after week 2. The selection of the best model was based on the Schwartz-Bayesian information criterion and on the Lo-Mendell-Rubin (LMR) likelihood ratio test...
That's nice... but they don't present the raw data. They don't tell us whether, looking at the individual trajectories of people on antidepressants, you'd actually see two classes. What I want is a graph of how likely people are to get better by a certain amount. If Gueorguieva et al are right, I want it to look like this i.e. bimodal -


We're not shown this graph. I'll eat my hat if it does look like that, frankly, because if it did people would have noticed the bimodality in antidepressant trials ages ago.

True, statistical models can tell us things that aren't obvious by inspection, so even if this isn't what the data look like, they might still be right. It could be that the two "peaks" are so broad, and there's so much random noise, that they blur into one.

However, it's also true that you can fit an infinite number of models to any set of data and at some point you have to step back and say - am I making this more complicated than it needs to be?

It could be that a 2-class model is better than a 1-class model for the people on antidepressants, but only because they're both crap, and really, every patient has a different, unpredictable trajectory which is poorly captured by such models.

Let's assume however that this is true. What would it mean?

Firstly, the fact that one class of people on antidepressants does worse than people on placebo doesn't mean that antidepressants are harming them. The authors miss this point, when they say
there are 2 trajectories for patients treated with antidepressants and 1 trajectory for patients treated with placebo [so] some patients would seem to be more effectively treated with placebo than with a serotonergic antidepressant.
But that's fallacious. It treats a purely statistical entity as representing individual people. Suppose that what antidepressants do is to take people who, on placebo, would have improved a bit, and make them improve a bit more than they otherwise would have. You'd then end up with more people doing well, but also fewer people doing moderately because they'd have been "moved up" out of the middle ground.

That "nudging people off the fence" could lead to a bimodal distribution and two distinct classes. But in this case the people doing badly would have done badly either way. The drug didn't make them do badly, it just made doing-badly into a class. On the other hand it's consistent with antidepressants doing real harm. We can't tell.

We do know that other randomized controlled trials show very convincingly that in a small minority of people, mostly but not exclusively young people, antidepressants do worsen suicidal thoughts and behaviours. So it's plausible. But we just don't know yet.

What worries me is that this paper is the latest in a series of attempts  to use, well, creative statistical approaches to antidepressant trial data. This one is nowhere near as dodgy as the Cherrypicker's Manifesto I discussed last year, but it cites that paper and others by the same group. The first sentence of the Abstract of this paper makes the intention clear:
The high percentage of failed clinical trials in depression may be due to high placebo response rates and the failure of standard statistical approaches to capture heterogeneity in treatment response.
In other words, the reason clinical trials of new antidepressants often fail to show a benefit over placebo is not because the drugs are crap but because the statistics aren't subtle enough. And you can see where this is going: if only we could use statistical models to find the people who do benefit from antidepressants, and compare them to placebo, there'd be no problem...

ResearchBlogging.orgGueorguieva R, Mallinckrodt C, and Krystal JH (2011). Trajectories of depression severity in clinical trials of duloxetine: insights into antidepressant and placebo responses. Archives of General Psychiatry, 68 (12), 1227-37 PMID: 22147842