Showing posts with label FixingScience. Show all posts
Showing posts with label FixingScience. Show all posts

Thursday, 15 February 2024

How Not To Do Great Science (The Lost Post)

This post was originally published on Discover Magazine on September 16th 2013, but has since vanished (although most of my other Discover posts are still available). Luckily, I saved a backup. So here's the original "How Not To Do Great Science".

---

This post is a bit special. For the first time ever, I've collaborated with an artist, Erene Stergiopoulos. Her webcomic is here and she's on Twitter here. I think you'll agree that the artistic standard is a little higher than I usually achieve. Anyway, here's what we did:



It would be silly to expect that every architect should finish buildings at a certain rate. That would make it impossible to anyone to build certain things. Some things take longer to build than others, and most great things take a great deal of time. Faced with a sufficiently demanding quota, builders might be reduced to rushing out follies that might look impressive from a distance, but that are no more than hollow shells. Yet, as silly it would be to make uniform demands of architects, this is what is happening to scientists.

Rather than build, scientists are expected to publish - and publish fast - or perish. My worry (and that of many others) is that the pressure to publish often fundamentally changes not just how much scientists write, but what they can write about. It turns researchers into prolific doers of small deeds, but it leaves them little time to think about, let alone complete, great works. Though the mills of God grind slowly...

Yet the problem is not just the speed of science today, but also the direction: go to a scientific conference and you'll see perfectly good data in the process of being oversold, misinterpreted, and p-hacked into a 'publishable' form.

Much has been said about how this leads to false positives - impressive follies that don't stand up to scrutiny. What's less discussed - and the point of this piece - is the opportunity cost. New theories come out of attempts to explain 'negative' data - negative from the perspective of the old theory. Null results are the foundations of future progress, but only if they are allowed to lie there awhile; not if they are torn up and used to prop up tottering old structures.

Sunday, 3 February 2013

Unilaterally Raising the Scientific Standard

For years, I and others have been arguing that the current system of publishing science is broken. Publishing and peer-reviewing work only after the study's been conducted and the data analysed allows bad practices - such as selective publication of desirable findings, and running multiple statistical tests to find positive results - to run rampant.

So I was extremely interested when I received an email from Jona Sassenhagen, of the University of Marburg, with subject line: Unilaterally raising the standard.

Sassenhagen explained that he's chose to pre-register a neuroscience study on a public database, the German Clinical Trials Register (DRKS).

His project, Alignment of Late Positive ERP Components to Linguistic Deviations ("P600"), is designed to use EEG to test whether the brain generates a distinct electrical response - the P600 - in response to seeing grammatical errors. The background here is that the P600 certainly exists, but people disagree on whether it's specific to language; Sassenhagen hopes to find out.

By publicly announcing the methods he'll use before collecting any data, Sassenhagen has, in my view, taken a brave and important step towards a better kind of science.

Already, most journals require trials of medical treatments to be publicly pre-registered, and the DRKS is one such registry. This study, however, is 'pure' neuroscience with nothing clinical about it, so it doesn't need to be registered - Sassenhagen just did it voluntarily.

Further, I should point out that he offered to pre-register his data analysis pipeline too by sending it to me. Unfortunately, I didn't reply to the email in time... but that was purely my fault.

I very much hope and expect that others will follow in his footsteps. Unilaterally adopting preregistration is one of the ways that I've argued reform could get started. As I said:
This would, at least at first, place these adopters at an objective disadvantage. However, by voluntarily accepting such a disadvantage, it might be hoped that such actors would gain acclaim as more trustworthy than non-adopters.
Pre-registration puts you at a disadvantage - insofar as it limits your ability to use bad practice to fish for positive results. It means you can't cheat, essentially, which is a handicap if everyone else can.

I don't know if this is the first time anyone's opted in to registering a pure neuroscience study, but it's certainly the first case I know of it being done for an entirely new experiment.

There have, however, recently been many pre-registered attempts to replicate previously published results e.g. the Reproducibility of Psychological Science; the 'Precognition' Replications; and an upcoming special issue of Frontiers in Cognition.

Replications are good, registered ones doubly so - but they're not enough to fix bad practice on their own. To do that we need to work on the source, original scientific research.

Thursday, 24 January 2013

Is Medical Science Really 86% True?

The idea that Most Published Research Findings Are False rocked the world of science when it was proposed in 2005. Since then, however, it's become widely accepted - at least with respect to many kinds of studies in biology, genetics, medicine and psychology.

Now, however, a new analysis from Jager and Leek says things are nowhere near as bad after all: only 14% of the medical literature is wrong, not half of it. Phew!

But is this conclusion... falsely positive?

I'm skeptical of this result for two separate reasons. First off, I have problems with the sample of the literature they used: it seems likely to contain only the 'best' results. This is because the authors:
  • only considered the creme-de-la-creme of top-ranked medical journals, which may be more reliable than others.
  • only looked at the Abstracts of the papers, which generally contain the best results in the paper.
  • only included the just over 5000 statistically significant p-values present in the 75,000 Abstracts published. Those papers that put their p-values up front might be more reliable than those that bury them deep in the Results.
In other words, even if it's true that only 14% of the results in these Abstracts were false, the proportion in the medical literature as a whole might be much higher.

Secondly, I have doubts about the statistics. Jager and Leek estimated the proportion of false positive p values, by assuming that true p-values tend to be low: not just below the arbitrary 0.05 cutoff, but well below it.

It turns out that p-values in these Abstracts strongly cluster around 0, and the conclusion is that most of them are real:

But this depends on the crucial assumption that false-positive p values are different from real ones, and equally likely to be anywhere from 0 to 0.05.
"if we consider only the P-­values that are less than 0.05, the P-­values for false positives must be distributed uniformly between 0 and 0.05."

The statement is true in theory - by definition, p values should behave in that way assuming the null hypothesis is true. In theory.

But... we have no way of knowing if it's true in practice. It might well not be.

For example, authors tend to put their best p-values in the Abstract. If they have several significant findings below 0.05, they'll likely put the lowest one up front. This works for both true and false positives: if you get p=0.01 and p=0.05, you'll probably highlight the 0.01. Therefore, false positive p values in Abstracts might cluster low, just like true positives.

Alternatively, false p's could also cluster the other way, just below 0.05. This is because running lots of independent comparisons is not the only way to generate false positives. You can also take almost-significant p's and fudge them downwards, for example by excluding 'outliers', or running slightly different statistical tests. You won't get p=0.06 down to p=0.001 by doing that, but you can get it down to p=0.04.

In this dataset, there's no evidence that p's just below 0.05 were more common. However, in many other sets of scientific papers, clear evidence of such "p hacking" has been found. That reinforces my suspicion that this is an especially 'good' sample.

Anyway, those are just two examples of why false p's might be unevenly distributed; there are plenty of others: 'there are more bad scientific practices in heaven and earth, Horatio, than are dreamt of in your model...'

In summary, although I think the idea of modelling the distribution of true and false findings, and using these models to estimate the proportions of each in a sample, is promising, I think a lot more work is needed before we can be confident in the results of the approach.

Friday, 18 January 2013

How (Not) To Fix Social Psychology

British psychologist David Shanks has commented on the Diedrik Stapel affair and other recent scandals that have rocked the field of social psychology: Unconscious track to disciplinary train wreck,


Lots of people are chipping in on this debate for the first time at the moment, but peoples' initial reactions often fall prey to misunderstandings that can stand in the way of meaningful reform - misunderstandings that more considered analysis has exposed.

For example, Shanks writes:
[despite claims that] social psychology is no more prone to fraud than any other discipline, but outright fraud is not the major problem: the biggest concern is sloppy research practice, such as running several experiments and only reporting the ones that work.
It's true that fraud is not the major issue, as I and many others have said. But bad practice, such as p-value fishing, is in no way "sloppy" as Shanks says. Running multiple experiments to get a positive results is a sensible and effective strategy for getting positive results; that's why so many people do it. And so long as scientists are required to get such findings to get publications and grants, it will continue.

Behavior is the product of rewards and punishments, as a great psychologist said. We need to change the reinforcement schedule, not berate the rats for pressing the lever.

Earlier, Shanks writes that evidence of unconscious influences on human behaviour - a popular topic in Stapel's work and in social psychology generally -
is easily obtained because it usually rests on null results, namely finding that people's reports about (and hence awareness of) the causes of their behaviour fail to acknowledge the relevant cues. Null results are easily obtained if one's methods are poor.
Thus journals have in recent years published extraordinary reports of unconscious social influences on behaviour, including claims that people are more likely to take a cleansing wipe at the end of an experiment in which they are induced to recall an immoral act [etc]...
...failures to replicate the effects described above have been reported, though often papers reporting such failures are rejected out of hand by the journals that published the initial studies. I await with interest the outcome of efforts to replicate the recent claim that touching a teddy bear makes lonely people more sociable.
Here Shanks first says that null results can easily result from poorly-conducted experiments, and then criticizes journals for not publishing null results that represent failures to replicate prior claims! But null replications are very often rejected because a reviewer says, like Shanks, "This replication was just poorly-conducted, it doesn't count." Shanks (unconsciously no doubt) replicates the problem in his article.

So what to do? Again, it's a systemic problem. So long as we have peer-reviewed scientific journals, and the peer-review takes place after the data are collected, it will be open to reviewers to spike results they don't like - generally although not always null ones. If reviewers had to judge the quality of a study before they knew what it was going to find, as I've suggested, this problem would be solved.

Other people have great ideas for fixing science of their own. The problem is structural, not a failing on the part of individual scientists, and not limited to social psychology.

Thursday, 22 November 2012

The Perils of Sharing Brain Scans

A fascinating paper by neuroscientists Van Horn and Gazzaniga chronicles their pioneering, but not entirely successful, attempt to get researchers sharing their brain scans: Why share data? Lessons learned from the fMRIDC.


It all started in 1999 when, along with some colleagues, they decided that the time was right for data sharing in neuroimaging. They got some public funding, and tried to get various major neuroscience journals to require that anyone publishing an fMRI study should make their data available to the fMRI Data Consortium (fMRIDC).

By making it mandatory, they'd ensure that there was no selection bias. Requirements to post raw data were already common in other fields of science like genetics and crystallography. So, they thought, why can't it happen here?

However, it didn't go down very well:
Upon becoming aware of our efforts and goals, fMRI researchers angered by journal requirements to provide copies of the fMRI data from their published articles began a letter writing campaign seeking to muster opposition  an effort which was featured in the news and editorial sections of several influential journals.
Editorials and commentaries over fMRI data sharing were aired in the pages of Science, Nature Neuroscience etc. expressing concern over the data sharing requirement, over what possession of the data implied, human subject concerns, and, if databasing was to be conducted at all, how it should be conducted “properly”.
This was all before my time, sadly. It sounds like a grand old academic street-fight. No doubt those on the other side remember it differently from how it's presented here, though.
The reactions of our colleagues caught us somewhat off guard. We were honestly surprised by the  negative and hostile response when we had believed that creation of a data archive would be of benefit to the neuroimaging community. Perhaps, they had a point.
Maybe the field wasn't ready for fMRI data sharing? Perhaps, it was too early to start such a project? We struggled with how best to move forward or whether to move forward at all.
Anyway, they decided they would continue, on a more modest scale. fMRIDC ended up with data from about 100 fMRI studies by the time the funders pulled the plug in 2006, only a fraction (and perhaps an unrepresentative one) of the papers published, but still, it's something.

As I said, I missed out on this debate, but if I had been in the field 10 years ago, I suspect I'd have been on the extreme wing of the pro-sharing faction, the Montagnard to the founders' Girondism. My view is that no researcher owns their data. The only person who owns the results of a brain scan is the person whose brain it is.

If I could wave a magic wand, I'd put a chip in every MRI scanner that automatically uploaded all scans to a public database as soon as they appeared (with the subject's consent, and/or with personal information about the subject stripped out). I'd fix all the software such that every time someone ran an analysis, it was publicly logged. Total transparency is best for science, I believe, and it would also make scientists lives  easier, once they got over the initial shock.

Sadly that's not possible... yet... but data sharing is a noble cause and, as Van Horn and Gazzaniga point out, even if fMRIDC is dead, the idea lives on with many new initiatives emerging, hydra-like, in its place.

ResearchBlogging.orgVan Horn JD, and Gazzaniga MS (2012). Why share data? Lessons learned from the fMRIDC. NeuroImage PMID: 23160115

Thursday, 8 November 2012

Blogging's First Academic Paper

In an historic achievement, I can announce that I have become (to my knowledge) the first blogger ever to publish in a peer-reviewed academic journal under a blogging pseudonym.
Skeptic, N. (2012) The Nine Circles of Scientific Hell Perspectives on Psychological Science 7 (6) 643-644
This is based on a post from two years ago (far and away the most popular post I've ever done).

Now as historic achievements go, this is fairly niche, but I do think it's important.

Most of the problems with the way science works today are problems of communication. We're trying to do 21st century work with a 19th century publishing model, and the cracks are showing.

Academic papers are a fine way of presenting the final results of research - they're here to stay. But scientists ought to be communicating (with each other and with the public) in many other ways as well, and I think that anything we can do to break down the hegemony of the 'final paper' - whether it be blogging, arxiv, the Open Science Framework or raw data sharing - is a step forward.

In that regard, I think it's great that the boundaries between 'real' papers and alternative forms of scientific communication have just become a little blurrier.

The piece appears in a special issue of Perspectives on Psychological Science full to bursting with other papers that readers of this blog are likely to find interesting... and it's all free to access (at least for now).

ResearchBlogging.orgNeuroskeptic (2012). The Nine Circles of Scientific Hell Perspectives on Psychological Science, 7 (6), 643-644 DOI: 10.1177/1745691612459519

Saturday, 20 October 2012

When Replication Goes Bad

How to ensure that results in psychology (and other fields) are replicated has become a popular topic of discussion recently. There's no doubt that many results fail to replicate, and also, that people don't even try to replicate findings as much as they should.


Yet psychologist Gregory Francis warns that replication per se is not always a good thing: Publication bias and the failure of replication in experimental psychology
Among experimental psychologists, successful replication enhances belief in a finding, while a failure to replicate is often interpreted to mean that one of the experiments is flawed. This view is wrong.

Because experimental psychology uses statistics, empirical findings should appear with predictable probabilities. In a misguided effort to demonstrate successful replication of empirical findings and avoid failures to replicate, experimental psychologists sometimes report too many positive results.

Rather than strengthen confidence in an effect, too much successful replication actually indicates publication bias, which invalidates entire sets of experimental findings...

Even populations with strong effects should have some experiments that do not reject the null hypothesis. Such null findings should not be interpreted as failures to replicate, because if the experiments are run properly and reported fully, such nonsignificant
findings are an expected outcome of random sampling... If there are not enough null findings in a set of moderately powered experiments, the experiments were either not run properly or not fully reported. If experiments are not run properly or not reported fully, there is no reason to believe the reported effect is real.
Say you took a pack of playing cards and removed half the red cards. Your pack would now be 2/3rds black, so if you took a random sample of cards, say a poker hand of 5 cards, then you'd expect more blacks than reds (a significant 'effect' of color). But you'd still expect some reds, and some random hands would in fact be entirely red, just by chance. If someone claimed to have drawn 10 random hands and they'd all been mainly black, that would be implausible - "too good".

Francis's approach is a bit like Uri Simonsohn's method for detecting fraudulent data - they both work on the principle that "If it's too good to be true, it's probably false" - but they differ in their specifics, and I believe that we should not conflate fraud with publication bias... so let's not get carried away with the parallels.

Earlier this year, Francis wrote a critical letter about a paper published in PNAS purporting to show that wealthier Americans are less ethical. He argued that the paper's results were "unbelievable" - it reported on the results of seven separate experiments, all of which showed a small, but significant, effect in favour of the hypothesis.

Even if rich people really were meaner, Francis said, the chance of 7/7 experiments being positive is very low: just by chance, you'd expect some of them to show no difference (given that the size of the difference in those seven was low, with a lot of overlap between the groups). Francis suggested that the authors may have run more than seven experiments, and only published the positive ones; the authors denied this in their Letter.

Anyway, in the new paper, Francis expands on this approach in much more detail, drawing from this 2007 paper, and suggests a Bayesian approach that might help mitigate the problem.

ResearchBlogging.orgFrancis G (2012). Publication bias and the failure of replication in experimental psychology. Psychonomic Bulletin and Review PMID: 23055145

Sunday, 14 October 2012

More on False Positive Neuroimaging

Back in June, I warned that the ever-increasing number of clever methods for analyzing brain imaging data could be a double-edged sword:
Recently, psychologists Joseph Simmons, Leif Nelson and Uri Simonsohn made waves when they published a provocative article called False-Positive Psychology - Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant.
It explained how there are so many possible ways to gather and analyze the results of a  simple psychology experiment that, even if there's nothing interesting really happening, it'll be possible to find some "significant" positive results purely by chance...
The problem's not just seen in psychology however, and I'm concerned that it's especially dangerous in modern neuroimaging research.
In a comment on that post, The Neurocritic pointed out that Michigan PhD student Joshua Carp had put forward the same argument in a conference presentation, several months previously.

Now Carp's published a paper on the topic: On the plurality of (methodological) worlds: estimating the analytic flexibility of fMRI experiments. It's free to access, so check it out.

Whereas I just talked the talk by listing lots of possible ways in which you could analyze a given set of data, Carp walked the walk, and actually did loads of analyses. He took a single dataset, the results of a simple experiment and looked at it in almost 7000 different ways. Each set of results was then thresholded to correct for multiple comparisons in 5 ways, for a grand total of 35,000 outputs.

The variants he considered ranged from how much smoothing to apply, to how to correct for head motion, and many more.

What happened? In a nutshell, the different options made a difference - and the variability was the largest in parts of the brain that were most activated (the "blobs" that lit up). In other words, analytic flexibility makes the most difference in the most interesting places. See the picture at the top.

The location of the maximum peak activation also varied. This is not unexpected, and not, in itself, that worrying - the great majority of the peaks clustered in a few small areas. However, it underlines that different options really can make a difference.

Carp concludes:
Nearly every voxel in the brain showed significant activation under at least one analysis pipeline. In other words, a sufficiently persistent researcher determined to find significant activation in virtually any brain region is quite likely to succeed...

If investigators apply several analysis pipelines to an experiment, and only report the analyses that support their hypotheses, then the prevalence of false positive results in the literature may far exceed the nominal rate. However, analytic flexibility only translates into elevated false positive rates when combined with selective analysis reporting. If researchers reported the results of all analysis pipelines used in their studies, then it would not be problematic.

To the author’s knowledge, there is no evidence that fMRI researchers actually engage in selective analysis reporting. But researchers in other fields do appear to pursue this strategy.
In my experience, fMRI researchers are actually fairly conservative in terms of using different analyses, and certainly I doubt anyone has ever run thousands of them just to get the result they want and I'd estimate that most published findings are not the result of more than a handful of 'attempts' at most.

However it's a serious concern that it could happen, and importantly it's getting ever-easier to do this, with the continuing increase in computer power making running an analysis quicker and cheaper than ever. As to what to do about it, Carp makes several suggestions, and here's one I made earlier...

ResearchBlogging.orgJoshua Carp (2012). On the plurality of (methodological) worlds: estimating the analytic flexibility of fMRI experiments Front. Neurosci. DOI: 10.3389/fnins.2012.00149

Wednesday, 3 October 2012

The Two Problems With Science


There's lots of concern at the moment over mistakes, misconduct and misbehaviour in science.
This concern is a good thing. There are serious, systemic problems with modern science as I and many others have long argued.

However, I worry that much of the recent discussion has failed to distinguish between two fundamentally distinct problems. On the one hand, we have outright fraud - i.e. making up data, or otherwise lying, breaking the basic rules of science.

On the other hand we have questionable practices such as: publication bias, p-value fishing, the File Drawer, sample size peeking, post-hoc storytelling, and all of the other dark arts that can lead to false positive science. These are permissible, even encouraged, by the current rules of doing and publishing science.

These two problems are similar in some ways - they're both "bad science", they both lead to failures to replicate, etc. - but in underlying essence they're very different, so much so that I'm not sure they can be usefully discussed in the same breath.

Fraud and questionable practices are different in terms of their harms. Fraud is a more serious act and it causes local harm, introducing major errors into the record. But in terms of its overall effects, I believe questionable practices are worse, as they systematically distort science: ensuring that, in some cases, it is difficult to publish anything but errors.

Fraud and questionable practices call for different solutions. Broadly speaking, fraudsters break the rules, so to stop them we need to enforce those rules, via deterrence, detection, and punishment - like with any criminal act. With questionable practices, it's the opposite: here the problem is the rules (or the lack of them), and the solution is to reform the system.

It's been suggested that fraud and questionable practices share a common cause in the "pressure to publish", the "publishing environment", the "culture" of modern science etc. But while this is a good explanation for questionable practices, I don't think this can explain fraud, any more than, say, the desire for money can explain theft.

Yes, thieves desire money, and yes they steal in order to get money, but everyone else wants money as well, yet most of us don't steal, so that's not an explanation. Frauds fake data to produce publications. But all scientists are under pressure to produce good publications and they always have been - which is why fraud is not new - what's changed recently is the criteria for a 'good' publication.

Now in retrospect, I blurred these distinctions somewhat with my own 9 Circles of Scientific Hell, in which I placed 6 questionable practices and 2 forms of misconduct on the same scale of "sinfulness". In fact there are two distinct hierarchies. In my defence though, that was a cartoon.

I think finance offers a great analogy here.


In finance, you have some people who break the rules. Bernie Madoff is the current poster boy for this. Such people harm others by outright criminal acts. But then we have the people who play by the rules, and still cause harm. The global financial crisis was in essence caused by all of the major American banks going all-in on a bet, and losing. Yet no-one broke the rules: the regulations allowed banks to gamble. The problem was not rule-breaking, but the rules (or lack thereof).

Here's the curious thing: the financial crisis did more harm than Madoff's scam, even though what Madoff did - theft by fraud - was more immoral than what the bankers did - gambling unwisely.

That's confusing to our ethical sense and our emotions (who should we feel more angry at? Who's 'worse'?) but it's really no surprise: precisely because what the banks did was above board, everyone did it so the damage was huge. If it had been illegal for banks to gamble all their money at once, individual banks might still have broken that rule, locally, but it's unlikely that the system would have been threatened.

Maybe you can see where I'm going with this: everyone following bad rules is often worse than individuals breaking good rules.

Science has its share of fraud. Hauser, Smeesters, Fujii - they broke good rules against such deceit. They are the Bernie Madoffs of science. But then there's 'questionable practices' like publication bias, p-value fishing, the File Drawer, and all the rest, which are allowed, but which are universally acknowledged to be bad for science. Scientists using these dark arts (and I don't know any who never do) may be the Lehman Brothers of science.

Sunday, 23 September 2012

Publication Bias in Animal Research

Publication bias has historically been thought of mostly in the context of clinical trials. But I have been banging on for the past 4 years about how it's a problem for more 'basic' science as well.


I'm not alone in my concerns as an interesting new paper reveals: Publication Bias in Laboratory Animal Research. The authors surveyed the approximately 3,000 Dutch scientists involved in research on laboratory animals. The response rate was about 20%.

When asked how much animal research ends up being published, university researchers estimated about half, but industrial scientists put it at only about 10% - which, if true, suggests that publication bias in Pharma animal work is extremely serious.

In terms of solutions, the survey considered two ideas which Neuroskeptic readers will be familiar with - public pre-registration of studies:
Mandatory anonymous publication of research protocols of all ethics-approved animal research experiments in a publicly available database
and also open access to all data:
Mandatory anonymous publication of a brief structured form in a publicly available database, that gave main results or explained why an experiment could not be completed
On average the surveyed researchers felt that these measures would aid scientific progress; improve the validity of the literature; and prevent wasteful duplication of effort - but they also worried that it would increase bureaucracy.

Now, bureaucracy is second only to bias on my list of Things I Hate About Science, so I share their concern - but I really think registration wouldn't have to involve any extra paperwork. In many cases, it could be implemented simply by making existing data public.

For instance, grant applications, and requests for ethical approval, already contain detailed a priori protocols in most cases. They could so easily be published (perhaps with certain details removed for confidentiality reasons) and turned into a powerful weapon against publication bias.

Having said that though - it easily could end up being needlessly complicated and obstructive, as so much of the scientific process unfortunately is today. It will all depend on how it's implemented.

This is why I think it's so important that, as scientists, we reform science ourselves, and get it right, rather than leaving it to the bureaucrats, who won't.

ResearchBlogging.orgTer Riet G, Korevaar DA, Leenaars M, Sterk PJ, Van Noorden CJ, Bouter LM, Lutter R, Elferink RP, and Hooft L (2012). Publication bias in laboratory animal research: a survey on magnitude, drivers, consequences and potential solutions. PloS one, 7 (9) PMID: 22957028

Saturday, 25 August 2012

Replication Alone Is Not Enough

Psychology has lately been hit by high-profile fraud scandals, and broader concerns over questionable research practices. Now the Society for Personality and Social Psychology (SPSP) has released a statement on "Responsible Conduct", and a task force has produced a report.

This is a start, and the SPSP is to be commended for facing up these problems (which affect many other fields) relatively early. However, neither of their documents contains much meat in my view.

Point One on the task force report is that "Replication is the key to building our science" and they suggest a "web site for depositing replications and failures to replicate" - but don't mention that various enterprising researchers have already made one. Nor do they tip their hats to the Open Science Initiative addressing just this issue. This makes me worried that they're planning to reinvent the wheel.

More fundamentally I disagree that replication is key to psychology or any field. Our goal should be replicability. Failure to replicate findings is a symptom of problems with those original findings, rather than being a problem in and of itself. Good results replicate; we want better results to be published.

In other words, we should strike at the root cause of invalid research, namely, the perverse incentives towards publishing as many eye-catching positive results with p values below 0.05 as possible by any means necessary. P-value fishing, selective reporting, post-hoc "prior hypotheses" and other questionable practices are a large part of what make unreplicable results.

We should encourage replication, but it's no panacea.

An overemphasis on replication, without addressing the incentives, could actually harm science. It could lead to scientists spending all their time worrying about the political drama of who's replicating who and why, and which questionable practices they can use to replicate their friends' data - rather than actually doing science.

This is why we shouldn't be satisfied with any reform effort that puts replication before replicability. If you can fudge a result, you can fudge the data a replication. How to fight questionable practices is another question but I've proposed reforms that I think would work, namely pre-registration of hypotheses, methods, and statistical analyses. Others have their own ideas.

A lesson from clinical medicine here. Clinical trials of new drugs adopted pre-registration, but only after they tried replication and it didn't work. Pharmaceutical regulators have long required multiple demonstrations of drug efficacy. One trial was not enough. Sounds good - but the problem was that drug companies just did lots of trials and analyses, picked the positive ones, and used them.

So in summary: replication is important, and we don't do enough of it, but replication alone is not enough to fix psychology.

Saturday, 28 July 2012

Catching Fraud: Simonsohn Says

Everyone's been talking about psychologist Uri Simonsohn and his role in the downfall of two scientific fraudsters.


When the story first broke, the methods Simonsohn used that allowed him to spot the dodgy data were mysterious - which only added to the buzz. The paper revealing the approach is now up online and it's a must-read. It's not often a statistics paper offers the train-wrecky schadenfreude of watching two fraudsters' careers come to a well-deserved end.

What's rather disturbing about the article, however, is that it doesn't really contain much that's new, in principle. Simonsohn used statistics to spot data in published papers that was, in effect, 'too good to be true'. He then followed up seemingly dodgy cases with some more stats, using simulations of what real data ought to look like, to verify that it was in fact made up. A simple idea in retrospect but one that's never been tried before. I don't think there's a single "Simonsohn method", rather, the paper uses multiple techniques, each one tailored to the particular data in question.

But it shouldn't have come to this. Someone else ought to have spotted that the data looked dodgy.

Take this table from one of Simonsohn's conquests, a soon-to-be-retracted paper by Lawrence J Sanna et al:

We now know that the data from Studies 2,3 and 4 were all made up. Each study compared 3 conditions, and what makes these data dodgy is that the standard deviations of the 3 sets of results for each study were almost identical. The chances of that happening are very low and it suggests that someone has (clumsily) made the data up.

I'm going to say that these data are obviously suspicious, at least to anyone who has worked with real data. Maybe you'll say that hindsight is 20/20, but Simonsohn didn't need hindsight and the stats he used were nothing remarkable. I'm not saying that to belittle his achievements, he deserves plenty of credit. But other people deserve blame.

Namely, whoever peer reviewed this paper should have spotted that these data looked unusual - and they should not have needed any special statistical tools to do so.

Simonsohn calls for journals to require that the raw data be made available for all published work, on the grounds that. That's a great idea - and not just because it would help catch bad science: it would facilitate proper research and teaching no end. But Simonsohn didn't need the raw data to detect these cases of fraud - he only checked the raw results to confirm the suspicions based on the published data.

Checking that the data are valid is the job of peer reviewers, and they dropped the ball. Instead Simonsohn had to conduct his own private crusade against fraud... a bit like Batman. Batman is awesome, but the point about Batman is that he's only needed because the police can't or won't cope on their own. He's not a superhero, he's just a guy with the will.

Peer reviewers are the police of science, but all too often, they're asleep on the job. Not just in psychology. Retraction Watch provides plenty of examples of published results in biology that were faked, often in comically crude fashion, and should have been obvious to anyone paying attention.

Peer reviewers are usually anonymous. I wonder if a policy of retrospectively naming and shaming the reviewers when a paper turns out to have been fraudulent, might help motivate them...?

Saturday, 21 July 2012

A Case Study in Voodoo Genetics

A new review of published studies looking at the relationship between a gene and brain structure offers a sobering lesson in how science goes wrong.

Dutch neuroscientists Marc Molendijk and colleagues took all of the studies that compared a particular variant, BDNF val66met, and the volume of the human hippocampus. It's a long story, but there are various biological reasons that these two things might be correlated.

It turns out that the first published reports found large genetic effects, but that ever since then, the size of the effects has dropped, with the latest studies finding no effects at all -


A cumulative meta-analysis confirms that as more studies on BDNF val66met have appeared, the overall effect estimate has steadily declined -


Finally, the authors found signs of publication bias: there were three small, imprecise studies that reported very large effects of the gene, but no such studies finding no effect (or a reverse effect). You'd expect that small and noisy studies would have a lot of random variation so they wouldn't all be positive even if there was a true effect; that all of the published ones were positive, suggests that null findings are out there, unreported.

Overall, this suggests that val66met probably isn't associated with hippocampus volume after all, and that the early studies showing that it was, were misleading. There's no reason to think that the early studies were wrong as such - they may have accurately reported an effect in the small sample of people they looked at, but it was only a chance finding.

I know a lot of neuroscientists who are now fairly skeptical of this whole genre of candidate gene studies; there was much excitement 5 or 10 years ago, but in retrospect, most of these studies were too small, and the publication process meant that it was the random chance findings that were most likely to get published.

However, we need to avoid any sense of complacency. Until we fix the scientific process, the same thing will happen again.

ResearchBlogging.orgMolendijk ML, Bus BA, Spinhoven P, Kaimatzoglou A, Voshaar RC, Penninx BW, van Ijzendoorn MH, and Elzinga BM (2012). A systematic review and meta-analysis on the association between BDNF val(66) met and hippocampal volume American journal of medical genetics B Neuropsychiatric genetics PMID: 22815222

Saturday, 30 June 2012

False Positive Neuroscience?

Recently, psychologists Joseph Simmons, Leif Nelson and Uri Simonsohn made waves when they published a provocative article called False-Positive Psychology


The paper's subtitle was "Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant". It explained how there are so many possible ways to gather and analyze the results of a (very simple) psychology experiment that even if there's nothing interesting really happening, it'll be possible to find some "significant" positive results purely by chance. Then you could publish those 'findings' and not mention all the other things you tried.

It's not a new argument, and the problem has been recognized for a long time as the "file drawer problem", "p-value fishing", "outcome reporting bias", and by many other names. But not much has been done to prevent it.

The problem's not just seen in psychology however, and I'm concerned that it's especially dangerous in modern neuroimaging research.

*

Let's assume a very simple fMRI experiment. The task is a facial emotion visual response. Volunteers are shown 30 second blocks of Neutral, Fearful and Happy faces during a standard functional EPI scanning. We also collect a standard structural MRI as required to analyze that data.

This is a minimalist study. Most imaging projects have more than one task, commonly two or three and maybe up to half a dozen, as part of one scan. If one task failed to show positive results, it need never be reported at all, so additional tasks would compound the problems here.

Our study is comparing two groups: people with depression, and healthy controls.

How many different ways could you analyze that data? How much flexibility is there?

*
First some ground rules. We'll stick to validated, optimal approaches. There are plenty of commonly used less favoured approaches, like using uncorrected thresholds (and then, which ones?) or voodoo stats, but let's assume we want to stick to 'best practice'.

As far as I can see, here's all the different things you could try. Please suggest more in the comments if you think I've missed any:

First off, general points:
  • What's the sample size? Unless it's fixed in advance, data peeking - checking whether you've got a significant result after each scan, and stopping the study when you get one - gives you multiple bites at the cherry.
  • Do you use parametric, or nonparametric analysis?
Now, what do you do with the data?
  • Preprocessing
    • How much smoothing?
    • Do you reject subjects for "too much head movement"?  If so, what's too much?
  • Straightforward whole-brain general linear model (GLM) analysis followed by a group comparison.
    • What's the contrast of interest? You could make a case for Fear vs Neutral, Happy vs Neutral, Happy vs Fear, "Emotional" vs Neutral.
    • Fixed effects or random effects group comparison?
    • Do you reject outliers? If so, what's an 'outlier'?
    • Do you consider all of the Fear, Happy and Neutral blocks to be equivalent, or do you model the first of each kind of block seperately? etc.
  • "Region of Interest" (ROI) GLM analysis. Same options as above, plus:
    • Which ROI(s)?
      • How do you define a given ROI?
  • Functional connectivity analysis
    • Whole-brain analysis, or seed region analysis?
      • If seed region, which region(s)?
    •  Functional connectivity in response to which stimuli?
  • Dynamic Causal Modelling?
    • Lots and lots of options here.
  • MVPA?
    • Lots and lots of options here.
But remember we also got structural MRIs, and while they may have been intended to help analyze the functional data, you could also examine structural differences between groups. What method?
    • Manual measurement of volume of certain regions.
      • Which region(s)?
    • VBM.
    • Cortical morphometry.
      • What measure? Thickness? Curvature...?
That's just the imaging data. You've almost certainly got some other data on these people as well, if only age and gender but maybe depression questionnaire scores, genetics, cognitive test performance...
  • You could try and correlate every variable with every imaging measure discussed above. Plus:
    • Do you only look for correlations in areas where there's a significant group difference (which would increase your chances of finding a correlation in those areas, as there'd be fewer multiple comparisons)?
  • You could define subgroups based on these variables.
*

So even a very straightforward experiment could give rise to hundreds or thousands of possible analyses. 1 in 20 of these would give a statistically significant result at p=0.05 by chance alone, and even if you throw out half those for being "in the wrong direction" (and that's subjective in most cases) you've got plenty of false positives.

This problem is growing. As computation power continues to expand, running multiple analyses is cheaper and faster than ever, and new methods continue to be invented (DCM and MVPA were very rarely used even 5 years ago.)

I want to emphasize that I am not saying that all fMRI studies of this kind are in fact junk. My worry is that it's hard to be confident that any given published study is sound, given that papers are written only after all the data has been collected and analyzed.

I'll also point out that some imaging research, especially what might be called "pure" neuroscience investigating brain function per se rather than "clinical" studies looking at differences between groups, has many fewer variables to play with, but still quite a lot.

As to how to solve this problem, the one solution I believe would work in practice is to require pre-approval of study protocols.
ResearchBlogging.orgSimmons JP, Nelson LD, and Simonsohn U (2011). False-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological science, 22 (11), 1359-66 PMID: 22006061

Wednesday, 23 May 2012

Rich People May Not Be So Unethical

There was quite the stir a few weeks back about a psychology paper claiming that rich people aren't very nice: Higher social class predicts increased unethical behavior.

The article, in PNAS, reported that upper class individuals were more likely to lie, cheat, and break traffic laws.

However, these results have been branded "unbelievable" in a Letter to PNAS just published. Psychologist Gregory Francis notes that the paper contains the results of 7 seperate experiments, and they all found statistically significiant socioeconomic effects on unethical behaviour.

Those 7 replications of the effect "might appear to provide strong evidence for the claim" - one study good, 7 studies better, right? - but Francis says that actually, it's too good to be believed.

Each of the studies was fairly small, and the effects they found were modest, and only just significant. So the observed power of the studies - the probability that a study of that size would detect the effect that they did, in fact, find - was only about 50-88% in each case.

Think of it this way: if you took a pack of cards and discarded half of the black ones, then shuffled the remainder, a random card from the deck would most likely be red. But even so, it would be unlikely that you'd pick seven reds in a row.

The chances of all 7 studies finding a positive result - even assuming that the effect claimed in the paper was real - is just 2%, by Francis's calculations.

Ow.

He concludes "The low probability of the experimental findings suggests that the data are contaminated with publication bias. Piff et al. may have (perhaps unwittingly) run, but not reported, additional experiments that failed to reject the null hypothesis (the file drawer problem), or they may have run the experiments in a way that improperly increased the rejection rate of the null hypothesis (4)".

What might have happened? Maybe there were more than 7 studies and only the positive ones were published. Maybe the authors peeked at the early data before settling on the sample size, or took other outcome measures that showed no effect and went unreported. See also the 9 Circles of Scientific Hell.

Or maybe not. Piff et al respond in their own Letter, firmly denying that they ran any other unpublished experiments, and saying that they "scrutinized our data collection procedures, coding protocols, experimental methods, and debriefing responses. In no case have we found anything untoward." They go on to criticize the method Francis used to get his magic 2% figure, which they point out relies on some debatable assumptions.

Even if you buy the 2% figure, it doesn't mean that the true effect is zero; it might be real, but exaggerated. Ultimately it all becomes rather murky and subjective, which is why I think we need preregistration of research, which would prevent any possibility of such data fiddling, and also remove the possibility of false accusations of it... but that's another story.

ResearchBlogging.orgFrancis, G. (2012). Evidence that publication bias contaminated studies relating social class and unethical behavior Proceedings of the National Academy of Sciences DOI: 10.1073/pnas.1203591109

Tuesday, 24 April 2012

Bias in Studies of Antidepressants In Autism

There's little evidence that antidepressants are useful in reducing repetitive behaviors in autism - but there is evidence of bias in the published literature. That's according to Carrasco, Volkmar and Bloch in an important report just out in Pediatrics: Pharmacologic Treatment of Repetitive Behaviors in Autism Spectrum Disorders: Evidence of Publication Bias

They looked at all of the published trials examining whether antidepressant drugs (mostly SSRIs, like Prozac) were better than placebo in reducing repetetive behaviours in children or adults with an autism spectrum disorder (ASD).

A meta-analysis showed that there was a statistically significant benefit of the drugs overall, but it was marginal, with a small effect size d=0.22, and it was driven mainly by two old, very small studies that found big benefits. One of them only had 12 subjects. By far the largest study, King et al (2009) with 149 people, showed zero effect.

This plot shows all of the studies, with the red line being no benefit of drug vs placebo. The further to the right of the line, the bigger the benefit, but the grey horizontal lines show the uncertainty. As you can see, two small, messy studies found big effects, the others didn't.


Worse yet, although there were 5 published studies, the authors also found that there had been 5 studies that had been completed, but never published. Carrasco, Volkmar and Bloch wrote to the people in charge of those studies and asked for data; only one out of 5 replied. The data showed no benefit.

We don't know what the other 4 unpublished studies found, but the way science works means they probably came out negative. If we assume that they did, then even the small benefit seen in the published studies disappears.

This paper, incidentally, is great example of why trial registration is a great thing. Without mandatory pre-registration on clinicaltrials.gov, no-one would know about the 6 unpublished trials at all. It would have been even better if the researchers had been forced to make public the results, as well as the existence, of the unpublished trials; but it's a lot better than nothing.

Finally, the authors of this paper stress that this doesn't mean antidepressants don't help at all in autism - just that they probably don't help with repetitive behaviors.


ResearchBlogging.orgCarrasco, M., Volkmar, F., and Bloch, M. (2012). Pharmacologic Treatment of Repetitive Behaviors in Autism Spectrum Disorders: Evidence of Publication Bias Pediatrics, 129 (5) DOI: 10.1542/peds.2011-3285

Saturday, 14 April 2012

Fixing Science - Systems and Politics

There is increasing concern that the structure of modern science is flawed and that most published research findings may be false.
Commonly cited problems with how science works today include:
  • Publication bias and the file drawer problem.
  • "Result fishing", data dredging etc. - analyzing data in different ways to "get a finding"
  • The privileging of "positive" results over "negative" ones.
I have previously argued that, to solve these, problems we need a way to ensure that scientists publicly announce which studies they are going to run, what methods they will use, and how they will analyze the data, before running their studies.

We already have such a registration system in place for clinical trials. It's a good system. It's not perfect but it's helped. I propose we extend it to all science. But how would that work in practice?

I'm not sure. So what follows is a series of ideas. These are intended to spark debate.

Here are some options for systems:
  1. There could be a central registry, free and open to the public, where protocols are pre-registered. Call this the 'clinicaltrials.gov option' because we already have one for clinical trials. This registry could also serve as a repository of results and raw data, but it wouldn't have to.
  2. Academic journals could require studies to be pre-registered in order to be considered for publication: you submit the Introduction and Methods, these are peer reviewed, and if accepted, the journal is bound to publish the results when they arrive; the authors for their part are bound to follow their protocol (secondary analyses could take place, but they would be explicitly flagged as such.) and submit the results.
  3. Scientific funding bodies could make all successful scientific grant applications public via an open database. These applications already contain pre-specified methods, hypotheses, and statistical analyses, in most cases; part of this plan could be to make these more detailed.
  4. Authors could have the individual responsibility to publicly announce their methods, hypotheses and plans before starting studies on their own websites.
How can we actually make this happen? That's a question of politics:
  1. Governments could introduce legislation to force this. This is the most extreme option. It is probably unviable, because it would place researchers in different jurisdictions under different rules. Science is a global enterprise, and we don't have a global legislature. (The USA did this for clinical trials, but for various reasons these are a special case and more 'international' than others.)
  2. A consortium of major scientific journal editors could announce that they'll only publish research that complies with the system. Notably, this was how clinical trial registration started.
  3. A consortium of major funding bodies could refuse to finance research that doesn't adhere to the system.
  4. Individual scientists, journals, and funding bodies could unilaterally adopt the system. This would, at least at first, place these adopters at an objective disadvantage. However, by voluntarily accepting such a disadvantage, it might be hoped that such actors would gain acclaim as more trustworthy than non-adopters.
My own preference would be for System 1 via a combination of Politics 2 and 3. Yet any combination of these options would be better than the current system.

Some possible objections:
  1. Pre-registration of all science would be impractical. What about pilot studies and 'tinkering'? - I'm only proposing that any research which might be published, should be publicly registered. This leaves anyone free to tinker away all they like - in private. We just need to be clear, from the outset, whether we're tinkering or doing 'proper' publishable research, a line which is currently very murky.
  2. Many interesting results are unexpected. Post-hoc analysis or interpretation of data is important. - There's nothing wrong with post-hoc analysis or interpretation, so long as everyone knows it was post-hoc. The problem is when it is passed off as being a priori. Registration doesn't seem to have discouraged legitimate post-hoc analyses in the case of clinical trials: there are lots of excellent post-hoc analyses coming out, clearly labelled as such.
  3. It would be unfair to scientists to make them 'tip off' their rivals about what they're working on in advance. It would penalize originality. - My gut instinct here is that this is not a big problem; everyone would be in the same boat so it would be a fair system. However, if this were felt to be a concern, there's an easy solution - just build in a delay to the publication of registered protocols. Put them in a 'sealed envelope' to be opened after a 12- or 24- month 'grace period', and that would give people a head start while ensuring that their original protocol was eventually revealed.
  4. This wouldn't solve all of the other problems with science. - No, it wouldn't, and it's not intended to. However, I do feel that we'll struggle to make progress in other areas without something like this happening. The current system of post-results publication is not the only problem, but it is a large part of it.
On that note, here's a sketch of how I see this relates (or not) to some other issues in science today:

Replication - there's been much discussion of late around ensuring the replicability of results in certain fields e.g. neuroimaging studies and psychology too. My view is that most published false (i.e. unreplicable) findings are a product of publication bias and positive result fishing. Solving those problems, as outlined here, would increase the replicability of science. It wouldn't be a panacea. There will always be dodgy results due to fraud, incompetence, and bad luck, but the current system too often rewards scientists for fiddling around until they get a positive one.

Careers - There is a widespread complaint that the current system of science is unsatisfactory. Our jobs, promotions, funding and tenure depend on our ability to generate high impact papers - which means, in effect, novel and interesting positive results. Pre-registration of science would change the game. Scientists would be judged on their ability to design and run interesting experiments, rather than on their ability to generate 'good papers'.

Open Access - The issue of free open access to scientific papers is an important one. It's a separate question to the one I've discussed here. But I see a spiritual overlap. In both cases, the  fundamental question is who owns science? At the moment, scientists own their work until and unless they decide to publish parts of it. When they do, they sell it to a publisher, who sells it to the world. In my view, the world should be told about science, from the beginning.

*

Fundamentally, this will only happen if a critical mass of scientists want it to happen. It will not be easy, but whereas four years ago I was, deep down, skeptical that it would ever be possible, today I really think it might.

Already we're seeing signs of hope, from informal pre-registration  to calls for preregistration in particular fields in major journals. 10 years ago, the idea was being written off as impractical, and with the technology available at the time, it probably was. I do not think that is true today.

Change can happen. All it needs is will.