Whose job is it to detect scientific fraud?
You've probably heard of Diederik Stapel, a Dutch psychologist who's just admitted to scientific fraud on a grand scale, with dozens and maybe over 100 papers published based on made-up data. This comes just months after Harvard's Marc Hauser resigned over unspecified data-meddling activities.
What disturbs me is not just that this fraud happened, but the way it was detected. Both Stapel and Hauser were busted by their own junior lab members. Browsing Retraction Watch and reading over other fraud cases reveals that
fraud is almost always detected either by 1) By readers of
published papers who notice oddities in the data, or 2) by
internal whistleblowers, almost always junior lab members
But these are both ad hoc methods. They rely heavily on
individual vigilance and courage in speaking out (especially in the
latter case). It seems to me that there's no working mechanism for catching fraud. If
there were, such acts of individual heroism wouldn't be needed.
So whose job is it to catch fraud? At the moment, it's all the work of private investigators. Where are the police?
First off, is it the job of journals? That seems plausible. Journals publish scientific papers and by doing so they are saying, implicitly, that the papers are good quality. The way this works is meant to be through peer review.
But peer review is failing to catch many cases of fraud. I guess we don't know how many fraudlent papers are caught at the peer review stage and never published. But one would hope that such cases would come to light anyway because reviewers who suspect fraud ought to alert the relevant authorities. I can't think of any recent cases in which fraud investigations were started by peer reviewers, or at least not that we know about.
Maybe it's up to the institution that employs the fraudster? It's the institution that carries out investigations, "convicts" the fraudster and enacts the punishment. Clearly it's in their interests to do this because they don't want to be seen as soft on misconduct. But rarely do they go out and proactively try and catch or prevent fraud. It's not in their interests to do that.
Undetected fraud does no harm to anyone's reputation. On the contrary fraudsters are often the "stars" of their faculty until they get caught. Hauser and Stapel were. Plus, a department that got a reputation for hard-hitting anti-fraud measures might struggle to recruit people, even perfectly innocent ones who just found it annoying.
So what we see is departments who perform (fairly) good investigations into fraud, but only when someone else tells them to.
Maybe it's the funding bodies? They're paying for the
research, so they clearly have an interest in making sure their money is
well spent. At present, though, they lack the mechanisms to investigate it.
So those are the three possibilites as I see them - journals, institutions and grant awarders. While all of these organizations have policies for investigating and punishing fraud when it comes to light, they rarely (if ever) actually catch it, leaving this hazardous and stressful job to individuals.
Is there a better way?
18 comments:
maybe to make it obligatory to put all raw data available online, so it would be hazardous to make them up? and not only hazardous, hard to produce too. is easy to fake and hard to detect if only final output is presented.
it looks, now, there are small risks and high advantages to cheat. for all players. authors and "regulators" too. maybe any new tool should increase risks, so people would think twice. public control of data (or maybe just such a possibility) may increase risks (real or just psychological) and reduce cheating profits.
Interesting post.
Personally, I think the Journals - not the peer reviewers, but the journals themselves - should be the ones to address this issue.
They are getting large sums of money from university libraries (and sometimes authors), free content from authors and free labor from peer reviewers. With things moving to digital, they are providing less and less value and there is less reason for them to profit given how little they do (e.g., compared to the PLoS model).
They also have a reputation that they should presumably be trying to defend and appropriate expertise in the relevant areas.
So why not have a set of people on staff, much like copy editors, that work with the author to vet the data at a specific stage of the publishing process?
That seems like the least intrusive, most logical solution to me. Plus it has the added bonus of providing additional job opportunities for academics!
Outright fraud is probably less of a problem than flawed methodology or plain old bias, all of which are much harder to detect without (all) the raw data. So I second Anon@12:32. (Ben Goldacre has been pushing for pre-registration of clinical trials for years.)
Ultimately it's the job of all of us in a field - manuscript reviewers, journal editors, bloggers, journal readers, etc. - to ask the awkward questions. We're all part of the problem and also part of the solution. What matters is the evidence we have available on which to take the claimed results to task. So feed us better information and we can do a better collective job of policing.
(Incidentally, were this fraud to have happened in the US with federal funding involved, I would expect federal charges to be filed. Lying to the feds - whether at a Congressional hearing or in a grant application - has been shown to be a fast track to prison. Al Capone went to prison for screwing the tax man, not for murder.)
I agree with mettle, this would be hardly different from editors of newspapers being expected to perform fact-checking for information they post. From both an economic and sensible stance, this seems to make the most sense.
I agree with practiCalfMRI. It is the duty of all of us.
~~~~~~ Metaphor On ~~~~~~
But instead of accepting this duty by kicking the scientific bandwagon's wheels and checking the horse's hooves too often folks are eager to jump on a scientific bandwagon for fear of being left behind. And far too often it turns out that the wagon was predicatbly destined to circle the town square dropping you off where you started minus the fare and time.
~~~~~~ Metaphor Off ~~~~~~
Over and over again I see the manipulation (either intentionally or due to shear ignorance) of ill-posed problems (in the Hadamard sense) leading to sexy publications which later turn out to be not so hot after all. It is my opinion that many reviewers are willing to accept the pseudo-inverses of ill-posed problems in order to breath vitality into their field of study. That is, they accept and often encourage wishful thinking in lieu of rigor. Once a certain level of tolerance for such solutions is established in the literature then ...
~~~~~~ Metaphor On ~~~~~~
... it is often difficult to get the horse back in the barn.
Thanks for the comments. I agree with mettle that the journals ought to be taking a lead on this.
Out of the three organizations I mentioned, they are the only ones who already have an infrastructure (peer review) that could be adapted to combat fraud. In theory peer review should catch fraud, the problem is that it doesn't.
The idea of having people on staff to do this would only work at big journals I think. Most just couldn't afford it. But the publishing consortia could.
One could imagine a kind of "audit team" based at the publishers who randomly select submitted articles and check the raw data.
But even this would be expensive and might be perceived as unfair. Making everyone post raw data online would be an alternative (and would be good in other ways too).
I do agree that we all have a part to play in this. But I think we're doing a pretty good job on a grassroots level and I don't see how we could get much better, without institutional change.
Data is the property of the researchers who bust their asses day in and day out to collect it to advance their careers and protect their jobs. Making it publicly available opens the door for people to use the data and present/publish it as they see fit.
Bust their asses? You mean squabbling over squat. Who can BS the best, adding no intrinsic value to societies today. Does anyone care about obscure datasets even if it was publicly available? Yay a bunch of numbers! I'm going to replicate and analyze it better! I say go for it who cares. Everyone bust their asses to put bacon on the table. Academics less I say, politicians even more less.
@omg If you don't believe research has any intrinsic value, why are you spending time on this site?
I think fraud - outright making-up data and massaging the data - occur quite frequently, but little is ever discovered.
To me, this means that heavy skepticism is in order for all of our science.
Making the data publicly available or even providing it to journals wouldn't change a whole lot in my opinion. I feel the culprits do not fabricate entire data sets, more commonly it is that they will change a few values (or do some curious outlier removal) that changes their tests and results to p<.05
@anon, the last one: haha. if true, problem is impossible to solve. and we should just sit and believe in scientific honesty. what should be more massive tool to uncover cheating then check of all raw data provided?
a sudden assault of auditor commando confiscatig all computers&totebooks? .-)
One way to reduce fraud would be for us to conduct and value exact replication studies more than we do. The incentive to commit fraud would go down, and the incentive to avoid errors and be rigorous would go up, if you knew that people were going to be running and publishing replication attempts, and failures to replicate would be hung around your neck (or your published article's neck).
Journals and databases could help in this regard: journals by automatically publishing exact replication attempts in online supplements, and databases by including "replication attempted by" links analogoous to "cited by" links. I wrote about this in the context of the Bem ESP paper: http://hardsci.wordpress.com/2011/05/10/how-should-journals-handle-replication-studies/
data fraud is extremely frustrating because it leads you down blind ends! very very annoying!
getting caught should be a big enough deterrent; i nearly had a heart attack when i caught some genuine errors in the proof stage, i can imagine falling over dead if i ever did commit fraud & got caught.
here's a tip - if the data are weird, the results don't really fit any current theories, it's most likely the data are valid - speaking from personal experience. refraining from 'massaging' strange results can lead to some pretty cool insights, even if you can't really explain them fully.
i guess the point of all that rambling is that humans, being humans will cheat. The shame of being caught is a deterrent for most of us; so make it unprofitable for the rest. Reward negative findings, don't penalize them & force people to 'find' p<0.05!
There needs to be a sea-change in academia - the publish or perish attitude may be driving a lot of fraud.
In the end though, quis custodiet ipsos custodes? Each of us should be our own guardian- idealistic, impractical, yes... but there's still hope.
I think that much of the fraud that occurs is caught prior to publication. In an ideal world, those who have earned terminal or professional degrees have been through a rigorous curriculum that ensures that they not only know how to conduct proper, ethical research, but also catch those who are likely to go so far as to make up data. The small number of high profile cases that we see of researchers losing their credibility to unethical behavior is but a small drop in the bucket compared to the vast number of researchers who are carrying out proper experiments and faithfully reporting the data to peer-reviewed journals.
Having been in academia for some time myself, I would like to think that if one of my colleagues were publishing data that was suspect, I would have the notice irregularities when their data seemed to be magically appearing. But, by that time, they should have been weeded out of science anyways, or taught the error of their ways so that this never became an issue in the first place.
Journals everybody involved has the responsibility to be vigilant against fraud. There is no scientific investigator, no experiment, no publication that may simply be trusted. As scientists and lay people alike, we must all question what we are told, what we hear and what we read. Does the new data make sense with regard to the current scientific thought? Are there other experiments independently that show similar data? If we all continue to critically analyze the information that we are given, there need not be a single authority calling out scientists for lying.
Finally, why would we want a single authority? No single group knows everything or can critically analyze each topic to the same degree. Even journals which rely on peer-review cannot be perfect as they are reviewed by imperfect humans. The journals can easily catch content which does not fit our current understanding of the universe, but they cannot be held to verify every statement in every paper that is published. They may only do what is reasonable to ensure that the journal publishes scientifically valid papers. We need to know how to digest the information ourselves, then we will be much less likely to be caught by lies. Sure, there will always be people that can game the system, but as our (the whole of the planet's) grasp of scientific knowledge increases, those who will be able to cheat and lie will become fewer and further between. The students who caught this individual in his lies, while definitely noble, were not heros by any stretch. They were doing what scientists are supposed to be doing at all times: they were questioning the data presented, the conclusions that followed and noting inconsistencies up their chain of command. Definitely, they deserve credit for risking their own hides by going against a senior scientist in their institution, but if other people had given the same scrutiny to the data he was producing then he probably would have been caught much earlier.
Anonymous who said "Data is the property of the researchers who bust their asses day in and day out to collect it"
I agree that they ought to get 'first use rights' to it. But I don't think they should own it. If you publish some conclusions based on some data then anyone should be able to see it.
Maybe there ought to be a kind of copyright-style period during which everyone else can "look but not touch" (i.e. not publish their own analyses without permission), but after that, copyright expires, and it's a free for all. Say 5 years.
Unfortunately the current journal system means that as soon as you raise difficult questions about a paper the authors withdraw it and take it somewhere else.
Once it's withdrawn the journal doesn't care and the peer reviewer will likely not know about the new destination until it is published. By which time they will have no power to request further clarification from the authors, and not enough evidence to allege fraud.
Hence the sotto voce hints and insinuations about this or that lab or researcher that go on at conferences.
"Hence the sotto voce hints and insinuations about this or that lab or researcher that go on at conferences."
This could be a potentially important restraint on fraud, although it could catch a few innocent people too.
I think someone could design a good survey of scientists over this - have you suspected any of your colleagues of fraud, have you avoid citing/using their research, have you discouraged funding for them/collaborating with them - to see how important a role it plays in the culture of science.
Post a Comment