Tuesday, 19 June 2012

Beware Stimulus Effects in Psychology

Recently I've blogged about methodological problems in neuroscience research but just to even things out a bit, here's a paper that highlights a potentially serious issue for psychologists - Treating Stimuli as a Random Factor in Social Psychology: A New and Comprehensive Solution to a Pervasive but Largely Ignored Problem

Suppose you want to find out whether people react differently to stimuli from two different groups. The reactions, stimuli, and groups could be anything: maybe you want to see if people prefer listening to sound clips of cats as opposed to dogs. Or maybe you show people photos of blonde men vs dark-haired men and see whether people judge guys with one colour as less trustworthy.

A lot of psychology studies amount to this.

Going with the blonde vs. dark example, suppose you take 1000 volunteers, show them some pictures of blonde guys and dark guys, and get them to rate them on trustworthiness. You find a significant difference between the two groups of stimuli. You conclude that your volunteers are hair-bigots and submit it as a paper. The reviewers think, 1000 volunteers? That's a big sample size. They publish it.

Now that study I just described might be perfectly valid. But it might be seriously flawed. The problem is that while your sample size may be large in terms of volunteers, it might be very small in another way. Suppose you have just 10 photos per group. Your 'sample size', as regards the sample of stimuli, is only 20. And that sample size is just as important as the other one.

It might be that there's no real hair difference in perceived trustworthiness, but there are individual differences - some men just look dodgy and it's nothing to do with hair - and in your stimuli, you've happened to pick some dodgy looking blonde guys. Or whatever.

Now you can run your statistical analyses taking these possible stimulus variation effects into account. But according to Judd, Westfall and Kenny, authors of this paper, this is rarely done. They show with both real and hypothetical data, that unless you take care of this, you can find "statistically significant" differences from pure random noise. This is not a new argument, but they say it's been ignored for too long.

The worst part is that increasing the number of volunteers actually makes it more likely that you'll fall foul of this, not less. Only increasing the stimulus sample size can prevent it.

The paper goes into lots of detail, and tackles various hot potatoes, including one of Daryl Bem's notorious precognition "retroactive priming" experiments. Bem claimed that college students were able to predict the future - they responded differently to different pictures... before the pictures appeared on the screen. The effect was statistically significant and he published it. But Judd et al say that accounting for stimulus variation removes the effect.

ResearchBlogging.orgJudd CM, Westfall J, and Kenny DA (2012). Treating Stimuli as a Random Factor in Social Psychology: A New and Comprehensive Solution to a Pervasive but Largely Ignored Problem. Journal of Personality and Social Psychology PMID: 22612667

29 comments:

Anonymous said...

Really, looking for item effects is rarely done? I know I don't see it reported a lot, but do researchers not check?

Anonymous said...

A better title would be: Beware of Psychology, Subtitle: It is not science.

An abstract concept thought up by a drug addict during one of his more delirious episode added and compounded over the decades by well meaning but mostly inept researchers.

As in this little piece, anything can be proven since the whole concept is vague.

Why people still bother with this stuff beats me.

Anonymous said...

Item analyses have been standard procedure in psycholinguistics for decades. Perhaps people have somehow been less bothered by random variation in faces than by random variation in words.

DS said...

Keep your stimuli equal in every respect (or at least paired with another such stimulus) except the one tested and the problem is solved. In the case of the hair stimulus the images could have been digitally altered to give such stimuli.

"All else equal": Why is this axiom forgotten so often in psychology? Yeh, sometimes its hard to fulfill this axiom. Then don't do the experiment!

Anonymous said...

"Perhaps people have somehow been less bothered by random variation in faces than by random variation in words."

When results can reproduced by random walks the results are invalid.

There is a lot of wishful thinking in psychology studies if not plain forgery.

Recently a psychologist
http://retractionwatch.wordpress.com/2011/09/07/dutch-university-investigating-psych-researcher-stapel-for-data-fraud/
was discovered to have fabricated all of the research data but still his papers where accepted as government policy guidelines without question.

He only got busted by accident btw.

This "incident" has triggered lots of withdrawals all over the place, not a few from colleagues incorporating his 'research' in their work.

All in all it is plain to see that soft psychology as a concept is just much to insubstantial to take seriously.

Observational studies gave us the book of horrors in the form of the DSM, when can we finally accept it's just a very error prone method?

The observer already confines the results by defining the study parameters. That what's NOT included could be the more important.

Anonymous said...

I'm neither a psychologist nor any kind of scientist so I'm genuinely asking and not being facetious: why aren't correct procedures mandatory and universally understood? This appears to be basic stuff, no?

CM said...

If you have correctly balanced your stimuli, including item effects just reduces your power. For things like randomly chosen images from the internet, it makes sense to include an item effect. But if your stimuli are instances on a well formed, validated scale (for example, artificial images), putting in a random effect for item only reduces your power. These kind of things are debated in psychology and in other sciences and for good reason. There is certainly a grey area and that is why there is some confusion about what the right thing to do is. To be honest, I see this problem the most in (i) business school studies and (ii) poorly motivated second tier social neuroscience.

The argument about Bem is weird. Bem's thing doesn't replicate, is wrong and is an embarrassment to (social) psychology. But it is not because of item effects. For the item to have an effect, it would need to be presented before the response. This is because of physics. Item effects have very little to do with it.

As to saying the person psychology is not a science and was invented by a drug addict... swing by my lab sometime and see how careful we are with experiments and with statistics and I'll show you the effects we have that replicate across cultures, across items and across experiments. You see colors on your computer monitor largely due to psychologists. Much of modern measurement theory was invented by psychologists. There are countless other examples. Conversely, statistical problems also exist in physics, in biology and across the board in science. If you don't know this, more fool you.

Ase said...

Well, I largely agree with CM. You are well aware (and get extensively trained) in the fact that stimuli matter, and can trip you up, and you spend a lot of time thinking about and considering the tradeoffs (and validating data-sets, and selecting validated datat sets, and considering what you can safely ignore).

I'll check out the paper, just to check the analysis ideas they have. But, it really is standard to consider and control for (and discuss, and problematize) stimulus selection.

Asha ten Broeke (Dutch science writer) said...

The Dutch ex-psychologist Diederik Stapel didn't get busted by accident. His actions were secretly and carefully checked over the period of several months by soms suspicious phd-students/postdocs, who then bravely came forward with accusations of fraud.

Professor Keith R Laws said...

Some of this is to do with 'cultures' that develop in specific sub areas of Psychology. For example in psycholinguistics (and some areas of cog neuro), it is fairly common practice to check for item effects - analyse across subjects and analyse across items and report e.g. F stats values for both - and/or minF
In other areas of psycholgy, other (possibly less stringent) practices have evolved- testifies to the sprawling nature of psychology and the way that poor culture evolves because of the 'in-bred' nature of sub disciplines

Stuart Brown said...

Sociolinguistics has attempted to handle a similar problem with attitudinal speech perception tests with the so-called "matched guise" test. In the traditional version, participants are apparently played a number of people with different dialects all reading the same passage and are asked to judge them on attitudinal scales. In fact, other than the confounds, all of the recordings are made by a single, skilled phonetician, who can reproduce all the dialects in question. This is supposed to factor out issues such as voice quality, depth of voice, and so forth.

However it's not without its problems: obviously it depends on the phonetician being able to perfectly, naturally and accurately reproduce another accent, about which I am somewhat skeptical, and there are a few other similar problems, such as the presumption that, say, voice pitch has the same significance across dialects.

However what is now happening is that digital manipulation techniques are being used to micro-modify files, tweaking solely the variable under examination. So, for instance, a study looking at the pronunciation of terminal -ing ("-in" vs "-ing") took naturalistic recordings and simply removed the actual spoken termination from all instances, and spliced in "-ing" or "-in." Participants were played one or the other for each speaker. This approach I think is quite promising.

Anonymous said...

@Asha ten Broeke
I stand corrected

@The Rest.

With every form of semi-scientific endeavor where one pursues a moving goal with approximating methods one is bound to loose track.

Add in a good dose of confirmation bias, sloppy work, poorly designed studies and you get results without value.

Psychology is a model. A model of various mental processes common amongst a species. It takes parameters based on observations and tries to fit them into a single unifying description.

When the processes are simple, the model works. Pavlov's dog salivates on cue.

But the error made assuming that if the model works for simple things it by extrapolating also works for complex things.

Which it doesn't. Which it never can. That's the big problem with all models, once you start to add layers of complexity the predictive value of them decreases exponentially.

At least that's something that came out of the AGW mess. 100's of billions of $ wasted but at least we know now you can't model self modifying parameters.

Anonymous said...

This is well known in psycholinguistics (following Coleman and later Clark). However, until fairly recently psycholinguists still often got it wrong. Looking at item effects doesn't solve the problem. You need to have a model that combines item/stimulus and participant/variablity or an experimental design that samples variability among items and subjects in the correct way (essentially a nested design rather than a crossed one).

Anonymous said...

Haven't read the article, but i really don't see how this could be a critique of the Bem studies. Imagine having one specific deck of cards that people could predict ahead of time better than chance the next card coming up...although odd, this would still certainly be amazing.

Jake Westfall said...

Maybe I can help clarify a few points brought up in the reader responses. (Although I can see now that some of my comments turned out to be a bit technical... so hopefully I at least do not make things worse!)

The retroactive priming experiment conducted by Bem, and reanalyzed here, employed two different sets of items: primes (which were printed words) and targets (which were images). As some here have noted, observing "item effects" of primes/words on response latencies would still be remarkable because, in the retroactive priming condition, responses were collected prior to presentation of the primes. However, there may well be (and in fact, there was evidence of) item effects for the targets/images. As discussed in the paper, the result of this target variability would be an anticonservative bias in the standard F test, possibly a severe one.

Critically, and as tsbaguley hints, correctly addressing this requires that this item variability be *included in the statistical model*, not simply loosely acknowledged by examining separate subject- and item-level F tests, as is apparently common in some fields. Alas, two biased analyses do not add up to a single unbiased analysis.

Anyway, as it turns out, in the Bem data we find significant item effects for targets/images in both the forward-priming and retro-priming conditions, BUT we only find item effects for primes/words in the forward-priming condition (where primes preceded responses), not in the retro-priming condition (where primes followed responses). Similarly, we find significant variation in the size of the congruency vs. incongruency priming effects across subjects in the forward-priming condition (i.e., some subjects are more influenced by the congruency between primes and targets than are other subjects), but we do not find significant variation in the priming effects in the retro-priming condition (i.e., if there is no such thing as retroactive priming, it cannot be that some subjects are "better at it" than others). So these results are pretty much what you would expect.

One final thought about statistical versus experimental control of item effects. Certainly an experimenter should try as much as is feasible to carefully balance stimuli with respect to features that might influence the outcome variable. And of course, most labs routinely do this. However, the extent to which one is ultimately successful in achieving this balance is an empirical, statistical question. One that is easily answerable by referring to the estimated random effect variances from a mixed model of the data. It is not sufficient for one to simply say that they attempted to use a balanced stimulus set and that therefore the relevant variance components for items are probably small in the data. At best, they must present statistical evidence that this is the case. And there are many, including us, who feel that the only truly compelling excuse for not modelling item effects is model nonconvergence or model nonidentifiability when item effects are included in the model (the latter of which would be the case in particular experimental designs).

Matt Craddock said...

I started off doing both by-participants and by-items analyses but was always a bit suspicious of how useful that was, given that it often means relying on means made up of tiny numbers of trials or switching variables that were within-participants to being between-items with concomitant lower statistical power. There is actually a paper somewhere against using separate F tests or minF, but I can't remember enough details to track it down.

Sure, in many cases you can control the stimuli so that only the variable of interest is different, but this is typically only the case for low-level factors such as contrast or luminance. I cannot eradicate physical differences between a set of animals and a set of tools, even if I can (and do) try to match them on low-level factors as far as possible.

There are some great guides to doing mixed-effects analyses kicking around - I can recommend Baayen's book, Analyzing Linguistic Data for a good intro using R. Mixed-effects models are the way to go, definitely.

Anonymous said...

More detail on my earlier comment at:

http://psychologicalstatistics.blogspot.co.uk/2012/06/stimuli-as-fixed-effect-fallacy.html

I have just now seen Westfall's comments (which seem to confirm my speculation about their approach - as I've not seen the full paper yet).

@Matt Craddock: Your suspicions about the by-item and by-subject analysis are sort of correct. However, the multilevel (linear mixed) model approach doesn't solve this issue - it merely brings it into focus. Whenever your number of items or subjects is low it is problematic estimating the variance components and therefore also the standard errors. Variances have a skewed sampling distribution so you need large numbers of items for the asymptotics to kick in (and sometimes very large). You generally want 30 to 50 items and 30-50 subjects minimum ... It is possible to get away with fewer (but the lower you go the more cautious you should be). It is also worth noting that the limiting factor is usually going to be the smaller of the sample sizes (though it will also depend on how variable items and subjects are). There is pretty much no compensatory effect ... so 5 items and 1000 participants is worse than 20 and 20 ...

Controlling the stimuli is a tricky issue as you may be reducing your analysis to a case study of those obscure items. The more carefully you control stimuli the smaller pool you draw from ... thus it becomes less clear what population they are sampled from. However, it is still worth trying to model the extra variability from items IMV (and I think that is also what Westfall was suggesting). Including or at least exploring the extra variability seems like an appropriately cautious approach.

Anonymous said...

Westfall says: "Similarly, we find significant variation in the size of the congruency vs. incongruency priming effects across subjects in the forward-priming condition (i.e., some subjects are more influenced by the congruency between primes and targets than are other subjects), but we do not find significant variation in the priming effects in the retro-priming condition (i.e., if there is no such thing as retroactive priming, it cannot be that some subjects are "better at it" than others). So these results are pretty much what you would expect."

So there was less variance in the effect for the backward condition, and that's considered a bad thing? I could see how if it was the other way around, one could say there was more variance because it's not a stable effect...

Anonymous said...

@Anonymous: the point about the item and subject effects is that the variability is systematic. The retro-priming effect could be highly variable (noisy) but it it doesn't show consistent variability - so there is no evidence that that variance arises from specific items or people being consistently good performers. It is odd that standard priming effects show individual differences between items or people and that the retro ones don't (assuming their the same kind of effect).

On the other hand this is probably the kind of pattern you'd expect if the retroactive-priming were an artifact of some noisy process being filtered to show an average effect. In fact, if the retroactive priming were the same effect going backwards you'd perhaps expect similar patterns in the individual differences (with the same participants and items tending to be better).

omg said...

I can't concentrate at all with those sets of bedroom eyes. They want me.

dalmeida said...

This whole "X-as-a-fixed-effect-fallacy" needs to be put to rest. Yes, in experimental designs where you RANDOMLY SAMPLED participants AND materials, you probably should treat them as random effects in your statistical model.

Now show me where this kind of design is used in experimental psychology. Samples of participants are generally convenience samples, and materials are generally carefully picked and/or generated, often with explicit matching put in place.

In designs like these, why should we treat items as random samples when they are anything but random?

Also, why care so much about population inferences, if what in fact what psychologists generally want are causal inferences?

Matt Craddock said...

@dalmeida umm, most object recognition experiments? For example, any naming experiment featuring familiar objects, or any experiment with superordinate classification (something along the lines of living/non-living). Pretty much anything featuring non-exhaustively sampled categories, in fact. Why would we be interested in population inferences? If I want to claim there's something different about living versus non-living objects, I need to know whether effects from my subset of living objects are likely to generalize to other living objects. Otherwise, my causal inferences only apply to my specific set of objects. Same reason we treat subjects as random effects.

Neuroskeptic said...

"and materials are generally carefully picked and/or generated, often with explicit matching put in place."

If they really are picked or generated such that they're exactly matched except on the variable of interest then that's fine; in some fields of psychology they often are, but in others this is not common, sometimes impossible - if your stimuli are all photos of real world objects, say.

dalmeida said...

@MattCraddock: The problem with the "stimuli-as-fixed-effect-fallacy" is this idea that if items are not exhaustively sampled from a category they should be treated as random. This idea is, at the very least, very controversial. Look at the response to Clark's 1973 article that appeared in 1976 (Journal of Verbal Learning and Verbal Behavior) by Wike and Church and the reponses that followed. The conceptual issue is how can we justify treating as randomly sampled items that were never really randomly sampled. As for the issue of "generalizability across items" you mentioned, you can really only achieve that through replication, especially if you never really sampled your items nor your subject randomly (which no one ever does).

dalmeida said...

@Neuroskeptic: Unfortunately, the idea that unless items are *perfectly* matched they should be treated as random has really caught on, despite there being very little justification for it. For starters, the main assumption of treating items as random is that they were randomly sampled. That is virtually never the case. Items are chosen often with some degree of control, even if perfect matching is seldom possible. This is a far cry from a random sampling procedure, so why use a statistical model that assumes something that was never really done?

Matt Craddock said...

@dalmeida Actually, yes, good point about the somewhat non-random sampling; but putting it the way you do, wouldn't the same apply to subjects? Should we treat them as fixed too? It's more the case that we should consider our design before blanketly applying fixed- or random-effects. I think the Raaijmakers papers (1999 and 2003) are quite helpful; I now remember it was reading the 1999 one that prompted me to stop doing F1 (by-subjects) and F2 (by-items) tests and just do F1, since I was using counterbalanced lists in a Latin-square-stylee (i.e. list 1 of items appeared in condition 1 for one set of subjects, condition 2 for the next set of subjects and so on), with which the problem (such as it is) should be minimal and by-subjects alone should be fine (if I'm reading the 1999 paper correctly).

dalmeida said...

@MattCraddock: Yes, the same reasoning should also apply to treating subjects as random-effects. I know of no experiment in experimental psychology in which the authors have actually used a simple random sample of subjects. If such experiments exist, they are exceedingly rare. The norm is to use convenience samples.

So we really need to ask what justifies using a statistical model that assumes a simple random sampling procedure that virtually never takes place in our experiments. And here, there are two schools of thought, as far as I can tell:

(1) Even though subjects are not technically randomly sampled, the sampling procedure is haphazard enough for us to consider them as a good approximation of a random sample (Clark, 1976 makes this argument in response to his critics). However, materials are often not sampled in the same way, because researchers often strive for some degree of either "representativeness" or control over their materials. This is the argument of Wickens & Keppel (1983, Journal of Verbal Learning and Verbal Behavior) for disagreeing with Clark (1973) that materials should always be treated as a random-effect if they do not exhaust a population of materials. There is a nice discussion of this topic in the Keppel & Wickens textbook (Design and Analysis: a Researcher's Handbook, 4th ed)

(2) There is no real justification for us to treat subjects as random-effects in a way that allow us to make population inferences (ie, we have no real external validity). There is, however, a good justification for us to have high confidence in the internal validity of this kind of analysis. And this is because these kinds of parametric tests provide a really good approximation to the results of the correct permutation test. Therefore, parametric tests using subjects as random-effects are a valid statistical model insofar as they are a good proxy to a valid permutation test. What is important here is that while there is a clear rationale for treating subjects as a random effect, this very same rationale disallows the population inference that is intended when we treat subjects as random effects in a parametric test; the permutation test framework only allows for a causal inference to the sample of subjects (and materials) tested. External validity here can only really be achieved through replication. This is the tack taken by most of Clark's 1973 critics (Wike and Church, Cohen, Keith Smith and to some extent Keppel) as well as Onghena and Edgington (2007), Siemer (1997, 2003) and Maxwell and Delaney's handbook.

I tend to regard the arguments such as (2) much more persuasive than the ones in (1), mainly because I cannot see anything beyond mere assertion backing the claim that "haphazard enough" should be considered for all intents and purposes as equivalent to "truly random".

At any rate, I see no real rationale for materials to *always* be treated as a random factor when they do not exhaust the population of materials. These arguments should only extend to cases where materials are in fact randomly sampled, which is rarely the case in psychology experiments.

The arguments presented by Coleman (1964) and expanded by Clark (1973) have in fact been extremely controversial and criticized by other statisticians, but they somehow gained so much traction with researchers that they are seen as canonical. That is extremely unfortunate, because it makes researcher care more about arcane statistical issues that do little to increase the validity of their experiments while at the same time it gives them the false impression that they don't have to worry so much with replication because hey, the math says the results generalize to the population, nevermind that this population has never been really randomly sampled.

Neuroskeptic said...

Thanks for the excellent comments everyone.

dalmeida: What about the argument that, although it would be wrong to only analyze data treating stimuli as random effects, you should at least check to see whether your key results still hold under such an analysis?

dalmeida said...

@Neuroskeptic: A proposal similar to what you are suggesting has been advocated by Coleman (1979, Journal of Verbal Learning and Verbal Behavior). He suggested that every researcher should provide three F tests on their papers (F1, F2 and MinF'), so that everyone can have the information they deem relevant. Those who do not buy Coleman's and Clark's arguments can just look at F1, while those who do can look at MinF' (or can re-calculate it using F1 and F2).

On the face of it, I think this is a reasonable suggestion, but there are some problems with it.

First, just presenting the three statistics as equally plausible tests does ignore the important conceptual differences behind the interpretation of the p-values (causal inference restricted to the sample or descriptive inference about a population parameter, etc).

Second, if the concern is to see whether the crucial results "hold" when you treat the items as random effects, then you run into a problem: Wickens and Keppel (1983) already showed that under an adequate level of experimental control over items, MinF' tests are less powerful (sometimes considerably so) than the correct F1 tests. In light of their finding, it is interesting to notice that all the demonstrations of the "stimulus-as-fixed-effect-fallacy", from Clark 1973 on, rely on washing away results that were originally reported significant via regular F1-type tests. If these tests are significantly less powerful under conditions in which the items were not randomly sampled, but rather carefully selected, then simply washing away the original results shows very little: it could be that the original results were false positives, or it could be that the new results are false negatives, due to the loss of statistical power. There is no real solution here, because the logic of the argument is really mutually exclusive: either the materials can be considered a random sample and treating them as fixed inflates the type I error rate (as argued by Clark 1973), or they cannot be considered a random sample and treating them as random will inflate your type II error rate (as shown by Wickens and Keppel, 1983).

This is an issue that is not solvable by math alone, but rather by careful consideration of which model is a better match to the experimental conditions.