Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Saturday, August 20, 2011

Bayesian truth serum, grading and student evaluations

In one of my last posts, I examined some proposals for making university grading more equitable and less prone to grade inflation. Currently, professors are motivated to inflate grades because high grades correlate with high student evaluations, and these are often the only metrics of teaching effectiveness available. Is there a way to assess professors' teaching abilities independent of the subjective views of students? Similarly, is there a way to get students to provide more objective evaluation responses?

It turns out that one technique may be able to do both. Drazen Prelec, a behavioral economist at MIT, has a very interesting proposal for motivating a person to give truthful opinions in face of knowledge that his opinion is a minority view. In this technique, awesomely named "Bayesian truth serum"*, people give two pieces of information: the first is their honest opinion on the issue at hand, and the second is an estimate of how the respondent thinks other people will answer the first question.

How can this method tell if you are giving a truthful response? The algorithm assigns more points to responses to answers that are "surprisingly common", that is, answers that are more common that collectively predicted. For example, let's say you are being asked about which political candidate you support. A candidate who is chosen (in the first question) by 10% of the respondents, but only predicted as being chosen (the second question) by 5% of the respondents is a surprisingly common answer. This technique gets more true opinions because it is believed that people systematically believe that their own views are unique, and hence will underestimate the degree to which other people will predict their own true views.

But, you might reasonably say, people also believe that they represent reasonable and popular views. They are narcissists and believe that people will tend to believe what they themselves believe. It turns out that this is a corollary to the Bayesian truth serum. Let's say that you are evaluating beer (as I like to do), and let's also say that you're a big fan of Coors (I don't know why you would be, but for the sake of argument....) As a lover of Coors, you believe that most people like Coors, but feel you also recognize that you like Coors more than most people. Therefore, you adjust your actual estimate of Coors' popularity according to this belief, therefore underestimating the popularity of Coors in the population.

It also turns out that this same method can be used to identify experts. It turns out that people who have more meta-knowledge are also the people who provide the most reliable, unbiased ratings. Let's again go back to the beer tasting example. Let's say that there are certain characteristics of beer that might taste very good, but show poor beer brewing technique, say a lot of sweetness. Conversely, there can be some properties of a beer that are normal for a particular process, but seem strange to a novice, such as yeast sediment. An expert will know that too much sweetness is bad and the sediment is fine, and will also know that a novice won't know this. Hence, while the novice will believe that most people agree with his opinion, the expert will accurately predict the novice opinion.

So, what does this all have to do with grades and grade inflation? Glad you asked. Here, I propose two independent uses of BTS to help the grading problem:

1. Student work is evaluated by multiple graders, and the grade the student gets is the "surprisingly common" answer. This motivates graders to be more objective about the piece of work. We can also find the graders who are most expert by sorting them according to meta-knowledge. Of course, this is throwing more resources after grading in an already strained system.

2. When students evaluate the professor, they are also given BTS in an attempt to elicit an objective evaluation.

* When I become a rock star, this will be my band name.

Tuesday, April 19, 2011

Gender and scientific success

I didn't want to write this post.  I really don't want to touch this with a ten foot pole. What follows is messy and complicated and guaranteed to make everyone mad at least some of the time. (Ask Larry Summers).

We need a sane approach to how we deal with gender in the sciences.

Women are making measurable representation gains in the sciences. This is an undisputed good. Everyone benefits when the right people are doing the right job. However, despite the fact that the majority of bachelor's degrees are now being awarded to women, women only make up about 20% of professorships in math and the sciences. Why?

The three basic alternative answers: 1.) women tend not to choose careers in math or science (either willingly or due to life/family circumstances); 2.) women are barred from achievement in math and science through acts of willful discrimination; or 3.) women do not have the same aptitude for achievement in math and sciences as men.

This is a difficult issue to study as people's careers cannot be manipulated experimentally, and we are left to mostly correlational evidence. An exception are CV studies where identical CVs are given to judges with either a woman or man's name on the top. Judges are asked to determine the competence of the candidates. These studies typically find that the "male candidates" are judged to be more competent than the "female candidates". As no objective differences exist between them, this is a measure of sex discrimination.

Reviewing the correlational evidence for gender discrimination in the sciences, Ceci and Williams find that when examining researchers with equal access to resources (lab space, teaching loads, etc), that no productivity difference is found between male and female scientists. Female scientists are, on average, less likely to have as many resources as male scientists as they are more likely to take positions with heavier teaching loads. How to reconcile the CV studies showing discrimination and the correlational evidence suggesting none? In an excellent analysis of the Ceci and Williams paper, Alison Gopnik asserts a possible hypothesis: "Women, knowing that they are subject to discrimination, may work twice as hard to produce high-quality grants and papers, so that the high quality offsets the influence of discrimination".

It's possible. But Gopnik also admits that it is also possible that policy changes could be responsible. In other words, that affirmative action-style policies that give women advantages could counteract the subconscious gender discrimination seen in the CV studies.

There's a darker side to these policies, though. Some worry about the discounting of a female professor's abilities, assuming she rose to the position via policy rather than talent. Furthermore, some policies designed to given women more voice actually end up give them more work - if a certain number of women need to be on a committee, then female professors are doing more service work than their male counterparts.

And then there's the matter of why female faculty find themselves in low-resource situations to begin with. Stated eloquently by Gopnik, "the conflict between female fertility and the typical tenure process is one important factor in women's access to resources. You could say that universities don't discriminate against women, they just discriminate against people whose fertility declines rapidly after 35."


And well-meaning policies also interact with the fertility issue in insidious ways. For example, many universities offer to "pause" the tenure clock for a year for a faculty member who gives birth before tenure. Sounds great, right?  It could be, except that there is a tremendous amount of pressure to not take this credit for fear of seeming weak. This is especially true in departments that have faculty members who have already chosen not to take the time.

So... we have unconscious discrimination, conscious policies to counter said unconscious discrimination, conscious and unconscious backlash against the policies, and a structural problem for female fertility. In other words, it's a complicated picture and I don't know what the answer is. I do, however agree with Shankar Vedantam's assessment: "It is true that fewer women than men break into science and engineering careers today because they do not choose such careers. What isn't true is that those choices are truly "free.""

Saturday, January 15, 2011

What are you writing that will be read in 10 years?

This was a question asked to an acquaintance during a job interview for a professorship in the humanities. It's one hell of a question, and one that I find unfortunately unasked in the sciences.

In my other life, I submitted a paper this week. It's not a bad paper - it shows something new,  but like too many papers being published today, it's incremental and generally forgettable. It's not something that will be read much in 10 years.

I love reading old papers. They are from a time when authors were under less pressure to produce by volume. They are consequently more theoretical, thoughtful and broad than most papers published today because the authors had the luxurious time to sit and think about the results, and place them in context.

As I've pointed out earlier, the competitive academic environment tends to foster bias in publications: when trying to distinguish oneself amongst the fray of other researchers, one looks for sexy and surprising results. So do the journals, who want to publish things that will get cited the most. And so do media outlets, vying for your attention.

Jonah Lehrer's new piece on the "decline effect" in the New Yorker almost gets it right. The decline effect, according to Lehrer, is the phenomenon of a scientific finding's effect size decreasing over time. Lehrer dances around the statistical explanations of the effect (regression to the mean, publication bias, selective reporting and significance fishing), and seems all-too-willing to dismiss these over a more "magical" and "handwave-y" explanation:

"This is largely because scientific research will always be shadowed by a force that can’t be curbed, only contained: sheer randomness" 

But randomness (along with the sheer number of experiments being done) is the underlying basis of the other effects he wrote about and dismissed. The large number of scientists we have doing an even larger number of experiments is not unlike the proverbial monkeys randomly plunking keys on a typewriter. Eventually, some of these monkeys will produce some "interesting" results: "to be or not to be" or "Alas, poor Yorick!" However, it is unlikely that the same monkeys will produce similar astounding results in the future.

Like all analogies, this one is imperfect as I am not trying to imply that scientists are only shuffling through statistical randomness. What I am saying is that given publication standards of large, new, interesting and surprising results, it is very likely that any experiment meeting these standards is an outlier and that its effect size will regress to the mean. This cuts two ways: although some large effects will get smaller, some experiments that were shelved for having small effects will probably have larger effect sizes if repeated in the future.

 This gets us back to my penchant for old papers. With more time, a researcher could do several replications of the study, and find the parameters under which the effects could be elicited. And often, these papers are from the pre-null-hypothesis significance testing days, so the effects tend to be larger as they need to be visually obvious from a graph. (A colleague once called this the JFO statistical test for "just f-ing obvious". It's a good standard) This standard guards against many of the statistical sins outlined by John Ioannidis.

This is also why advances in bibiometrics are going to be key for shaping science in the future. If we can formalize what makes a paper good, and what makes a scientist's work "good", then (hopefully) we can go about doing good, rather than voluminous, science.

Friday, January 7, 2011

"Hooking up" and depression

Do attachment-free sexual encounters increase or decrease depression? Let me quote from this abstract:

"Young adults who reported more depressive symptoms and feelings of loneliness at Time 1 and subsequently engaged in penetrative hook ups reported fewer depressive symptoms and lower feelings of loneliness at Time 2 as compared to young adults who did not hook up. However, young adults who reported fewer depressive symptoms and were less lonely at Time 1 and engaged in penetrative hook ups over the 4 month period reported more depressive symptoms and greater feelings of loneliness at Time 2 as compared to young adults who did not hook up."

Holy regression to the mean, batman!

Wednesday, December 29, 2010

Why brain-based lie detection is not ready for "prime time"

We are in a new and interesting legal world. Although to date, no US court cases have used brain-based lie detection techniques as evidence, several cases have sought such evidence and settled out of court. fMRI is the most frequent type of brain-based lie detection technology, with two companies, Cephos and No Lie MRI providing this service in the legal domain. There have also been attempts made to use EEG for deception detection. Notably, such a technique was used in part to prosecute a young woman for murder in India in 2008.

I am far from the first to point out that this technology is highly exploratory and not accurate enough to be used in the court of law. My goal here is to outline a good number of the reasons this is the case.

9. We do not know how accurate these techniques are. Although the two aforementioned companies boast lie detection accuracy rates of 90%+, these cannot be independently verified by an independent lab as the methods used by these companies are trade secrets. For example, there are few peer-reviewed studies of the putative EEG-based marker of deception, the P300, and most come from the lab that is commercially involved with a company trying to sell the technique as a product. Interestingly, an independent lab studying the effect of countermeasures on the technique found an 82% hit rate in controls (not the 99% accuracy claimed by the company), and this was reduced to 18% when countermeasures were used!

8. In the academic literature, where we do have access to methodology, we are limited to testing typical research participants: undergraduate psychology majors (although see this). For a lie detection method to be valid, it would need to be shown as accurate in a wide variety of populations, varying in age, education, drug use, etc. This population is not likely to be skilled in deception as a career criminal might, and it has been shown that the more often one lies, the easier it is to lie. Most fMRI-based lie detection techniques are based on the assumption that lying is hard to do, and thus requires the brain to use more energy. If frequent lying makes lying easy, then it could be the case that practiced liars don't have this pattern of brain activity.
     Although a fair amount has been made lately about WEIRD subjects, participants in these studies are actually beyond WEIRD: they are almost exclusively right handed, and predominantly male.

7. Along this same line, the "lies" that are told in these studies rarely have an impact on the lives of the student participants. Occasionally, an extra reward is given if the participant is able to "trick" the system, but in the real world, with reputations and civil liberties at stake, one might imagine that one might do a better job at tricking the scanner. However, being instructed to lie about a low-stakes laboratory situation is not the same as the high-stress situations where this technology would be used in real-life. Occasionally, a study will try to ameliorate this situation by using a mock crime (such as a theft) as the deceptive stimuli. However, these are also of limited use as participants know that the situation is contrived.

6. Like traditional polygraph tests, it is possible to fool brain-based lie detection systems with countermeasures. Indeed, in an article in press at NeuroImage, Ganis and colleagues found that deliberate countermeasures on the part of their participants dropped deception detection from 100% to 30%. Most studies of fMRI lie detection have found more brain activation for lies than truth, suggesting that it is more difficult for participants to lie. However, is this still the case with well-rehearsed lies? What about subjects performing mental arithmetic during truth to fool the scanner?
    
5. A general lack of consistency in the findings in the academic literature. To date, there are ~25 published, peer-reviewed studies of deception and fMRI. Of these studies there are at least as many brain areas implicated in deception, including the anterior prefrontal area, ventromedial prefrontal area, dorsolateral prefrontal area, parahippocampal areas, anterior cingulate, left posterior cingulate, temporal and subcortical caudate, right precuneous, left cerebellum, insula, putamen, caudate, thalamus, and various regions of temporal cortex! Of course, we know better than to believe that there is some dedicated "lying region" of the brain, and given the diversity of deception tasks (everything from "lie about this playing card" to "lie about things you typically do during the day"), the diversity of regions is not surprising. However, the lack of replication is a cause for concern, particularly when we are applying science to issues of civil liberties.

4. An additional issue surrounds the fact that many of these studies are not properly balanced. In other words, participants are instructed to lie more or less often than they are instructed to tell the truth.

3. There is a large difference between group averages and finding deception within an individual. Knowing that on average, brain region X is significantly more active in a group of subjects during deception than during truth does not tell you than for subject 2 on trial 9 than deception was likely to occur due to the differences in activation. Of course, some studies are trying to study this level of analysis, but right now they are the majority.

2. Some things that we think that are not true are not necessarily lies. Most of us believe we are above-average drivers, and smarter and more attractive than most even when these beliefs are not true. Memories, even so-called "flash-bulb" memories are not fool proof.

1. Are all lies equivalent to the brain?  Are lies about knowledge of a crime the same in the brain as white lies such as "no, honey those pants don't make you look fat" or lies of omission or self-deceiving lies?

Monday, December 20, 2010

Why no one bats .299 in late September

This paper shows that people strive for round-number goals, showing evidence from Major League baseball players, high school students taking the SAT, and from laboratory subjects answering hypothetical surveys of behavior.

As can be seen in the figure, baseball players are 4 times more likely to end the season with a 0.300 batting average than a 0.299 average! How does this happen? Players that are at 0.298 or 0.299 are more likely to have at-bats (rather than having a pinch hitter), they are slightly more likely to have hits at those at bats, and once a batter hits the magic 0.300 point, batters often take walks and sit out for pinch hitters.

The SAT takers were 10-20 percentage points more likely to re-take the test if they had an exam ending in -90 (e.g. 1190) than one ending in -00 (e.g. 1200).


Last, the authors gathered laboratory participants and asked them how they would react given certain situations. To give an example situation, imagine running laps around a track and you are getting tired. You have run either 28, 29, 30 or 31 laps (depending on what condition you are in). Do you want to run one more lap? They found that participants in the just under a round number condition (29 laps) were more likely to run one more, and participants in the just over the round number (31) were less likely to do one more.

Are these round number goals rational? In other words, is a baseball player more likely to get a lucrative contract with a 0.300 batting average than a 0.299? Do highly selective colleges have round number cut-offs for admissions? The authors examined data from university admissions that showed no discontinuities in the probability of admission as a function of SAT score, suggesting that such round number goals are not, in fact, rational.

Pope D, & Simonsohn U (2010). Round Numbers as Goals: Evidence From Baseball, SAT Takers, and the Lab. Psychological science : a journal of the American Psychological Society / APS PMID: 21148460

Friday, December 17, 2010

Roundup of epic visualizations

A map of the world revealed by networks of Facebook friendships.

Play with the New York Times' tool for visualizing census data.

Spatial distributions of tourists and locals in various cities.

How did your schools compare with others within district, state, country?

Asteroid discoveries from the last 30 years.

Sunday, October 17, 2010

How many published studies are actually true?


I’d like to point readers to this excellent new article in The Atlantic on meta-researcher John Ioannidis. Ioannidis is building quite the career on exposing the multiple biases in medical research. He has taken a field to task publishing papers with shy titles such as “Why most research findings are false”. He is rapidly becoming a personal hero of mine.

Ioannidis has examined and formally quantified research biases at all levels of “production”: in which questions are being asked, in the design of experiments, in the analysis of these experiments, and in the presentation and interpretation of the results. “At every step in the process, there is room to distort results, a way to make a stronger claim or to select what is going to be concluded,” says Ioannidis in the article. “There is an intellectual conflict of interest that pressures researchers to find whatever it is that is most likely to get them funded.”

While I have examined some of these biases for both general research and fMRI experiments, it’s worth noting that in the context of medical research, the stakes are even higher as they affect patient care. It is also unfortunate that medical studies are, according to Ioannidis, more likely to contain bias as there are stronger financial interests vested in the results, compared to cognitive neuroscience. 

An unfortunate result of the competitive research environment is a lack of replication of scientific results. Although replication is the gold standard of a result’s truth, there is little acknowledgment, and thus little motivation for researchers to do this, except for the most bold of claims. Without replication, bias in research increases. However, even when a failure to replicate a major study is published, it often gets very little attention. A case in point is the failure to replicate the “Mozart effect”: the finding that listening to 10 minutes of a Mozart sonata significantly increased participants’ performance on a spatial reasoning test. A quick Googling of “Mozart effect” will show you several companies selling you Mozart recordings to increase your child’s IQ, despite the failure to replicate.

It is very easy to get discouraged by this, after all, science should be a science, right? Ioannidis seems less discouraged, and reminds us of the following: “Science is a noble endeavor, but it’s also a low-yield endeavor… I’m not sure that more than a very small percentage of medical research is ever likely to lead to major improvements in clinical outcomes and quality of life. We should be very comfortable with that fact.”

Wednesday, September 29, 2010

Where does bad fMRI science writing come from?


A perennial favorite topic on science blogs is the examination of badly designed, or badly interpreted fMRI data. However, little time is spent on why there is so much of this material to blog about! Here, I’m listing a few reasons why mistakes in experimental design and study interpretation are so common.

Problems in experimental design and analysis
These are problems in the design, execution and analysis of fMRI papers.

Reason 1: Statistical reasoning is not intuitive
Already, I have mentioned non-independence error in the context of clinical trial design. In fMRI, researchers often ask questions about activity in certain brain areas. Sometimes these areas are anatomically defined (such as early visual areas), but more often they are functionally defined, meaning they are areas that cannot be distinguished from surrounding areas by the physical appearance of the structure, but are rather defined by responding more to one type of stimulation than another. One of the most famous functionally defined areas is the Fusiform Face Area (FFA), which responds more to faces than objects. Non-independence error often comes from these functionally defined areas. It is completely kosher to run a localizer scan containing examples of stimuli known to drive your area and stimuli known to contrast with it (faces and objects, in the case of the FFA), and then run a separate experimental block containing whatever experimental stimuli you want. Then, when you analyze your data, you test your experimental hypothesis using the voxels (“volumetric pixels”) defined by your localizer scan. What is not acceptable is to run one long block that defines your region of interest in the context of the experimental block. 

A separate, but frequent error in fMRI data analysis is the failure to correct for multiple comparisons. There are hundreds of thousands of voxels in the brain, so it is probable that high activation in any particular voxel could be due to random chance. Making this point in a memorable way was Craig Bennett and colleagues who found a 3-voxel sized area in the brain of a dead salmon that responded to photographs of emotional situations. Of course, the deceased fish was not thinking about highly complex human emotions, the area was due to chance.

Now, it is all too easy to read about these problems and feel very smug about the retrospectively obvious. But it’s not that these researchers are misleading or dumb. But the non-independence problem stated another way is “we found voxels in the brain that responded to X, and then correlated this activation with Y”. Part of the controversy surrounding “voodoo correlations” surrounds the fact that, intuitively, there doesn’t seem to be much difference between correct and incorrect data analysis. Another important factor affecting the persistence of incorrect analysis is the fact that statistical reasoning is not intuitive, and that our intuitions have systematic biases.

Reason 2: It is both too easy and too hard to analyze fMRI data
There are many steps to fMRI data analysis, and there are several software packages available to do this, both free and non-free. Data fresh out of the scanner need to be pre-processed before any data analysis takes place. This pre-processing takes out small movements made by subjects, smoothes the data to take out noise, and often warps each individual’s brain to a standard brain. For brevity, I will refer the reader to this excellent synopsis of fMRI data analysis. The problem is that it is altogether too easy to “go through the motions” of data analysis without understanding how decisions made about various parameters affect the result. And although there is wide consensus about the statistical parameters used by analysis packages, this paper shows that differences in statistical decisions made by software developers have big effects in the overall results of the study. It is, in other words, too easy to go through analysis motions that are too hard to understand.

Problems in stretching the conclusions of studies
In contrast, these are problems in the translation of a scientific study to the general public.

Reason 3: Academics are under pressure to publish sexy work
As I discussed earlier, academia is a very competitive, and it is widely believed that fMRI publications have higher impact in hiring and tenure decisions than do behavioral studies. (Note, I have not found evidence of this, but it seems like someone should have computed it). Sexy fMRI work makes great sound-bites for university donors. (See “neuro-babble and brain-porn”). Here, slight exaggerations of the conclusions may be formed (and noisy peer review does not catch it).

Reason 4: Journals compete with one another for the sexiest papers
Science and Nature each have manuscript acceptance rates below 10%. If we assume that more than 10% of all papers submitted to these journals have sufficient quality to be accepted, then it is likely that some other selection criteria is being applied during the editorial process, such as novelty. It is also of note that these journals have word limits of < 2000 words, making it impossible to fully describe experimental techniques. Collectively, these situations make it possible for charismatically expressed papers with dubious methods to be accepted.

Reason 5: Pressure on the press to over-state scientific findings
Even for well-designed, well-analyzed and well-written papers, things can get tricky in the translation to the press. Part of the problem is the fact that many scientists are not completely effective communicators. But equally problematic is the pressure placed on journalists to express every study as a revolution or break-through. The truth is that almost no published papers will turn out to be revolutionary in the fullness of time; science works very slowly. Journalists perceive that their audience would rather hear about the newly discovered “neural hate circuit” or the “100% accurate Alzheimer’s disease test” than the moderate-strength statistical association found between a particular brain area and a behavioral measure.

Reason 6: Brain porn and Neuro-babble
(I would briefly like to thank Chris Chabris and Dan Simons for putting these terms into the lexicon).  “Brain porn” refers to the colored-blob-on-a-brain style photographs that are nearly ubiquitous in popular science writing. “Neuro-babble” is often a consequence of brain porn: when viewing such a science-y picture, one’s threshold for accepting crap explanations is dramatically lowered. There have been two laboratory demonstrations of this general effect. In one study, a bad scientific explanation was either presented to participants by itself, or with one of two graphics: a bar graph or a blobby brain scan image. Participants who viewed the brain, but not the bar graph or no image were more likely to say that the explanation made sense. In the other, a bad scientific explanation was given to subjects, either alone or with the preface “brain scans show that….”. Non-scientist participants as well as neuroscience graduate students were more likely to rank the prefaced bad explanations as better, even though the logic was equally un-compelling in both cases. These should serve as cautionary tales for the thinking public.

Reason 7: Truthiness
When we are presented with what we want to believe, it is much harder to look at the world with the same skeptical glasses.

Sunday, September 26, 2010

Double-dipped sundae with a picked cherry on top

Let’s say that your hypothesis is that pitching in the 2010 baseball season is much stronger than 2009 pitching. In support of your hypothesis, you take a sample of excellent starting pitchers, say the ten pitchers with the most complete games. The average ERA for this group (at writing) is 3.17, and when you compare this to the 2009 MLB average ERA of 4.45, you say “see, I told you that we’re entering a new era of pitcher-dominated baseball!”.


Not so fast. Pitching a complete game is correlated with a low ERA (if batters were hitting you, you’d be taken out for a relief pitcher). This logic is circular: you are taking the best pitchers to prove that pitchers are great. These best pitchers are not representative of all pitchers.

Unfortunately, this statistical mistake is not uncommon in science, and a couple of recent papers have addressed this “voodoo” or “double dipping”. 

The Neuroskeptic just pointed out a particularly egregious case of a paper advocating double dipping as a way of getting better results from clinical drug trials. Briefly, their method is to run clinical trials at many centers, and then discount the centers that show a strong placebo effect. As the effect of any drug is measured by the amount of benefit that participants in the drug condition get over the participants in the placebo condition, centers with a strong placebo effect have a weaker drug effect.

Not all placebos are created equal, and not all types of patients respond to placebos in the same way. For example, severely depressed people have very little placebo effect in antidepressant trials, so antidepressants only have a strong effect in this population. 

There have been many recent, hard-hitting criticisms of several practices of big pharma, and they have been known to cherry pick studies for publication. Although only 50% of government-funded clinical drug trials find that a particular drug works, over 85% of industry-funded studies do.



Monday, September 20, 2010

Should we crowd-source peer review?


Peer review has been the gold standard for judging the quality of scientific work since World War II. However, it is a time consuming and error-prone process. Now, both lay and academic work is questioning whether the peer review system should be ditched in favor of a crowd-sourced model. 

Currently, a typical situation from an author’s perspective is to send out a paper, and receive 3 reviews about three months later. Typically, the reviewers will not completely agree with one another, and it is up to the editor to decide what to do with, for example, two mostly-positive and one scathingly negative review. How can the objective merit of a piece of work accurately be judged on such limited, noisy data? Were all of the reviewers close experts in the field? Were they rushed into doing a sloppy job? Did they feel the need for revenge against an author that unfairly judged one of their papers? Did they feel like they were in competition with the authors of the paper? Did they feel irrationally positive or negative towards the author’s institution or gender?

And from the reviewer’s point of view, reviewing is a thankless and time-consuming job. It is often a full day’s work to read, think about, and write a full and fair review of a paper. It requires accurate judgment on all matters from grammar and statistics to a determination of future importance to the field. And the larger the problems the paper has, the more time is spent in the description of and prescription for these problems. So, at the end of the day, you send your review and feel 30 seconds of gratitude that it’s over and you can go on to the rest of your to-do list. In a couple of months, you’ll be copied on the editor’s decision, but you almost never get any feedback about the quality of the review from an editor, and very little professional recognition of your efforts.

The peer review process is indeed noisy. A study of reviewer agreement of conference presentations found that the rate of reviewer agreement was not different from chance. In a study described here, women’s publications in law reviews were shown to have more citations than mens’. A possible interpretation of this result is that women are treated harsher in the peer review process, and as a consequence publish (when they can publish) better quality articles than men who do not have the same level of scrutiny. 

In peer review, one must also worry about competition and jealousy. In fact, a perfectly "rational" (Machiavellian) reviewer might reject all work that is better than his own for the purpose of advancing his career. In a simple computational model of the peer review process, it was found that the ratio of either "rational" or random reviewers needed to be kept below 30% for the system to beat chance. It also concludes that the refereeing system works the best when only the best papers are published. One can easily see how the “publish or perish” system hurts science.

It is a statistical fact that averaging over many noisy measurements provides a more accurate answer than any one answer. Francis Galton discovered this when asking individuals in a crowd to estimate the weight of an ox. Pooling over noisy estimates works when you ask for one measurement from many people, or when you ask the same person to estimate multiple times. A salient modern example of the power of crowd-sourcing is, of course,Wikipedia.

In a completely crowd-sourced model of publication, everything that is submitted gets published, and everyone who wants to can read and comment. Academic publishing would be quite similar to the blogosphere, in other words. The merits of a paper could then be determined by the citations, track backs, page views, etc.

On one hand, there are highly selective journals such as Nature who reject more than half of submitted papers before they even get to peer review and finally publish 7% of submissions. In this system, too many good papers are getting rejected. On the other hand, a completely crowd-sourced model means that there are too many papers for any scientist in the field to keep up with, and too many good papers won’t be read because it’s not worth one’s time to find diamonds in the rough. Furthermore, although the academy far from settled on the matter of how to rate professors for hiring and tenure decisions, it is more unclear what a “good” paper would be in this system as more controversial topics would get more attention.

The one real issue I see is that without editors seeking out reviewers to do the job, I worry that the only people reviewing a given paper will be the friends, colleagues and enemies of the authors, and this could make publication a popularity contest. Some data bear out this worry. In 2006, Nature conducted an experiment on the addition of open comments to the normal peer review process. Of the 71 papers that took part in the experiment, just under half received no comments at all, and half of the total comments were on only eight papers!

So, at the end of the day, I do believe that with good editorial control over comments, that a more open peer-reviewing system would be of tremendous benefit to authors, reviewers and science.


Sunday, September 19, 2010

New York Times' retraction of Alzheimer's test

Last month, an article in the New York Times proudly proclaimed a new 100% accurate test for the prediction of future Alzheimer’s disease. (This is the part where you should scroll down and read the retraction.).

Confused about the issue? Don’t feel bad – many doctors also fail at this type of reasoning.

Let’s say for the sake of argument that I tell you that I have a brand-new medical test. It’s called the “sleep gives you cancer” test, and it is one question long: Do you sleep? If you answered “yes”, then you will get cancer! As everyone who has ever had cancer has slept, this test is (according to the logic of the New York Times) 100% accurate. But, as a smart reader you will now tell me that my test isn’t so great because there are plenty of people in the world who sleep all the time and have never had cancer.

A test must do 2 things in order to be accurate: it must predict which people will get the disease (this is called sensitivity in the medical literature, and hit rate in psychology), and it also must predict which people won’t (the specificity in medical-speak, correct rejections in psych-speak). The test in question had a 100% sensitivity (everyone in their sample who later got Alzheimer’s tested positive), but 36% of people in the sample who didn’t get Alzheimer’s also tested positive.

So, how good is this test really?  Fortunately, some useful math exists to help us figure this out. Let’s say we have 1000 55-year olds. We know that 10% of them will develop Alzheimer’s by age 60. We give all 1000 people this test, and wait 5 years.  Looking at our sample, we’ll find that all 100 patients with Alzheimer’s tested positive for the test, as well as 324 (36%) of the non-Alzheimer group. Therefore, if one participant tested positive for the test there is only a 100/424 chance that s/he will have AD.

We also need to examine how useful an Alzheimer’s prediction test would be because, as of this writing, there isn’t a whole lot that can be done for AD. As pointed out here, the test described in the New York Times is based on a painful and invasive spinal tap, which makes the cost-benefit ratio quite large. However, there exist several predictive tests for AD based in neuroimaging that are less invasive.  However, given the high degree of uncertainty in the tests, coupled with the lack of meaningful therapeutic options spells years of needless anxiety for patients and families, in the opinion of this writer.