One of the central assumptions of modern neuroscience research is that the brain’s shape and structure affect behavior. Scientists have linked a thicker cortex (the brain’s outer layer) to higher intelligence, certain brain wave patterns to better volleyball ability, and higher connectivity between parts of the brain to chess-playing skills. There’s a vast number of these so-called brain-wide association studies (BWAS).
But when experts repeat these studies, they can’t replicate the results.
This “replication crisis” suggests that a substantial amount of brain-behavior imaging research rests on a shaky foundation. There are myriad problems affecting this area of research, said Randy Ellis, a senior scientist at Oracle who wrote about these issues while working as a biomedical informatician at the Icahn School of Medicine at Mount Sinai in New York.
Some of these issues aren’t unique to neuroscience. They begin with what Ellis called the “original sin” of science: Academics are put under huge pressure to publish positive research findings.
But some of these factors do apply specifically to this field.
The problems are big enough that they are “halting the growth and development of science and the curing of diseases,” Ellis told Live Science.
Tiny differences, small samples
Several brain studies that could not be reproduced in follow-up work show how problems can creep in. Each of the original papers has over 500 citations, with the number of citations reflecting how much these studies may influence thinking in the field.
Get the world’s most fascinating discoveries delivered straight to your inbox.
It can be difficult for researchers to get access to MRI scanners for their studies, and the scans are expensive. This has historically meant that many brain-imaging studies included relatively few participants.
(Image credit: Shutterstock)
For example, a landmark study in 2007 found that the brains of kids with attention-deficit/hyperactivity disorder took longer to mature. But a study published earlier this year found that this link vanished once the different rates of aging between boys and girls were taken into account. “That’s what made the whole house of cards topple,” Matthew Albaugh, co-author of the replication paper and a clinical neuroscientist at the University of Vermont, previously told Live Science.
In 2011, researchers at University College London published a study that found that the density of gray matter, the part of the brain where brain cell bodies are found, in several brain areas was linked to participants’ number of Facebook friends. (In 2011, this was a vital social metric.)
A 2015 paper failed to replicate this finding, but the authors of the original paper argued that was because the follow-up study didn’t follow the same protocol.
“It’s hard to do exactly what was done in the original studies,” said Sarah Genon, a neuroscientist at the Research Center Jülich in Germany who was not involved in either the Facebook study or its replication.
So, in 2019, Genon and her colleagues tried a different type of replication analysis.
They used a vast database and conducted more than 10,000 analyses of MRI scans from hundreds of volunteers. The median neuroimaging study sample size was around 25, so they divided this larger dataset into smaller ones and then looked for significant associations within each.
Most neuroscience studies use MRI scans to study the brain and deduce findings from these images.
(Image credit: Image Editor, Flickr)
If one of these micro-studies reported a significant association between a behavior and a brain structure, the team tried to reach the same findings using a different subset of the data. Very few of the repeat studies replicated the results of the first. “It was clear evidence that the replicability of those brain-behavior associations is relatively weak,” Genon said.
One of the key problems is that in healthy volunteers, variation among individual brains is relatively small, which means you need a large number of participants to detect average differences that tie to behavior, Genon said. That’s similar to genome-wide association studies, which pick up the very tiny effects of individual genes, and they must sample tens or hundreds of thousands of people to find robust effects. To detect meaningful differences in brain-wide imaging studies, only sample sizes in the thousands would do the trick, a 2022 paper suggested. Because brain scans are expensive and time-consuming to capture, most studies have relied on much smaller samples.
Everybody always says that replication is great. But then when I suggest that their specific studies could be replicated, often they are a bit more skeptical
Luca Kämmer, doctoral student at the Max Planck Institute for Human Cognitive and Brain Sciences
Additionally, the behavioral tests used to establish psychological variables can be inconsistent, Genon said. What’s more, some of these studies may be operating on the outdated assumption that small brain regions control complex characteristics like intelligence, she added. In reality, however, “those types of abilities are usually relatively distributed across the brain,” she said. When researchers analyze multiple small brain areas with that assumption in mind using separate statistical tests, it can raise the risk of producing mirage associations.
Some of these studies may be flawed at the outset. A 2017 brain imaging study by researchers at Weill Cornell Medical College explored whether MRI data could be used to stratify patients with depression. The researchers found that brain activity clustered into four “biotypes” of depression, each of which featured distinct, unusual activity patterns in key brain networks.
But a follow-up study found that the differences among the clusters weren’t statistically significant, meaning they could have occurred by chance. Richard Dinga, a neuroscientist at the Friedrich Schiller University Jena in Germany who worked on the follow-up study, said the main problems lay in its design.
Its statistical approach, he said, essentially guaranteed that significant links would be detected. That’s because of a problem called overfitting, in which a model is so well tuned to one set of data that it struggles to find real patterns in other datasets. This design correlated 17 clinical features of depression with 30,000 brain imaging features, but Dinga said the authors could have used “30,000 coin flips” and still produced the same strong correlation.
Conor Liston, a psychiatrist at Weill Cornell Medicine and co-author of the original paper, acknowledged the model’s susceptibility to overfitting. “That was not at all our intention, but that was the result,” Liston told Live Science
A brain scan from the re:vision project.
(Image credit: Luca Kämmer and Martin Hebart)
Still, he argued that the biotype finding they identified was real, pointing to later studies his team had published to back up their original finding. When asked why his team hadn’t updated the original paper to acknowledge that their analysis method had serious flaws, Liston said their subsequent publications were sufficient to correct the record. “I think the message is out there, and I think the solutions to those overfitting issues are very clear,” he said.
Solving the problem
The results of many studies that share the same issues as these papers will never be replicated and yet will be left to stand. For one thing, there’s no money in replication efforts. “Grant agencies wouldn’t be really excited” to fund a study whose goal was to confirm past work, Genon said. And even when these studies get done, they aren’t as likely to get published when they undermine prior findings.
For Dinga, the way forward is to collect more data. “That’s what solved the same irreproducibility problem in genetics,” he said, referring to genome-wide association studies that have produced important findings over the past two decades by using sample sizes that have reached the millions. But it’s also important to improve study design, he added, since bigger datasets won’t fix studies with incorrect analysis approaches. “The methods just don’t have a chance to produce something useful,” he said.
Some researchers are leveraging large brain imaging datasets. Martin Hebart, a group leader at the Max Planck Institute for Human Cognitive and Brain Sciences in Germany, and Luca Kämmer, a doctoral student in Hebart’s lab, are developing re:vision, a project that is collecting imaging data — the largest initiative yet to capture how the brain responds to visual stimuli. They plan to make this information available to other researchers, who will then explore whether hypotheses tested in other papers can be replicated in re:vision’s dataset.
Hebart and Kämmer said they have already received over a dozen applications for their project. Reception has been particularly positive among younger researchers in the field. But convincing senior scientists to retest their data might take more work because they have more to lose if their long-standing finding is ultimately not replicated, Kämmer said.
“Everybody always says that replication is great,” he told Live Science. “But then when I suggest that their specific studies could be replicated, often they are a bit more skeptical.”
See how much you know about the most complex organ in the human body with our brain quiz!













