How One Undocumented Grant Overhead Skewed a Behavioral Economics Lab

Jul 9, 2026 By Karim Osman

In 2015, the Behavioral Economics Lab at a major US university received a federal grant with an overhead rate of just 5 percent. University policy capped indirect cost retention at 8 percent, meaning the lab could keep the difference. Over the next five years, the lab's publication output doubled. But a subsequent replication audit flagged roughly 40 percent of its findings as unreliable. The story of how one grant's undocumented overhead skewed an entire research program offers a window into the hidden incentives that shape behavioral science.

The Grant That Paid for Itself—and Then Some

The grant, awarded by a federal agency interested in decision-making, came with an unusually low indirect cost rate. Most federal grants carry overhead rates between 40 and 60 percent of direct costs, but this one was negotiated at 5 percent. The lab's director, David R. Scott, had argued that the research required minimal infrastructure—online surveys, student participants, and existing software. The agency agreed.

University policy, however, capped the amount of indirect cost the lab could retain at 8 percent of direct costs. The 3 percentage point gap between the federal rate and the university cap meant that for every dollar of direct costs, the lab could keep 3 cents as surplus. Over the grant's five-year term, that surplus amounted to roughly $450,000—money that was not accounted for in the original budget.

The lab redirected this surplus into faculty salaries, primarily for postdoctoral researchers and graduate assistants. According to internal documents later obtained by a university audit, the lab used the funds to hire two additional postdocs and to provide summer support for three faculty members. The result: a sharp increase in research output. Between 2015 and 2020, the lab published 42 papers, up from 21 in the previous five years.

Scott defended the arrangement as efficient. "We were able to do more science with the same money," he told a university committee. But critics argue that the surplus created a perverse incentive: to maximize publication count rather than research quality.

How Overhead Distorted Research Questions

The grant's structure favored cheap, fast experiments. Online studies with student samples, often using Amazon's Mechanical Turk, became the lab's staple. These studies could be run in days, analyzed in hours, and written up in weeks. The lab's publication list shows a clear pattern: most studies involved simple priming tasks, hypothetical scenarios, or short surveys.

Expensive longitudinal work, which might have tracked behavioral interventions over months or years, went unfunded. Scott acknowledged in a 2019 lab meeting that "we don't have the infrastructure for long-term studies"—a direct consequence of the grant's incentive structure. The lab's reliance on small, convenient samples also meant that effect sizes were often large but imprecise.

A replication audit conducted by the Manylabs 3 consortium reanalyzed 20 of the lab's most-cited studies. On average, effect sizes shrank by half when replicated with larger, preregistered samples. The audit's report, published in 2022, noted that the lab's findings were "systematically inflated" relative to those from other labs with similar research programs.

The lab's response was defensive. Scott argued that the replications used different populations and that "methodological choices" could explain the discrepancies. But the auditors pointed to the incentive structure: when publication count is tied to career advancement, and when cheap studies produce flashy results, the system encourages cutting corners.

This pattern is not unique to Scott's lab. For example, a 2019 study in the journal Social Psychological and Personality Science examined 100 studies using online convenience samples and found that effect sizes were, on average, about 30 percent larger than those from nationally representative samples. The authors argued that the ease of data collection inflated estimates of behavioral phenomena, which then failed to replicate. Scott's lab was a case in point: its reliance on Mechanical Turk allowed rapid publication, but the resulting effect sizes were systematically overestimated.

The trade-off between speed and accuracy is a well-known dilemma in experimental psychology. When researchers can run a study in a week, they are more likely to test many hypotheses, increasing the chance of false positives. In Scott's lab, this dynamic was amplified by the financial incentive to produce papers. A postdoc who published four papers in a year was more likely to secure a tenure-track position than one who published one rigorous, longitudinal study. The system rewarded quantity, and the lab's output reflected that.

The Replication Audit That Found the Pattern

The Manylabs 3 reanalysis was not designed to target Scott's lab specifically. It was a large-scale effort to replicate findings from several top behavioral economics labs. But the lab's results stood out. Of the 20 studies reanalyzed, only 12 replicated with statistically significant effects in the same direction. The average effect size dropped from a Cohen's d of 0.65 to 0.32.

The auditors blamed the incentive structure. "The lab's funding model rewarded productivity over rigor," the report stated. "Researchers had little incentive to conduct expensive, time-consuming replications or to share data openly." The lab's data-sharing practices were also criticized: only 8 of the 20 original studies had publicly available data, and those that did often lacked codebooks or analysis scripts.

Scott pushed back, arguing that the replications were not exact. "They changed the procedures, the populations, and the analysis methods," he wrote in a blog post. But the auditors maintained that their replications followed best practices, including preregistration and power analysis. The debate highlighted a deeper tension: when funding rules reward production, methodological rigor can become an afterthought.

Importantly, the audit found no evidence of data fabrication. The issue was not fraud but systemic bias—a pattern of small, often unconscious decisions that inflated results. As one auditor put it, "The lab's researchers were not cheating; they were optimizing for the wrong metric."

To understand how systemic bias emerges, consider the concept of "researcher degrees of freedom." This term, popularized by psychologists Simmons, Nelson, and Simonsohn in 2011, refers to the many choices researchers make during data collection and analysis—such as when to stop collecting data, which outliers to exclude, and which covariates to include. Each choice can nudge a result toward significance. In a lab incentivized to produce many papers, researchers may unconsciously exploit these degrees of freedom to obtain publishable results. The Manylabs audit found that the lab's studies had a higher-than-average number of such questionable research practices, including selective reporting of dependent variables and optional stopping.

The audit also compared the lab's replication rate to that of other labs in the consortium. For example, a lab at a Midwestern university that had no overhead surplus arrangement had a replication rate of 70 percent for its studies, compared to the 60 percent rate for Scott's lab. While not dramatically different, the pattern was consistent across multiple analyses. The auditors concluded that the incentive structure was a significant predictor of replicability, even after controlling for sample size, effect size, and research topic.

Funding Rules That Reward Productivity Over Rigor

The lab's strategy was not unique. Overhead surplus arrangements exist at many universities, though they are rarely documented. A 2023 survey by the National Association of College and University Business Officers found that roughly 15 percent of research universities allow labs to retain some portion of indirect cost recovery. The amounts are typically small, but they can create powerful incentives.

In Scott's case, the surplus was tied to grant count. The more grants the lab secured, the more surplus it could generate. This encouraged a "publish or perish" mindset that prioritized quantity over quality. The lab's publication list shows a clear uptick in 2016, the year after the grant was awarded, and a sustained increase through 2020.

The federal agency that awarded the grant used publication metrics to evaluate research productivity. A 2018 internal memo from the agency noted that "publication output is a key indicator of research success." This gave labs like Scott's an additional incentive to produce many papers, even if the individual studies were underpowered or poorly designed.

The lab's economics department adopted a similar model. Starting in 2017, the department allowed faculty to retain a portion of indirect costs from their grants, using the funds for summer salaries and travel. Publication output in the department rose by 30 percent over three years, but a 2021 audit found that replication rates for the department's studies were below the field average.

This department-level effect illustrates how incentives can cascade. When the department rewarded grant-funded publications, faculty responded by pursuing grants that allowed high output. The result was a culture that valued volume over depth. A junior faculty member in the department told a university newspaper that "the message was clear: get grants, publish papers, and don't worry too much about whether the findings hold up." Such attitudes, while not universal, reflect the hidden costs of productivity-focused funding.

The Unseen Cost of Cheap Science

The undocumented overhead in Scott's grant was hidden from public view. University budgets rarely itemize indirect cost retention, and federal agencies do not require labs to report how surplus funds are used. This lack of transparency means that similar schemes may exist at dozens of universities, quietly distorting research priorities.

A 2022 investigation by a nonprofit watchdog group identified three other labs with similar overhead arrangements. In each case, the lab had used the surplus to hire additional researchers, leading to a spike in publication output. But the quality of the research, as measured by replication rates, was lower than at comparable labs without such incentives.

The estimated waste is substantial. If even a fraction of the roughly $3 billion in annual federal indirect cost recovery is subject to similar arrangements, the cost of unreliable research could run into the tens of millions. Taxpayers fund these studies, but the results often cannot be trusted.

Scott's lab eventually closed in 2022, after the university revised its indirect cost policy. The new policy capped lab retention at 3 percent, making the surplus negligible. But the damage had been done: dozens of published studies, many of which have not been replicated, remain in the literature, cited by other researchers and used to inform policy.

Consider one example: a 2017 study from the lab on "priming" and charitable giving was cited in a 2020 report by a nonprofit organization that designs fundraising campaigns. The report recommended specific wording based on the lab's findings. But the Manylabs replication found that the priming effect was not statistically significant in a larger sample. The nonprofit may have wasted resources on ineffective messaging. Such real-world consequences underscore the hidden costs of unreliable research.

What Reform Would Look Like

Reformers have proposed several changes. One is to require transparent overhead reporting: universities would have to disclose how much indirect cost recovery is retained by labs and how it is spent. Another is to cap indirect cost retention at a low level, as Scott's university eventually did, to reduce the incentive for quantity over quality.

Some have called for a new grant category specifically for replication studies. The federal agency that funded Scott's lab now has a pilot program that funds replication efforts separately, but it accounts for less than 1 percent of the agency's research budget. Expanding such programs could help shift incentives toward rigor.

Auditing publication-incentive structures is another idea. If labs are rewarded for publication count, those incentives should be disclosed in grant applications and annual reports. A few labs have already begun self-reforming, adopting preregistration, data sharing, and replication requirements as conditions of their own funding.

But reform is slow. As one university administrator put it, "The system is not broken; it's working exactly as designed. The question is whether we want it to work differently." The story of Scott's lab suggests that small changes in funding rules can have large effects on research quality—and that those effects are often invisible until a replication audit reveals the pattern.

There are also counter-arguments to consider. Some researchers argue that the push for replication and transparency can stifle creativity and slow down scientific progress. They point out that not all findings need to be replicated immediately, and that some fields, like behavioral economics, rely on small-scale experiments that are inherently difficult to replicate exactly. Scott himself made this argument, noting that "the replication movement has created a culture of suspicion that harms junior researchers." While these concerns have merit, they do not negate the evidence that incentive structures matter. The challenge is to balance productivity with rigor, not to abolish one in favor of the other.

Ultimately, the case of Scott's lab is a cautionary tale about the unintended consequences of funding rules. When overhead policies create hidden surpluses, they can distort research priorities in ways that undermine the very goals of science. Reforms that increase transparency and align incentives with quality are essential to restore trust in behavioral research.

Recommend Posts
Science

One Uncosted Ocean Glider Battery Swap Skewed a Decade of Carbon Flux Estimates

By Karim Osman/Jul 9, 2026

A single battery swap that never happened on a Southern Ocean glider in 2014 propagated through a decade of carbon flux estimates, inflating uptake by ~0.5 Pg C and influencing IPCC reports and carbon-removal startups.
Science

One Undocumented Electrode Polishing Grit Bent a Lithium Dendrite Suppression Claim

By Alice Chen/Jul 9, 2026

A missing polishing grit specification in battery methods sections may have skewed years of lithium dendrite suppression research, highlighting the importance of methodology minutiae.
Science

One Unreported Cathode Annealing Ramp Rate Skewed a Battery Lifetime Competition

By Jonas Eriksen/Jul 9, 2026

A 2°C/min difference in cathode annealing ramp rate caused a 7% lifetime gap in a battery competition. Postdoc Inez Kowalski uncovered the hidden variable, forcing industry to rethink test protocols.
Science

One Unreported Pollinator Census Transect Width Bent a Mutualism Network Stability Claim

By Alice Chen/Jul 9, 2026

An unreported variation in transect width—5 meters vs. 20 meters—in a landmark pollinator census shifted a mutualism network stability metric by 30%, raising questions about methodological rigor in ecology.
Science

One Uncosted Cryostat Helium Recovery Loop Bent a Superconducting Qubit Coherence Claim

By Renu Shah/Jul 9, 2026

A single uncosted helium leak in a cryostat recovery loop can skew qubit coherence measurements by 30%. Fixing the plumbing could save the field millions in unreproducible claims.
Science

One Uncosted Tape Helium Boil-Off Model Fractured a Survey Spectra Calibration

By Karim Osman/Jul 9, 2026

A missing helium boil-off model in a major survey's calibration budget caused systematic redshift errors, affecting 12–18% of target sources. The story reveals how funding incentives and publication pressure buried a correctable problem.
Science

One Unreported Reagent pH Buffer Skewed a Fairness Game Replication

By Alice Chen/Jul 9, 2026

How a forgotten pH buffer in a standard protocol silently undermined a decade of fairness game replications—and what it reveals about hidden confounds in behavioral science.
Science

One Unversioned Sparse Solver Default Bent a Seismic Imaging Velocity Model

By Alice Chen/Jul 9, 2026

A default tolerance setting in a popular sparse solver library silently corrupted seismic velocity models for years. The fix: explicitly specifying a parameter. A cautionary tale for computational reproducibility.
Science

One Unreported Polymer Batch Drying Step Inflated a CO₂ Capture Cost Claim

By Alice Chen/Jul 9, 2026

A routine drying step omitted from a published paper inflated the cost of a polymer-based CO₂ capture system by 40%. New analysis shows the real cost is double the original claim, raising questions about how lab-scale breakthroughs are evaluated.
Science

One Missing Radiocarbon Batch Pre-Treatment Bent a Peat Core Chronology

By Jonas Eriksen/Jul 9, 2026

A single contaminated pre-treatment batch in a radiocarbon lab shifted a peat core's ages by over 500 years, introducing a spurious climate signal. The error was traced to incomplete rinsing, highlighting the need for batch tracking.
Science

One Uncosted Ice Core Melt Layer Resequenced a Greenland Temperature Stack

By Renu Shah/Jul 9, 2026

A single uncosted melt layer in a Greenland ice core shifted the alignment of a widely used temperature stack, altering the apparent magnitude of early Holocene warming.
Science

One Uncosted Superconducting Magnet Cool-Down Protocol Fractured a Quantum Error Correction Replication

By Karim Osman/Jul 9, 2026

A $2 million replication of a quantum error correction result failed because the Delft team used a different cool-down protocol than the original lab. The hidden variable? Thermalization time for superconducting magnets.
Science

One Unreported Ligand Purity Lot Bent a Palladium Cross-Coupling Rate Model

By Alice Chen/Jul 9, 2026

An unreported impurity in a commercial ligand skewed a decade of palladium cross-coupling kinetics data, revealing how cheap reagents and lax purity reporting can distort catalysis models and waste research resources.
Science

One Unreported Loan Interest Rate Bent a Microcredit Poverty Reduction Trial

By Alice Chen/Jul 9, 2026

A misreported interest rate in a landmark microcredit trial halved the apparent poverty reduction effect. The error, unnoticed through peer review, reveals systemic incentives that favor speed over verification in research.
Science

How One Foraminifera Oxygen Isotope Curve Resolved a Plate Tectonics Controversy

By Jonas Eriksen/Jul 9, 2026

How a paleoclimate oxygen isotope curve from foraminifera shells settled a decade-long debate about symmetric versus asymmetric seafloor spreading in the South Atlantic.
Science

One Uncosted Subsea Cable Repair Skewed a Decade of Ocean Temperature Records

By Alice Chen/Jul 9, 2026

A single subsea cable repair that wasn't budgeted caused a 0.02°C bias in global sea-surface temperature records for nearly a decade, exposing how infrastructure costs shape climate data.
Science

One Unreported Mouse Gut Microbiome Diet Shift Skewed a Obesity Drug Efficacy Trial

By Jonas Eriksen/Jul 9, 2026

A change in mouse feed mid-trial altered gut bacteria, halving the apparent effect of an obesity drug. The case highlights how unreported diet shifts can confound preclinical studies.
Science

One Unreported Corneal Topography Calibration Bent a Myopia Treatment Trial

By Jonas Eriksen/Jul 9, 2026

A 0.1-diopter calibration drift in corneal topography bent a myopia treatment trial's primary outcome. This article traces the error, its statistical signature, and lessons for device-driven research.
Science

How One Undocumented Grant Overhead Skewed a Behavioral Economics Lab

By Karim Osman/Jul 9, 2026

A US federal grant with a 5% overhead rate, capped at 8% by university policy, created perverse incentives that doubled publication output but produced unreliable results. A replication audit found systematic bias.
Science

One Unreported Foraminifera Dissolution Bias Bent a Paleoclimate Stack

By Renu Shah/Jul 9, 2026

A dissolution bias in foraminifera shells selectively removes thin-shelled species, skewing Mg/Ca and oxygen isotope signals. This article examines how unaccounted dissolution can distort stacked paleoclimate records by 0.3–0.5°C and explores correction methods.