Managing Research Projects

A practical system for getting research finished

John M. Drake

Odum School of Ecology & Center for the Ecology of Infectious Diseases

CEID Professional Development Workshop

Center for the Ecology of Infectious Diseases, Athens, GA

October 9, 2026

Every lab has a graveyard of projects that have been 80% done for five years.

Projects fail from entropy, not stupidity

  • Undefined endpoints — nobody said what “done” is
  • Mission creep — the question moved after the data arrived
  • Entropy — files, decisions, and intentions decay

Everything today is a cheap, mechanical defense against all three failure modes.

Roadmap

  1. What a project is (and isn’t)
  2. Choosing the project
  3. The protocol — why written plans change outcomes
  4. Structure: repositories, folders, reproducible code
  5. Git and GitHub, the concept
  6. Driving to done
  7. You draft a protocol + discussion

What a project is

What are you working on?

A project is the work aimed at one product

  • Project = all the documents, data, and code aimed at one product
  • Product = something you can deliver: a paper, a chapter, a package, a data archive
  • One project = one product = one repository

If you can’t write a plan for it, it isn’t a project

  • A dissertation is not a project; each chapter is
  • Peer reviews, proposals, and presentations are not projects
  • An NSF or NIH grant is not a project either — it’s an umbrella funding mechanism that pays for several
  • A project is over when the product is delivered — the second paper from the same data is a new project, not scope growth

Choosing the project

The highest-leverage decision you will make

The stakes: only about 57% of PhD students finish within ten years

  • ~57% ten-year completion (~63% in the life sciences), students entering 1992–95 at ~30 universities
  • Project choice is the highest-leverage, least-taught decision in a research career
  • Time spent choosing — weeks, even months — is not procrastination

Source: Council of Graduate Schools, PhD Completion Project (2008). Alon (2009): “Do not commit to a problem before 3 months have elapsed.”

Plot every candidate project on two axes


“If you do not work on an important problem, it’s unlikely you’ll do important work.”

— Richard Hamming, “You and Your Research” (1986)

Candidate projects plotted by feasibility and interest (gain in knowledge); adapted from the same two axes in Fig. 1 of Alon (2009), Mol Cell 35:726.

Where I part ways with Alon: everyone works easy/high-gain

  • Alon: first problems are easy and modest; postdocs need feasible and high-gain; new PIs may take on a hard grand challenge — if it splits into tractable projects
  • My advice: difficulty is not a virtue — it is a cost that discounts everything downstream
  • For almost every hard problem there is an easier project that teaches the same lesson — find the easier route before you commit

Most published science is conservative — for a reason

  • Across ~6.5 million biomedical abstracts, 86% of the chemical relationships reported were already known; 14% were new
  • Papers introducing a new chemical were cited more — 12.9 citations on average versus 8.4 — but a project that fails produces no paper at all, and by the authors’ calculation the extra citations don’t make up for that risk
  • Highly novel papers had ~57% higher odds of a top-1% hit — and fatter tails at both ends
  • In a field experiment, expert reviewers gave the most novel proposals markedly lower scores

Sources: Foster, Rzhetsky & Evans (2015) Am Sociol Rev 80:875; Wang, Veugelers & Stephan (2017) Res Policy 46:1416; Boudreau et al. (2016) Manag Sci 62:2765.

The highest-impact papers are conventional work with one unusual combination

  • 17.9 million papers, scored by how usual or unusual each pair of cited journals is
  • Papers that were mostly conventional but included one rare combination reached the top 5% of citations 9.1% of the time — versus 5.8% for purely conventional papers and 5.3% for purely novel ones
  • Only 7% of papers achieve that combination
  • Teams of three or more included an unusual combination in half their papers; solo authors in a third

Source: Uzzi, Mukherjee, Stringer & Jones (2013) Science 342:468.

Donald Stokes and Pasteur’s Quadrant

  • Donald E. Stokes (1927–1997) — political scientist; co-author of The American Voter (1960); dean of Princeton’s Woodrow Wilson School, 1974–1992
  • Pasteur’s Quadrant (Brookings, 1997) was published shortly after his death
  • His target: the old linear model — basic research upstream, applied downstream, purity and usefulness as opposite ends of one axis
  • Stokes split that into two independent questions — and the quadrant falls out

Point the project with Pasteur

Stokes (1997): two axes, not one. Use-inspired basic research — Pasteur’s quadrant — is where understanding and consequence advance together.

Discussion

Where does your current project sit on the feasibility–interest diagram?

Could an easier project teach the same lesson?

“Here to Help” — Randall Munroe, xkcd.com (CC BY-NC 2.5)

The protocol

Write the plan before the data exist

After trial registration arrived, positive results fell from 57% to 8%

  • 55 large NHLBI trials of drugs and supplements, all with cardiovascular outcomes
  • Published before 2000: 17 of 30 positive on their primary outcome
  • Published 2000 or later, with outcomes registered in advance on ClinicalTrials.gov: 2 of 25 positive

Source: Kaplan & Irvin (2015) PLoS ONE 10:e0132382.

Four innocuous choices, left open, turn 5% false positives into 61%

  • Two outcome variables instead of one; peeking before adding observations; an optional covariate; dropping a condition
  • Each inflates error modestly; combined: a 60.7% false-positive rate at nominal p < .05
  • No fraud required — it’s enough that you would have analyzed different data differently (the “garden of forking paths”)

Sources: Simmons, Nelson & Simonsohn (2011) Psych Sci 22:1359; Gelman & Loken (2014) Am Sci 102:460.

This is us: 174 ecology teams, same data, opposite conclusions

  • Two unpublished datasets (blue tit sibling competition; Eucalyptus recruitment) analyzed independently by 174 teams
  • Effects ranged from significantly negative to significantly positive
  • Peer-rated analysis quality did not predict who landed where
  • In a survey of ecologists and evolutionary biologists, 64% admitted cherry-picking and 51% HARKing — Hypothesizing After the Results are Known: presenting an unexpected finding as if it had been predicted from the start

Sources: Gould et al. (2025) BMC Biology 23:35; Fraser et al. (2018) PLoS ONE 13:e0200303.

The Drake Lab protocol: a contract with your future self

A real protocol: Tredennick’s 2018 measles early-warning study — background, research questions, and study design, set down at the outset and kept current as the work moved.

Seven parts, one page is enough to start

  1. Title and researchers — an originator with contact info
  2. Background and rationale — 2–3 succinct paragraphs
  3. Study design — steps + explicit research questions (for models: design it like an experiment — Drake 2025)
  4. Data sources — with permissions needed
  5. Analysis — including the anticipated figures, sketched before the data exist (for models: treatments, levels, and responses — Drake 2025)
  6. Checklist — tasks with dates and owners
  7. Changelog — the plan may change, but only in writing, with the reason

For computational projects, the study design is an experimental design

  • Treatments — what you vary (parameters, scenarios, interventions)
  • Levels — the values each treatment takes, chosen in advance
  • Responses — the outputs you will measure, named before the first run
  • Replication — how many stochastic realizations, and the seed policy
  • Control — the baseline scenario everything is compared against

After Drake (2025), “Modelling Like an Experimentalist,” Ecology Letters 28:e70251.

Discussion

If your current project had a changelog, what would its first entry say?

“The General Problem” — Randall Munroe, xkcd.com (CC BY-NC 2.5)

Structure

Repositories, folders, and code that reruns

One project = one repository, named so a stranger can find it

  • Create the repository the day the protocol exists
  • Naming: surname-description (individual) or drakelab-description (team)
  • README.md is the front door: project name, originator, collaborators, the protocol — and a map of how the repository is organized
  • Private project, public product: the repo is the full internal history; the product (paper + archive) is the curated public face

A standard layout means never asking “where does this go?”

project/
├── README.md        protocol, or pointer to it
├── data/            raw, read-only, never edited
├── scripts/         01-clean.R, 02-fit.R, 03-figures.R
├── output/          derived data + figures, regenerable
└── documents/       manuscript, notes, correspondence

Two rules do most of the work: raw data are sacred; everything derived is disposable.

After Noble (2009) PLoS Comput Biol 5:e1000424 and Wilson et al. (2017) PLoS Comput Biol 13:e1005510.

Undocumented projects decay

  • Information entropy: detail is lost from the moment of publication — specifics first, then generalities, with step losses at retirement, accidents, and death
  • Odds a dataset still exists fall 17% per year after publication
  • Even under strong archiving policies, 56% of ecology datasets were incomplete and 64% archived in ways that partly or wholly prevented reuse

Sources: Michener et al. (1997) Ecol Appl 7:330; Vines et al. (2014) Curr Biol 24:94; Roche et al. (2015) PLoS Biol 13:e1002295.

The test: if output/ was lost tonight, could one command rebuild it?

  • When published R code was re-run in a clean environment: 74% of scripts failed — library and working-directory errors among the most common
  • In ecology journals with code-sharing policies, 79% of articles shared data but only 27% shared code
  • Everything by script; relative paths; record seeds and package versions

Sources: Trisovic et al. (2022) Sci Data 9:60; Culina et al. (2020) PLoS Biol 18:e3000763; Sandve et al. (2013) PLoS Comput Biol 9:e1003285; Noble (2009).

Version control is a lab notebook you cannot lose

Commit per logical change; push at least once a day.

Commits are dated snapshots of the whole project; branches let ideas fail safely; GitHub holds the shared copy. Mechanics — branching, merging, pull requests — get their own hands-on workshop: October 20, 2:55 PM.

Driving to done

The checklist is the engine; dependencies are where projects stall

The Tredennick protocol’s checklist, mid-project: struck-through tasks, surviving blockers, and a changelog recording why the plan changed.

Done means delivered — then you stop

  • A project is over when the product is delivered: paper out, code and data archived (e.g., Dryad) with the repository linking to them
  • New ideas from old projects become spinoffs with new protocols and new repos
  • The portfolio view: one fast/safe, one core, one stretch — chapters as separate projects makes this natural

Exercise

Start your own protocol — right now

Fifteen minutes: draft the skeleton for one real project

  1. Name the product — one sentence: what do you deliver?
  2. Background — two sentences
  3. Research question — and what would an answer look like? (a p-value? a pattern in a graph? a table of numbers?)
  4. Three checklist items — each with a date
  5. One dependency — the thing most likely to block you

Then two minutes: tell your neighbor your product in one sentence.

Discussion

What did you discover drafting it? Where do you disagree with the system?

“Standards” — Randall Munroe, xkcd.com (CC BY-NC 2.5)

The starter kit

  • Protocol template + worked example: github.com/DrakeLab/Wiki → lab-docs → project-protocols
  • Folder skeleton: README + data + scripts + output + documents
  • One repo per project, created the day the protocol exists
  • Definition of done: the product is delivered, the archive is public, you stop

You don’t need more discipline or better software. You need a plan you can print, a repo with five folders, and a definition of done. Write the protocol this week.

Backup

Does better documentation of analytic choices fix the many-analysts problem?

  • 161 researchers, 73 teams, one dataset (ISSP) and one hypothesis — does immigration erode support for social policy? — 1,261 models submitted
  • Of 89 team-level conclusions: ~61% reject, ~26% support, ~13% “not testable”; estimates ran from large negative to large positive
  • After coding 107 analytic decisions plus expertise and prior beliefs, 95.2% of the variance in results remained unexplained
  • Expertise didn’t shrink it, and no team changed its model after seeing the others’ — the divergence is idiosyncratic, not bias

Source: Breznau et al. (2022) PNAS 119:e2203150119.

What do honest base rates look like?

  • Standard psychology literature: 96% of first hypotheses “confirmed” (n=152 papers)
  • Registered Reports: 44% (n=71)
  • Across 113 Registered Reports in biomedicine and psychology, ~61% of 296 hypotheses were not supported

Sources: Scheel, Schijen & Lakens (2021) AMPPS 4(2); Allen & Mehler (2019) PLoS Biol 17:e3000246.

How much do effects shrink on replication?

  • When an experiment is repeated, the measured effect is almost always smaller than the original report — the gap is the shrinkage
  • Reproducibility Project: Cancer Biology — for positive effects, median replication effect 85% smaller than original; 46% of effects replicated by a majority of five criteria
  • The project planned 193 experiments and completed 50 — missing data, under-described methods, and unhelpful authors blocked the rest
  • Nature survey of 1,576 researchers: >70% failed to reproduce another’s experiment; >50% failed to reproduce their own

Sources: Errington et al. (2021) eLife 10:e71601 and 10:e67995; Baker (2016) Nature 533:452.

Why doesn’t the system just reward good methods?

  • Model: when institutions reward publication volume, low-effort methods out-reproduce careful ones — no misconduct required
  • Mean statistical power to detect small effects in the social and behavioral sciences: 0.24, unchanged 1960–2011

Source: Smaldino & McElreath (2016) R Soc Open Sci 3:160384.

Doesn’t “work on important problems” contradict easy/high-gain?

  • “If you do not work on an important problem, it’s unlikely you’ll do important work.”
  • But: “It’s not the consequence that makes a problem important, it is that you have a reasonable attack.”
  • Great scientists keep 10–20 important problems waiting for an attack to appear

Source: Hamming (1986), “You and Your Research,” Bell Communications Research Colloquium.

References

References (1 of 2)

Allen C & Mehler DMA (2019) PLoS Biology 17:e3000246. doi:10.1371/journal.pbio.3000246

Alon U (2009) Molecular Cell 35:726. doi:10.1016/j.molcel.2009.09.013

Baker M (2016) Nature 533:452. doi:10.1038/533452a

Boudreau KJ et al. (2016) Management Science 62:2765. doi:10.1287/mnsc.2015.2285

Breznau N et al. (2022) PNAS 119:e2203150119. doi:10.1073/pnas.2203150119

Council of Graduate Schools (2008) PhD Completion and Attrition. Washington, DC

Culina A et al. (2020) PLoS Biology 18:e3000763. doi:10.1371/journal.pbio.3000763

Drake JM (2025) “Modelling Like an Experimentalist.” Ecology Letters 28:e70251. doi:10.1111/ele.70251

Errington TM et al. (2021) eLife 10:e71601. doi:10.7554/eLife.71601

Errington TM et al. (2021) eLife 10:e67995. doi:10.7554/eLife.67995

Foster JG, Rzhetsky A & Evans JA (2015) Am Sociol Rev 80:875. doi:10.1177/0003122415601618

Fraser H et al. (2018) PLoS ONE 13:e0200303. doi:10.1371/journal.pone.0200303

Gao X et al. (2024) Scientometrics 129:909. doi:10.1007/s11192-023-04925-w

Gelman A & Loken E (2014) American Scientist 102:460. doi:10.1511/2014.111.460

Gould E et al. (2025) BMC Biology 23:35. doi:10.1186/s12915-024-02101-x

Hamming RW (1986) “You and Your Research,” Bell Communications Research Colloquium

Kaplan RM & Irvin VL (2015) PLoS ONE 10:e0132382. doi:10.1371/journal.pone.0132382

Kerr NL (1998) Pers Soc Psychol Rev 2:196. doi:10.1207/s15327957pspr0203_4

References (2 of 2)

Michener WK et al. (1997) Ecological Applications 7:330. doi:10.1890/1051-0761(1997)007[0330:NMFTES]2.0.CO;2

Noble WS (2009) PLoS Comput Biol 5:e1000424. doi:10.1371/journal.pcbi.1000424

Perez-Riverol Y et al. (2016) PLoS Comput Biol 12:e1004947. doi:10.1371/journal.pcbi.1004947

Ram K (2013) Source Code Biol Med 8:7. doi:10.1186/1751-0473-8-7

Roche DG et al. (2015) PLoS Biology 13:e1002295. doi:10.1371/journal.pbio.1002295

Sandve GK et al. (2013) PLoS Comput Biol 9:e1003285. doi:10.1371/journal.pcbi.1003285

Scheel AM, Schijen MRMJ & Lakens D (2021) AMPPS 4(2). doi:10.1177/25152459211007467

Silberzahn R et al. (2018) AMPPS 1:337. doi:10.1177/2515245917747646

Simmons JP, Nelson LD & Simonsohn U (2011) Psych Science 22:1359. doi:10.1177/0956797611417632

Smaldino PE & McElreath R (2016) R Soc Open Sci 3:160384. doi:10.1098/rsos.160384

Stokes DE (1997) Pasteur’s Quadrant. Brookings Institution Press

Tredennick AT et al. (2022) J R Soc Interface 19:20220123. doi:10.1098/rsif.2022.0123

Trisovic A et al. (2022) Scientific Data 9:60. doi:10.1038/s41597-022-01143-6

Uzzi B et al. (2013) Science 342:468. doi:10.1126/science.1240474

Vines TH et al. (2014) Current Biology 24:94. doi:10.1016/j.cub.2013.11.014

Wang J, Veugelers R & Stephan P (2017) Research Policy 46:1416. doi:10.1016/j.respol.2017.06.006

Wilson G et al. (2017) PLoS Comput Biol 13:e1005510. doi:10.1371/journal.pcbi.1005510