Study Selection #5: The Lord of the reviews, the two towers of screening

Screens slow now. Go fast later.
Introduction
Study screening looks simple from the outside: search, sift, include, exclude. But after more than a decade working on systematic reviews — first in academia, where I learned the hard way by making my own mistakes across several published reviews, and then six years at the National Guideline Alliance (NGA) in London — I realised that study selection is one of the most fragile parts of the whole evidence process. At the NGA, I worked as both a systematic reviewer and a health economist on NICE guidelines across clinical, public health, qualitative, and economic evidence. I screened thousands of titles, abstracts, and full texts for developing guidelines on may subjects (including supporting adult carers [NG150]; early & locally advanced breast cancer [NG101]; epilepsies [NG217]; pancreatic cancer [NG85]; self-harm [NG225]; cerebral palsy; [NG119]; cystic fibrosis [NG78]; and endometriosis [NG73]). Across these years — and through my own academic reviews — I learned that study screening is where reviews succeed or fail.
Good study selection supports three things that the PRISMA 2020 statement and Cochrane repeat again and again:
- Transparency — readers can see how decisions were made
- Reproducibility — another team could repeat the process
- Consistency — decisions follow the same rules from start to end
But screening is also a very human task. It involves judgment. Fatigue. Interpretation.
And without good systems, things fall apart fast.
In this final post of the series related to the screening process in systematic reviews, I look back at what I learned: the methods, the reality, and the things that can quietly go wrong. This post brings together the four pillars discussed in earlier posts (dual screening, eligibility criteria, pilot testing, full-text and snowballing) and adds what I learned across my own published reviews and guideline work.
What often goes wrong!
Even with good intentions and solid methods, the same problems appear again and again.
I saw them in my early academic work and continued to see them while working on NICE guidelines.
Below are the recurring issues — and what they taught me.
1. Vague or overly broad eligibility criteria
In some of my early academic reviews — for example, my systematic reviews on:
- Primary care efficiency measurement (Pelone 2015)
- Economic impact of childhood obesity (Pelone 2012)
—I initially adapted criteria from other reviews with too little tailoring.
The result?
Too many borderline studies, too many uncertainties, and a lot of unnecessary debate.
Lesson learned:
- Write the criteria for your question, not for someone else’s.
- If you cannot explain a criterion in one sentence, it is not ready.
2. Screening fatigue and “drift”
During guideline work, I sometimes screened 3,000–5,000 records for a single evidence review. It is very easy to lose consistency over time — something every reviewer experiences but rarely admits. This “drift” means reviewers unconsciously change how they apply criteria.
Lesson learned:
- Pilot testing before screening, and scheduled “check-ins” during large batches, stop drift before it becomes a problem.
3. Poor tracking of decisions
In some early systematic reviews — especially:
- Interprofessional education (Reeves 2016)
- Diabetes EBM tools (de Belvis 2009)
—I sometimes recorded exclusion reasons in inconsistent ways. Nothing major, but enough to slow down writing the Methods section and reduce clarity.
And during guideline development, lost PDFs or missing notes happened more often than people think.
Lesson learned:
- Record everything — clearly, consistently, and immediately.
- Future-you will thank present-you.
4. Skipping supplementary searches
Time pressure is constant in guideline settings.
The risk: skipping citation chasing, related-article searches, or grey literature.
But I have seen snowballing produce 5–15% of the final included studies — especially in:
- qualitative syntheses (e.g. Richardson 2016)
- systematic reviews of economic evaluations (e.g. Pelone 2022)
Lesson learned:
- What feels optional usually finds key papers.
5. Inadequate reporting
Sometimes teams forget to report:
- how many reviewers screened each stage
- how disagreements were resolved
- how many records came from snowballing
- which lists/databases were used for citation chasing
Both the PRISMA 2020 statement and Cochrane are clear: these details must be reported.
Lesson learned:
- If it happened, report it.
- If you cannot report it, it probably was not done properly.
After 6+ years of guideline work, several academic reviews and many mistakes, my study selection workflow goes through 7 steps:
- Define eligibility criteria early
- Use a pilot screening round
- Screen in pairs, independently
- Use tools to reduce errors
- Always run snowballing
- Document everything
- Complete PRISMA reporting
This approach has saved time, reduced errors, and strengthened every review I’ve worked on.
Study screening is not just a methodological step.
Most screening problems are preventable — but only if you catch them early. And the tension in academia and guideline work remains the same:
- speed is rewarded
- quality is required
Good screening sits between the two. By combining clear criteria, calibration, dual screening, full-text review, and snowball searches, we build reviews that are complete, defensible, and trustworthy.