Study selection #3: Pilot screening — the calm before the storm!
Why pilot screening matters
Without a pilot screening round, several problems tend to appear later in the review process:
- Sometimes eligibility criteria are too vague because they were drafted quickly or adapted from another review. When reviewers begin screening, they realise the definitions are unclear. A pilot round makes this visible early, while there is still time to adjust the wording.
- Reviewers may also interpret the same criteria differently. Each person brings their own experience, which influences judgment. The pilot round allows the team to compare decisions and clarify how the criteria should be applied in practice.
- Another issue is that screening decisions become inconsistent, especially for borderline studies. Without prior calibration, reviewers may apply rules differently. The pilot round creates shared interpretation and agreed rules for edge cases.
- Finally, when disagreement is discovered too late, reviewers must go back and re-screen all studies. This is inefficient and can slow the entire review – and may be boring and painful!
After defining eligibility criteria and setting up documentation procedures, the next key step is to run a pilot screening round before starting full screening. A pilot screening round means that reviewers use a small sample of citations to practice applying the inclusion/exclusion criteria. The purpose is not to “start screening early,” but to check if reviewers interpret the criteria the same way, and adjust wording or guidance when they do not. When it is done well, it increases consistency, efficiency, e transparency throughout the review.
During screening, two people can read the same title and abstract and still reach different decisions — especially when concepts are broad or unclear. A pilot screening round helps to:
- Check agreement between reviewers early
- Identify unclear eligibility criteria
- Avoid repeated disagreements later
- Clarify how to handle borderline cases
- Save time in full screening
- Most importantly, it prevents systematic errors from being carried into the main screening phase.
How to get it right
A typical pilot screening round includes:
- Selecting a sample of titles/abstracts/full-text articles
- Reviewers screen independently
- Disagreements are compared
- Reasons for decisions are discussed
- Eligibility criteria are clarified or refined
- If needed, a second pilot round is run
TIP: If disagreement remains high after the pilot, the criteria may need rewriting or simplifying.
Pilot screening works best when reviewers use structured screening forms, which make decisions consistent and comparable. A structured screening form usually includes:
- A yes/no question for each eligibility criterion
- A final include/exclude decision
- A dropdown list of exclusion reasons
These forms can be created in:
PRISMA does not prescribe how to run pilot screening, but it emphasizes:
- Transparent documentation
- Consistent application of inclusion/exclusion criteria
- Reporting how many records were excluded, and why
Pilot screening supports these goals by ensuring that decisions are consistent before screening begins.
The Cochrane Handbook recommends pilot screening as part of the standard workflow. Key points include:
- Training reviewers before full screening
- Testing the clarity of eligibility criteria
- Using structured decision rules
- Recording uncertainties and revisiting decisions as a team
My experience: screen slow now, go fast later!
This approach saves time, avoids late disagreements, and produces more consistent screening decisions:
- In my own systematic reviews, pilot screening has consistently reduced confusion, re-work, and disagreement later in the process.
- Using tools like Excel, Rayyan, DistillerSR, and EPPI-Reviewer has helped make pilot screening more structured and efficient, because every decision is logged in a consistent format and differences can be reviewed clearly.
- These tools also help calibrate reviewers, especially when distinctions matter — for example, telling the difference between cost-effectiveness studies e cost analyses in health economic reviews.
- In more complex reviews, such as qualitative syntheses, I often run two or three pilot rounds instead of one. Qualitative eligibility criteria often involve conceptual interpretation rather than simple study design or population rules. Multiple pilot rounds allow the team to discuss meaning, not just classification.
- In short:
- Clinical or economic reviews → usually one pilot round
- Qualitative or theory-based reviews → often two or three pilot rounds
Pilot screening is sometimes seen as a small step, or even a delay. But it shapes how decisions are made throughout the review. When teams skip calibration, disagreements happen late, exclusions are uneven, and justification becomes harder — which can weaken the credibility of the review. Yet in academic environments, there is often pressure to “keep moving” to meet deadlines and publication expectations. The tension is simple: speed is rewarded, quality is required.
Choosing to run a pilot screening round — even a short one — protects the reliability of the review and strengthens trust in the final evidence.
