The validity of conclusions drawn from systematic reviews depends fundamentally on the reliability of the results from the studies they synthesize. Over more than six decades, instruments for assessing the risk of bias in randomized controlled trials (RCT) have evolved considerably, from informal five-item checklists applied to general medical research to algorithmically structured tools specifically designed for RCT.
The concept of methodological assessment has also evolved over time. Alongside the development of these instruments, three distinct concepts have been more clearly defined: (1) methodological quality as the application of specific methodological safeguards in the conduct of a study to avoid or reduce systematic errors; (2) risk of bias as the risk that the study result over- or underestimates the true effect; and (3) reporting quality as the completeness and transparency with which a study is described in its scientific publication (Faggion 2023).
This commentary provides an expanded historical account of that evolution, tracing the development of risk of bias assessment from its earliest precursors in the early 1960s through to the launch of the ROBUST-RCT instrument in 2025. We published a shorter version of this history, focusing specifically on the Cochrane Collaboration’s contribution to the field, in the Journal of Evaluation in Clinical Practice (Pacheco & Riera 2025). The present text expands that to include the pre-Cochrane landscape and greater detail on each iterative development.
Throughout this history, a recurring tension emerges: the trade-off between methodological comprehensiveness and practical usability. Each new/updated tool responded to challenges of its predecessor, yet in doing so often introduced new challenges. Understanding this is valuable not only for historians of methodology but for anyone who regularly uses these instruments in systematic reviews.
1. The Pre-Cochrane Era: Early Checklists and Scales
1.1 An Annotated Bibliography of Scales and Checklists for Assessing the Quality of RCT (Moher, 1995)
In 1995, Moher and colleagues published a review (Moher et al. 1995) of different tools to assess the quality of RCT. They identified 25 scales and 9 checklists, defined as follows:
“To be considered as a scale, the construct under consideration should be a continuum, with quantitative units that reflect varying levels of a trait or characteristic. There had to be evidence that the scale was developed to measure quality and each item had to have a numeric score attached to it with an overall summary score. For a checklist to be included the author(s) intentions were not to attach a quantitative score to each of the questions or to have an overall numeric score.”
The review found that while the checklists and scales varied considerably in their structure, item content, and scoring approaches, most shared fundamental weaknesses. Critically, Moher and colleagues concluded that none of the scales available at that time could be recommended without reservation (Moher et al. 1995). This conclusion set the stage for the Cochrane Collaboration’s subsequent efforts to build guidance from first principles rather than adapting existing instruments.
1.2 The First Identified Checklist (Badgley, 1961)
The earliest identified checklist for assessing research methodology in medical studies was published in 1961 by Robin F. Badgley in the Canadian Medical Association Journal (Badgley 1961). The author applied a five-item checklist to 103 articles published in the journal between January and July 1960.
The checklist was not designed exclusively for RCT but introduced a formal and reproducible evaluation framework for study quality assessment. It did not produce a numerical quality score, instead each item was rated qualitatively.
The checklist contained the following questions:
- Are the terms defined in such a way that it is possible to replicate the study?
- Who (or what) does the population (or sample) being studied represent? Are the criteria given by which cases were selected or rejected?
- What type of control group was used in the study?
- If the results were not analyzed statistically, could statistical analysis have provided additional descriptive and analytical measures?
- Are the generalizations induced in the conclusions of the study limited to the findings of the study?
This first checklist preceded the systematic review movement in medicine and was not designed to inform evidence synthesis. However, it created a precedent that research methodology could and should be evaluated systematically, using explicit criteria that could be applied consistently across studies.
1.3 The First Identified Quality Scale for RCT (Chalmers, 1981)
The first quality scale for RCT appeared in 1981 in work by Chalmers and colleagues (Chalmers et al. 1981). This introduced a weighted composite quality index across three domains: trial design and protocol (assigned a weight of 0.60), statistical analysis (0.30), and presentation of results (0.10). Scores for individual items within each domain were summed and divided by the total possible score, producing a normalized composite index between 0 and 1.
This approach offered an appealing simplicity for comparing RCT but had several assumptions that would later attract criticism. The weighting scheme was largely arbitrary and the use of a single composite score meant that serious deficiencies in one domain could be masked by other domains.
1.4 The Jadad Scale (Jadad, 1996)
Among the multiple scales proposed for assessing RCT, the Jadad scale deserves special mention given its widespread use. Formally published in 1996 (Jadad et al., 1996), the scale originally targeted RCT in pain research and organized the quality assessment in three scoring domains: randomization, blinding, and handling of withdrawals. The scale was conceptualized to be a direct and rapid assessment: “It should not take more than 10 minutes to score a report and there are no right or wrong answers” (Jadad et al. 1996).
Jadad scale questions, to be answered as yes/no, with 1 point for each yes:
- Was the study described as randomized (this includes the use of words such as randomly, random, and randomization)?
- Was the study described as double blind?
- Was there a description of withdrawals and dropouts?

Figure 1. Scoring system from the Jadad scale, as described in the Appendix of Jadad et al. (1996).
Considering the instructions in Figure 1, the Jadad scale produces a total score from 0 to 5. In practice, this score was dichotomized and many systematic reviews used a threshold of 3 out of 5 to classify trials as high or low quality.
The scale’s strengths were its simplicity, ease of application, and explicit focus on three domains with established theoretical links to bias.
However, the Jadad scale attracted substantial criticism. The most persistent concern was its conflation of reporting with methodological quality: both the randomization and blinding items award a point for description of the method, independent of whether the conduct itself was robust. A trial that reports an adequate randomization method receives the same score regardless of whether the method was actually implemented correctly. Conversely, a trial with rigorous but unreported procedures scores lower than one that reports inadequate methods.
A second criticism concerned the treatment of blinding. The scale assigns up to two points for double blinding regardless of whether blinding was feasible in the trial context. Trials of surgical interventions, physical therapies, or behavioral treatments, where participant and provider blinding is often infeasible, are systematically penalized relative to pharmacological trials. This introduced a structural bias against certain types of interventions.
A third limitation was the absence of allocation concealment as an explicit domain. As the literature had already noted, concealment of the allocation sequence prior to assignment is the domain with the strongest empirical evidence for an association with effect size inflation. The Jadad scale does not assess this directly.
Despite these limitations, the scale’s brevity and the clarity of its scoring rules contributed to its rapid diffusion during the late 1990s and 2000s, where it was used in thousands of systematic reviews before domain-based tools became the common approach.
2. Early Cochrane Collaboration Guidance (1994–2006)
2.1 The 1994 Toolkit
In March 1994, the Cochrane Collaboration issued its first formal guidance on systematic review methodology, embedded within a section of its Toolkit titled ‘Section VI: Preparing and Maintaining Systematic Reviews’ (Oxman 1994). This guidance drew directly on the work of Moher and colleagues (Moher et al. 1995) to justify a departure from existing scales and checklists.
The guidance organized methodological quality assessment around four ‘categories’ of potential bias, derived from the conceptual model of sources of systematic error in clinical trials (Figure 2).

Figure 2. Criteria for assessing the methodology quality of RCT as described in the Cochrane’s 1994 Toolkit.
Selective reporting of statistically significant results was noted as a potentially relevant consideration in certain contexts, but was not included as a core domain.
The evaluation system produced an overall judgment: low risk of bias (all criteria met), moderate risk of bias (one or more criteria partially met), or high risk of bias (one or more criteria unmet).
A notable feature of this guidance was its explicit caution against complexity. The document stated directly that more elaborate scoring systems had not been shown to produce more reliable assessments and carried a greater risk of conflating reporting quality with actual trial conduct. This position, that simplicity should be preferred when it does not sacrifice validity, would prove both prescient and contested throughout the subsequent three decades of tool development.
The 1994 guidance acknowledged the near absence of empirical evidence supporting most of the criteria proposed. It stated that among all the domains, only adequate allocation concealment had strong empirical support for its association with bias in trial results.
2.2 The Cochrane Collaboration Handbook (1996 and 1999)
In October 1996, Section VI was revised and published as a standalone document: the Cochrane Collaboration Handbook (Mulrow & Oxman, 1996). The terminological shift in this version was notable: the language of ‘methodological quality assessment’ was replaced by ‘critical appraisal,’ signaling a conceptual distinction between evaluating whether a trial was conducted appropriately (quality) and evaluating whether its results were trustworthy.
The four-domain structure was preserved with a terminology revision: ‘exclusion bias’ was renamed ‘attrition bias’. Additionally, the three-tier overall risk of bias rating system (low, moderate, high) remained unchanged.
An interesting feature of the 1996 Handbook is the description of a code for allocation concealment to studies entered into Cochrane’s systematic review software (Review Manager): “When reviewers enter studies into Review Manager (RevMan) they are required to whether allocation concealment was adequate (A), unclear (B), inadequate (C), or that allocation concealment was not used (D) as a criterion to assess validity.”
The 1999 fourth edition of the Handbook maintained these core elements without substantive modification (Clarke & Oxman 1999). Subsequent sub-versions (4.2.1, labeled the ‘Cochrane Reviewers’ Handbook,’ and 4.2.6, the first to hold the current title ‘Cochrane Handbook for Systematic Reviews of Interventions’) introduced minor updates and clarifications but preserved the same conceptual framework (Alderson et al. 2003; Higgins & Green 2006).
3. The Cochrane Risk of Bias Tool (2007–2011)
In June 2007, the Cochrane Collaboration Methods Groups Newsletter (Volume 11) published a preview by Julian Higgins and Doug Altman of a new tool under development since 2005 (Cochrane 2007).
The tool adopted a formal distinction between methodological quality and risk of bias. This distinction implied that a trial could be methodologically sound without necessarily being at high risk of bias for a given outcome, and conversely that a well-conducted trial might still have important sources of bias.
The new tool proposed six domains of assessment:
- Sequence generation (the method used to generate the allocation sequence).
- Allocation concealment (the method used to conceal the random allocation sequence).
- Blinding of participants, personnel, and outcome assessors.
- Incomplete outcome data (including attrition and exclusions from analysis).
- Selective outcome reporting.
- Other sources of bias.
Each domain would receive one of three judgments: low risk, high risk, or unknown risk of bias. The intermediate ‘moderate risk’ category of earlier frameworks was eliminated.
The new tool recognized that some sources of bias, particularly blinding and incomplete outcome data, might operate differently for different outcomes within the same trial, leading to a recommendation that these domains be assessed separately for each main outcome. This represented an important shift from trial-level to outcome-level evaluation.
The tool also emphasized the requirement for transparent documentation: for each domain, reviewers were expected to provide a description of what happened in the trial, using verbatim quotes from the report where possible.

Figure 3. Cover of the Cochrane Collaboration Methods Groups Newsletter, Volume 11, June 2007, which announced the development of the Cochrane risk of bias tool.
3.1 Incorporation into the Cochrane Handbook (2008–2011)
In February 2008, the Cochrane Handbook was updated to its fifth version and incorporated the new tool, formally named the ‘Cochrane Risk of Bias Tool’ (Higgins & Green 2008). One terminological change was made from the 2007 preview: ‘unknown risk of bias’ became ‘unclear risk of bias’.
In this first published version, the risk of bias in each domain was judged by answering an overarching question with ‘yes’ or ‘no’, where ‘yes’ indicated low risk. For example, the sequence generation domain was assessed with the question ‘Was the allocation sequence adequately generated?’ This format provided explicit guidance while maintaining a relatively low burden of application.

Figure 4. The Cochrane Collaboration’s tool for assessing risk of bias as published in the Cochrane Handbook for Systematic Reviews of Interventions version 5.0 (2008) (Higgins and Green 2008), showing the six domains and their corresponding directive questions.
In March 2011, version 5.1 of the Handbook introduced a structural change with the blinding domain split into two distinct components (Higgins & Green 2011). ‘Blinding of participants and personnel’ addressed performance bias, the risk that knowledge of treatment allocation affected the behavior of participants or those caring for them. ‘Blinding of outcome assessment’ addressed detection bias, the risk that knowledge of allocation affected how outcomes were measured or recorded. This split was important because it acknowledged that the mechanisms and consequences of these two forms of unblinding are distinct.
At the same time, version 5.1 removed the questions that had guided judgments in the 2008 version. Judgments were now expressed directly as low, high, or unclear risk of bias, without an intermediate question format. The rationale for this change was not fully elaborated in the Handbook, but it appears to have reflected a desire to encourage genuine methodological reasoning rather than a mechanical response to a binary decision.
The Cochrane Risk of Bias Tool was formally published in The BMJ in October 2011 (Higgins et al., 2011), in a paper by Higgins and colleagues. The paper also proposed a method for generating a summary risk of bias judgment for each outcome across all domains:
- Low risk of bias (low risk of bias for all key domains): bias, if present, is unlikely to alter the results seriously.
- Unclear risk of bias (low or unclear risk of bias for all key domains): a risk of bias that raises some doubt about the results.
- High risk of bias (high risk of bias for one or more key domains): bias may alter the results seriously.
Following its 2011 publication in the BMJ, several meta-research studies examined how the tool was being used in practice. Two studies found poor inter-reviewer agreement on domain judgments (Könsgen et al. 2020; Hartling et al. 2013). Another study concluded that risk of bias judgments for RCT included in more than one Cochrane review often differed (Bertizzolo et al. 2019).
Other studies demonstrated that reviewers adopted inconsistent approaches to generating an overall risk of bias summary across domains (Babic et al. 2020) and reported inadequate application across all domains of the tool (Barcot et al. 2019a; Barcot et al. 2019b; Babic et al. 2019a; Babic et al. 2019b; Saric et al. 2019). Jørgensen and colleagues concluded that the tool was often implemented in a non-recommended way (Jørgensen et al. 2016), although another study presented qualitative evidence of positive experiences and perceptions of the tool (Savović et al. 2014). These problems were not confined to Cochrane reviews. Puljak and colleagues found that the tool was used inadequately in most non-Cochrane systematic reviews examined (Puljak et al. 2020). Collectively, these findings suggested that the Cochrane Risk of Bias Tool was not achieving reliable implementation in practice.
4. The Cochrane Risk of Bias Tool 2.0 (RoB 2)
In July 2019, the sixth version of the Cochrane Handbook (Higgins et al. 2019) introduced a substantially revised instrument: the Cochrane Risk of Bias Tool 2.0 (RoB 2) (Sterne et al. 2019). The tool was redesigned to address the documented shortcomings of Cochrane risk of bias table, particularly the problem of inconsistent and inadequately justified judgments.
RoB 2 reorganized assessment around five domains:
- Bias arising from the randomization process.
- Bias due to deviations from intended interventions.
- Bias due to missing outcome data.
- Bias in measurement of outcomes.
- Bias in selection of the reported result.
The ‘other bias’ domain was removed, reflecting a judgment that this catch-all category had been applied inconsistently. Baseline differences between groups, previously handled under ‘other bias’, were incorporated into the domain ‘bias arising from the randomization process’ through the signaling question ‘Did baseline differences between intervention groups suggest a problem with the randomization process?’ This illustrates how RoB 2 uses structured assessment within domains to avoid, for example, the common mistake of judging chance imbalances as high risk of bias.
The domain addressing selective reporting was also reconceptualized to focus specifically on selection of reported results (e.g., reporting a different analysis than pre-specified), rather than the broader question of selective non-reporting of entire outcomes.
Two features from the pre-2011 era were deliberately reintroduced. The first was an intermediate judgment category: ‘some concerns’ replaced the old ‘moderate risk’ designation, acknowledging that the binary low/high structure of the original tool had led reviewers to default disproportionately to ‘unclear’ when evidence was ambiguous.
The second was the return of judgment-aiding questions, but in a substantially more sophisticated form than the yes/no questions of 2008. RoB 2 introduced a first attempt to structure assessments within domains with the concept of signaling questions that could be answered as ‘yes,’ ‘probably yes,’ ‘probably no,’ ‘no,’ or ‘no information’. Responses to signaling questions were incorporated in explicit algorithms to domain-level judgments.
The tool also formally distinguished between two conceptually different questions of interest: the effect of assignment to the intervention (the intention-to-treat effect) and the effect of adhering to the intervention (the per-protocol effect). Domain judgments, particularly for the deviations from intended interventions domain, can differ depending on which effect was being estimated.
4.1 Implementation Challenges
RoB 2 encountered significant implementation difficulties almost immediately after its release. We participated in a RoB 2 workshop at the 25th Cochrane Colloquium in Edinburgh, where a group of experienced methodologists was unable to reach consensus on how to approach the first domain after two hours of discussion. This is illustrative of a wider pattern documented in the literature. Early empirical investigations confirmed poor adherence to RoB 2 in both Cochrane and non-Cochrane reviews ((Martimbianco et al. 2023; Babić et al. 2024). Minozzi and colleagues found low interrater reliability in a prospective study of its application (Minozzi et al. 2020). A study in the Journal of Clinical Epidemiology in 2023 concluded that while RoB 2.0 was useful in principle, it was challenging and resource-intensive to apply, requiring considerably more time and reviewer expertise than its predecessor (Crocker et al. 2023).
5. The ROBUST-RCT (Wang, 2025)
In March 2025, a new instrument was published in The BMJ: the ROBUST-RCT (Risk of Bias Instrument for Use in Systematic Reviews: For Randomized Controlled Trials), developed by Wang and colleagues, including founders and seniors’ members of the GRADE Working Group (Wang et al. 2025). The primary motivation for developing yet another tool, as stated explicitly in the publication, was the call for a return to simplicity and practicability, qualities that the authors argued had been sacrificed in the pursuit of methodological rigor in the development of RoB 2.
ROBUST-RCT is organized around six core domains:
- Random sequence generation.
- Allocation concealment.
- Blinding of participants.
- Blinding of healthcare providers.
- Blinding of outcome assessors.
- Outcome data not included in the analysis.
Each domain receives one of four judgments: low, probably low, probably high, or definitely high risk of bias. This four-level judgment is a notable departure from the three levels of most predecessor tools, introducing a probabilistic gradation that acknowledges uncertainty more explicitly.
The assessment process returns to a format of straightforward directive questions, similar to the 2008 version of the original Cochrane tool, such as ‘Was the allocation sequence adequately generated?’ and ‘Was the allocation adequately concealed?’
ROBUST-RCT also includes optional domains that reviewers may incorporate if deemed relevant, including selective outcome reporting and several items that partially overlap with the ‘other bias’ domain from the original Cochrane tool.
The separation of blinding into three distinct domains (participants, providers, and outcome assessors) allows for a more granular evaluation of performance and detection bias than either the original or the RoB 2 framework.

Figure 5. The ROBUST-RCT, as published in 2025 by Wang et al. (2025), showing the six domains and their corresponding directive questions.
As we noted in the Journal of Evaluation in Clinical Practice (Pacheco and Riera 2025), ROBUST-RCT shares many structural characteristics with the original 2008 version of the Cochrane Risk of Bias Tool. This raises a question that only future research will resolve: whether the simplicity that in theory makes the tool more usable will simultaneously reintroduce the reliability and adequacy problems that motivated the development of RoB 2.
From a practical standpoint, we hypothesize that ROBUST-RCT is likely to achieve a swifter uptake than RoB 2, given its simplified application format. Whether this uptake will be accompanied by adequate application and comprehensive discussion of risk of bias remains to be assessed.
6. Summary and Reflections
The structured assessment of risk of bias of RCT has shifted between simplicity and complexity across successive generations of tools. The 1994 Cochrane framework emphasized simplicity explicitly, having reviewed the evidence and found that more complex instruments did not produce more reliable assessments. The 2019 RoB 2 tool moved substantially toward a more complex and detailed judgment system in pursuit of greater methodological rigor and consistency. The 2025 ROBUST-RCT instrument returned to simplicity in response to documented implementation failures.
The term simplicity, used throughout this period to characterize and justify different tools, deserves attention. Most publications, especially in the case of ROBUST-RCT, seem to use the term as an indicator of usability or time required to apply the instrument, that is, favoring a shorter tool. Simplicity, however, may also be understood as the quality of a tool that is clear and supported by understandable guidance. The distinction matters because reliability is often determined less by how few items an instrument displays than by how consistently assessors can apply its underlying judgments. As we discussed, the structural similarities between the 2008 Cochrane tool and ROBUST-RCT are substantial but whether simplicity alone is sufficient to avoid the reliability problems that the former tool exhibited remains uncertain.
The conceptual framework for understanding bias in RCT has deepened considerably across this period but the tension between developer intent and user implementation has been a persistent theme. Meta-research consistently shows that tools are applied inadequately, even by experienced systematic reviewers. This suggests that the development of tools cannot be evaluated in isolation from the implementation systems within which they operate.
Finally, as we observed in the Journal of Evaluation in Clinical Practice (Pacheco & Riera 2025), the history of these tools raises an important forward-looking challenge: our field may have accumulated more tools without learning to implement any single one reliably. The future requires not only the development of methodologically sound instruments, but sustained attention to how they are implemented, taught, and applied in practice.
Acknowledgements
We acknowledge Mike Clarke, Julian Higgins and Robert Wolff for their contributions during the editorial and peer-review stages. Versions of the Cochrane Handbook ranging from 4.2.1 to the most recent edition are publicly available on the Cochrane Collaboration website (https://www.cochrane.org/authors/handbooks-and-manuals/handbook#previous-versions). Earlier versions, including those from 1994 and 1996, along with the 2007 reference document, were obtained from personal archives and the Library of the Evidence-Based Medicine Discipline at the Universidade Federal de São Paulo (Unifesp).
References
Alderson P, Green S, Higgins JPT (editors) (2003). Cochrane Reviewers’ Handbook 4.2.1. Available from https://www.cochrane.org/authors/handbooks-and-manuals/handbook#previous-versions. Accessed 15 June 2026.
Babić A, Barcot O, Visković T, Šarić F, Kirkovski A, Barun I, Križanac Z, Ananda RA, Fuentes Barreiro YV, Malih N, Dimcea DA, Ordulj J, Weerasekara I, Spezia M, Žuljević MF, Šuto J, Tancredi L, Pijuk A, Sammali S, Iascone V, von Groote T, Poklepović Peričić T, Puljak L (2024). Frequency of use and adequacy of Cochrane risk of bias tool 2 in non-Cochrane systematic reviews published in 2020: Meta-research study. Research Synthesis Methods 15(3):430-440. doi: 10.1002/jrsm.1695.
Babic A, Pijuk A, Brázdilová L, Georgieva Y, Raposo Pereira MA, Poklepovic Pericic T, Puljak L (2019a). The judgement of biases included in the category “other bias” in Cochrane systematic reviews of interventions: a systematic survey. BMC Medical Research Methodology 19(1):77. doi: 10.1186/s12874-019-0718-8.
Babic A, Tokalic R, Amílcar Silva Cunha J, Novak I, Suto J, Vidak M, Miosic I, Vuka I, Poklepovic Pericic T, Puljak L (2019b). Assessments of attrition bias in Cochrane systematic reviews are highly inconsistent and thus hindering trial comparability. BMC Medical Research Methodology 19(1):76. doi: 10.1186/s12874-019-0717-9.
Babic A, Vuka I, Saric F, Proloscic I, Slapnicar E, Cavar J, Pericic TP, Pieper D, Puljak L (2020). Overall bias methods and their use in sensitivity analysis of Cochrane reviews were not consistent. Journal of Clinical Epidemiology 119:57–64. doi: 10.1016/j.jclinepi.2019.11.008.
Badgley RF (1961). An assessment of research methods reported in 103 scientific articles from two Canadian medical journals. Canadian Medical Association Journal 85(5):246-50.
Barcot O, Boric M, Dosenovic S, Poklepovic Pericic T, Cavar M, Puljak L (2019a). Risk of bias assessments for blinding of participants and personnel in Cochrane reviews were frequently inadequate. Journal of Clinical Epidemiology 113:104-113. doi: 10.1016/j.jclinepi.2019.05.012.
Barcot O, Boric M, Poklepovic Pericic T, Cavar M, Dosenovic S, Vuka I, Puljak L (2019b). Risk of bias judgments for random sequence generation in Cochrane systematic reviews were frequently not in line with Cochrane Handbook. BMC Medical Research Methodology 19(1):170. doi: 10.1186/s12874-019-0804-y.
Bertizzolo L, Bossuyt P, Atal I, Ravaud P, Dechartres A (2019). Disagreements in risk of bias assessment for randomised controlled trials included in more than one Cochrane systematic reviews: a research on research study using cross-sectional design. BMJ Open 9(4):e028382. doi: 10.1136/bmjopen-2018-028382.
Chalmers TC, Smith H Jr, Blackburn B, Silverman B, Schroeder B, Reitman D, Ambroz A (1981). A method for assessing the quality of a randomized control trial. Controlled Clinical Trials 2(1):31-49. doi: 10.1016/0197-2456(81)90056-8.
Clarke M, Oxman AD, editors (1999). Cochrane Reviewers’ Handbook 4.0 [updated July 1999]. In: Review Manager (RevMan) [Computer program]. Version 4.0. Oxford: The Cochrane Collaboration.
Cochrane Collaboration Methods Groups Newsletter (2007). Volume 11. Oxford: The Cochrane Collaboration; June 2007. Published by the UK Cochrane Centre, United Kingdom. ISSN 1500-3825.
Crocker TF, Lam N, Jordão M, Brundle C, Prescott M, Forster A, Ensor J, Gladman J, Clegg A (2023). Risk-of-bias assessment using Cochrane’s revised tool for randomized trials (RoB 2) was useful but challenging and resource-intensive: observations from a systematic review. Journal of Clinical Epidemiology 161:39-45. doi: 10.1016/j.jclinepi.2023.06.015.
Faggion CM Jr (2023). Methodological quality, risk of bias, and reporting quality: A confusion persists. Journal of Evidence Based Medicine 16(3):261-263. doi: 10.1111/jebm.12550.
Hartling L, Hamm MP, Milne A, Vandermeer B, Santaguida PL, Ansari M, Tsertsvadze A, Hempel S, Shekelle P, Dryden DM (2013). Testing the risk of bias tool showed low reliability between individual reviewers and across consensus assessments of reviewer pairs. Journal of Clinical Epidemiology 66(9):973-81. doi: 10.1016/j.jclinepi.2012.07.005.
Higgins JP, Altman DG, Gøtzsche PC, Jüni P, Moher D, Oxman AD, Savovic J, Schulz KF, Weeks L, Sterne JA; Cochrane Bias Methods Group; Cochrane Statistical Methods Group (2011). The Cochrane Collaboration’s tool for assessing risk of bias in randomised trials. BMJ 343:d5928. doi: 10.1136/bmj.d5928.
Higgins JPT, Green S (editors) (2006). Cochrane Handbook for Systematic Reviews of Interventions 4.2.6. Available from https://www.cochrane.org/authors/handbooks-and-manuals/handbook#previous-versions. Accessed 15 June 2026.
Higgins JPT, Green S (editors) (2008). Cochrane Handbook for Systematic Reviews of Interventions Version 5.0.0. Available from https://www.cochrane.org/authors/handbooks-and-manuals/handbook#previous-versions. Accessed 15 June 2026.
Higgins JPT, Green S (editors) (2011). Cochrane Handbook for Systematic Reviews of Interventions Version 5.1.0. Available from https://www.cochrane.org/authors/handbooks-and-manuals/handbook#previous-versions. Accessed 15 June 2026.
Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA (editors) (2019). Cochrane Handbook for Systematic Reviews of Interventions version 6.0. Available from https://www.cochrane.org/authors/handbooks-and-manuals/handbook#previous-versions. Accessed 15 June 2026.
Jadad AR, Moore RA, Carroll D, Jenkinson C, Reynolds DJ, Gavaghan DJ, McQuay HJ (1996). Assessing the quality of reports of randomized clinical trials: is blinding necessary? Controlled Clinical Trials 17(1):1-12. doi: 10.1016/0197-2456(95)00134-4.
Jørgensen L, Paludan-Müller AS, Laursen DR, Savović J, Boutron I, Sterne JA, Higgins JP, Hróbjartsson A (2016). Evaluation of the Cochrane tool for assessing risk of bias in randomized clinical trials: overview of published comments and analysis of user practice in Cochrane and non-Cochrane reviews. Systematic Reviews 5:80. doi: 10.1186/s13643-016-0259-8.
Könsgen N, Barcot O, Heß S, Puljak L, Goossen K, Rombey T, Pieper D (2020). Inter-review agreement of risk-of-bias judgments varied in Cochrane reviews. Journal of Clinical Epidemiology 120:25–32. doi: 10.1016/j.jclinepi.2019.12.016.
Martimbianco ALC, Sá KMM, Santos GM, Santos EM, Pacheco RL, Riera R (2023). Most Cochrane systematic reviews and protocols did not adhere to the Cochrane’s risk of bias 2.0 tool. Revista da Associação Médica Brasileira 69(3):469-472. doi: 10.1590/1806-9282.20221593.
Minozzi S, Cinquini M, Gianola S, Gonzalez-Lorenzo M, Banzi R (2020). The revised Cochrane risk of bias tool for randomized trials (RoB 2) showed low interrater reliability and challenges in its application. Journal of Clinical Epidemiology 126:37-44. doi: 10.1016/j.jclinepi.2020.06.015.
Moher D, Jadad AR, Nichol G, Penman M, Tugwell P, Walsh S (1995). Assessing the quality of randomized controlled trials: an annotated bibliography of scales and checklists. Controlled Clinical Trials 16(1):62-73. doi: 10.1016/0197-2456(94)00031-w.
Mulrow CD, Oxman AD, editors (1996). Cochrane Collaboration Handbook [updated 21 October 1996]. The Cochrane Collaboration; Issue 3. Oxford.
Oxman AD, editor (1994). Preparing and maintaining systematic reviews. Section VI of The Cochrane Collaboration Handbook. Oxford.
Pacheco RL, Riera R (2025). Risk of Bias Assessment of Randomized Controlled Trials: From the 1994 Cochrane Collaboration Toolkit to ROBUST-RCT. Journal of Evaluation in Clinical Practice 31(8):e70328. doi: 10.1111/jep.70328.
Puljak L, Ramic I, Arriola Naharro C, Brezova J, Lin YC, Surdila AA, Tomajkova E, Farias Medeiros I, Nikolovska M, Poklepovic Pericic T, Barcot O, Suarez Salvado M (2020). Cochrane risk of bias tool was used inadequately in the majority of non-Cochrane systematic reviews. Journal of Clinical Epidemiology 123:114-119. doi: 10.1016/j.jclinepi.2020.03.019.
Saric F, Barcot O, Puljak L (2019). Risk of bias assessments for selective reporting were inadequate in the majority of Cochrane reviews. Journal of Clinical Epidemiology 112:53-58. doi: 10.1016/j.jclinepi.2019.04.007.
Savović J, Weeks L, Sterne JA, Turner L, Altman DG, Moher D, Higgins JP (2014). Evaluation of the Cochrane Collaboration’s tool for assessing the risk of bias in randomized trials: focus groups, online survey, proposed recommendations and their implementation. Systematic Reviews 3:37. doi: 10.1186/2046-4053-3-37.
Sterne JAC, Savović J, Page MJ, Elbers RG, Blencowe NS, Boutron I, Cates CJ, Cheng HY, Corbett MS, Eldridge SM, Emberson JR, Hernán MA, Hopewell S, Hróbjartsson A, Junqueira DR, Jüni P, Kirkham JJ, Lasserson T, Li T, McAleenan A, Reeves BC, Shepperd S, Shrier I, Stewart LA, Tilling K, White IR, Whiting PF, Higgins JPT (2019). RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ 366:l4898. doi: 10.1136/bmj.l4898.
Wang Y, Keitz S, Briel M, Glasziou P, Brignardello-Petersen R, Siemieniuk RAC, Zeraatkar D, Akl EA, Armijo-Olivo S, Bassler D, Gamble C, Gluud LL, Hutton JL, Letelier LM, Ravaud P, Schulz KF, Torgerson DJ, Guyatt GH (2025). Development of ROBUST-RCT: Risk Of Bias instrument for Use in SysTematic reviews-for Randomised Controlled Trials. BMJ 388:e081199. doi: 10.1136/bmj-2024-081199.
