On the other hand, for trials with few clusters say 10 or 20 per arm , minimum detectable differences become large. So, for example for continuous outcomes, with say 10 clusters per arm and an ICC in the region of 0. For binary outcomes Figure 4 with 10 clusters per arm and ICC in the region of 0. In a real example, a CRCT is to be designed to evaluate the effectiveness of lay support workers to promote breastfeeding initiation and sustainability until 6 weeks postpartum.
Due to fears of contamination, whereby new mothers indivertibly gain access and support from the lay workers, the intervention is to be randomised over cluster units. Cluster randomisation will also ensure that the trial is logistically simpler to run, as randomisation will be carried out at a single point in time, and midwives will have the benefit of remaining in either the intervention or control arm for the duration of the trial.
The cluster units to be used are midwifery teams, which are teams of midwives who visit a set number of primary care general practices to deliver antenatal and postnatal care. The trial is to be carried out within a single primary care trust within the West Midlands. The nature of this design therefore means that the number of clusters available is fixed at the number of midwifery teams delivering care within the region. It was known that 40 clusters are available i.
Estimates of ICC range from 0. Firstly, the feasibility check is implemented to determine whether the 20 available clusters per arm are sufficient to detect the 10 percentage point change assuming the lower estimated ICC 0.
This therefore means that 20 clusters per arm will be sufficient for this design provided an adequate number of individuals are recruited in each cluster. Secondly, the feasibility check is evaluated to determine whether the 20 available clusters per arm is sufficient to detect the 10 percentage point change assuming the higher estimated ICC 0. Therefore, 20 clusters per arm is not a sufficient number of clusters, however many individuals are included within each cluster, to detect the required effect size at the pre-specified power and significance.
Since this latter design is not feasible, formulae at equation 25 allow determination of the minimum detectable difference or maximum achievable power from equation In health care service evaluation cluster RCTs, pre-specifying the numbers of clusters available, are frequently used. That is, trials are designed based on a limited number of cluster units e. GP practices willing or able to participate [ 6 , 7 , 9 , 10 ].
In contrast, sample size methods are almost exclusively based on pre-specified average cluster sizes, as opposed to number of clusters available [ 1 , 4 ]. Whilst mapping sample size formulae from one method to the other is straightforward, a limit on the precision of estimates in such designs leads to a maximum available power that is, a limit on the power available irrespective of how large the clusters are and minimum detectable differences that is, a limit on the difference detectable irrespective of how large the clusters are.
For example, with just 15 clusters available per arm and an ICC of 0. Cluster trials with just 15 clusters available per arm are not uncommon and a 10 percentage point change not an unrealistic goal in many settings. Re-formulation of the problem in terms of minimum detectable difference can thus be used to compare the difference which is statistically detectable at acceptable power levels to that which is clinically, or managerially, important.
Should the situation arise in which the postulated ICC suggests that it is not possible to detect the required difference at pre-specified power , it might be tempting to lower the estimated ICC. Such an approach should be strongly discouraged, since loss of power will most likely result, potentially leading to a non-significant finding [ 12 ].
Rather, formulae here allow sensitivity of the design to be explored in light of possible variations in the ICC. However, other avenues to increase available power might reasonably be considered. For example, it may be plausible to consider relaxing alpha and even to set alpha and beta equivalent [ 17 ]. Or alternatively, incorporating prior information in a Bayesian framework may lead to increases in power.
It might further be argued that studies of limited power are of importance as they contribute to the evidence framework by ultimately becoming part of future systematic reviews [ 18 ], and the methods presented here thus allow for the achievable power to be computed. Before-and-after type studies offer a further avenue of exploration, as by their very nature induce smaller intra-cluster correlations. Methodological limitations of the work presented here include the assumption of equal sized arms; equal standard deviations; Normality assumptions which might not be tenable for small numbers of clusters as well as small numbers of individuals ; and lack of continuity correction for binary variables.
Furthermore, CRCTs with a small number of clusters are controversial, primarily because the small number of units randomised open results to the possibility of bias and approximations to Normality become questionable. However, despite this, CRCTs with a small number of clusters are frequently reported. The Medical Research Council, for instance, has issued guidelines that cluster trials with fewer than 5 clusters per arm are inadvisable [ 19 ]. Others have considered some of the issues involved in community based intervention trials with a small number of clusters, but have focused on issues of restricted randomisation and whether the analysis should be at the individual or cluster level [ 20 ].
Evaluations of health service interventions using CRCTs, are frequently designed with a limited available number of clusters. Sample size formulae for CRCTs, are almost exclusively evaluated as a function of the average cluster size. Where no formal limits exist on the number of individuals enrolled within each cluster, increasing the numbers of individuals leads to a limited increase in the study power.
This in turn means that for a trial with a fixed number of clusters, some designs will not be feasible, and we have provided simple guidelines to evaluate feasibility. A simple rule is that the number of clusters k will be sufficient provided:. For infeasible designs to retain acceptable levels of power, detectable difference might not be as small as desired, leading to the notion of a minimum detectable difference.
Useful aidese memoires are that the detectable difference in a CRCT is that of an individual RCT inflated by the square root of the variance inflation factor; and the power is that under individual randomisation with the standardised effect size deflated by the square root of the variance inflation factor.
Google Scholar. London: Arnold. Computers in Biology and Medicine. Article PubMed Google Scholar. American Journal of Epidemiology. Health Technology Assessment. Article Google Scholar.
International Journal of Epidemiology. Annual Review of Public Health. Part 2. Study design. Qual Saf Health Care. Chapter Google Scholar. Kerry SM, Bland MJ: Sample size in cluster randomised trials: effect of coefficient of variation of cluster size and cluster analysis method. Clinical Trials. New England Journal of Medicine. Medical Research Council: Cluster randomsied trials: methodological and ethical considerations; Yudkin PL, Moher M: Putting theory into practice: a cluster randomised trial with a small number of clusters.
Statistics in Medicine. Download references. Hemming, R. Lilford and A. The authors would like to express their gratitude to Monica Taljaard and Sandra Eldridge for review comments which helped to develop the material.
You can also search for this author in PubMed Google Scholar. Correspondence to Karla Hemming. KH wrote the first and subsequent drafts. AG and RJL helped developed the ideas. All authors read and approved the final manuscript.
This article is published under license to BioMed Central Ltd. Reprints and Permissions. Hemming, K. Sample size calculations for cluster randomised controlled trials with a fixed number of clusters. Download citation. Received : 15 February Accepted : 30 June Published : 30 June Anyone you share the following link with will be able to read this content:.
Sorry, a shareable link is not currently available for this article. Provided by the Springer Nature SharedIt content-sharing initiative. View archived comments 1.
Skip to main content. Search all BMC articles Search. Download PDF. Methods We systematically outline sample size formulae including required number of randomisation units, detectable difference and power for CRCTs with a fixed number of clusters, to provide a concise summary for both binary and continuous outcomes. Introduction Cluster randomised controlled trials CRCTs , in which clusters of individuals are randomised to intervention groups, are frequently used in the evaluation of service delivery interventions, primarily to avoid contamination but also for logistic and economic reasons [ 1 — 3 ].
CRCTs of fixed size: fixed number of clusters each of fixed size Where a CRCT is to be designed with a completely fixed size, that is with a fixed number of clusters, each of a fixed size although this size may vary between clusters , then it is possible to evaluate both the detectable difference and the power, as would be the case in a design using individual randomisation. CRCTs with fixed number of clusters but flexible cluster size Standard sample size formulae for CRCTs, by assuming knowledge of the cluster size m and determining the required number of clusters k , implicitly assume that the number of clusters can be increased as required.
CRCTs with a fixed number of clusters: sample size per cluster The standard sample size formulae for CRCTs assumes knowledge of cluster size m and consequently determines the number of clusters k required. Figure 1. Full size image. Figure 2.
Figure 3. Figure 4. A simple design effect described by Donner, Birkett and Buck 12 can be used for parallel-group trials when the cluster size is assumed constant and the outcome is continuous, binary, count or time-to-event.
Design effects have been derived for more complex designs including: variable cluster sizes; individual level attrition; cross-over trials; stepped-wedge designs; inclusion of baseline measurements; analysis by GEE; and three levels of clustering.
These design effects are relatively straight forward to calculate. However, the opportunity to use them may depend upon the availability and quality of estimates of the parameters required for the calculation. When incorporating variable cluster size, the choice of methods depends upon whether every cluster size is known in advance, or just information on cluster size distribution.
In the case of incorporating stratification, the only method available requires knowledge about the proportion of individuals in the stratum as well as the success probabilities in each, information which is unlikely to be available at the beginning of the trial. The intracluster correlation coefficient featured more frequently as a measure of within-cluster correlation than the coefficient of variation, in our assessment of the sample size literature. The majority of papers specify binary or continuous outcomes; few deal with other types of outcome.
Simple approaches for alternative outcomes data potentially warrant future development. Sample size by simulation is an alternative to using an analytical formula. Although the procedure may be computationally intensive, in some cases it may be preferable to complex numerical procedures and was used in four papers identified in the literature.
However, the type I error is often inflated when the number of clusters is small, the cluster size is variable and for particular analyses such as the frailty model, and this should be taken into consideration during the planning and interpretation of simulations. We have provided a comprehensive description of sample size methodology for cluster randomized trials, presented in a simple way to aid researchers designing future studies.
With the increasing availability of more advanced methods to incorporate the full complexity that can arise in the design of a cluster randomized trial, the researcher may feel overwhelmed by the volume of methods presented. However it should be noted that in some situations a simple formula may perform reasonably well in comparison with a more complex methodology. For example, when the coefficient of variation in cluster size is less than 0.
For continuous outcomes with equal cluster sizes, the cluster-level and individual-level analyses are equivalent.
Therefore a sample size calculation assuming either of these with the same measure of correlation should produce equivalent results. When cluster size is variable, an individual-level analysis is more efficient than a cluster-level analysis weighted by cluster size; therefore a sample size calculation based upon cluster-level analyses will be somewhat conservative if an individual analysis is then conducted.
For binary outcomes, if the intervention is designed to reduce the outcome proportion use of the coefficient of variation 27 will produce marginally smaller sample sizes than using the ICC. When several methods may be used, the choice between them is also a question of practicality. The distribution of the outcome and whether required estimates are available should be considered. Further work is required to formally compare the resulting sample sizes calculated under competing methods, when alternative analyses are conducted, and to evaluate the situations in which the simple methods can provide reasonable results over the more complex.
This was beyond the scope of this paper. A limitation of this paper is that a full critique and comparison of the sample size methods were difficult due to the lack of consistency in reporting across the papers. No guidelines exist at present to judge the quality of methodological papers and guide authors in clear and transparent reporting. We hypothesize that the way in which these methods are reported can also be a barrier to their uptake. We hope that their presentation in this article will improve uptake and research in the performance of these methods.
We are planning further work looking at developing guidelines for the reporting of methodology papers. There is often a large amount of uncertainty associated with the estimate of the ICC, and the appropriateness of any of the methods described here will depend upon the level of uncertainty.
In the case of a large amount of uncertainty, we recommend that at a minimum the sample size sensitivity to a range of ICC values be explored. We recommend that, at the design stage, an appropriate simple formula be used in the first instance to provide the researcher with a benchmark figure upon which the impact of incorporating further complexities can be assessed. We would also like to thank Sally Kerry and two anonymous reviewers for their comments on this paper, which significantly improved its development.
National Center for Biotechnology Information , U. Int J Epidemiol. Published online Jul Author information Article notes Copyright and License information Disclaimer. E-mail: ku. Accepted Jun 2. This article has been cited by other articles in PMC.
Abstract Background: The use of cluster randomized trials CRTs is increasing, along with the variety in their design and analysis.
Keywords: Sample size, cluster randomization, design effect. Key Messages. Introduction Cluster randomized trials In a cluster randomized trial, groups or clusters, rather than individuals, are randomly allocated to intervention groups.
A simple approach to sample size calculation A consequence of clustering is that the information gained is less than that in an individually randomized trial of the same size, making randomization by cluster less efficient. Measuring variability between clusters A key parameter common to all sample size calculations for cluster randomized trials is the extent of similarity between units within a cluster.
Comparison of ICC and coefficient of variation Sample size calculations often make the assumption that the measure of correlation, be it the ICC or k, is the same in each treatment group.
Trial design features that impact on sample size The most common and simplest design choice for a cluster randomized trial is the completely randomized, two-arm parallel-group design with fixed cluster sizes. Results: sample size methods Where possible, sample size formulae have been re-expressed to use consistent terminology for ease in comparability.
Standard parallel-group, two-arm design Continuous and binary outcomes Table 1 summarizes the methodology available for the standard parallel-group trial with equal sized clusters. Table 1. Open in a separate window. Ordinal outcomes A method for correlated ordinal outcomes assuming a GEE analysis has been proposed.
Time-to-event outcomes Methods have been suggested for time-to-event outcomes that adapt the formulae for individual randomization provided by Schoenfeld. Variations to the standard parallel-group design Table 2 provides a summary of all sample size methodology for variations to the standard parallel group trial.
Table 2. Uncertainty around the estimate of the ICC There is often large uncertainty around the estimate of the ICC, leading to wide confidence intervals. Variable cluster sizes The use of the standard design effect assumes that the number of observations from each cluster to be included in the analysis is the same.
Methods that require only the mean and standard deviation of the distribution of cluster size: It is not common to have knowledge about each cluster size at the design stage. Internal pilots For trials that recruit a relatively large number of clusters over a fairly long period of time, it may be appropriate to re-estimate the sample size during the trial once information has been gained on the ICC and other nuisance parameters.
Allocation ratio Design efficiency is maximized with equal allocation to treatment groups, and this has been assumed in the majority of the methodology presented here. Small number of clusters The majority of the methods assume that a relatively large number of clusters is to be recruited, making the approximation to the normal distribution in the formulae appropriate. Equivalence and non-inferiority Non-inferiority and equivalence designs are less commonly used in cluster randomized trials.
Attrition In a cluster randomized trial, individuals within a cluster may withdraw from the trial or an entire cluster may withdraw or not recruit any participants.
Non-compliance Sample size requirements increase as the level of non-compliance increases. Inclusion of baseline measurements Sample size calculations can be adapted to allow covariates in the analysis, as this may increase power by explaining variability and reducing the between-cluster variation, which is particularly important when the number of available clusters is limited or the cost of recruiting each additional cluster is high.
Pre-post design Inclusion of the baseline measurement of the primary outcome into the analysis is referred to as a pre-post design. Inclusion of other covariates Although the inclusion of covariates can reduce the sample size requirements, there are costs associated with taking additional measurements. Inclusion of repeated measurements Multiple time points introduce additional components of correlation, as the observations for each cluster will be correlated over time. Alternative designs The above methods are described for the parallel group trial and small variations to this standard design.
Table 3. Sample size methodology for alternative designs. Stratification and matching Cluster randomized trials in general recruit a smaller number of units than an individually randomized trial. Cross-over designs Cross-over designs require a smaller number of clusters than a parallel-group trial and are therefore useful when the availability of clusters is limited.
Stepped-wedge design The stepped-wedge design is similar to the cross-over design, except that the cross-over of treatments is all in one direction and staggered over time. Three-level cluster randomized trials Additional levels of clustering may occur due to the choice of cluster.
Discussion Sample size calculations for individually randomized trials must be inflated in order to be used for cluster randomized trials, to account for the inefficiency introduced by the correlation of outcomes between members of a cluster. Funding This work was supported by the Medical Research Council.
Supplementary Material Supplementary Data: Click here to view. Conflict of interest: None declared. References 1. Eldridge S, Kerry S. Chichester, UK: Wiley, Donner A, Klar N. Murray D. Design and Analysis of Group-Randomized Trials.
Hayes R, Moulton L. Cluster Randomised Trials. Methods for evaluating area-wide and organisation-based interventions in health and health care: a systematic review.
Health Technol Assess ; 3 : iii— On design considerations and randomization-based inference for community intervention trials. Stat Med ; 15 : — Issues in the design and interpretation of studies to evaluate the impact of community-based interventions. Trop Med Int Health ; 2 : — Campbell MJ.
Cluster randomized trials in general family practice research. Stat Methods Med Res ; 9 : 81— Selected methodological issues in evaluating community-based health promotion and disease prevention programs. Annu Rev Public Health ; 13 — Design and analysis of group-randomized trials: a review of recent methodological developments. Am J Public Health ; 94 — Cornfield J.
Randomization by group: a formal analysis. Am J Epidemiol ; : — Randomization by cluster- sample size requirements and analysis. Statistical considerations in the design and analysis of community intervention trials.
J Clin Epidemiol ; 49 : — Incorporation of clustering effects for the Wilcoxon rank sum test: a large-sample approach. Biometrics ; 59 : — Austin PC. A comparison of the statistical power of different methods for the analysis of cluster randomization trials with binary outcomes. Stat Med ; 26 : — Donner A.
A review of inference procedures for the intraclass correlation-coefficient in the one-way random effects model. Int Stat Rev ; 54 : 67— Estimating intraclass correlation for binary data. Biometrics ; 55 : — Patterns of intra-cluster correlation from primary care research to inform study design and analysis.
J Clin Epidemiol ; 57 : — Determinants of the intracluster correlation coefficient in cluster randomized trials: the case of implementation research. Clin Trials ; 2 — Components of variance and intraclass correlations for the design of community-based surveys and intervention studies: data from the Health Survey for England Intracluster correlation coefficients and coefficients of variation for perinatal outcomes from five cluster-randomised controlled trials in low and middle-income countries: results and methodological implications.
Trials ; 12 : Paediatr Perinat Epidemiol ; 22 : — Parameters to aid in the design and analysis of community trials: intraclass correlations from the Minnesota Heart Health Program. Epidemiology ; 5 : 88— The worksite component of variance: design effects and the Healthy Worker Project.
Health Educ Res ; 8 : — School-level intraclass correlation for physical activity in adolescent girls. Med Sci Sports Exerc ; 36 : — Murray DM, Short B. Intraclass correlation among measures related to alcohol use by young adults: estimates, correlates and applications in intervention studies. J Stud Alcohol ; 56 : — Hayes R, Bennett S. Simple sample size calculation for cluster-randomized trials.
Int J Epidemiol ; 28 : — Developments in cluster randomized trials and statistics in medicine. Stat Med ; 26 : 2— Shih W. Sample size and power calculations for periodontal and other studies with clustered samples using the method of generalized estimating equations. Biometr J ; 39 : — Kerry S, Bland J. Trials which randomize practices II: sample size. Fam Pract ; 15 : 84— Connelly LB. Balancing the number and size of sites: an economic approach to the optimal design of cluster samples.
Control Clin Trials ; 24 : — Hsieh F. Sample-size formulas for intervention studies with the cluster as unit of randomisation. Stat Med ; 7 : — Rosner B, Glynn R. Power and Sample size estimation for the clustered Wilcoxon test. Biometrics ; 67 : — Sample size determination for clustered count data. Stat Med ; 32 : — Sample-size calculations for studies with correlated ordinal outcomes. Stat Med ; 24 : — Campbell M, Walters S. How to design, analyse and report cluster randomised trials in medicine and health related research.
Wiley, Chichester, Whitehead J. Sample size calculations for ordered categorical data. Stat Med ; 12 : — Schoenfeld D. Sample-size formula for the proportional-hazards regression model. Biometrics ; 39 : — Gangnon R, Kosorok M. Sample-size formula for clustered survival data using weighted log-rank statistics. Biometrika ; 91 : — Sample size in cluster-randomized trials with time to event as the primary endpoint.
Byar DP. The design of cancer prevention trials. Recent Results Cancer Res ; : 34— Xie T, Waksman J. Design and sample size estimation in clinical trials with clustered survival times as the primary endpoint. Stat Med ; 22 : — Manatunga A, Chen S. Sample size estimation for survival outcomes in cluster-randomized studies with small cluster sizes. Biometrics ; 56 — Spiegelhalter D. Bayesian methods for cluster randomized trials with continuous responses.
Stat Med ; 20 : — Prior distributions for the intracluster correlation coefficient, based on multiple previous estimates, and their application in cluster randomized trials. Clin Trials ; 2 : — Allowing for imprecision of the intracluster correlation coefficient in the design of cluster randomized trials.
Stat Med ; 23 : — Feng Z, Grizzle JE. Correlated binomial variates: properties of estimator of intraclass correlation and its effect on sample size calculation. Stat Med ; 11 : — Exploratory cluster randomised controlled trial of shared care development for long-term mental illness. Br J Gen Pract ; 54 : — Mukhopadhyay S, Looney S. Quantile dispersion graphs to compare the efficiencies of cluster randomized designs.
J Appl Stat ; 36 : — Sample size for cluster randomized trials: effect of coefficient of variation of cluster size and analysis method. Int J Epidemiol ; 35 : — Unequal cluster sizes for trials in English and Welsh general practice: implications for sample size calculations.
Pan W. Sample size and power calculations with correlated binary data. Control Clin Trials ; 22 : — Liu G, Liang K. Sample size calculations for studies with correlated observations. Biometrics ; 53 : — Sample size calculation for dichotomous outcomes in cluster randomization trials with varying cluster size.
Drug Inform J ; 37 : — Sample size estimation in cluster randomized studies with varying cluster size. Biometr J ; 43 : 75— Relative efficiency of unequal versus equal cluster sizes in cluster randomized and multicentre trials.
Sample size adjustments for varying cluster sizes in cluster randomized trials with binary outcomes analyzed with second-order PQL mixed logistic regression. Stat Med ; 29 : — Sample size re-estimation in cluster randomization trials. Stat Med ; 21 : — Yin G, Shen Y. Adaptive design and estimation in randomized clinical trials with correlated observations. Biometrics ; 61 : — Liu X. Statistical power and optimum sample allocation ratio for treatment and control having unequal costs per unit of randomization.
J Educ Behav Stat ; 28 : — Hoover D. Power for t-test comparisons of unbalanced cluster exposure studies.
J Urban Health ; 79 : — Statistical Methods. Some aspects of the design and analysis of cluster randomization trials. Lui K, Chang K. Test non-inferiority and sample size determination based on the odds ratio under a cluster randomized trial with noncompliance. J Biopharm Stat ; 21 — Accounting for expected attrition in the planning of community intervention trials. Sample size determination for hierarchical longitudinal designs with differential attrition rates. Biometrics ; 63 : — Sample size determination for testing equality in a cluster randomized trial with noncompliance.
J Biopharm Stat ; 21 :1— J Clin Epidemiol ; 62 : — Design effects for binary regression models fitted to dependent data. A simple sample size formula for analysis of covariance in cluster randomized trials. Stat Med ; 31 : — Planning for the appropriate analysis in school-based drug-use prevention studies. J Consult Clin Psychol ; 58 : — The importance and role of intracluster correlations in planning cluster trials. Epidemiology ; 18 : — An integrated population-averaged approach to the design, analysis and sample size determination of cluster-unit trials.
Cohort versus cross-sectional design in large field trials: precision, sample size, and a unifying model. Stat Med ; 13 : 61— McKinlay S. Cost-efficient designs of cluster unit trials. Prev Med ; 23 : — Raudenbush S. Statistical analysis and optimal design for cluster randomized trials.
Psychol Methods ; 2 : — Design issues for experiments in multilevel populations. J Educ Behav Stat ; 25 : — Optimal experimental designs for multilevel logistic models.
Optimal experimental designs for multilevel models with covariates. Commun Stat Theor Stat ; 30 : — Moerbeek M, Maas C. Optimal experimental designs for multilevel logistic models with two binary predictors. Commun Stat Theor Stat ; 34 : — Moerbeek M. Power and money in cluster randomized trials: When is it worth measuring a covariate? Stat Med ; 25 : — Data-analysis and sample size issues in evaluations of community-based health promotion and disease prevention programs — a mixed-model analysis of variance approach.
J Clin Epidemiol ; 44 : — Heo M, Leon A. Sample size requirements to detect an intervention by time interaction in longitudinal cluster randomized clinical trials. Stat Med ; 28 : — Sizing a trial to alter the trajectory of health behaviours: methods, parameter estimates, and their application. Sample size and power determination for clustered repeated measurements. Sta Med ; 21 — Sample size requirement to detect an intervention effect at the end of follow-up in a longitudinal cluster randomized trial.
Stat med ; 29 — Assessing the gain in efficiency due to matching in a community intervention study. Stat Med ; 9 — The merits of breaking the matches: a cautionary tale. The design and analysis of paired cluster randomized trials: An application of meta-analysis techniques. Stat Med ; 16 — Calculation of power for matched pair studies when randomization is by group.
0コメント