Guidelines for Reporting Reliability and Agreement Studies (GRRAS) were proposed
Journal of Clinical EpidemiologyPublished 18 June 2010
Jan Kottner, Laurent Audigé, Stig Brorson, Allan Donner, Byron Gajewski, Asbjørn Hróbjartsson
Citations2,097
SJR quartileQ1
SJR score3.15
SNIP2.66
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The objective was to develop guidelines for reporting reliability and agreement studies and the proposed guidelines intend to improve the quality of reporting.
Abstract
The proposed guidelines intend to improve the quality of reporting.
Keywords
Decision SciencesMedicine
BiometricsThe Measurement of Observer Agreement for Categorical Data
78,934 Citations1977J. Richard Landis, Gary G. Koch
A general statistical methodology for the analysis of multivariate categorical data arising from observer reliability studies is presented and tests for interobserver bias are presented in terms of first-order marginal homogeneity and measures of interob server agreement are developed as generalized kappa-type statistics.
Psychological BulletinIntraclass correlations: Uses in assessing rater reliability.
23,020 Citations1979Patrick E. Shrout, Joseph L. Fleiss
Wiley series in probability and statisticsStatistical Methods for Rates and Proportions
18,048 Citations2003Joseph L. Fleiss, Bruce Levin +1 more
Journal of Clinical EpidemiologyQuality criteria were proposed for measurement properties of health status questionnaires
10,786 Citations2006Caroline B. Terwee, Sandra D.M. Bot +6 more
The criteria can be used in systematic reviews of health status questionnaires, to detect shortcomings and gaps in knowledge of measurement properties, and to design validation studies.
Health Measurement Scales: A Practical Guide to Their Development and Use
9,651 Citations1989David L. Streiner, Geoffrey R. Norman +1 more
This chapter discusses the design of items, their properties, and methods of administration, and some of the considerations that went into designing the items.
Statistical Methods in Medical ResearchMeasuring agreement in method comparison studies
8,740 Citations1999John M. Bland, Douglas G. Altman
The 95% limits of agreement, estimated by mean difference 1.96 standard deviation of the differences, provide an interval within which 95% of differences between measurements by the two methods are expected to lie.
Nursing Research - Generating And Assessing Evidence For Nursing Practice
8,521 Citations2016Denise F. Polit, Cheryl Tatano Beck
Integrating Research Evidence: Meta-Analysis and Metasynthesis 26: Disseminating Evidence: Reporting Research Findings 27: Writing Proposals to Generate Evidence Methodologic References Glossary.
Psychological MethodsForming inferences about some intraclass correlation coefficients.
6,685 Citations1996Kenneth O. McGraw, S. P. Wong
Nurse Education in PracticeNursing Research Generating and Assessing Evidence for Nursing Practice
5,593 Citations2013Fiona Timmins
ScienceOn the Theory of Scales of Measurement
4,682 Citations1946S. S. Stevens
The current issues will remain at 32 pages until a more adequate supply of paper is assured, due to a shortage of paper for Bacto-Agar research.
Journal of Clinical EpidemiologyHigh agreement but low Kappa: I. the problems of two paradoxes
2,886 Citations1990Alvan R. Feinstein, Domenic V. Cicchetti
In a fourfold table showing binary agreement of two observers, the observed proportion of agreement, p0, can be paradoxically altered by the chance-corrected ratio that creates kappa as an index of concordance.
Statistics in MedicineSample size and optimal designs for reliability studies
2,174 Citations1998Stephen D. Walter, Michael Eliasziw +1 more
A functional approximation to earlier exact results is shown to have excellent agreement with the exact results and one can use it easily without intensive numerical computation.
American Journal of Public HealthConsensus methods: characteristics and guidelines for use.
2,131 Citations1984Alexander Fink, Jacqueline Kosecoff +2 more
The characteristics of several major methods (Delphi, Nominal Group, and models developed by the National Institutes of Health and Glaser) are surveyed and guidelines for those who want to use the techniques are provided.
BMJConfidence intervals rather than P values: estimation rather than hypothesis testing.
2,038 Citations1986Michael Gardner, Douglas G. Altman
Some methods of calculating confidence intervals for means and differences between means are given, with similar information for proportions, and the paper also gives suggestions for graphical display.
British Journal of Mathematical and Statistical PsychologyComputing inter‐rater reliability and its variance in the presence of high agreement
1,932 Citations2006Kilem L. Gwet
This paper explores the origin of these limitations, and introduces an alternative and more stable agreement coefficient referred to as the AC1 coefficient, and proposes new variance estimators for the multiple-rater generalized pi and AC1 statistics, whose validity does not depend upon the hypothesis of independence between raters.
Journal of Clinical EpidemiologyWhen to use agreement versus reliability measures
1,727 Citations2006Henrica C. W. de Vet, Caroline B. Terwee +2 more
If the research question concerns the distinction of persons, reliability parameters are the most appropriate, but if the aim is to measure change in health status, which is often the case in clinical practice, parameters of agreement are preferred.
Journal of Chronic DiseasesA methodological framework for assessing health indices
1,563 Citations1985Bram Kirshner, Gordon Guyatt
This work explores the implications of index purpose for each stage of instrument development: selection of the item pool, item scaling, item reduction, determination of reliability, of validity, and of responsiveness.
Epidemiology: Beyond the Basics
1,457 Citations1999Moysés Szklo, F. Javier Nieto
Pure Amsterdam UMCThe STARD statement for reporting studies of diagnostic accuracy: explanation and elaboration
1,233 Citations2003
Statistical Methods in Medical ResearchMeasurement reliability and agreement in psychiatry
1,173 Citations1998Patrick E. Shrout
Psychiatric research has benefited from attention to measurement theories of reliability, and reliability/agreement statistics for psychopathology ratings and diagnoses are regularly reported in empirical reports.
Medical TeacherThe Delphi technique in health sciences education research
951 Citations2005Marietjie de Villiers, Pierre J.T. De Villiers +1 more
The authors used the Delphi technique to assist with making recommendations regarding education and training for medical practitioners working in district hospitals in South Africa to obtain consensus opinion on content and methods relating to the maintenance of competence of these doctors.
Statistics in MedicineSample size requirements for estimating intraclass correlations with desired precision
944 Citations2002Douglas G. Bonett
A method is developed to calculate the approximate number of subjects required to obtain an exact confidence interval of desired width for certain types of intraclass correlations in one-way and two-way ANOVA models.
American Journal of EpidemiologyMISINTERPRETATION AND MISUSE OF THE KAPPA STATISTIC
816 Citations1987Malcolm Maclure, WC Willett
Statistics in MedicineA critical discussion of intraclass correlation coefficients
719 Citations1994Reinhold Müller, Petra Büttner
Statistics in MedicineSample size requirements for reliability studies
701 Citations1987Allan Donner, Michael Eliasziw
This paper provides exact power contours to guide the planning of reliability studies, where the parameter of interest is the coefficient of intraclass correlation rho derived from a one-way analysis of variance model.
Computers in Biology and MedicineA note on the use of the intraclass correlation coefficient in the evaluation of agreement between two methods of measurement
679 Citations1990Martin Bland, Douglas G. Altman
It is shown that neither technique is appropriate for assessing the interchangeability of measurement methods, and an alternative approach based on estimation of the mean and standard deviation of differences between measurements by the two methods is described.
Annals of the Rheumatic DiseasesClinimetric evaluation of shoulder disability questionnaires: a systematic review of the literature
539 Citations2004Sandra D.M. Bot, C B Terwee +4 more
The DASH, SPADI, and ASES have been studied most extensively, and yet even published validation studies of these instruments have limitations in study design, sample sizes, or evidence for dimensionality.
Statistics in MedicineKappa coefficients in medical research
521 Citations2002Helena C. Kraemer, Vyjeyanthi S. Periyakoil +1 more
Development and definitions of the K (categories) by M (ratings) kappas (K×M) are recapitulate, and what they are well‐ or ill‐designed to do are discussed, and where k appas now stand with regard to their application in medical research are summarized.
Scandinavian Journal of Work Environment & HealthCommentary
511 Citations2000Gustav Wickström, Theo. Bendix
Instead of referring to the ambiguous and disputable Hawthorne effect when evaluating intervention effectiveness, researchers should introduce specific psychological and social variables that may have affected the outcome under study but were not monitored during the project, along with the possible effect on the observed results.
Computers in Biology and MedicineStatistical evaluation of agreement between two methods for measuring a quantitative variable
472 Citations1989James Lee, David Koh +1 more
It is suggested that two methods for measuring a quantitative variable can be judged interchangeable provided all of the following conditions are met: first the methods must not exhibit marked additive or nonadditive systematic bias; second the difference between the two mean readings is not "statistically significant"; third, the lower limit of the 95% confidence interval of the intraclass correlation is at least 0.75.
Statistical Methods in Medical ResearchSample size requirements for the design of reliability study: review and new results
468 Citations2004Mohamed M. Shoukri, Musa Hakan Asyalı +1 more
The reliability of continuous or binary outcome measures is usually assessed by estimation of the intraclass correlation coefficient (ICC), and the optimal allocation for the number of subjects k and thenumber of repeated measurements n that minimize the variance of the estimated ICC is discussed.
Educational and Psychological MeasurementReliability Generalization: Exploring Variance in Measurement Error Affecting Score Reliability Across Studies
446 Citations1998Tammi Vacha‐Haase
Physical TherapyUse of the Standard Error as a Reliability Index of Interest: An Applied Example Using Elbow Flexor Strength Data
426 Citations1997Paul W. Stratford, Charlie H. Goldsmith
Using actual elbow flexor make and break strength measurements, this article illustrates a method for estimating a confidence interval for the SEM, shows how an a priori specification of confidence interval width can be used to estimate sample size, and provides several approaches for comparing error variances.
Statistics in MedicineAssessing intrarater, interrater and test–retest reliability of continuous measurements
423 Citations2002Valentin Rousson, Théo Gasser +1 more
It is argued that the usual notion of product-moment correlation is well adapted in a test-retest situation, whereas the concept of intraclass correlation should be used for intrarater and interrater reliability.
Age and AgeingReliability of the Barthel Index when used with older people
378 Citations2005Anita Sainsbury, Gudrun Seebass +2 more
There was evidence that the BI might be less reliable in patients with cognitive impairment and when scores obtained by patient interview are compared with patient testing, and there remain important uncertainties concerning its reliability when used with older people.
Journal of Clinical and Experimental NeuropsychologyMethodological Commentary The Precision of Reliability and Validity Estimates Re-Visited: Distinguishing Between Clinical and Statistical Significance of Sample Size Requirements
357 Citations2001Domenic V. Cicchetti
This critique will stress the inappropriateness of considering precision solely in the context of increasing N, or using sample sizes of 400 and more, as appears to be Charter's main objective or desideratum.
PsychometrikaRamifications of a Population Model for <i>κ</i> as a Coefficient of Reliability
295 Citations1979Helena C. Kraemer
Statistics in MedicineA goodness‐of‐fit approach to inference procedures for the kappa statistic: Confidence interval construction, significance‐testing and sample size estimation
258 Citations1992Allan Donner, Michael Eliasziw
A new procedure for constructing a confidence interval about the kappa statistic in the case of two raters and a dichotomous outcome is proposed, based on a chi-square goodness-of-fit test as applied to a model frequently used for clustered binary data.
Statistics in MedicinePlanning a reproducibility study: how many subjects and how many replicates per subject for an expected width of the 95 per cent confidence interval of the intraclass correlation coefficient
248 Citations2001Bruno Giraudeau, J. Y. Mary
An approximation of the expected width of the 95 per cent confidence interval of the ICC is derived and is shown to be of good accuracy and can therefore be used reliably in reproducibility studies.
Developmental Medicine & Child NeurologyGoal attainment scaling in paediatric rehabilitation: a critical review of the literature
240 Citations2007Duco Steenbeek, Marjolijn Ketelaar +2 more
The literature supports promising qualities of GAS in paediatric rehabilitation, and GAS is a responsive method for individual goal setting and for treatment evaluation, but current knowledge about its reliability when used with children is insufficient.
The Journal of Nervous and Mental DiseaseRating Scales, Scales of Measurement, Issues of Reliability
232 Citations2006Domenic V. Cicchetti, Richard A. Bronen +5 more
These issues are the critical reassessment of S. S. Stevens’ quadripartite conceptualization of scales of measurement; the application of criteria to determine the clinical significance of reliability estimates; the detection of subsets of reliable and unreliable raters, when the overall level is of little clinical import.
Annals of The Royal College of Surgeons of EnglandMedical Statistics: A Guide to Data Analysis and Critical Appraisal
226 Citations2006Jim Lewsey
Journal of Orthopaedic TraumaA Concept for the Validation of Fracture Classifications
210 Citations2005Laurent Audig�, Mohit Bhandari +2 more
The classification of fractures from an epidemiological and clinical decision-making perspective is discussed and a standardized methodological concept for their development and scientific validation is proposed.
BMJUse of consensus development to establish national research priorities in critical care
203 Citations2000Keryn Vella
A nominal group technique is feasible and reliable for determining research priorities among clinicians and suggests that clinicians perceive research into the best ways of delivering and organising services as a high priority.
Journal of Manipulative and Physiological TherapeuticsManual Examination of the Spine: A Systematic Critical Literature Review of Reproducibility
189 Citations2006Mette Jensen Stochkendahl, Henrik Wulff Christensen +6 more
Critically analyzes the literature pertaining to the inter- and intraobserver reproducibility of spinal palpation to investigate the consistency of study results and assess the level of evidence for reproducecibility.
Acta Orthopaedica ScandinavicaHow reliable are reliability studies of fracture classifications?A systematic review of their methodologies
179 Citations2004Laurent Audigé, Mohit Bhandari +1 more
A search in MEDLINE and EMBASE for fracture classi- fication reliability studies found a wide variation of methodologies, and investigators should consider alternative methods that focus upon the accuracy of the classification systems.
Journal of Clinical EpidemiologyThe dependence of Cohen's kappa on the prevalence does not matter
178 Citations2005Werner Vach
Journal of Psychopathology and Behavioral AssessmentMeasures of interobserver agreement: Calculation formulas and distribution effects
178 Citations1981Alvin E. House, Betty J. House +1 more
StrokeA Reappraisal of Reliability and Validity Studies in Stroke
174 Citations1996L D'Olhaberriague, Irene Litvan +2 more
The identification of the most reliable stroke classifications and scales should encourage their use in selection of homogeneous populations of patients for clinical research studies and to improve communication among scientists.
Journal of Clinical NursingInter‐rater reliability of the EPUAP pressure ulcer classification system using photographs
169 Citations2004Tom Defloor, Lisette Schoonhoven
The inter-rater reliability of the European Pressure Ulcer Advisory Panel classification appears to be good for the assessment of photographs by experts and can be used as a practice instrument to learn to discern pressure ulcer from incontinence lesions and to get to know the different grades of pressure ulcers.
Journal of Manipulative and Physiological TherapeuticsAre chiropractic tests for the lumbo-pelvic spine reliable and valid? A systematic critical literature review
166 Citations2000Lise Hestœk, Charlotte Leboeuf‐Yde
The detection of the manipulative lesion in the lumbo-pelvic spine depends on valid and reliable tests, and because such tests have not been established, the presence of the manipulation lesion remains hypothetical.
Journal of Manipulative and Physiological TherapeuticsIntertester Reliability and Diagnostic Validity of the Cervical Flexion-Rotation Test
161 Citations2008Toby Hall, Kim Robinson +3 more
The cervical flexion-rotation test can be used accurately and reliably by inexperienced examiners and may be a useful aid in CeH evaluation.
WorkReliability of work-related assessments
161 Citations1999Ev Innes, Leon Straker
It is indicated that a number of commercially available work-related assessments have insufficient evidence of reliability, and clinicians will be able to examine their options with regard to the reliability of the assessments they choose to use.
Journal of General Internal MedicineHow reliable are assessments of clinical teaching?
153 Citations2004Thomas J. Beckman, Amit Kumar Ghosh +3 more
Characteristics of teacher evaluations vary between educational settings and between different learner levels, indicating that future studies should utilize more narrowly defined study populations.
Journal of Pediatric OrthopaedicsDevelopment and Validation of the AO Pediatric Comprehensive Classification of Long Bone Fractures by the Pediatric Expert Group of the AO Foundation in Collaboration With AO Clinical Investigation and Documentation and the International Association for Pediatric Traumatology
142 Citations2005Theddy Slongo, Laurent Audigé +3 more
Disagreement and misclassification of fractures were overall very low; hence, experienced and trained surgeons can classify pediatric long bone fractures using the proposed system with high accuracy based on standard radiographic views.
PubMedThe development of a national registration form to measure the prevalence of pressure ulcers in The Netherlands.
136 Citations1999G J Bours, Ruud J.G. Halfens +2 more
The pilot study concluded that it is possible to collect accurate and reliable data on the scope and severity of pressure ulcers with a uniform instrument in different healthcare settings.
Journal of Psychopathology and Behavioral AssessmentMeasures of interobserver agreement: Calculation formulas and distribution effects
129 Citations1982Alvin E. House
Measures of interobserver agreement
127 Citations2004M. M. Shoukri
Journal of Health Services Research & PolicyA comparison of formal consensus methods used for developing clinical guidelines
126 Citations2006Andrew Hutchings, Rosalind Raine +2 more
The advantages of nominal groups (more consensus; greater understanding of reasons for disagreement) could be combined with the greater reliability of the Delphi approach by developing a hybrid method.
International Archives of Occupational and Environmental HealthReliability and validity of Functional Capacity Evaluation methods: a systematic review with reference to Blankenship system, Ergos work simulator, Ergo-Kit and Isernhagen work system
118 Citations2004Vincent Gouttebarge, Haije Wind +2 more
Statistical Evaluation of Measurement Errors: Design and Analysis of Reliability Studies
116 Citations2004Graham Dunn
This chapter discusses methods for categorical (binary) data and sources of variation, as well as method comparison 1 and 2, which compare the methods for paired observation and categorical data.
Statistical Methods in Medical ResearchMeasurement of reliability for categorical data in medical research
113 Citations1992Helena C. Kraemer
The problem of measuring reliability of categorical measurements, particularly diagnostic categorizations, is addressed and a general model is proposed, leading to definition of reliability indices.
Clinical RheumatologyReliability and validity of clinical outcome measurements of osteoarthritis of the hip and knee — A review of the literature
111 Citations1997Y. Sun, Til Stürmer +2 more
Overall, knowledge on reliability and validity of clinical scores of hip and knee osteoarthritis is limited, underlining the need for further properly designed and conducted studies.
Biological PsychiatryPenny-wise and pound-foolish: the impact of measurement error on sample size requirements in clinical trials
108 Citations2000Diana O. Perkins, R J Wyatt +1 more
The relationship between reliability and sample-size requirements is model and the potential tangible cost savings resulting from the decreased number of subjects needed when reliability of raters is improved or multiple ratings are used are considered.
The LancetWhy we need large, simple studies of the clinical examination: the problem and a proposed solution
100 Citations1999Finlay A. McAlister, Sharon E. Straus +1 more
Since the results of the initial clinical assessment allow us to modify the authors' pretest calculations of probability of disease in patients and to tailor subsequent investigations appropriately, a thoughtful examination can improve efficiency and lower the costs of care.
Australian Journal of StatisticsCategory Distinguishability and Observer Agreement
97 Citations1986J. N. Darroch, P. McCloud
Journal of Clinical EpidemiologyThe statistical analysis of kappa statistics in multiple samples
93 Citations1996Allan Donner, Neil Klar
Partitioning methods allow a variety of hypotheses to be tested, including an assessment of the degree of agreement within each sample, a testing procedure based on the pooled data, and a test of heterogeneity that may be used to assess the validity of pooling across samples.
Statistics in MedicineEffective number of subjects and number of raters for inter‐rater reliability studies
88 Citations2005Yuki Saito, Takashi Sozu +2 more
A reliability study in which multiple raters evaluate multiple subjects was assumed in order to confirm the inter-rater reliability of rating scales, and it was concluded that the use of identical numbers of raters and subjects minimizes variance when the relative ratio is substantially large and variance is maximized.
Journal of Clinical EpidemiologyThe inter-rater agreement of retrospective assessments of adverse events does not improve with two reviewers per patient record
85 Citations2009Marieke Zegers, Martine C. de Bruijne +4 more
A record review process with two physicians per record including a consensus procedure to assess AEs is not more reliable than a record reviewprocess with one physician.
Medical CarePeer Review of Medical Care
85 Citations1972Fred MacD. Richardson
It was shown that individual judges consistently differed in the degree of harshness or permissiveness of their respective judgments, and that the number of independent judges required to reach a stable judgment of care quality exceeded the number logistically available to meet the probable future demands of third-party payors.
Journal of Clinical EpidemiologyClinimetrics vs. psychometrics: an unnecessary distinction
79 Citations2003David L. Streiner
It is shown that the clinimetric approach is neither new nor unique, but is rather a subset of psychometrics, and use of the term "clinimetric" cuts people off from a rich source of information.
Journal of Nursing Care QualityReliability Testing of the National Database of Nursing Quality Indicators Pressure Ulcer Indicator
78 Citations2006Sara Hart, Sandra Bergquist +2 more
Findings suggest that nurses can accurately differentiate pressure ulcer from other ulcerous wounds in Web-based photographs, reliably stage pressure ulcers, and reliably identify community versus nosocomial pressure Ulcers.
PubMedPressure ulcer prevalence, incidence and associated risk factors in the community.
74 Citations1993Oot-Giromini Ba
This study examined the prevalence and incidence of pressure ulcers as well as associated risk factors in the community by using the Web of Causation and the Braden conceptual schema, and barriers to effective interventions were identified and analyzed.
SpineComparison of Observer Variation in Conventional and Three Digital Radiographic Methods Used in the Evaluation of Patients With Adolescent Idiopathic Scoliosis
70 Citations2008James M. Mok, Sigurd Berven +4 more
The results suggest that different observers will obtain similar measurements when viewing the same image, but care should be taken when interpreting images printed on 2 unstitched films.
Methods of Information in MedicineThe Kappa Coefficient and the Prevalence of a Diagnosis
70 Citations1988T Gjørup
The kappa coefficient is a widely used measure of agreement between observers’ independent recording of diagnoses and means that kappa does not give a general statement of the reproducibility of a diagnosis.
Academic MedicineThe Reported Validity and Reliability of Methods for Evaluating Continuing Medical Education: A Systematic Review
69 Citations2008Neda Ratanawongsa, Patricia A. Thomas +9 more
The data indicate that reporting about internal structure validity exceeded reporting about other categories of validity evidence, and Educators should devote more attention to the development and reporting of high-quality CME evaluation methods and to emerging guidelines for establishing the validity of CME evaluated methods.
Archives of Physical Medicine and RehabilitationImpact of quality scales on levels of evidence inferred from a systematic review of exercise therapy and low back pain
69 Citations2002F. Colle, François Rannou +3 more
Two of the 3 main results of the systematic review (conflicting evidence on the effectiveness of exercise therapy compared with inactive treatments; strong evidence that exercise therapy is more effective than usual care by a general practitioner) were influenced by the scale used.
Wound Repair and RegenerationSubepidermal moisture differentiates erythema and stage I pressure ulcers in nursing home residents
68 Citations2008Barbara M. Bates‐Jensen, Heather McCreath +2 more
SEM may assist in predicting early PU damage, allowing for earlier intervention to prevent PUs, and was responsive to visual assessment changes, differentiated between erythema and stage I PU, and higher SEM predicted greater likelihood of ery thema/stage I PU at the sacrum the next week.
BiometricsStatistical Implications of the Choice between a Dichotomous or Continuous Trait in Studies of Interobserver Agreement
64 Citations1994Allan Donner, Michael Eliasziw
The effect of the decision to measure the trait of interest on a continuous or dichotomous scale on the overall number of subjects required to test whether a population reliability coefficient equals a specified criterion value is measured.
International Journal of Nursing StudiesAn interrater reliability study of the Braden scale in two nursing homes
62 Citations2008Jan Kottner, Theo Dassen
Although the calculated interrater reliability coefficients for the total Braden score were high in some cases, several clinically relevant differences occurred between the nurses, and it is doubtful if their assessment contributes to any valid results.
Statistics in MedicineA comparison of methods for calculating a stratified kappa
61 Citations1991William E. Barlow, Mei‐Ying Lai +1 more
BMC Medical ImagingObserver variation in chest radiography of acute lower respiratory infections in children: a systematic review
60 Citations2001George Swingler
A systematic review of agreement between and within observers in the detection of radiographic features of acute lower respiratory infections in children and described the quality of the design and reporting of studies, whether included or excluded from the review.
Journal of Clinical NursingA systematic review of interrater reliability of pressure ulcer classification systems
58 Citations2009Jan Kottner, Kathrin Raeder +2 more
There is at present not enough evidence to recommend a specific pressure ulcer classification system for use in daily practice and on the basis of this review there are no recommendations as to which system is to be given preference.
The Journal of Applied Behavioral ScienceUsing Behavioral Science Strategies for Defining the State-of-the-Art
56 Citations1980Edward M. Glaser
The author describes an innovative, iterative review paradigm for synthesizing the knowledge base of a given subject, leading to a state-of-the-art consensus document regarding "best practice" in that field, namely, chronic obstructive pulmonary diseases (COPD).
International Journal of Nursing StudiesInter- and intrarater reliability of the Waterlow pressure sore risk scale: A systematic review
55 Citations2008Jan Kottner, Theo Dassen +1 more
Empirical evidence is rare regarding reliability and agreement among nurses when using the Waterlow scale in clinical practice and evaluation of the applicability of the waterlow scale to clinical practice are limited.
International Journal of Nursing StudiesInterpreting interrater reliability coefficients of the Braden scale: A discussion paper
53 Citations2007Jan Kottner, Theo Dassen
It is shown that the intraclass correlation coefficient is an appropriate statistical approach for calculating the interrater reliability of the Braden scale and it is recommended to present intrusion correlation coefficients in combination with the overall percentage of agreement.
American Journal of PsychiatryInterrater Reliability in Clinical Trials of Depressive Disorders
52 Citations2002Benoit H. Mulsant, Kari B. Kastango +4 more
Few published reports of clinical trials of treatments for depressive disorders document adequately the number of raters, rater training, assessment of interrater reliability, and rater drift.
Journal of Manipulative and Physiological TherapeuticsAre chiropractic tests for the lumbo-pelvic spine reliable and valid? A systematic critical literature review
50 Citations2000Lise Hestbœk, Charlotte Leboeuf‐Yde
Journal of Tissue ViabilityEPUAP statement on prevalence and incidence monitoring of pressure ulcer occurrence
50 Citations2005Tom Defloor, Michael Clark +5 more
This statement upon pressure ulcer monitoring issued in May 2005 by the European Pressure Ulcer Advisory Panel has not been subject to the peer review process undertaken by this journal.
Journal of Clinical EpidemiologyTraining improves agreement among doctors using the Neer system for proximal humeral fractures in a systematic review
49 Citations2007Stig Brorson, Asbjørn Hróbjartsson
A consistently low level of observer agreement was found among doctors classifying proximal humeral fractures according to the Neer system, and the widely held belief that experts disagree less than nonexperts could not be supported.
Journal of Clinical PsychopharmacologyRater Training in Multicenter Clinical Trials
48 Citations2004Kenneth A. Kobak, Nina Engelhardt +2 more
Despite the lack of empirical data supporting acquisition of a specific set of rater skills, it is believed that the following general skills are essential to conduct a competent clinical interview using clinician-administered symptom rating scales.
Journal of Clinical and Experimental NeuropsychologySample Size Requirements for Increasing the Precision of Reliability Estimates: Problems and Proposed Solutions
44 Citations1999Domenic V. Cocchetti
Journal of Clinical PsychopharmacologyA New Approach to Rater Training and Certification in a Multicenter Clinical Trial
42 Citations2005Kenneth A. Kobak, Joshua D. Lipsitz +3 more
Journal of Clinical NursingInterrater reliability using Modified Norton Scale, Pressure Ulcer Card, Short Form‐Mini Nutritional Assessment by registered and enrolled nurses in clinical practice
41 Citations2008Carina Bååth, Marie‐Louise Hall‐Lord +3 more
The Modified Norton Scale and Short Form Mini-Nutritional Assessment were reasonably understandable and easy to utilize in clinical care and it seems possible for nurses to accomplish assessment using these tools.
The Journal of Trauma: Injury, Infection, and Critical CareShould an Allen Test Be Performed Before Radial Artery Cannulation?
41 Citations2006James E. Barone, Robert Madlinger
Performance of an Allen test before radial artery cannulation should not be considered a "standard of care" and the significance of an equivocal or abnormal test is unclear.
International Journal of Nursing StudiesExamining the validity of pressure ulcer risk assessment scales: a replication study
40 Citations2003Dinah Gould, L.A. Goldstone +2 more
A simulation study was conducted in which clinical nurses were asked to identify the degree of risk experienced by four patients employing the three RASs discussed most frequently in the literature (Norton, Braden and Waterlow Scores), and nurses' clinical judgment agreed much more closely with expert opinion than any of the Rass.
Journal of Wound Ostomy and Continence NursingSensitivity and specificity of the Braden scale in the cardiac surgical population*1, *2
40 Citations2000Linda J. Lewicki
Findings illustrate that optimum prediction of pressure ulcer risk can only be accomplished with reassessments and determination of the Braden cutoff score or scores that are reflective of the patient's changing clinical condition throughout the hospitalization.
Statistics in MedicineInter‐rater reliability of pressure ulcer staging: ordinal probit Bayesian hierarchical model that allows for uncertain rater response
40 Citations2007Byron Gajewski, Sara Hart +2 more
This article describes a method for estimating the inter‐rater reliability of pressure ulcer staging (stages I–IV) from raters in National Database of Nursing Quality Indicators (NDNQI) participating hospitals and allows for an unstageable PU rating to be included in the analysis.
Statistics in MedicineA general goodness‐of‐fit approach for inference procedures concerning the kappa statistic
37 Citations2001Mekibib Altaye, Allan Donner +1 more
A new procedure for constructing inferences for the kappa statistic is proposed that is based on a chi-square goodness-of-fit test as applied to the Dirichlet multinomial model, and is a natural extension of previously proposed procedures that apply to more restricted cases.
…
