Scinovex
articleTop 1% cited

Measurement of Observer Agreement

Radiology · 2003 · Vol. 228(2) · pp. 303–308
Harold L. KundelMarcia Polansky

Abstract

Statistical measures are described that are used in diagnostic imaging for expressing observer agreement in regard to categorical data. The measures are used to characterize the reliability of imaging methods and the reproducibility of disease classifications and, occasionally with great care, as the surrogate for accuracy. The review concentrates on the chance-corrected indices, kappa and weighted kappa. Examples from the imaging literature illustrate the method of calculation and the effects of both disease prevalence and the number of rating categories. Other measures of agreement that are used less frequently, including multiple-rater kappa, are referenced and described briefly.

Reliability and Agreement in MeasurementMeta-analysis and systematic reviewsRadiology practices and educationKappaMedicineCategorical variableCohen's kappaReliability (semiconductor)ReproducibilityInter-rater reliabilityStatisticsObserver (physics)Medical physics

MeSH terms

Data Interpretation, StatisticalDiagnostic ImagingHumansReproducibility of ResultsObserver Variation
Citations
1,492
FWCI
52.51
field-weighted impact
References
28
Percentile
100%
vs. same field & year
Citations per year
Cited by
References
High agreement but low Kappa: I. the problems of two paradoxes
Journal of Clinical Epidemiology · 1990 · 2,818 citations
Categorical Data Analysis
Technometrics · 1991 · 9,086 citations
High agreement but low kappa: II. Resolving the paradoxes
Journal of Clinical Epidemiology · 1990 · 1,815 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

Measurement of Observer Agreement · Scinovex